Short answer: 4D Gaussian splatting records a moving real scene so a viewer can move around inside it freely, at photographic quality, in a browser. It adds real value in one specific situation: when the motion is the thing being examined and the viewpoint has to be the viewer’s choice rather than a director’s. That covers technique and movement analysis, procedure training where hands and angles matter, incident and process review, and a narrow band of premium storytelling. Everywhere else something cheaper wins — 360° video when the viewpoint is fixed, real-time 3D when the content must respond or change, and a static 3D splat when the scene does not move. The reason is structural: 4D is a new scene at every instant, so its production, storage and bandwidth costs sit an order of magnitude above everything it competes with. Treat it as a clip medium, pilot it on one clip, and decide with a real file rather than a showreel.

What the fourth dimension actually costs you.

A 3D Gaussian splat is one scene: millions of small, oriented, coloured blobs that together reconstruct a static place with photographic fidelity, renderable in real time and deliverable to a browser. That technology is now ordinary production work.

4D adds time, so the contents can move — people, vehicles, machinery, hands, water. The obvious way to do that is also the wrong way, and understanding why explains the entire commercial picture.

The shape of the problem, not a quotation A static 3D splat scene one scene, prepared once, streamed with LOD chunking -> a download The naive version of 4D a complete new scene for every frame, thirty times a second -> 30x that, per second of runtime ... which is why nobody ships this How practical 4D is actually built one canonical scene, plus a learned deformation field describing how it moves over time -> scene + a much smaller motion field What that means commercially 4D is a CLIP medium. Seconds to a couple of minutes, one subject, one capture volume. It is not an environment you wander for an hour.

Every practical 4D system takes the second shape — a canonical scene with motion described on top of it — because the first is arithmetically hopeless. That single design fact is what you are buying into, and it has three consequences that show up in every real project.

Length is bounded. Deformation fields describe change from a reference state. The further the scene drifts from that reference, the worse and heavier the representation becomes. Short clips are cheap; long continuous capture is not, and often is not possible at acceptable quality.

The capture volume is bounded. Multi-camera rigs define a physical space. Motion outside it is not recorded, which means the subject is a subject, not a world.

Editing is limited. A splat is a reconstruction, not a model. You can crop it, clean it, place it and light it approximately. You cannot restage the performance, change what someone did, or fix a missed angle without recapturing.

If the detail of how these scenes are produced is what you need, we have covered the pipelines and tooling side separately. This article is about whether to pay for it.

The one question that decides it.

Strip away the format debate and one question does almost all the work:

Does the viewer need to choose the viewpoint while the subject is moving?

If the answer is no — if a well-chosen camera, or two, tells the story — you want video. Ordinary video, or 360° video if a sense of place matters. It is an order of magnitude cheaper to produce, universally supported, trivially distributed, and every editor on earth can cut it.

If the subject does not move, you want a static 3D splat. Full six-degree-of-freedom exploration of a real place, at a fraction of the cost, using tooling that is mature today.

If the content needs to respond — be configured, carry live data, have things clicked, change after launch — you want real-time 3D. A splat of any dimension is a recording, and recordings do not take instructions.

4D is the answer only in the remaining corner: real motion, of a real subject, examined from viewpoints the viewer picks. That corner is genuinely valuable and genuinely narrow, and most briefs that arrive asking for volumetric belong in one of the other three boxes. Working out which is a decision worth making deliberately — we have set out the full comparison between volumetric, 360° and real-time 3D in its own guide.

Where it genuinely adds value.

Four areas where that corner is exactly the shape of the problem.

Movement and technique analysis. Sport is the obvious case — a coach circling a swing, a scrum, a landing, choosing the angle the point requires rather than the angle the camera happened to have. The commercial logic is that the alternative is a multi-camera shoot repeated every time somebody asks a new question, and here the same capture answers questions nobody had thought of yet.

Procedural training where hands and angles matter. Some procedures cannot be taught from a fixed camera because the informative view is over the practitioner’s shoulder, or from underneath, or from where a second person would have to stand. Surgical and clinical technique, complex maintenance, assembly sequences with awkward access. When the expert who could demonstrate it is scarce or expensive, capturing them once and letting many learners choose their own view has a clear payback.

Incident and process review. Something happened, and the argument afterwards is about what was visible from where. A free-viewpoint record of the motion answers that in a way a fixed camera cannot. This has obvious application in industrial safety, logistics and operations — and an equally obvious governance requirement, because a photoreal record of identifiable people is personal data and needs to be treated as such from the first day of the project, not the last.

Premium storytelling and heritage performance. Dance, ceremony, craft, a performance in a specific place. This is real, it is beautiful, and it is a marketing budget rather than an operations budget. Judge it on attention and reach like any other campaign asset, and be suspicious of anyone who tells you the format alone will earn the engagement.

Where it does not, and what to use instead.

The five briefs that most often arrive labelled volumetric, and what each one actually wants.

What the brief saysWhat it usually needsWhy
“A volumetric tour of our site or building”A static 3D splatThe building is not moving. You want six degrees of freedom, which a static splat gives you at a fraction of the cost.
“An immersive brand film”360° or 180° videoThe viewpoint is the director’s choice and should be. Cheaper, longer, better looking, and it plays anywhere.
“A volumetric product demonstration”Real-time 3DPeople want to spin it, configure it and see variants. A recording cannot do any of that.
“A virtual presenter or spokesperson”Video, on a plane, in a 3D sceneNobody walks around a presenter. The extra freedom is paid for and unused.
“A volumetric twin of the production line”A real-time 3D twin, with a static splat for contextA twin needs live data and a model underneath. A capture is context, not a system.

None of that is an argument against the technology. It is an argument against buying six degrees of freedom for content nobody moves around in, which is the most common and most expensive mistake in this category.

The delivery constraints that decide feasibility.

A capture that will not stream is a research result, not a product. Four constraints do most of the deciding, and they should be settled before a camera is hired.

Streaming and level of detail. Serious viewers use chunked loading, progressive detail and aggressive culling — the same discipline any large 3D web application needs, and the same one people skip when they are excited about a capture. The difference between a demo and a product is almost always whether it loads on a phone on a normal mobile connection.

The decode budget. A moving splat is being reconstructed continuously while it renders. That is CPU and GPU work happening in addition to drawing, on a device that may also be running a browser, a video call and a corporate security agent. Test on the worst device in your audience, not the best one on the team.

Memory, which does not degrade gracefully. Browser 3D does not slow down when it runs out of memory; it dies. Every decision about capture resolution and clip length is a memory decision before it is a creative one.

Where it will be watched. A controlled venue with good Wi-Fi is a different product from a link sent to a mailing list. Decide which one you are making, because it sets the entire quality budget.

The practical upshot: specify the target device and the acceptable wait before the shoot, then let those two numbers constrain resolution, clip length and capture volume. Doing it the other way round produces a beautiful file nobody can open.

What a sensible pilot looks like.

One clip. Not a format decision, not a content strategy — one clip, scoped to answer three questions with a real file on real devices.

  1. Does the free viewpoint actually change the outcome? Put the same content in front of two groups, one with the 4D version and one with well-shot conventional video. If the people who could move cannot articulate what moving told them, you have your answer, and it is a cheap answer.
  2. Does it load and hold up on the device the audience really has? Measured, on a real network, on a mid-range phone as well as a laptop.
  3. What does the second one cost? The first capture always carries setup, rig time and learning. The number that matters for a programme is the marginal one, and a pilot is the only honest way to establish it.

Scope this deliberately. Our Idea Validation Sprint at £495 is built for settling whether the question is even the right one, and prototype-tier work starts from £2,800. That is the right order of magnitude for finding out, and it is a great deal less than the cost of discovering the same thing halfway through a campaign.

Where we actually stand on this.

Worth being exact, because this is a field where a lot of showreels imply more delivery experience than exists behind them.

What we run in production: static Gaussian splat capture, cleanup, editing and browser delivery. That includes our own browser splat editor, a gallery of published scenes you can open without installing anything, and client delivery — a photoreal capture used as a working wayfinding tool on a large site where the drawings had fallen behind reality.

What we deliver rather than produce: volumetric and 360° media playback. Simam Immerse carries volumetric titles alongside 360° and 180° ones and streams them to a headset from a link, so the delivery half of this is something we run.

What we treat as pilot-stage: the capture half. We have not run a 4D Gaussian splatting capture-to-delivery programme for a client. Everything above about production is drawn from the published research on how these systems are constructed, from the streaming and memory constraints that apply to any large 3D web application — which we deal with constantly — and from what we see arriving in briefs.

If someone tells you 4D volumetric is a routine, de-risked production format in 2026, ask them to send you a URL that opens on your phone. That is the whole test, and it is the same test we would want applied to us.

What this means for a buyer.

Start with the business decision, audience, and evidence the project must produce. Simam Digital can turn that into a focused discovery, prototype, MVP, or production roadmap across AI applications, SaaS platforms, digital twins, real-time 3D, XR, and interactive systems.

Sources and further reading