Short answer: For a decade, volumetric and Gaussian splat content has been delivered as a large file you download or a bespoke player you install, because it had none of the plumbing that made ordinary video cheap: a standard container, a standard codec, adaptive bitrate, and a CDN that did not need to know what it was carrying. That is what changed in 2026. Splat content is being packaged into MP4, compressed with MPEG work aimed specifically at Gaussian data, and delivered over DASH — the same adaptive protocol behind most streaming video — while a parallel run of research papers has solved the layered, progressively decodable representations that adaptive delivery needs. The commercial consequence is not better pictures. It is that one capture can serve a phone and a headset at different quality tiers, over infrastructure a broadcaster already pays for, at a cost per viewer that can finally be forecast. What has not changed: capture is still expensive, decode still costs the client device more than video does, and a stand demonstration is not proof of internet-scale economics.

The part of video nobody finds interesting is the part that made it cheap.

Ask why you can watch a football match on a phone on a train for a few pounds a month and the answer has almost nothing to do with cameras. It is a stack of unglamorous agreements.

A codec everyone implemented. A container everyone can read. A manifest that lists the same content at several quality levels. A player that measures your bandwidth and quietly moves between them. A CDN that caches the segments without knowing or caring what is in them.

None of that is exciting, and all of it is load-bearing. Between them those pieces turned video from a bespoke delivery project into a line item. Volumetric media has spent most of its life without any of them.

Every volumetric project so far has had to re-solve delivery from scratch. That is why the cost per viewer was never a number anyone could quote you.

The practical symptoms will be familiar to anyone who has commissioned this work. A capture arrives as an enormous file. It plays in one viewer, on one class of device, over a connection nobody tested. The quality is fixed at whatever the producer chose, so a phone on 4G gets the same payload as a laptop on fibre and one of them fails. And when you ask what it would cost to put it in front of fifty thousand people, the honest answer is a shrug.

What actually changed in 2026.

Two things moved at once, and they matter in combination rather than separately.

Standards bodies started treating splats as media. Nokia’s MPEG standardisation work has been aimed at defining an end-to-end system for delivering volumetric video over DASH, the same adaptive protocol that carries most commercial streaming video, alongside the V3C family of coding standards for volumetric content. At IBC 2026 the company listed a streaming demonstration of dynamic 3D Gaussian splats built on standards-compliant components rather than a proprietary stack.

The research caught up on the hard part. Adaptive streaming is not simply a matter of putting a file behind a manifest. It requires a representation that can be decoded usefully at several quality levels, which a naive splat file cannot. That problem has been attacked steadily: layered progressive representations, sliding-window approaches that allow arbitrary clip length, fine-grained scalable encodings, and DASH systems built specifically for multi-layer dynamic splat scenes. The papers in the sources below are worth a skim even if you never read research, because they describe the exact shape of the constraint you will be buying into.

The pipeline that made video cheap camera -> H.264 / HEVC / AV1 -> MP4 segments -> manifest at several bitrates -> ordinary CDN -> DASH / HLS player picks a tier -> anything with a browser What volumetric delivery looked like rig or scan -> proprietary representation -> one enormous file -> direct download or bespoke server -> a specific viewer -> one quality, take it or leave it -> the devices you tested What is now being assembled rig or scan -> Gaussian representation -> MPEG splat coding -> MP4 segments -> manifest at several tiers -> ordinary CDN -> DASH player picks a tier -> phone gets fewer Gaussians, headset gets more

The third diagram is the whole story. It is the same shape as the first one, which is precisely the point: it means the infrastructure is already built, already paid for, and already understood by the people who run it.

Adaptive bitrate is the load-bearing idea, not the container.

It is tempting to read “volumetric in MP4” as the headline. It is not. MP4 is a filing decision. The idea that changes the economics is one capture serving many devices at different weights.

Think about what a fixed-quality volumetric asset forces on you. You must choose, at production time, the worst device and connection you are willing to support, and then every viewer pays that cost — or you choose a higher quality and simply lose the viewers who cannot carry it. Either way one capture equals one audience.

Adaptive delivery breaks that. The same scene can be served as a coarse representation to a phone on a poor connection and a dense one to a headset on good Wi-Fi, from the same origin, chosen automatically, segment by segment. That has three consequences a commercial buyer should care about.

Device reach stops being a production decision. You are no longer picking an audience when you pick a quality setting. This is the difference between a piece of content you can put in an email to everyone and one you can only show at a controlled venue.

Cost per viewer becomes forecastable. Not low, necessarily — forecastable. Bandwidth billed through a CDN at tiers you can model is a fundamentally different procurement object from a bespoke streaming arrangement with a specialist vendor.

Failure becomes degradation rather than collapse. The characteristic failure of fixed-quality 3D on the web is not slowness, it is a blank screen or a dead tab. Adaptive delivery converts that into a softer picture, which is the difference between an experience and an incident.

What this does not fix.

This is where a lot of coverage will overreach, so it is worth being blunt about the parts that are unchanged.

Capture is still the expensive half. Nothing about a delivery standard makes a multi-camera volumetric rig cheaper, a capture volume larger, or a clip longer. If the economics of your project were broken at the shoot, they are still broken.

The client device still has to decode it. A streamed splat is being reconstructed continuously while it is drawn. That is real work on the viewer’s hardware, in addition to rendering, on a device that may also be running a browser with thirty tabs. Video decode has dedicated silicon in every phone sold; splat decode does not. Test on the worst device in your audience.

Memory does not degrade gracefully. Browser 3D does not get slower when it runs out of memory, it dies. Adaptive tiers help, but only if the player is genuinely dropping detail rather than buffering more of it.

A stand demonstration is not proof of scale. This deserves saying plainly, because it is the trap in every emerging-format story. Demonstrating an architecture at a trade show establishes that the pieces fit together. It does not establish what happens at fifty thousand concurrent viewers, what the egress bill looks like, or how the encoder behaves on content nobody curated. Ordinary video took years to go from “this works” to “this is a line item”.

Volumetric media does not need another impressive demo. It needs the boring infrastructure that made ordinary video commercially scalable — and in 2026 those pieces are finally appearing.

The questions worth asking a volumetric vendor now.

The arrival of standards changes what a sensible procurement conversation sounds like. A year ago most of these questions had no good answer, so asking them was pointless. Now the answers are diagnostic.

Ask thisWhat a good answer sounds likeWhat it tells you
What format do the assets leave in?A named, documented representation, ideally one with a published specificationWhether you own the capture or merely rent access to it
Can this be served from our own CDN?Yes, as segments behind a manifestWhether delivery cost is yours to negotiate or theirs to set
How many quality tiers, and how are they produced?Several, generated from one master, automaticallyWhether device reach costs you extra production
What happens on a mid-range phone on 4G?A measured answer with a number in itWhether they have tested outside the studio
What is the decode cost on the client?An honest budget, not a claim that it is negligibleWhether they have shipped to real audiences
If you disappeared tomorrow, what could we still open?Files, in a format with other implementationsYour actual exposure

That last question is the one most worth asking and least often asked. A photoreal capture of a place, a product or a performance is an asset with a long life. Buying one that can only be opened by one company’s player is a choice, and it should be a deliberate one rather than an accident of procurement.

What we would actually do about this today.

Nothing here justifies restructuring a content plan. It justifies changing two things.

Change what you write into contracts. If you are commissioning volumetric or splat capture in the next year, specify the deliverable as files in a documented representation, with the right to serve them yourself. That costs nothing today and is worth a great deal if the standards-based path matures as it looks like it might.

Change what you pilot. The interesting pilot in 2026 is no longer “can we make a volumetric thing”. That is settled. It is “does a volumetric thing survive contact with our actual audience, on their actual devices, over their actual connections”. That is a cheap thing to find out and an expensive thing to discover late.

A sensible first step is one clip, delivered adaptively, measured on a mid-range phone and a headset, with the numbers written down. Our Idea Validation Sprint at £495 exists to settle whether that is even the right question for you, and prototype-tier work starts from £3,250. Both are a great deal cheaper than finding out during a campaign.

Where we actually stand on this.

Worth being exact, because this is a field where showreels routinely imply more delivery experience than sits behind them.

What we run in production: static Gaussian splat capture, cleanup, editing and browser delivery. That includes our own browser splat editor, Simam 3D Studio, a gallery of published scenes you can open without installing anything, and client delivery — including a photoreal capture used as a working wayfinding tool on a site where the drawings had fallen behind reality.

What we deliver rather than produce: volumetric and 360° playback. Simam Immerse carries volumetric titles alongside 360° and 180° ones and streams them to a headset from a link, so the delivery half is something we operate rather than something we have read about.

What we are watching rather than shipping: standards-based adaptive volumetric delivery. We have not built a DASH volumetric pipeline for a client. Everything above is drawn from the published standards work and research listed below, and from the streaming, decode and memory constraints that apply to any large 3D web application — which we deal with constantly.

If someone tells you standards-based volumetric streaming is a solved, de-risked production choice in 2026, ask them for a URL that opens on your phone and holds up. That is the whole test, and it is the one we would want applied to us.

What this means for a buyer.

Start with the business decision, audience, and evidence the project must produce. Simam Digital can turn that into a focused discovery, prototype, MVP, or production roadmap across AI applications, SaaS platforms, digital twins, real-time 3D, XR, and interactive systems.

Sources and further reading