Spatial audio is half the immersion, and most installations get it wrong

Every immersive room budget over-indexes on pixels and under-indexes on sound. Here's how spatial audio actually works in domes and installations - formats, speaker layouts, sync, and the geometry problems nobody warns you about.

LISTEN
0:00 / -:-
MP3
Studio monitors and a multichannel audio interface staged inside a geodesic projection dome during a spatial audio calibration session

Spatial audio installations are the half of immersion that budgets forget. A dome or a mapped room gets specified projector-first: lumens, blend, resolution, content hours. Sound shows up in the last two weeks as a line item called "AV package" and arrives as a pair of powered speakers pointed at the audience. The result is a room that looks wraparound and sounds flat, and visitors describe the experience as impressive rather than convincing.

That gap is not a taste problem. Localization - the brain's ability to place a sound in space - is fast and involuntary. If a bird flies across the dome overhead and the sound of it stays pinned to two boxes at the front of the room, the visual is contradicted before the visitor has consciously registered anything. Immersion is a consensus between senses, and audio breaks the consensus faster than any other channel.

This post covers what we actually specify on dome and installation projects: the formats worth using, how many speakers a room really needs, why curved rooms are acoustically hostile, and how to keep audio locked to real-time visuals that are not playing back a fixed timeline.

What spatial audio means in an installation context

The term covers at least four distinct technologies that solve different problems. Getting the vocabulary right matters, because vendors use the words loosely and the wrong choice is expensive to reverse after the truss is up.

Channel-based audio
Sound is mixed to a fixed speaker layout such as 5.1 or 7.1. Simple and universally supported, but the mix only works if the playback layout matches the mix layout exactly.
Ambisonics
A full-sphere format that encodes the sound field itself rather than speaker feeds, then decodes to whatever speaker array the room actually has. First-order ambisonics uses 4 channels; order n uses (n+1) squared channels, so third order is 16.
Object-based audio
Each sound is stored as an audio file plus positional metadata, and a renderer computes speaker feeds live. Dolby Atmos, L-Acoustics L-ISA, and d&b audiotechnik Soundscape all work this way.
Wave field synthesis
A dense speaker array physically reconstructs a wavefront so sources have a stable position anywhere in the room rather than at one sweet spot. Highest speaker count and cost of any approach.
Sweet spot
The listening area where the intended spatial image holds together. Every format except wave field synthesis has one, and installation design is largely the practice of making it big enough to cover the audience.

For most installation work the real choice is between ambisonics and object-based rendering. Ambisonics is layout-agnostic, which is exactly what you want when the speaker positions are dictated by a dome frame or a fabrication constraint rather than a standard. It is also cheap to author: the IEM Plugin Suite from the Institute of Electronic Music and Acoustics in Graz is free, runs in any DAW, and handles encoding, rotation, and decoding to arbitrary arrays. Object-based systems are better when the piece has discrete sources that need to track precisely - a voice that follows a projected character across the wall - and when there is budget for a proprietary renderer.

Reference material worth reading before you commit: the IEM Plugin Suite documentation for practical ambisonic workflow, and ambisonic.net for the underlying theory going back to Michael Gerzon's work in the 1970s.

Why curved rooms are acoustically hostile

A projection dome is a concave reflector. The same geometry that makes it a good screen makes it a bad room. Sound reflecting off a dome shell converges toward the center of curvature instead of scattering, which produces audible hot spots, long flutter, and a low-frequency buildup that no amount of EQ will fix because it is a geometry problem, not a frequency-response problem.

Three mitigations, in order of effectiveness. First, put the speakers behind the projection surface, which is standard practice in planetariums: a perforated or micro-woven dome screen passes sound while still holding an image, so speakers sit outside the reflective volume. Second, treat the floor, which is the surface that completes the focusing path and is usually the only one you are allowed to touch in a rented venue. Third, use directional speakers aimed inward and downward so less energy ever reaches the shell in the first place.

The room geometry also sets your delay budget. Sound travels roughly 343 meters per second, or about 3 milliseconds per meter. In a 12-meter dome, a visitor standing near one edge is about 12 meters from the opposite speaker, so that speaker's contribution arrives 35 milliseconds late relative to the near one. That is well past the precedence-effect window, which is why per-speaker delay alignment - not just level matching - is the calibration step that separates a coherent room from a smeared one. We go deeper on the visual side of the same geometry problem in designing for curved surfaces in fulldome.

Choosing a format and a speaker count

The honest version of this decision is a table. Speaker counts below assume a room in the 8 to 15 meter range, which covers most dome and gallery installations.

ApproachTypical speakersLayout flexibilityBest for
Stereo or LCR2 to 3RigidLobby loops, kiosks, anything with a single viewing direction
Channel-based surround6 to 8RigidSeated theaters where the layout is standard and fixed
First-order ambisonics8 to 12HighDomes and irregular rooms on a normal budget
Third-order ambisonics16 to 24HighDomes where overhead precision matters and the budget allows
Object-based rendering12 to 40HighShows with discrete tracked sources; requires a proprietary renderer
Wave field synthesis60+LowPermanent flagship venues only

For a typical dome we specify first-order ambisonics decoded to a ring of 8 speakers at ear height plus 3 to 4 in the upper hemisphere, with a single subwoofer. That configuration gives real overhead imaging, survives being decoded to whatever positions the frame allows, and costs a fraction of an object-based deployment. Third order is the upgrade path, and because ambisonics is layout-agnostic the content does not have to be re-authored when you add speakers.

Getting the signal there

Above 8 channels, analog cabling stops being reasonable. Audio-over-IP is the answer: Dante carries hundreds of channels over standard gigabit Ethernet with sample-accurate clocking, and AES67 provides interoperability when the room mixes vendors. Practical consequence for the install: one Cat6 run to a dome instead of a 24-pair loom, which also means the audio and show-control networks can share infrastructure if the VLANs are set up properly.

Syncing audio to real-time visuals

This is where installations that are otherwise well-designed fall apart. If the visuals are pre-rendered playback, sync is a solved problem: run SMPTE linear timecode, lock both machines, done. But most of the interesting work now is generative and real-time, as we've argued in the case for real-time over pre-rendered content, and a real-time visual system has no fixed timeline to lock to.

Three patterns that work, in increasing order of complexity:

  • Audio as master clock. The audio engine plays a fixed piece and emits timecode or Ableton Link; the visual system follows. Simplest and most robust, but the visuals cannot alter the pacing.
  • Shared event bus. Both systems subscribe to OSC or MIDI messages from a show controller, so a cue fires audio and visuals together. Good for interactive pieces where sections are triggered rather than timed.
  • Parameter-level coupling. The visual system streams continuous values - a particle system's energy, a tracked visitor's position - over OSC into the spatial renderer, so sound sources move with what is on the surface. Most convincing and most work to build.

Whichever pattern you use, the latency budget is tighter than people expect. ITU-R BT.1359 puts the detectability threshold for audio-video sync error at roughly 45 milliseconds when audio leads video and 125 milliseconds when it lags. Audio leading is the perceptually worse failure, and it is also the more common one in installations, because projector processing adds 1 to 3 frames of video latency that nobody accounts for. Measure the projector's actual delay and push the audio back to match it.

If a visitor turns their head toward a sound and finds the right thing there, you've built an immersive room. If they turn and find a speaker, you've built a screening room with better lighting.

What this costs, roughly

Budget bands we plan against, hardware only, excluding content authoring and labor. An 8 to 12 speaker ambisonic dome rig with a multichannel interface, amplification, and a sub lands in the low five figures. Moving to third order and 16 to 24 speakers roughly doubles that. Object-based systems from L-Acoustics, d&b, or Meyer Sound carry renderer licensing and vendor-specific hardware, which puts a permanent install into the high five to low six figures before content.

Content authoring is the cost people miss entirely. A spatial mix takes meaningfully longer than a stereo one, and it cannot be finished off-site: the final calibration pass has to happen in the room, with the speakers in their real positions, which means schedule and site access, not just budget. Put two days of on-site audio time in the plan and defend it, because it is the first thing cut and the most visible thing lost.

Frequently asked questions

Do I need spatial audio for a projection dome, or is stereo enough?

If the dome shows content with any sense of direction - anything flying, moving, or coming from a place - stereo actively fights the image. A minimum useful spatial rig is 8 speakers in a ring plus 3 or 4 overhead, decoded from first-order ambisonics. Below that, honest stereo is better than bad surround.

What is the difference between ambisonics and Dolby Atmos?

Ambisonics encodes the whole sound field and decodes to whatever speaker layout the room has, so it is layout-agnostic and free to author with tools like the IEM Plugin Suite. Atmos is object-based: sounds are stored with positional metadata and rendered live by a licensed renderer. Ambisonics suits irregular installation rooms; Atmos suits standardized cinema and broadcast layouts.

How many speakers does an immersive installation need?

Eight to twelve for a typical 8 to 15 meter room using first-order ambisonics, and sixteen to twenty-four if you want precise overhead imaging with third order. Speaker placement and per-speaker delay alignment matter more than raw count - a well-aligned 8-speaker rig outperforms a badly placed 16.

How do you keep spatial audio in sync with real-time generative visuals?

Pick a clock master. Either the audio engine emits timecode or Ableton Link and the visual system follows, or both subscribe to OSC cues from a show controller. Then measure the projector's video latency, which is typically 1 to 3 frames, and delay the audio to match. Audio arriving early is detectable at around 45 milliseconds.

Can we add spatial audio to an existing installation later?

Yes, and ambisonic content is the reason to start there - it decodes to a larger array without re-authoring. What is hard to retrofit is cabling and rigging points, so pull spare Cat6 and plan speaker mounts during the original build even if the budget only covers stereo on day one.

Does a fabric touring dome need different audio treatment than a hard dome?

Yes. A fabric shell is acoustically transparent, so ambient venue noise leaks in and you lose the option of hiding speakers behind a perforated screen. Compensate with more speakers at lower individual output, directional placement aimed inward, and floor treatment where the venue allows it.

We design audio alongside the projection system rather than after it on every dome we build. If you're scoping a room, the projection domes capability page covers how the two systems get planned together, and The Aurora Experience is a good example of a piece where the sound field and the visuals were authored as one thing. For the perceptual side of why sensory contradictions break an experience so quickly, the first thirty seconds of VR comfort covers the same failure mode in a headset.

/// RELATED TRANSMISSIONS

Working on something immersive?