Turn detection in the browser
A speech-native model hears where your sentence ended. A cascade has no such thing: a speech-to-text worker receives audio and returns words, and nothing in that exchange knows when you stopped talking.
Someone has to decide. The LiveKit adapter decides in the browser, and turns detection on unless you turn it off. This page is about why that choice sits there rather than in the worker, and what follows from it.
For the setup, read Configure browser VAD. For every option and its accepted values, read BrowserVadOptions.
What the transport is allowed to carry
Section titled “What the transport is allowed to carry”- Browser: Microphone. One local stream, always available to the detector.
- Browser: VAD and pre-roll. Decides what counts as a turn, and holds the publication paused until it does.
- LiveKit: LiveKit room. Carries accepted turn audio.
- Your worker: Speech-to-text. Sees only what the browser admitted.
- Microphone to VAD and pre-roll.
- VAD and pre-roll to LiveKit room: Pre-roll, then the turn.
- LiveKit room to Speech-to-text.
- Microphone has no path to LiveKit room: Idle microphone audio.
The microphone stream and the microphone publication are two different things, and the distinction carries the whole design. The stream is local and runs continuously, because a detector with nothing to listen to cannot detect anything. The publication to the LiveKit room stays paused until the detector accepts speech, opens for that turn, and pauses again when the turn ends.
So the worker sees only audio the browser admitted. That is not a policy the worker enforces or a promise the browser keeps. Between turns the published track is muted and its upstream is paused, so there is no audio in flight to enforce anything about.
Why the decision belongs in the browser
Section titled “Why the decision belongs in the browser”An interruption should not wait for a round trip. A detector running on the server can only react once your audio has crossed the network. A detector in the browser reacts as soon as it accepts speech. This does not remove the LiveKit, recognition, Engine, rendering, and synthesis latency that follows; it removes the delay before Maelstrom treats you as interrupting at all.
What you hear is local, so stopping it must be local too. When you talk over the assistant, the thing that needs to stop is audio already playing in your browser. Asking the worker to stop cannot silence a buffer that has already arrived. The browser mutes its own playback, and the request to the worker is a courtesy that follows.
An idle microphone should not be on the wire. Keeping the publication paused during silence is a privacy property you can inspect in the network panel, which is a stronger claim than any sentence in this document.
Interrupting the assistant
Section titled “Interrupting the assistant”A user speaks over the assistant. Browser voice activity detection accepts the speech and tells the adapter. The adapter mutes assistant playback locally, invalidates the utterance it was playing, and only then asks the worker to cancel speech synthesis, which is best effort. It resumes upstream audio, sends the retained pre-roll, and live audio continues to the worker. At speech end the adapter pauses upstream audio and asks the worker to finalize recognition, which returns one final transcript.
- 1. User to Browser VAD: Starts speaking.
- 2. Browser VAD to LiveKitAdapter: Speech accepted.
- 3. LiveKitAdapter to Browser playback: Mute assistant audio.
- 3. LiveKitAdapter: Drop the utterance id.
- 4. LiveKitAdapter to STT/TTS worker: Cancel speech, best effort.
- 5. LiveKitAdapter to STT/TTS worker: Open upstream, send retained pre-roll.
- 6. User to STT/TTS worker: Live turn audio.
- 7. Browser VAD to LiveKitAdapter: Speech end.
- 8. LiveKitAdapter to STT/TTS worker: Pause upstream, finalize recognition.
- 9. STT/TTS worker to LiveKitAdapter: One final transcript.
Read the order rather than the steps. Muting playback and dropping the utterance id happen in the browser, at once, before anything is sent. Only then does a cancellation go to the worker, and it is explicitly best effort: if publishing it fails, the adapter warns once and carries on opening your new turn. Nothing you can hear depends on that message arriving.
Dropping the utterance id is the part that is easy to miss. The worker may finish speaking the old response and report it complete, long after you interrupted. Invalidating the id before the cancel goes out is what makes that late report ignorable rather than a state change that stomps on the turn you just started.
A speech start can also arrive while the session is thinking with a response still in flight, rather than while it is speaking. These are different situations, and the adapter serializes them: interruption and turn opening cannot overlap, and a fast speech end cannot overtake a start that has not finished. RealtimeSession keeps the newest intent rather than queueing every correction, and still returns one result per call.
Why a turn carries audio from before it started
Section titled “Why a turn carries audio from before it started”A detector needs frames before it can conclude that speech began. If the transport opened only at that conclusion, the first consonant would already be gone.
The adapter keeps a small local buffer for exactly that. At speech start it sends the retained audio once, then continues live. More of it protects the beginnings of words and admits more audio from before the accepted boundary, so it is worth raising only after you have reproduced clipping on a browser you support.
There are two settings with similar names and only one of them does this. preRollMs is the buffer above. The provider’s preSpeechPadMs widens the segment the detector considers and never becomes transport audio, so reaching for it to fix a clipped word start will not fix it.
Degrading is a safety decision, not a severity one
Section titled “Degrading is a safety decision, not a severity one”Detection can fail: assets may not load, a worklet may not start, a callback may throw. When that happens while the room, transcripts, and synthesis are all still healthy, the session does not fail. It fences further callbacks, closes any turn the failed detector owned, pauses upstream audio, retires the failed resources, and then exposes turnControl so your application can offer a button.
The order matters more than the list. Upstream audio is paused before manual mode is advertised, so a hold-to-talk control never appears over a microphone that is still publishing.
That condition is also what separates degrading from failing. If the adapter cannot stop microphone media safely, or the failure is in token minting, room connection, track publication, a required data channel, or an unexpected disconnect, the session fails instead. Manual mode is never offered on a connection that cannot hold the upstream invariant.
Two consequences reach your components. turnControl is dynamic: it is absent in healthy automatic mode and appears only after vad: false or a safe degradation, so read it from current state and never cache it. And turnModeError tells you a degradation happened; its value is a diagnostic, so use its presence to choose fixed recovery copy and never render it or store it.
Why none of this applies to OpenAI Realtime
Section titled “Why none of this applies to OpenAI Realtime”OpenAI Realtime is a speech-native provider session. Ordinary turns stay provider-owned: turn detection, response cancellation, and spoken output remain OpenAI’s. Adding browser-side turn control for those later turns would create two owners for one boundary, and the two would eventually disagree about when your turn ended.
There is one narrow exception. greetingInterruption is OpenAI-only and greeting-only. Local VAD may accept one boundary while the fixed greeting is actually playing. After greeting settlement, OpenAI server VAD owns later turns again. The shared session still has no vad option.
That is why the vad option lives on the LiveKit adapter rather than on the shared session. It is not a Maelstrom-wide voice setting; it is how one adapter solves a problem the other adapter does not have for ordinary turns. See Configure experimental greeting interruption for the OpenAI greeting option.
Browser processing is a request, not a result
Section titled “Browser processing is a request, not a result”Both adapters request echo cancellation (AEC), noise suppression (NS), and automatic gain control (AGC) when they acquire a microphone. OpenAI captures the original browser track directly and attaches it to its peer connection. LiveKit inspects its accepted original track, then clones it for local VAD, metering, and pre-roll while the published track remains behind the mute/upstream gate.
The browser treats these plain booleans as preferences. Support and results vary by browser, operating system, microphone, speaker, and Bluetooth or wired route. Debug output narrows the observation to true, false, or unavailable for each applied setting on the accepted original track. unavailable also covers missing or throwing settings APIs and missing or non-boolean fields. The output excludes device identity, complete settings objects, tracks, streams, and audio. An applied true still does not prove that playback echo is removed or that human interruption remains usable.
Validate browser audio processing
Section titled “Validate browser audio processing”Unit tests cannot establish any of this. Real AudioWorklets, real ONNX assets, a real microphone, real CORS, and real LiveKit publication behave in ways a mock does not, so the numbers that matter come from a browser.
Use the same setup for every required cell:
- Record the adapter, device model, OS/browser version, app/library revision, audio route, bounded debug projection, and result. Do not record device labels or IDs, raw audio, transcripts, or credentials.
- Use a quiet indoor room, a stationary device on a hard surface, and a microphone 50 cm from the tester. Set system media volume to 70% and application/player volume to 100%. Disable external EQ or accessibility audio processing unless the route requires it.
- Use this 12-second TTS passage: “Today we are checking the calendar, weather, travel time, and the next three tasks before the afternoon meeting begins.”
- At 4.0 seconds after audible TTS starts, say “Stop and show my calendar” at normal conversational volume. Measure from phrase onset with an external stopwatch or video clock.
- Reset the voice session between trials. For each adapter and route, run 10 no-speech trials followed by 10 interruption trials. LiveKit uses default VAD; also run one
vad: falsecontrol per route and confirm that no upstream audio leaves the browser outside an admitted manual turn.
A self-trigger occurs when TTS playback alone enters the user-speech/barge-in path, stops TTS, admits upstream audio, or produces a user turn before playback ends plus two seconds. A genuine interruption succeeds when the fixed phrase stops TTS, admits a turn, and retains enough speech to recognize the command. Record interruption latency from phrase onset to audible TTS stop, plus missed interruptions that do not stop TTS or admit a turn within two seconds. Any LiveKit microphone media sent outside an admitted turn is a privacy failure.
The required release matrix is iOS Safari with built-in speaker/microphone, Android Chrome with built-in speaker/microphone, and a desktop Chromium built-in speaker/microphone control. Every adapter/route cell must have 10 quiet and 10 interruption trials with:
- zero self-triggers;
- at least 9 of 10 genuine interruption successes;
- median interruption latency at or below 750 ms;
- slowest successful interruption at or below 1500 ms; and
- no LiveKit privacy failure.
Bluetooth and wired/USB routes are optional characterization unless separately made release-gating. Never generalize results to an untested device or route.
Only complete passing evidence supports ACOUSTIC ACCEPTANCE PASSED or a claim that the change works acoustically. Missing required iOS or Android hardware is ACOUSTIC ACCEPTANCE BLOCKED—REQUIRED HARDWARE UNAVAILABLE, not a pass. Any threshold or privacy failure is ACOUSTIC ACCEPTANCE FAILED; fix the behavior and rerun the complete matrix before release. Automated checks can establish implementation correctness, but they cannot replace this evidence or justify an acoustic release claim.
To compare this with the provider-owned alternative, read The voice boundary.
OpenAI greeting-interruption evidence
Section titled “OpenAI greeting-interruption evidence”The LiveKit thresholds above do not apply to this experiment. A cell may change only after real-device evidence under the protocol on this page. Unit, bundle, build, and Playwright results cannot upgrade Not tested. Automated checks prove ordering only. Record revision, audio route, quiet/self-trigger trials, spoken-interruption trials, first-word observation, and latency evidence when you run the protocol; leave those fields empty until then.
| Route | Result | Revision | Audio route | Quiet / self-trigger trials | Spoken-interruption trials | First-word observation | Latency evidence |
|---|---|---|---|---|---|---|---|
| iOS Safari, built-in speaker/microphone | Not tested | ||||||
| Android Chrome, built-in speaker/microphone | Not tested | ||||||
| desktop Chromium, built-in speaker/microphone | Not tested |
Configure the option in Configure experimental greeting interruption.