LiveKit cascade
The LiveKit cascade adds browser-detected voice without adding another planning model. A worker transcribes speech with Deepgram and synthesizes settled results with Cartesia, while the browser sends each final transcript through RealtimeSession as an ordinary update_ui intent. If automatic detection is disabled or safely degrades, the adapter exposes manual turn controls instead.
You mint a room grant on your server, run a model-free worker, connect the browser adapter, then harden the route before it ships.
New to the cascade? Run the cascade demo locally first to hear one working turn, then come back here to build the deployable version.
Before you start
Section titled “Before you start”Provide:
- a Maelstrom React instance with its normal Engine Connector
- a LiveKit server and credentials for both the Minter and worker
- Deepgram and Cartesia credentials for the worker
- a browser secure context with microphone permission
- the
@maelstrom-co/realtime-livekit,@maelstrom-co/realtime-livekit-server,@maelstrom-co/realtime-livekit-agent, and@maelstrom-co/react-realtimepackages
Three things decide whether this integration is safe, and none of them are optional.
What you own
Section titled “What you own”The cascade splits into three roles you deploy, arranged around a room you may not operate. Each has its own lifecycle and its own credential boundary.
- Browser: Voice session. RealtimeSession, the LiveKit adapter, useVoice. Lives with the page.
- Your server: Mint route. Authenticates, scopes the room, signs. Holds LIVEKIT_API_SECRET. Ends after each request.
- Your server: Engine. The only planning model.
- LiveKit: LiveKit room. Relays audio and transcripts. Cloud or self-hosted.
- Your worker: Voice worker. Deepgram STT, Cartesia TTS. No model, no UI grounding. Holds Deepgram and Cartesia keys. Long-running.
- Mint route to Voice session: Room grant.
- Voice session and LiveKit room: Audio and session data.
- LiveKit room and Voice worker.
- Voice session and Engine: Intent and plan.
Two absences in that picture matter. Audio reaches the room over the top of your server rather than through it, which is why the mint route can end after each request. And nothing joins the worker to the Engine: the worker publishes each final transcript to the room, the browser adapter turns it into an ordinary Maelstrom intent through RealtimeSession, and the Engine plans against the client-built Chart snapshot exactly as it does for typed intent.
Routing every transcript back through the browser is also what decides when the worker speaks. A successful UI update waits for rendering to settle before the result returns through LiveKit for TTS. Malformed tool arguments and genuine submission failures return immediately, without waiting for render settlement.
One detected turn goes: microphone → LiveKit room → worker STT → final transcript → update_ui intent → Engine → settled UI → worker TTS. Automatic mode detects the boundary in the browser; manual mode uses the same path after the user ends the turn. The voice boundary walks each stage.
Mint a room grant on your server
Section titled “Mint a room grant on your server”Create the authenticated endpoint that derives room scope and signs a grant before you write any browser code that consumes it. Keep LiveKit signing credentials in server-only code. The minter validates the request, trims the room name, signs a token scoped to that room, and returns a typed Result.
This is the smallest route that is safe to run at all — it authenticates, and it refuses to cache the grant:
import { isOk } from '@maelstrom-co/protocol';import { mintFailureHttpStatus, mintLiveKitToken,} from '@maelstrom-co/realtime-livekit-server';
export async function POST(request: Request): Promise<Response> { const authenticated = await requireAuthenticatedVoiceSession(request); if (authenticated === null) { return Response.json({ error: 'unauthorized' }, { status: 401 }); }
const result = await mintLiveKitToken({ roomName: authenticated.roomName, identity: authenticated.identity, }); if (!isOk(result)) { return Response.json( { error: 'mint-failed' }, { status: mintFailureHttpStatus(result.error.reason) }, ); }
return Response.json(result.value, { headers: { 'Cache-Control': 'private, no-store' }, });}requireAuthenticatedVoiceSession is yours to implement. It is not exported by any Maelstrom package — mintLiveKitToken validates shape and signs a token, and deliberately has no opinion about who may ask. It must return null for an authentication or authorization denial, bind the principal to its server-side session, and derive roomName and identity from that trusted session alone. The browser never sends either.
The minter issues only the grants this browser role uses: roomJoin for the server-derived room, canPublish for microphone audio and data, and canSubscribe for worker audio and room events. It does not grant room administration, recording, ingress, egress, or permission to join any other room. If your application needs a narrower publication policy, wrap or replace the minter and test that policy against the media and data topics your integration actually publishes.
The server process requires LIVEKIT_URL, LIVEKIT_API_KEY, and LIVEKIT_API_SECRET. LIVEKIT_TOKEN_TTL is optional and defaults to 15m; supported values are a positive integer number of seconds or one <int><unit> value using s, m, h, or d. Unsupported or unsafe values fail closed as invalid-config.
These token and reconnect statements are scoped to the repository’s locked livekit-client and livekit-server-sdk 2.13.1 releases. Re-check the provider’s token generation guide before changing either dependency.
This route is enough to get voice working and is not enough to deploy. Harden the mint route names the origin, CSRF, rate-limit, and telemetry controls it is missing.
Fetch the grant in the browser
Section titled “Fetch the grant in the browser”Fetch the grant with same-origin credentials and a CSRF proof, then validate the unknown response before connecting. This runs in the browser bundle, so it holds no signing credentials:
import { type LiveKitTokenGrant, LiveKitTokenGrantSchema,} from '@maelstrom-co/realtime-protocol/livekit';
export async function fetchAuthenticatedLiveKitToken(): Promise<LiveKitTokenGrant> { const response = await fetch('/api/livekit-token', { method: 'POST', credentials: 'same-origin', headers: { 'X-CSRF-Token': getCsrfToken() }, });
if (!response.ok) { throw new Error(`LiveKit token mint failed with status ${response.status}`); }
return LiveKitTokenGrantSchema.parse(await response.json());}LiveKitTokenGrantSchema comes from the @maelstrom-co/realtime-protocol/livekit entry point, which carries the grant schema without pulling in the LiveKit browser transport. The request sends the app’s session credentials and CSRF token, and no client-selected room or participant identity. The schema requires a non-empty token, a ws: or wss: LiveKit URL, and an absolute expiry in Unix seconds. Shape validation does not replace endpoint authentication or authorization.
Start the STT/TTS worker
Section titled “Start the STT/TTS worker”Run the model-free agent as a separate long-running process, holding only its server and speech-provider credentials. Keep it out of the application server and the browser.
The worker has no planning LLM and receives no UI grounding. It runs only the Deepgram STT and Cartesia TTS paths; the browser converts final transcripts into intent, and the Maelstrom Engine plans from the client-built Chart snapshot.
import { runVoiceAgent } from '@maelstrom-co/realtime-livekit-agent';
runVoiceAgent();Run that entry with LiveKit’s start worker command. Node cannot execute a bare .ts file before v23, so on Node 20 or 22 either build the entry first and run the emitted .js, or run it through your own TypeScript loader:
# Node 23+node --env-file=.env voice-worker.ts start
# Node 20 or 22: run the compiled entry insteadnode --env-file=.env dist/voice-worker.js startTo run the checked-in demo instead, follow Run the cascade demo locally rather than launching only the worker role. Its Nx target starts apps/demo/voice-worker.ts together with the Engine host and browser app.
When a custom worker harness needs the underlying definition, export the factory result instead:
import { defineMaelstromVoiceAgent } from "@maelstrom-co/realtime-livekit-agent";
export default defineMaelstromVoiceAgent();Keep environment values grouped by the process that reads them:
| Process | Required environment | Optional environment |
|---|---|---|
| Token endpoint | LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET |
LIVEKIT_TOKEN_TTL (default 15m) |
| Agent worker | LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET, DEEPGRAM_API_KEY, CARTESIA_API_KEY |
MAELSTROM_STT_DRAIN_TIMEOUT_MS (default 5000) |
| Browser | None of these secrets | None |
Keep all provider credentials in worker-only or server-only configuration. MAELSTROM_STT_DRAIN_TIMEOUT_MS bounds how long the Deepgram stream may take to drain after the worker closes its input, so a stalled stream cannot block the next turn. It must be finite, positive, and no greater than 60000; any other value fails worker startup rather than changing turn behavior later.
The worker commands and APIs above are scoped to @livekit/agents, @livekit/agents-plugin-deepgram, and @livekit/agents-plugin-cartesia 1.4.8. Follow LiveKit’s voice AI quickstart when upgrading the worker packages, then re-run this guide’s compile-checked fixture.
Connect the adapter
Section titled “Connect the adapter”Create one adapter per Maelstrom session. The generic LiveKitAdapterOptions contract calls mintToken(roomName) so the package can support integrations in which the callback uses a configured room name. This hardened example deliberately ignores that browser-side callback argument: authenticated server policy derives authenticated.roomName and authenticated.identity, and the signed grant alone selects and authorizes the actual room.
import { createMaelstrom } from "@maelstrom-co/react";import { type CreateAdapterContext, maelstromRealtime,} from "@maelstrom-co/react-realtime";import { LiveKitAdapter } from "@maelstrom-co/realtime-livekit";import { fetchAuthenticatedLiveKitToken } from "./voice/livekit-token";import { vessels } from "./vessels";
export const m = createMaelstrom({ vessels }).with( maelstromRealtime({ createAdapter: ({ sessionId }: CreateAdapterContext) => new LiveKitAdapter({ roomName: `maelstrom-${sessionId}`, mintToken: (_roomName) => fetchAuthenticatedLiveKitToken(), }), }),);The roomName option remains necessary to satisfy the generic adapter API, and the adapter forwards it only to mintToken(roomName). It is not an authorization input in this integration. The browser fetch sends no room or identity in its body; the authenticated endpoint’s signed grant is the authority used when LiveKit joins the participant.
connect() joins the room and asks LiveKit to capture with echo cancellation (AEC), noise suppression (NS), and automatic gain control (AGC). These plain-boolean constraints are best-effort browser preferences, not guarantees for a device, operating system, or audio route. The adapter inspects the accepted original track, clones it for local metering, VAD, and pre-roll, mutes the LiveKit track, and only then publishes it and pauses it upstream. The order is deliberate: setMicrophoneEnabled(true) would publish an enabled track before the adapter could apply its gate, so the adapter does not use it.
With debug: true, the adapter logs only the three requested values and the original track’s applied values as true, false, or unavailable; it never treats the local clone as evidence for the transmitted track. A missing or throwing settings API, missing field, or non-boolean value becomes unavailable without failing connection. The diagnostic excludes tracks, streams, the complete settings object, device identity, and audio. It reports bounded browser settings, not acoustic success.
This implementation requests microphone access during connection, before the first turn. The local clone remains enabled for detection while the published track stays muted and upstream-paused outside an admitted turn. The browser retains microphone access until disconnect() leaves the room and stops the adapter’s media resources. Starting a turn changes the existing gated track; it does not reacquire permission. Passing vad: false changes only turn detection and keeps the same capture preferences and privacy gate. Turn boundaries are found in the browser, not by LiveKit or the worker — see Configure browser VAD next.
Token expiry is checked for each initial or fresh room connection, including an explicit connect() after disconnect(). It does not impact LiveKit SDK-managed reconnects for an already connected participant because LiveKit pushes refreshed tokens to connected clients, and the SDK uses the refreshed token for managed reconnects. Expiry alone does not disconnect an active participant or end the room. Each fresh connect() invokes mintToken; the adapter does not read expiresAt or schedule proactive renewal. Expiry does not impose a maximum room duration or schedule adapter cleanup, so production hosts must enforce their own maximum session duration, call disconnect() on policy termination, and clean up abandoned rooms and participants separately.
Configure browser VAD
Section titled “Configure browser VAD”The adapter detects turn boundaries in the browser, and does so unless you pass vad: false. The adapter above therefore runs in automatic mode: your users talk, and nothing has to be held.
Automatic mode needs asset files that this package does not publish for you. Copy them before your build, so the versions match the ones this package pins:
{ "scripts": { "prepare:vad": "maelstrom-vad-assets public/vad-web", "build": "bun run prepare:vad && vite build" }}Skip this step and the outcome is quiet rather than loud. The assets fail to load, the session degrades to manual turns, and your users are told to hold a button you may not have built. Serve the directory from the same origin, or from a credential-free HTTPS host you control — the browser executes worklet, JavaScript, and WebAssembly from it.
To tune detection, pass an object. An unsupported key fails to compile, and fails again at runtime with a TypeError, so it can never become a setting that silently does nothing:
new LiveKitAdapter({ roomName, mintToken, vad: { // Where the build step copied the assets. `/vad-web/` is the default. assetBaseUrl: '/vad-web/', preRollMs: 300, providerOptions: { positiveSpeechThreshold: 0.8, minSpeechMs: 100, }, },});Raising positiveSpeechThreshold cuts false interruptions and starts missing quiet speakers. preRollMs is the audio sent at the start of a turn so the first word survives; the provider’s preSpeechPadMs is a different setting and does not do this. BrowserVadOptions lists every key and its accepted values.
To take the boundary yourself instead, turn detection off:
new LiveKitAdapter({ roomName, mintToken, vad: false });That is one of the two ways turnControl appears; Delimit each manual turn builds the control. For why this decision sits in the browser at all, read Turn detection in the browser.
Use the protocol entry point
Section titled “Use the protocol entry point”The /protocol entry point is shared by the browser adapter and worker. It defines the topics without pulling browser dependencies into the worker.
import { SPEAK_TOPIC, TURN_TOPIC, TTS_TOPIC,} from "@maelstrom-co/realtime-protocol/livekit-wire";TURN_TOPIC carries turn-end and TTS-cancel data messages. SPEAK_TOPIC carries settled result text to the worker, and TTS_TOPIC carries playout-complete or synthesis-error status back to the browser. Application code normally lets the adapter and worker publish these messages.
Delimit each manual turn
Section titled “Delimit each manual turn”Build this only if you configured vad: false, or if you want a fallback ready for a session that degrades to manual mode. turnControl is undefined while automatic detection is healthy, so this control renders in exactly those two cases and stays hidden otherwise.
Disable hold-to-talk until the transport is connected. Prevent repeated starts and queue serialized, best-effort cleanup when a normal release is lost to pointer-capture loss, control or window blur, a hidden document, disconnect, or unmount. Browser termination can still prevent JavaScript cleanup, so server-side session limits remain necessary.
const LIVE_STATES = new Set(['connected', 'listening', 'thinking', 'speaking']);
export function HoldToTalk() { const { state, turnControl } = m.useVoice(); const pressRequested = useRef(false); // Which pointer owns the turn. A second finger landing on the same button // must not be able to end a turn it never started: its pointerdown bails out // early, but its pointerup would otherwise flush STT mid-sentence. const activePointer = useRef<number | null>(null); const requestSequence = useRef(0); const turnActive = useRef(false); const lifecycle = useRef<Promise<void>>(Promise.resolve()); const connected = LIVE_STATES.has(state);
// Every turn operation queues behind the previous one. A release can arrive // while startTurn() is still pending; without this, endTurn() would run first // and leave the microphone open for the rest of the session. const serialize = useCallback((operation: () => Promise<void>) => { lifecycle.current = lifecycle.current.then(operation).catch(() => {}); return lifecycle.current; }, []);
const begin = useCallback(async () => { if (pressRequested.current || !turnControl || !connected) return; pressRequested.current = true; requestSequence.current += 1; const request = requestSequence.current;
await serialize(async () => { // requestSequence guards against a later press reviving a start that an // earlier release already invalidated. if ( !pressRequested.current || requestSequence.current !== request || turnActive.current ) return; await turnControl.startTurn(); turnActive.current = true; }); }, [connected, serialize, turnControl]);
const stopTurn = useCallback( async (force: boolean) => { if (!force && !pressRequested.current) return; pressRequested.current = false; requestSequence.current += 1;
await serialize(async () => { if (!turnActive.current) return; try { await turnControl?.endTurn(); } finally { turnActive.current = false; } }); }, [serialize, turnControl], );
// Release can be lost to a dropped connection, a blurred window, a hidden // tab, or unmount. Each one has to end the turn, or the mic stays live. useEffect(() => { if (!connected) void stopTurn(true); }, [connected, stopTurn]);
useEffect(() => { const stopForFocusLoss = () => void stopTurn(true); const stopWhenHidden = () => { if (document.visibilityState === 'hidden') stopForFocusLoss(); };
window.addEventListener('blur', stopForFocusLoss); document.addEventListener('visibilitychange', stopWhenHidden);
return () => { window.removeEventListener('blur', stopForFocusLoss); document.removeEventListener('visibilitychange', stopWhenHidden); void stopTurn(true); }; }, [stopTurn]);
const endPress = () => void stopTurn(false); const cancelPress = () => { activePointer.current = null; void stopTurn(true); }; // Only the pointer that opened the turn may close it. const endOwnedPress = (event: PointerEvent<HTMLButtonElement>) => { if (activePointer.current !== event.pointerId) return; activePointer.current = null; void stopTurn(false); }; const isTurnKey = (key: string) => key === 'Enter' || key === ' ';
// LiveKit exposes turnControl only in manual mode. Automatic turn detection // leaves it undefined, and this control must not render at all. if (!turnControl) return null;
return ( <button type="button" disabled={!connected} onPointerDown={(event: PointerEvent<HTMLButtonElement>) => { if (pressRequested.current) return; activePointer.current = event.pointerId; event.currentTarget.setPointerCapture(event.pointerId); void begin(); }} onPointerUp={endOwnedPress} onPointerLeave={endOwnedPress} onPointerCancel={endOwnedPress} onLostPointerCapture={cancelPress} onBlur={cancelPress} onKeyDown={(event: KeyboardEvent<HTMLButtonElement>) => { if (!isTurnKey(event.key) || event.repeat || pressRequested.current) return; event.preventDefault(); void begin(); }} onKeyUp={(event: KeyboardEvent<HTMLButtonElement>) => { if (isTurnKey(event.key)) endPress(); }} > Hold to talk </button> );}turnActive is local ordering bookkeeping, not provider status. finally clears it even when endTurn() rejects, so the next press is not wedged behind stale state. The disconnect and focus-loss effects queue cleanup through the same serializer, so it runs after pending turn work. This is best-effort browser cleanup, not a guarantee against process termination — server-side session limits remain necessary.
An empty held turn produces no final transcript, no intent or Engine request, and no spoken response. After the bounded empty-turn fallback completes, the session returns to connected.
This control shows no errors, and a rejected startTurn() or endTurn() needs one. It is a recoverable local failure rather than a session failure — serialize already swallows it so the queue survives — so surface it separately and let a valid press or a reconnect clear it. Send the original value through your own redacted-telemetry boundary with only the lifecycle phase as public context. Session-level failures belong in the status control below.
Show session state
Section titled “Show session state”The hold-to-talk button says nothing about what the session is doing. Render the rest of useVoice() beside it: the lifecycle state, both transcripts, and the two failures a user can act on.
import type { VoiceSessionState } from '@maelstrom-co/client';
const STATE_LABEL: Record<VoiceSessionState, string> = { idle: 'Voice off', connecting: 'Connecting', connected: 'Voice on', listening: 'Listening', thinking: 'Thinking', speaking: 'Speaking', error: 'Voice unavailable',};
export function VoiceStatus() { const { state, error, userTranscript, assistantTranscript, turnMode, turnModeError, } = m.useVoice();
// Manual mode is either how you configured the adapter or where it landed // after automatic detection failed to start. turnModeError is set only in the // second case, so it is the one signal that explains why a hold-to-talk // button just appeared in a session that was meant to listen on its own. const degraded = turnMode === 'manual' && turnModeError !== undefined;
return ( <section aria-label="Voice status"> <p>{STATE_LABEL[state]}</p> {degraded && ( <p role="status"> Automatic voice detection is unavailable. Hold to talk instead. </p> )} {userTranscript && <p>You: {userTranscript.text}</p>} {assistantTranscript && <p>Assistant: {assistantTranscript.text}</p>} {error !== undefined && ( <p role="alert">Voice stopped. Start voice again to reconnect.</p> )} </section> );}turnMode is the piece unique to this mode. Manual is either how you configured the adapter or where it landed after automatic detection failed to start, and only the second case sets turnModeError. Without that check a hold-to-talk button appears in a session that was meant to listen on its own, and the user is given no reason for it.
A fatal adapter failure arrives as a RealtimeVoiceError on error and tears down the connection. Map the typed code to fixed recovery copy — never the raw message, stack, or cause, which can carry room tokens and provider detail. Send the original value through your own redacted-telemetry boundary instead.
Transcript entries accumulate: each one replaces the last until it is marked final. Render them directly rather than appending, or a partial phrase repeats as it grows.
m.VoiceOrb is the built-in alternative to a hand-written toggle. It renders a button that connects and disconnects the shared session, colours itself per state, and drives an audio-reactive meter. Take it when it fits and keep your own markup for the transcripts:
<m.VoiceOrb size={48} />It trades away the error surface built above: it discards connect() and disconnect() rejections rather than reporting them. Keep the hand-written toggle wherever a failed connection or a failed cleanup has to reach the user. size is a diameter in pixels, defaulting to 64; keep it at 24 or above so the control stays a usable touch target.
Verify
Section titled “Verify”Check one complete automatic turn across all three roles, then repeat in manual fallback:
- Token route: an authenticated request passes the configured origin, CSRF, and rate-limit controls; the returned grant validates with
LiveKitTokenGrantSchemaand every response carriesCache-Control: private, no-store. - Browser: after the adapter uses that grant, the shared voice state reaches
connectedwithout any browser-supplied room or identity becoming authoritative. - Worker: the separately deployed worker joins the authorized room and publishes one final transcript when browser VAD ends the turn.
- Lifecycle: a successful automatic turn progresses through the observable
listening,thinking, andspeakingstates; the final transcript becomes ordinary intent, the intent renders, and TTS follows settle rather than predicting the result. - Manual fallback: reconnect with
vad: false, confirmturnControlappears, and complete one held turn through the same lifecycle. An empty held turn produces no final transcript, intent, Engine request, or spoken response, then returns toconnectedafter the bounded fallback.
Unit tests can verify requested options, original-track inspection, and the mute/publish/pause order. They cannot establish mobile acoustics. Before claiming that AEC prevents self-triggering while preserving interruption, run the release protocol in Validate browser audio processing, including the LiveKit idle/upstream privacy checks.
Harden the mint route
Section titled “Harden the mint route”The route above authenticates and derives scope on the server. Add these before it is reachable from the internet. Each one closes a specific hole:
| Control | Without it |
|---|---|
| Exact-origin allowlist, checked first | Any site can drive a browser to mint rooms on your account |
| CSRF proof bound to the session | A logged-in user’s browser mints on an attacker’s behalf |
| Shared rate limit keyed by principal | One compromised account creates rooms without bound |
| Server-derived room and identity | The browser joins a room belonging to someone else |
| Short token lifetime | A leaked grant stays usable long after the session ends |
Cache-Control: private, no-store |
A shared cache serves one user’s room token to another |
| Maximum active-session duration | A held session bills against LiveKit and both speech providers until noticed |
| Redacted failure telemetry | Stack traces leak room tokens and provider keys into logs |
Every control is yours to implement. None are exported by a Maelstrom package or supplied by a server runtime, and each one must fail closed — a check that throws, times out, or cannot reach its policy store denies the request rather than falling through. None may accept a policy key or room scope from browser input.
The origin check runs first, before anything reads the body. It parses the request’s Origin, compares it with an exact configured allowlist, and returns false when the header or the policy is missing or invalid — not a prefix or suffix match.
Authentication and CSRF are separate checks. Authenticating the caller proves who is asking and supplies the server-derived room scope; it does not prove the request was intended. Validate CSRF separately against the authenticated session, then apply the shared rate limit, keyed atomically by the authenticated principal. The limit must deny on policy-store failure and return a bounded Retry-After.
Wrap the whole route so unexpected exceptions reach only your redacted-telemetry boundary, which must emit allowlisted metadata without raw messages, stacks, credentials, token contents, request bodies, or session data. If telemetry itself fails, the route still returns the same generic failure. Neither mapped mint failures nor unexpected exceptions expose internal failure reasons to the browser.
Every response sets Cache-Control: private, no-store. no-store prevents private and shared caches from retaining the bearer room token; private also documents that the response is user-specific. Keep this header even when an authenticated framework supplies its own default caching policy.
Token TTL does not bound a session. LiveKit pushes refreshed tokens to connected participants, so expiry only affects new connections — enforce duration separately and clean up abandoned rooms.
Monitor room creation, mint failures, and worker failures. Never log room tokens, LiveKit secrets, or provider API keys.
Recover from failures
Section titled “Recover from failures”| Failure | Recovery |
|---|---|
| Token request or grant validation fails | Keep voice disconnected, reject partial data, and correct caller policy or server configuration before retrying |
LiveKit URL uses http: or https: |
Configure a ws: or wss: URL; the grant schema rejects unsupported schemes |
| Microphone permission is denied | Explain how to allow microphone access, then reconnect before enabling hold-to-talk |
| Worker does not answer | Check the worker process, LiveKit connectivity, and the Deepgram and Cartesia keys; do not move planning into the worker |
| A press is cancelled or the connection drops | Call endTurn() from cancellation cleanup so the microphone does not remain active |
| TTS fails after rendering | Keep the session usable and report that spoken confirmation was lost; the UI result is already applied |
What you get without asking
Section titled “What you get without asking”Two behaviors are wired by the time a turn completes. Neither needs configuration, and both are easy to break by working around them.
Speak-after-render. Each final transcript becomes one update_ui call. For a successful UI update, RealtimeSession waits for rendering to settle before it calls sendToolResult(), so the Cartesia read-back describes the rendered result rather than a prediction. Malformed tool arguments and genuine submission failures return failure results immediately. Failed TTS is non-fatal, because the UI update has already applied.
A worker that cannot plan. Unlike the OpenAI mode, the LiveKit worker has no model session to ground. Settled structural summaries may reach LiveKitAdapter.injectGrounding(), but the method deliberately does nothing. Do not forward Chart structure or Vessel domain data to the worker.
Swap STT or TTS
Section titled “Swap STT or TTS”The checked-in worker currently constructs Deepgram for STT and Cartesia for TTS. That selection happens inside the worker, not in LiveKitAdapter or other browser configuration. defineMaelstromVoiceAgent() currently accepts no provider-injection options, so configuration such as sttProvider or ttsProvider would describe an API that does not exist.
Today, replacing either provider requires a custom worker. Preserve the integration contract when doing so: publish final transcripts in the form the browser adapter consumes, honor TURN_TOPIC for turn completion and cancellation, receive settled result text on SPEAK_TOPIC, and publish playout-complete or synthesis-error status on TTS_TOPIC. Keeping those semantics intact lets the browser continue turning one final transcript into one natural-language intent and keeps TTS after UI settlement. If a future worker API adds provider injection, prefer that supported seam instead of copying the worker implementation, but do not assume that seam exists until the package publishes it.
Teams may accept the custom-worker cost to specialize recognition by language or domain, improve medical-vocabulary handling, choose a different voice, meet provider or data-residency compliance requirements, or optimize usage cost. Separate provider stages also make later swaps possible without changing the Engine boundary. The tradeoff is additional integration, deployment, observability, and failure handling, and each network or processing stage may increase turn latency. Evaluate those costs against measured recognition quality, synthesis quality, and end-to-end latency for your own traffic.
API reference
Section titled “API reference”Use the generated reference for the public contracts used in this guide:
LiveKitAdapterfrom the@maelstrom-co/realtime-livekitrootSPEAK_TOPICfrom@maelstrom-co/realtime-protocol/livekit-wireLiveKitTokenGrantSchemafrom@maelstrom-co/realtime-protocol/livekitmintLiveKitTokenandLiveKitTokenRequestSchemafrom@maelstrom-co/realtime-livekit-serverdefineMaelstromVoiceAgentandrunVoiceAgentfrom@maelstrom-co/realtime-livekit-agent