OpenAI Realtime
OpenAI Realtime adds speech-native voice to a Maelstrom app. Your server mints a short-lived client credential, the browser connects directly to OpenAI over WebRTC, and RealtimeSession turns provider tool calls into ordinary Maelstrom intents.
By the end you will speak to your interface, watch it adapt, and hear the model describe what changed. You add a mint route, connect voice in the browser, render session state, then harden the route before it ships.
Before you start
Section titled “Before you start”Provide:
- a Maelstrom React instance with its normal Engine Connector
- a server runtime that can read
OPENAI_API_KEY - a browser secure context with microphone permission
- the
@maelstrom-co/react,@maelstrom-co/protocol,@maelstrom-co/realtime,@maelstrom-co/react-realtime,@maelstrom-co/realtime-openai, and@maelstrom-co/realtime-openai-serverpackages
Three things decide whether this integration is safe, and none of them are optional.
OpenAI manages voice activity and turn detection for this mode. OpenAIAdapter therefore does not expose turnControl; render an always-on connect/disconnect control rather than hold-to-talk controls.
What you own
Section titled “What you own”OpenAI Realtime uses two application-owned runtime roles: the browser session and a server mint route.
- Browser: Voice session. RealtimeSession, the OpenAI adapter, useVoice. Lives with the page.
- Your server: Mint route. Authenticates the caller, then mints. Holds OPENAI_API_KEY. Ends after each request.
- Your server: Engine. Unchanged by voice.
- OpenAI: Realtime session. Model inference and turn detection. Runs for the whole call.
- Mint route to Voice session: Client credential.
- Voice session and Realtime session: WebRTC audio and session data.
- Voice session and Engine: Intent and plan.
The arc is the part to remember: audio and session data go between the browser and OpenAI over the top of your server, not through it. That is why the mint route can end after each request while the call keeps running.
The same split decides which package goes where. The browser owns RealtimeSession, the OpenAI adapter, and the React binding. The application server owns the minter and the long-lived OpenAI API key. The Engine is unchanged and receives the same client-built request that typed intent uses.
Because audio travels between the browser and OpenAI rather than through Maelstrom infrastructure, OpenAI’s own terms govern it. Review the OpenAI Realtime WebRTC guide for the browser transport and the OpenAI Realtime guide for provider-managed session behavior and data handling.
Mint a credential on your server
Section titled “Mint a credential on your server”Build the minter once in server-only code, authenticate the caller, then map the minter’s Result to your framework’s HTTP response. This is the smallest route that is safe to run at all — it authenticates, and it refuses to cache the credential:
import { isOk } from '@maelstrom-co/protocol';import { createOpenAIMinter, mintFailureHttpStatus,} from '@maelstrom-co/realtime-openai-server';
const mint = createOpenAIMinter({ apiKey: process.env.OPENAI_API_KEY ?? '', model: 'gpt-realtime-2',});
export async function POST(request: Request): Promise<Response> { const authenticated = await requireAuthenticatedRealtimeSession(request); if (authenticated === null) { return Response.json({ error: 'unauthorized' }, { status: 401 }); }
const result = await mint({ provider: 'openai', sessionId: authenticated.sessionId, model: 'gpt-realtime-2', }); if (!isOk(result)) { return Response.json( { error: 'mint-failed' }, { status: mintFailureHttpStatus(result.error.reason) }, ); }
return Response.json(result.value, { headers: { 'Cache-Control': 'private, no-store' }, });}requireAuthenticatedRealtimeSession is yours to implement. It is not exported by any Maelstrom package — the minter deliberately has no opinion about who may call it. The server derives sessionId from trusted application state; the browser never sends one.
The credential is short-lived and scoped to the selected model. OpenAI returns the authoritative expiresAt in the grant.
This route is enough to get voice working and is not enough to deploy. Harden the mint route names the origin, CSRF, rate-limit, model-policy, and telemetry controls it is missing.
Fetch and validate the grant
Section titled “Fetch and validate the grant”Fetch the credential from the same origin and treat the route response as unknown data. Validate it with OpenAITokenGrantSchema before handing the opaque credential to OpenAIAdapter.
import { OpenAITokenGrantSchema, type OpenAITokenGrant,} from '@maelstrom-co/realtime-openai';import { getCsrfToken } from '../auth/csrf';
export async function mintOpenAIToken(): Promise<OpenAITokenGrant> { const response = await fetch('/api/realtime-token', { method: 'POST', credentials: 'same-origin', headers: { 'Content-Type': 'application/json', 'X-CSRF-Token': getCsrfToken(), }, body: JSON.stringify({}), });
if (!response.ok) { throw new Error(`Token mint failed with status ${response.status}`); }
return OpenAITokenGrantSchema.parse(await response.json());}The browser sends only its same-origin session credentials, the app-provided CSRF proof, and an empty JSON object required by this POST contract. It does not send sessionId, a model, or model-policy data. The server derives the session, principal, and allowed model from trusted application state before calling the minter. Schema validation proves only that the response has the expected wire shape.
Connect voice
Section titled “Connect voice”Give OpenAIAdapter the validated mint function, then extend the Maelstrom React instance with maelstromRealtime. Build one adapter per Maelstrom session: maelstromRealtime creates one shared RealtimeSession for all useVoice() consumers under that Maelstrom Provider.
The factory runs once per session, not per render, and receives a CreateAdapterContext carrying that session’s sessionId. Use it to scope provider-side resources you create in the browser. This mode creates none — the mint route derives the session server-side — so the factory takes no argument here.
import { createMaelstrom } from '@maelstrom-co/react';import { maelstromRealtime } from '@maelstrom-co/react-realtime';import { OpenAIAdapter } from '@maelstrom-co/realtime-openai';import { mintOpenAIToken } from './voice/mint-openai-token';import { vessels } from './vessels';
export function createOpenAIAdapter(): OpenAIAdapter { return new OpenAIAdapter({ mintToken: mintOpenAIToken, });}
export const m = createMaelstrom({ vessels }).with( maelstromRealtime({ createAdapter: createOpenAIAdapter, }),);Calling connect() mints a fresh credential and begins microphone capture. Once WebRTC connects, the browser sends that microphone track directly to OpenAI until disconnect() stops the track and closes the peer connection. Make this media lifecycle clear in consent and connection UI. A second session-level connect() while the session is already live is a no-op; the adapter itself rejects a direct duplicate connection.
Configure server VAD
Section titled “Configure server VAD”OpenAIAdapter uses OpenAI’s server voice activity detection (VAD). Configure its activation threshold and choose whether provider speech detection or input transcription confirms an interruption:
const adapter = new OpenAIAdapter({ mintToken: mintOpenAIToken, serverVad: { threshold: 0.65, bargeInMode: 'transcript-confirmed', },});The defaults are threshold: 0.5 and bargeInMode: 'immediate'. Immediate mode lets OpenAI create responses and interrupt response audio when its server VAD detects speech. Transcript-confirmed mode disables those provider flags: the adapter waits for a useful, non-whitespace provisional input-transcription delta before interrupting, then waits for a useful completed transcript and the matching audio commit before sending response.create for that item. Transcription arrival is nondeterministic, and non-whitespace text does not prove that a human produced the sound.
Transcript-confirmed mode requires input transcription. Combining bargeInMode: 'transcript-confirmed' with transcription: false throws a synchronous TypeError from the constructor, before token minting, microphone capture, peer creation, or connection. The hosted demo combines transcript-confirmed interruption with threshold: 0.8 so a raw server-VAD start cannot interrupt output by itself. That threshold is tuning for its own environment, not a production recommendation. Tune the threshold and confirmation latency against the browsers, devices, audio routes, microphone conditions, quiet speech, and intentional barge-in that your application supports.
The session also requests OpenAI’s near_field input noise reduction. During capture, the adapter asks the browser for echo cancellation (AEC), noise suppression (NS), and automatic gain control (AGC). These configurations are requests, not acoustic guarantees: OpenAI, the browser, selected device, operating system, and audio route may apply them differently or not at all. When debug: true is enabled, the adapter inspects the accepted original microphone track and logs only the three browser-requested values plus each applied value as true, false, or unavailable. Raw speech-start diagnostics also include a greetingActive boolean so device tests can distinguish greeting overlap from later assistant output. unavailable covers a missing or throwing settings API and missing or non-boolean fields. Diagnostics never include the track, stream, complete settings object, device identity, audio, or transcript content, and they do not prove that echo cancellation or noise suppression works acoustically.
OpenAI manages turn detection, model responses, and response audio for this mode. When OpenAI publishes audio, OpenAIAdapter attaches the remote audio stream to an autoplay audio element and exposes both microphone and playback streams through getAudioStreams(). disconnect() stops local tracks, pauses playback, clears both streams, and closes the data channel and peer connection. The adapter does not expose manual turn controls or an application-owned audio-response queue. LiveKit cascade instead uses browser VAD by default, exposes manual controls only when VAD is disabled or safely degrades, and leaves TTS to its worker.
When an application configures the optional OpenAI greeting synthesizer, a private greeting gate disables the adapter’s current microphone audio tracks while the greeting is synthesized, decoded, and played, then restores each track’s prior enabled state. Disabling sends silence without ending the tracks. That default gate does not accept spoken interruption. The experimental option below can add one playback-scoped local acceptance window; Skip greeting remains the reliable control.
Configure experimental greeting interruption
Section titled “Configure experimental greeting interruption”greetingInterruption is experimental and default-off. Omitted and false keep the current raw-track gate and do not load VAD or worklet runtime.
true selects shared defaults. An object is validated synchronously by the constructor; invalid shape, URL, pre-roll, or provider tuning throws TypeError naming greetingInterruption before mint, capture, or connection.
Enabled mode lazy-loads @maelstrom-co/realtime-vad only from connection preparation. Deploy matching VAD and ONNX assets at assetBaseUrl (default /vad-web/) with the package helper.
new OpenAIAdapter({ mintToken: mintOpenAIToken, greetingInterruption: { assetBaseUrl: '/vad-web/', },});Local acceptance is playback-scoped. Synthesis and decode are not interruption windows. Acceptance arms only after AudioBufferSourceNode.start(0) and closes synchronously at playback end or stop.
The first accepted local boundary stops greeting playback, then flushes one bounded 0–500 ms pre-roll, then live audio. Normal completion and Skip discard the retained tail and do not flush it.
Speech that begins before playback may not produce another accepted boundary. Skip greeting remains necessary.
After greeting settlement, OpenAI server VAD remains the sole turn and response owner. This option does not add turnControl or a LiveKit-style turn mode.
Loader or bridge-creation failure keeps voice connected and uses the old raw gate. Later VAD or worklet command failure recovers through generation-fenced raw-track replacement. An independent replaceTrack failure is fatal. No fallback publishes greeting-period audio.
Local acceptance may be playback echo, another sound, or speech. Requested AEC, NS, and AGC, and mock tests, do not prove those outcomes on a real device. See Turn detection in the browser for the real-device protocol; every required OpenAI greeting-interruption cell is Not tested.
Show session state
Section titled “Show session state”Use the extension’s hook for lifecycle state, connection controls, and the latest user and assistant transcript entries. The following custom control calls the public connect() and disconnect() functions from UseVoiceResult.
export function VoiceControl() { const { state, greetingState, error, userTranscript, assistantTranscript, connect, disconnect, skipGreeting, } = m.useVoice(); const [operationFailed, setOperationFailed] = useState(false); const canStart = state === 'idle' || state === 'error'; const hasFailure = operationFailed || error !== undefined; const isPreparingGreeting = greetingState === 'preparing'; const hasActiveTurn = state === 'listening' || state === 'thinking' || state === 'speaking'; const voiceLabel = isPreparingGreeting && !hasActiveTurn ? 'Preparing voice…' : state === 'connected' ? 'Listening' : `Voice: ${state}`; const canSkipGreeting = greetingState === 'preparing' && state !== 'idle' && state !== 'error';
return ( <section aria-label="Voice controls"> <button type="button" aria-describedby={hasFailure ? 'voice-error' : undefined} aria-busy={isPreparingGreeting} onClick={() => { setOperationFailed(false); // connect() and disconnect() both reject on failure. A rejected // disconnect must stay visible, so the user can retry cleanup. void (canStart ? connect() : disconnect()).catch(() => { setOperationFailed(true); }); }} > {canStart ? 'Start voice' : 'Stop voice'} </button> <p>{voiceLabel}</p> {canSkipGreeting && ( <button type="button" onClick={skipGreeting}> Skip greeting </button> )} {isPreparingGreeting && <progress aria-label="Preparing voice" />} {userTranscript && <p>You: {userTranscript.text}</p>} {assistantTranscript && <p>Assistant: {assistantTranscript.text}</p>} {hasFailure && ( <div id="voice-error" role="alert"> Voice is unavailable.{' '} {canStart ? 'Select Start voice to retry.' : 'Select Stop voice to retry cleanup.'} </div> )} </section> );}The session accumulates non-final transcript fragments and replaces them with the final entry. The control starts from idle or error; at every other point it remains a Stop voice control, including while a greeting is preparing, so the user can cancel a connection or an indefinite wait for Engine readiness.
greetingState belongs to the shared session rather than this component. When a fixed greeting is configured, preparing begins before the provider handshake and covers the Engine wait, caption, and speech attempt. The compact <progress> is indeterminate because the Engine wait has no deadline or meaningful percentage. Listening, thinking, and speaking labels take precedence when real voice activity overlaps preparation. settled changes an otherwise connected label to Listening, but it means only that greeting presentation has ended—not that speech succeeded or that the provider is currently producing audio. A retryable disconnect can return the lifecycle to pending; terminal proof remains settled if the shared provider session is recreated. A stale connection attempt cannot update the replacement session’s lifecycle.
Call skipGreeting() while greetingState is preparing to give the user immediate control. It is synchronous, shared by every useVoice() consumer, idempotent, and does not disconnect voice. A winning call sets the terminal skipped state, which survives reconnecting or recreating the provider session for the same Maelstrom session; a fresh session may greet normally. Skip remains the deterministic escape hatch, including for speech that starts too early or is not accepted. If it wins before or during the Engine wait, no acknowledgement, caption, provider item, synthesis, or playback starts. An acknowledgement or provider send already in flight can settle once, but no later irreversible side effect starts; after provider handoff, local work aborts and microphone state is restored without replaying history.
Two failure channels reach this component. A failed connect() rejects and moves the session to error, so the same failure can arrive twice — deduplicate before reporting. A rejected disconnect() only rejects; it never becomes a session error, which is why operationFailed exists rather than an empty catch that would hide a failed cleanup.
The alert copy is fixed on purpose. Before you add telemetry, decide what it may emit: allowlisted fields such as the operation and a typed error code, never raw messages, stacks, request or response bodies, credentials, transcript text, or session data. Wrap the reporting call so a synchronous throw and a rejected promise both land in one handled chain, and log a stable string if reporting itself fails.
m.VoiceOrb is the built-in alternative to a hand-written toggle. It renders a button that connects and disconnects the shared session, colours itself per state, and drives an audio-reactive meter. Take it when it fits and keep your own markup for the transcripts:
<m.VoiceOrb size={48} />It trades away the error surface built above: it discards connect() and disconnect() rejections rather than reporting them. Keep the hand-written toggle wherever a failed connection or a failed cleanup has to reach the user. size is a diameter in pixels, defaulting to 64; keep it at 24 or above so the control stays a usable touch target.
Verify
Section titled “Verify”Confirm each observable boundary before relying on the integration:
- Call the token route and confirm that
OpenAITokenGrantSchemaaccepts the response body. The successful response is a schema-valid grant withCache-Control: private, no-store. - Start voice and confirm that the session state reaches
connected. - Speak one utterance and confirm that it appears as a user transcript.
- Request a successful UI change and confirm that the intent and rendering settle before assistant narration begins. Then send malformed tool arguments or force a genuine submission failure and confirm that the failure result returns immediately without waiting for render settlement.
- Before making a mobile barge-in or echo-cancellation claim, run the device protocol in Turn detection in the browser. Mocked unit tests and debug settings verify the request and bounded observation, not the result heard by a real microphone.
Harden the mint route
Section titled “Harden the mint route”The route above authenticates, and that is all. Add these controls before it is reachable from the internet. Each one closes a specific hole:
| Control | Without it |
|---|---|
| Exact-origin allowlist, checked first | Any site can drive a browser to spend your quota |
| Pre-auth transport throttle | Unauthenticated floods reach your auth path |
| CSRF proof bound to the session | A logged-in user’s browser mints on an attacker’s behalf |
| Shared rate limit keyed by principal | One compromised account drains the account |
| Server-derived model allowlist | The browser selects an expensive model |
Cache-Control: private, no-store |
A shared cache serves one user’s credential to another |
| Redacted failure telemetry | Stack traces leak credentials and session data into logs |
| Bounded session limits | An abandoned open session bills until it is noticed |
Every control is yours to implement. None are exported by a Maelstrom package or supplied by a server runtime, and each one must fail closed — a check that throws or times out denies the request rather than falling through. createOpenAIMinter validates the final mint request and exchanges credentials; it does not decide who may ask.
Order is part of the control. The origin check runs first, before anything reads the body. The pre-auth throttle only matters when authentication and CSRF validation are not demonstrably cheap and constant-time, and then it runs before both — derive its client address from a trusted server or proxy, and reject browser-controlled forwarding headers rather than treating them as address data.
Authentication and CSRF are separate checks. Authenticating the caller supplies the server-derived sessionId and proves who is asking; it does not prove the request was intended. Validate CSRF separately, then apply the shared rate limit, keyed by principal and atomic across instances. Both throttles must deny on store failure, return a bounded Retry-After, and share one generic 429 body.
The model allowlist may read the parsed body only to pick from a list derived from server policy. Browser input must never define or widen that policy.
Every response — success, denial, malformed JSON, mapped mint failure, unexpected exception — carries Cache-Control: private, no-store and a stable public error. mintFailureHttpStatus maps the status and nothing else: no MintFailureReason, credential diagnostic, raw exception, request body, or session data reaches the response or the logs. Unexpected failures go only to your redacted-telemetry boundary, which must emit allowlisted metadata and fail without escaping the route.
Recover from failures
Section titled “Recover from failures”| Scoped outcome | Recovery |
|---|---|
400 Bad Request |
Fix the route-body or integration contract before retrying. |
401 Unauthorized |
Reauthenticate, obtain a fresh CSRF proof, and start a new mint attempt. |
403 Forbidden |
Do not retry until origin, CSRF, session authorization, or server-derived model policy is corrected. |
429 Too Many Requests |
Keep voice disconnected and wait for the bounded Retry-After interval before a deliberate retry. |
500 Internal Server Error |
Treat the mint service or its configuration as unavailable; alert operators without exposing diagnostics to the browser. |
502 Bad Gateway |
Treat OpenAI minting as temporarily unavailable and retry later with bounded backoff. |
| Grant validation fails | Reject the response; do not pass partial data to the adapter. |
| Microphone permission is denied | Explain how to enable site microphone permission, then let the user call connect() again. |
| WebRTC connection fails | Show the session’s error state, disconnect stale UI controls, and offer a deliberate reconnect. |
| A UI update is superseded | Let the latest voice correction proceed; the superseded call still receives a result, so no manual retry is required. |
What you get without asking
Section titled “What you get without asking”Two behaviors are already wired by the time voice connects. Neither needs configuration, and both are easy to break by working around them.
Speak-after-render. RealtimeSession exposes update_ui, get_ui_state, and get_capabilities to the provider. For a successful update_ui call, it withholds the successful tool result until the client reports that rendering has settled — so OpenAI describes a change the user can already see, rather than predicting one. Malformed tool arguments and genuine submission failures return failure results immediately, without waiting for render settlement.
Grounding that cannot plan. Settled UI changes produce a debounced, latest-wins structural summary on a separate, weaker channel. The adapter deletes the previous grounding item and recreates it at the conversation tail, while get_ui_state lets the model pull current structural state when it needs certainty. Grounding contains structure, not Vessel domain data, and never becomes Engine planning input.
API reference
Section titled “API reference”Use the generated reference for the public contracts used in this guide:
RealtimeSessionfrom@maelstrom-co/realtimeOpenAIAdapterfrom@maelstrom-co/realtime-openaiandOpenAITokenGrantSchemafrom@maelstrom-co/realtime-protocol/openaimaelstromRealtimeandUseVoiceResultfrom@maelstrom-co/react-realtimecreateOpenAIMinterfrom@maelstrom-co/realtime-openai-server