Overview
Voice lets people speak to a Maelstrom interface, and lets the interface speak back.
A person says what they want out loud. The UI adapts to it, the same way it adapts to typed input. Nothing about your Vessels changes.
Choose a voice mode
Section titled “Choose a voice mode”Maelstrom ships two voice modes. Pick one by deciding where the conversation, the speech-to-text (STT), and the text-to-speech (TTS) should run.
Realtime sends audio straight to a speech-native model. You get lower latency and fewer moving parts, and the provider owns the pipeline.
Cascade splits STT and TTS into a worker you operate. The shipped worker uses Deepgram and Cartesia, while the browser owns automatic turn detection and manual fallback. Replacing the worker enables other speech providers but requires a custom integration. The extra stages add operational work and may add latency.
Whichever you choose, the client — never the voice provider — builds the Chart snapshot the Engine plans from. Read The voice boundary for why that is structural rather than a rule you follow, and for the full comparison of latency, provider control, specialization, operations, observability, and privacy.
Wire voice into your UI
Section titled “Wire voice into your UI”Both modes reach your application the same way. Choose a provider adapter, then pass its factory to the provider-neutral extension. React is the binding available today, and your existing Vessels are untouched:
export const m = createMaelstrom({ vessels }).with( maelstromRealtime({ createAdapter }),);Every adapter then reaches your UI through the same m.useVoice() hook and m.VoiceOrb component. See the React voice API reference for the session contract those expose, and the Voice API reference for the provider-neutral session underneath. Your chosen guide below builds the full control.
Know the package boundaries
Section titled “Know the package boundaries”Each package runs in exactly one place. That is what keeps OPENAI_API_KEY, LiveKit secrets, and speech-provider keys out of browser code, and provider transport details out of the Engine. Your chosen guide lists the packages that mode actually needs.
| Package | Runtime | Responsibility |
|---|---|---|
@maelstrom-co/realtime |
Browser | Provider-neutral session and intent routing |
@maelstrom-co/realtime-openai |
Browser | OpenAI WebRTC transport |
@maelstrom-co/react-realtime |
Browser | React binding: useVoice(), transcripts, and VoiceOrb |
@maelstrom-co/realtime-openai-server |
Server | Mint a short-lived OpenAI client credential |
@maelstrom-co/realtime-livekit |
Browser | LiveKit room transport and turn control |
@maelstrom-co/realtime-livekit-server |
Server | Sign a short-lived grant for one LiveKit room |
@maelstrom-co/realtime-livekit-agent |
Worker | Run Deepgram STT and Cartesia TTS, with no LLM |
Work with a coding agent
Section titled “Work with a coding agent”Voice remains an optional input to the Engine’s normal intent flow. When voice is part of the product goal, use the voice integration prompt for a standalone task or select the optional voice lead skill from the portable skill bundle. The workflow keeps provider-specific behavior and the browser/server/worker boundaries explicit.
Integrate voice (optional)
Plan or implement voice for this Maelstrom application only if the requested
product goal requires it. Voice is an optional intent input, not a second planner.
Product goal: [USER GOAL]
Provider/deployment: [OPENAI REALTIME | LIVEKIT CASCADE | OTHER | UNDECIDED]
App/auth boundary: [DETAILS OR "inspect"]
Installed Maelstrom voice versions: [KNOWN OR "inspect"]
Implementation authorized: [YES / NO]
Inspect the app's current intent flow, React shell, server routes/deployment,
installed package versions, and the canonical voice architecture before proposing
changes. Read `docs/realtime-voice-architecture.md` and ADRs 0001, 0002, and 0003
from the app checkout when present. For a portable skill bundle, resolve the
installed skill root (commonly `.agents/skills/` or `.claude/skills/`) and read
`references/canonical-voice/docs/realtime-voice-architecture.md` and the ADRs
under `references/canonical-voice/docs/adr/`. For standalone use, retrieve the
canonical files from the Maelstrom source repository at
`https://github.com/maelstrom-co/maelstrom/tree/main/docs/` and
`https://github.com/maelstrom-co/maelstrom/tree/main/docs/adr/`. If none is
available, ask for the materials before implementation; do not infer their
contents or assume a public guide links the ADRs. Confirm selected
APIs against installed package exports/types/source/tests. Use a provider guide
as version-specific evidence only when its version matches the installed SDK; label
current external docs as unversioned when that match is unknown. Do not assume a
provider or add voice packages until the user goal and provider/deployment
decision justify them. If a decision is missing, return the bounded decision
needed rather than choosing a provider.
Keep the responsibilities distinct:
- **Browser/React:** voice UI, microphone interaction, session presentation, and
appropriate React binding. Voice controls are not automatically a Vessel.
- **Realtime:** provider adapter/session lifecycle and turn-taking, with provider
differences preserved rather than hidden behind invented common behavior.
- **Server/worker:** long-lived provider keys stay out of browser code. A
server or LiveKit worker may hold only the credentials its documented role
requires; for the supported LiveKit cascade, the worker uses LiveKit,
Deepgram, and Cartesia credentials while the browser receives a short-lived
room grant. Application auth, authorization, origin policy, and rate limits
remain app-owned; a package minter does not authenticate callers.
- **Planning:** speech becomes ordinary natural-language intent. The Engine stays
the planner and uses client-owned Chart state. Turn mode is runtime behavior,
not a fixed provider flag.
Honor the four-layer boundary: CONTRACT, BROWSER, SERVER, WORKER. No server or
worker package may depend on a browser package. A worker must never receive
Vessel/Chart structure or become a second planner. Distinguish actual provider support
and deployment requirements in the checked-in guides from hypothetical options.
For an authorized bounded implementation, preserve unrelated application flows,
add deterministic tests for locally testable lifecycle and failure behavior, and
run affected checks. Do not claim microphone, acoustic, live provider, or Engine
success unless that exact scenario was exercised. Return a layer-separated
integration map, provider/auth decisions and unresolved questions, version/API
anchors, changed files or plan, test commands/results, security and turn-mode
risks, and unrun manual scenarios.