Skip to content

Overview

Voice lets people speak to a Maelstrom interface, and lets the interface speak back.

A person says what they want out loud. The UI adapts to it, the same way it adapts to typed input. Nothing about your Vessels changes.

Maelstrom ships two voice modes. Pick one by deciding where the conversation, the speech-to-text (STT), and the text-to-speech (TTS) should run.

Realtime sends audio straight to a speech-native model. You get lower latency and fewer moving parts, and the provider owns the pipeline.

Cascade splits STT and TTS into a worker you operate. The shipped worker uses Deepgram and Cartesia, while the browser owns automatic turn detection and manual fallback. Replacing the worker enables other speech providers but requires a custom integration. The extra stages add operational work and may add latency.

Whichever you choose, the client — never the voice provider — builds the Chart snapshot the Engine plans from. Read The voice boundary for why that is structural rather than a rule you follow, and for the full comparison of latency, provider control, specialization, operations, observability, and privacy.

Both modes reach your application the same way. Choose a provider adapter, then pass its factory to the provider-neutral extension. React is the binding available today, and your existing Vessels are untouched:

src/maelstrom.ts
export const m = createMaelstrom({ vessels }).with(
maelstromRealtime({ createAdapter }),
);

Every adapter then reaches your UI through the same m.useVoice() hook and m.VoiceOrb component. See the React voice API reference for the session contract those expose, and the Voice API reference for the provider-neutral session underneath. Your chosen guide below builds the full control.

Each package runs in exactly one place. That is what keeps OPENAI_API_KEY, LiveKit secrets, and speech-provider keys out of browser code, and provider transport details out of the Engine. Your chosen guide lists the packages that mode actually needs.

Package Runtime Responsibility
@maelstrom-co/realtime Browser Provider-neutral session and intent routing
@maelstrom-co/realtime-openai Browser OpenAI WebRTC transport
@maelstrom-co/react-realtime Browser React binding: useVoice(), transcripts, and VoiceOrb
@maelstrom-co/realtime-openai-server Server Mint a short-lived OpenAI client credential
@maelstrom-co/realtime-livekit Browser LiveKit room transport and turn control
@maelstrom-co/realtime-livekit-server Server Sign a short-lived grant for one LiveKit room
@maelstrom-co/realtime-livekit-agent Worker Run Deepgram STT and Cartesia TTS, with no LLM

Voice remains an optional input to the Engine’s normal intent flow. When voice is part of the product goal, use the voice integration prompt for a standalone task or select the optional voice lead skill from the portable skill bundle. The workflow keeps provider-specific behavior and the browser/server/worker boundaries explicit.

Integrate voice (optional)

Plan or implement voice for this Maelstrom application only if the requested
product goal requires it. Voice is an optional intent input, not a second planner.

Product goal: [USER GOAL]
Provider/deployment: [OPENAI REALTIME | LIVEKIT CASCADE | OTHER | UNDECIDED]
App/auth boundary: [DETAILS OR "inspect"]
Installed Maelstrom voice versions: [KNOWN OR "inspect"]
Implementation authorized: [YES / NO]

Inspect the app's current intent flow, React shell, server routes/deployment,
installed package versions, and the canonical voice architecture before proposing
changes. Read `docs/realtime-voice-architecture.md` and ADRs 0001, 0002, and 0003
from the app checkout when present. For a portable skill bundle, resolve the
installed skill root (commonly `.agents/skills/` or `.claude/skills/`) and read
`references/canonical-voice/docs/realtime-voice-architecture.md` and the ADRs
under `references/canonical-voice/docs/adr/`. For standalone use, retrieve the
canonical files from the Maelstrom source repository at
`https://github.com/maelstrom-co/maelstrom/tree/main/docs/` and
`https://github.com/maelstrom-co/maelstrom/tree/main/docs/adr/`. If none is
available, ask for the materials before implementation; do not infer their
contents or assume a public guide links the ADRs. Confirm selected
APIs against installed package exports/types/source/tests. Use a provider guide
as version-specific evidence only when its version matches the installed SDK; label
current external docs as unversioned when that match is unknown. Do not assume a
provider or add voice packages until the user goal and provider/deployment
decision justify them. If a decision is missing, return the bounded decision
needed rather than choosing a provider.

Keep the responsibilities distinct:

- **Browser/React:** voice UI, microphone interaction, session presentation, and
  appropriate React binding. Voice controls are not automatically a Vessel.
- **Realtime:** provider adapter/session lifecycle and turn-taking, with provider
  differences preserved rather than hidden behind invented common behavior.
- **Server/worker:** long-lived provider keys stay out of browser code. A
  server or LiveKit worker may hold only the credentials its documented role
  requires; for the supported LiveKit cascade, the worker uses LiveKit,
  Deepgram, and Cartesia credentials while the browser receives a short-lived
  room grant. Application auth, authorization, origin policy, and rate limits
  remain app-owned; a package minter does not authenticate callers.
- **Planning:** speech becomes ordinary natural-language intent. The Engine stays
  the planner and uses client-owned Chart state. Turn mode is runtime behavior,
  not a fixed provider flag.

Honor the four-layer boundary: CONTRACT, BROWSER, SERVER, WORKER. No server or
worker package may depend on a browser package. A worker must never receive
Vessel/Chart structure or become a second planner. Distinguish actual provider support
and deployment requirements in the checked-in guides from hypothetical options.

For an authorized bounded implementation, preserve unrelated application flows,
add deterministic tests for locally testable lifecycle and failure behavior, and
run affected checks. Do not claim microphone, acoustic, live provider, or Engine
success unless that exact scenario was exercised. Return a layer-separated
integration map, provider/auth decisions and unresolved questions, version/API
anchors, changed files or plan, test commands/results, security and turn-mode
risks, and unrun manual scenarios.