Skip to content

Run the voice demo locally

By the end of this tutorial you will talk to the repository demo, watch the UI rearrange itself, and hear it describe what it just did.

This is the shortest path to hearing Maelstrom speak. Realtime voice is what the demo runs by default, so there is one command, and the only key you add is an OpenAI one.

This is a learning checkpoint, not a deployment. The demo mint route is development-only and is not deployable.

  • The repository cloned, with bun install already run.
  • An OPENAI_API_KEY with access to the Engine and realtime models.
  • A browser with microphone permission.
  1. Copy the root .env.example to .env. Fill in OPENAI_API_KEY. Leave everything else as it is — the Deepgram, Cartesia, and LiveKit values belong to the cascade tutorial, and this one never reads them.

  2. From the repository root, start the demo:

    Terminal window
    bunx nx dev demo

    This starts two processes: the Engine host and the browser app. Wait until both report ready.

Open http://localhost:3000 and connect voice. The browser asks for microphone permission — allow it.

Now just talk. There is no button to hold. OpenAI decides where your turn ends, so ask for something the demo can show and then stop speaking.

Three things should happen in order: your words appear as a transcript, the UI rearranges itself, and then the model describes the result out loud.

Watch that order. The model speaks after the screen changes, never before — it is describing a layout it has already been told about, rather than predicting one. That is the settle wait doing its job.

Try one more thing: start a new request while it is still speaking. The provider handles the interruption, and if your second intent arrives while the first is still being applied, the session keeps the newer one instead of queueing both.

If nothing is spoken back, check the terminal running nx dev demo. A rejected mint is reported there, and an expired or unauthorized OPENAI_API_KEY is the usual cause.

You did not write any of this, but three parts of it are worth knowing before you build your own.

The demo chose the provider for you. It selects OpenAI unless the client is built with VITE_VOICE_PROVIDER=livekit, which apps/demo/.env.livekit sets and the dev-livekit target uses. Both modes reach the same m.useVoice() hook and the same Vessels; only the adapter differs.

Your key never reached the browser. The browser posted to /api/realtime-token, and the server minted a short-lived credential from OPENAI_API_KEY and returned only that. The same server-held key also synthesizes the one fixed greeting through a development-only same-origin TTS route (gpt-4o-mini-tts / recommended cedar voice / WAV). Ordinary Realtime conversation audio still goes from the browser straight to OpenAI over WebRTC, so it never passed through the process you started. Both mint and greeting TTS routes are development-only — replace them before any deployment. The OpenAI Realtime guide shows the route you would deploy in place of the demo’s.

The greeting is AI-generated speech. The demo discloses that next to the voice orb. It is not a human voice recording, and it uses the same recommended built-in cedar voice as later model replies.

Nothing you said reached the Engine as structure. The model called one tool with one sentence, and the Maelstrom client attached the Chart snapshot itself. That is why speaking to the demo could not break a Vessel, and it is the subject of The voice boundary.