Skip to content

The Engine

The Engine is Maelstrom’s decision core. Given the current chart state and a user intent, it asks an LLM to produce a layout and an execution plan, then returns a structured response.

@maelstrom-co/engine is a reference implementation, not a mandate. Production systems may keep this package if its concurrency, history, logging, and safety trade-offs match their needs, or replace it with another implementation behind the same Engine.handle() contract. It ships so you can run Maelstrom locally, build demos, and understand how the contract works. It is not production-grade — production deployments supply their own implementation behind the same boundary.

Most React applications never import the engine directly. They use @maelstrom-co/react, which orchestrates the client and talks to whichever engine is configured. You’ll reach for the engine package when you’re:

  • Building a server that accepts chart requests over your own transport
  • Embedding Maelstrom into a non-React framework
  • Writing tests against engine behaviour
  • Implementing your own engine to swap in production

The engine exposes a single durable contract:

Engine.handle()
interface Engine {
handle(
message: EngineMessage,
signal?: AbortSignal,
): Promise<Result<EngineOutput, EngineError>>;
}

Inputs and outputs are tagged envelopes (see The Protocol). Every message carries a sessionId, id, and timestamp. Outputs carry sessionId and id so callers can correlate request and response. The wire format is the engine format — connectors and transports are pure passthrough.

This signature is the contract. The implementation behind it can change without anything else needing to know.

flowchart TD
  Request["ChartRequest"]
  Prompt["Build prompt"]
  Adapter["Call LLM adapter"]
  Parse["Parse structured JSON"]
  Validate["Validate protocol schema"]
  Heal["Heal recoverable<br/>layout/action issues"]
  Retry{"Would the user<br/>see a no-op?"}
  Return["Return EngineOutput"]

  Request --> Prompt
  Prompt --> Adapter
  Adapter --> Parse
  Parse --> Validate
  Validate --> Heal
  Heal --> Retry
  Retry -->|"yes, within budget"| Prompt
  Retry -->|"no"| Return

For a chart_request, the engine runs:

  1. Build the prompt from your vessels, the current chart state, and any update reasons
  2. Call the configured LLM adapter
  3. Parse and validate the structured response
  4. Heal recoverable problems (drop overflowed positions, prune unreachable layouts)
  5. Retry once if the result would be empty
  6. Record the turn in history and return the response

Errors short-circuit the pipeline; recoverable problems become diagnostics that consumers don’t see.

The engine validates inputs at three boundaries and the LLM’s output at a fourth:

  1. Engine config: createEngine() validates EngineConfig against a strict Zod schema; llmPlanner() (from @maelstrom-co/planner-llm) validates its own options the same way and eagerly builds system prompts at construction. Misconfigured system prompt customizers throw synchronously from llmPlanner(...), but user prompt customizer errors still surface on the first request as a ValidationError.
  2. Engine messages — every inbound EngineMessage envelope is parsed against the protocol Zod schema. Malformed payloads short-circuit before any LLM call with a RequestValidationError carrying the structured path of the failing field.
  3. LLM structured output — the JSON the model returns is parsed against the protocol schemas inside the adapter call. Schema-violating output surfaces as a ProviderError; a structurally valid response missing required content (such as an empty userMessage) surfaces as a ParseError.
  4. LLM output coherence — even a parseable response can be incoherent: positions overflowing the chart, unknown vessel/variant references, duplicate instance IDs, malformed tiling trees. This layer doesn’t fail; it heals.

For step 4, an empty UI is worse than a partial one. The heal layer drops invalid pieces and records each one with a structured reason — categorical labels such as position_overflow, unknown_vessel, unknown_variant, size_below_minimum, duplicate_instance_id, or tree_missing_child. Soft-failures like a missing rationale are recorded as warnings. Drop entries and warnings stay engine-internal: they feed the retry decision and engine logs, but the engine strips them at its boundary. Consumers receive only the surviving layout; rendering proceeds with whatever made it through.

If the model returns two positions and one overflows the chart, the engine can keep the valid position and drop the invalid one:

flowchart TD
  Model["Model output<br/>weather row 0<br/>calendar row 99"]
  Bounds["Chart size<br/>12 columns x 8 rows"]
  Heal["Heal layer"]
  Weather["weather survives"]
  Calendar["calendar drops<br/>position_overflow"]
  User["User sees partial useful layout<br/>instead of a failed turn"]

  Model --> Heal
  Bounds --> Heal
  Heal --> Weather
  Heal --> Calendar
  Weather --> User

The dropped entry is logged for diagnostics, but consumers receive only the surviving layout.

A chart resize is the one trigger that arrives with no user or vessel signal, so the chart state the LLM sees on such a turn is identical to the state before an action it already ran, and the engine never sees vessel contents, so it cannot tell a first run from a replay. On a turn whose every update reason is a chart_resize, an action runs only if its definition declares idempotent: true and the model emitted it with no parameters. Missing the flag is recorded as non_idempotent_on_resize; carrying parameters despite the flag is recorded as parameterized_on_resize.

The parameter half is what the flag alone cannot cover. idempotent promises that running an action twice with the same parameters matches running it once; it says nothing about running it with different ones. A resize asks for nothing, so any parameter value the model emits can only have come from the conversation history: showWeek() is a reshape the new size may call for, while setLocation("Tokyo") is either a replay of an earlier turn or an overwrite nobody asked for, even on an action marked idempotent. Parameters count as the model emitted them, so an action free to take parameters but invoked with none still runs.

One escape stays open by design: a resize batched together with a user intent or vessel mutation is an ordinary turn, and the guard never fires on it. A freshly declared instance id is not a second escape, because a vessel may back every instance of itself with one store, so an action replayed at a new id can still duplicate state an older id already holds. The same predicate drives a prompt-side instruction, so the LLM is told to return an empty plan before the guard ever has to refuse one.

After healing, the engine recommends retry only when the user would otherwise see a no-op:

  • Empty-when-needed layout — the request asked for instances but every position dropped (severity unrecoverable).
  • All actions skipped for fixable reasons — the proposed plan was non-empty but every action was rejected for a reason the LLM could plausibly fix on retry, e.g. unknown_vessel, params_parse_failed, action_not_valid (severity medium).

A partial layout or partial action plan does not trigger retry — partial output beats waiting for a re-roll. Non-fixable skip reasons (state_drift, self_transition_noop, non_idempotent_on_resize, parameterized_on_resize) don’t recommend retry on their own, but they don’t suppress retry either: if any other skip in the same response is fixable and no actions survived, retry still fires. The retry budget is EngineConfig.retry (default { max: 1 }).

Configuring the engine
import { geminiText } from '@tanstack/ai-gemini';
import { createEngine } from '@maelstrom-co/engine';
import { llmPlanner } from '@maelstrom-co/planner-llm';
const engine = createEngine({
planner: llmPlanner({ adapter: geminiText('gemini-2.0-flash') }),
// logger, history store, instrument, retry policy: all optional.
});

The pluggable surfaces are intentionally small:

  • planner — produces the candidate response for every chart_request. llmPlanner({ adapter, prompt?, modelOptions? }) from @maelstrom-co/planner-llm makes a single structured-output LLM call per attempt; adapter swaps the provider (Gemini, OpenAI, Anthropic, your own via TanStack AI), and prompt replaces, prepends, appends, or transforms individual prompt sections. pipeline, withFallback, and withRetry compose more than one stage into a planner — see the engine README.

  • history — opt-in conversation history with a store you supply (Redis, Postgres, SQLite). The reference engine ships an in-memory store for tests but never as a default.

  • instrument is the telemetry seam, defaulting to noopInstrument, which does nothing. pipeline wraps every stage with instrument.stage(name, stage) once per request, and stages report model calls and other events with instrument.event(name, attributes), a flat attribute bag keyed by OpenTelemetry GenAI names where they apply. A logging implementation is five lines:

    A logging instrument
    const instrument: Instrument = {
    stage: (name, stage) => async (input, ctx) => {
    ctx.logger.info('stage', { name });
    return stage(input, ctx);
    },
    event: (name, attributes) => console.info(name, attributes),
    };
  • retry — { max } re-runs the planner while healing recommends a retry; false runs it exactly once.

  • logger — any object with debug / info / warn / error works.

  • Multi-instance safe. The per-session mutex is in-process. Production deployments need session affinity or a distributed lock.
  • Persistent. The bundled in-memory store is for tests; supply your own store for anything that needs to survive a process restart.
  • Hardened against adversarial input. It’s a reference, not a security boundary.

If your use case crosses any of these lines, plan to replace this engine with one of your own. The Engine interface is the only thing you need to keep stable.