Skip to content

Why Maelstrom?

Traditional UIs are built for a world where developers know in advance what the user will need. In AI-powered applications, that assumption breaks: users ask for outcomes — “help me plan tomorrow around the rain,” “show the meetings that affect my launch checklist” — and the useful interface is the one that assembles the right pieces at the moment of need. Most UI systems treat arrangement as a build-time decision; AI interaction makes arrangement a runtime problem.

The industry has largely converged on two common approaches to that problem.

Bolt a chat box onto every app. The UI stays fixed; the LLM answers in text or fires a few predefined tools. This works well when the interaction is genuinely conversational, but the model never gets to use the interface — it can only describe the available actions, such as which buttons the user could click. As the number of things a user might want to do grows, walls of text replace working software and human reading speed becomes the bottleneck.

Let the model generate the interface. Generative UI invents components from a prompt, trading a team’s existing component library for on-the-fly generation. That’s fine for a single, simple answer, but every widget is a one-shot — no shared state, no layout system, no way to reason about whether two things fit on screen together. Coordinating more than one component at a time is where this approach runs into trouble.

Both miss the point. The interface doesn’t need to be retired into a chat log, and it doesn’t need to be generated from scratch. It needs to adapt — composed by AI in the moment, out of pieces you designed on purpose. Maelstrom exists for that gap: a real component system the LLM is allowed to drive — not by emitting markup, but by choosing among components, variants, and positions the application defined up front.

Three commitments separate Maelstrom from free-form generative UI:

  1. The component system stays yours. You declare Vessels — components with explicit state machines (variants and actions). The LLM never invents new components, never writes JSX. It chooses among the ones you shipped.
  2. The model reasons over a small, validated domain. Branded IDs. Discrete variants. Layout as integers or split tokens instead of pixels. A minSize per variant. Every input the LLM sees is constrained and typed; every output it returns is parsed against a schema before anything renders.
  3. Decisions and rendering are different jobs. The engine decides; the framework binding renders. The two halves can swap independently — replace the reference engine without touching the React app, replace React without touching the engine.

Those three together are the bet: if you give the model a closed world of well-typed components and a discrete surface to place them on, it can orchestrate UIs reliably enough to put in front of users.

Benefits of state machines and a discrete layout surface

Section titled “Benefits of state machines and a discrete layout surface”

Both choices look opinionated. Both are load-bearing.

State machines (The Vessel State Machine) because the LLM needs a closed set of legal moves. When the Calendar is collapsed, the model sees expand and nothing else — it cannot ask for showWeek and silently fail. Each variant also declares its minSize — the footprint that state needs — which the engine holds grid layouts to. That’s what turns “collapse Calendar to fit Weather” into a move the system can represent and enforce, rather than a guess about pixels.

A discrete surface because the smallest thing that still describes layout is not a pixel. The LLM never sees one. It works in whichever of two modes the app chose, and both are small enough to validate exhaustively:

  • Tiling (the default) — a split tree serialized as a preorder token array: H to split top/bottom, V to split left/right, ID: followed by an instance ID for a leaf, null for empty space. A bare split halves its region; a weighted one such as V:1,1,1 names both how many children it takes and the share each gets. There are no coordinates; the shape of the tree is the layout.
  • Grid — explicit integer placement: row, column, width, height, in whole cells.

The container computes its own pixel size in both cases; the model never has to.

The result is an architecture an LLM can drive with the failure modes you’d want: the schema rejects a response that’s malformed outright, and the heal pass drops any schema-valid layout element that still violates a domain constraint — like minSize or the grid bounds — so the rest of the turn still renders. For how a decision actually flows through the system, see How the LLM Orchestrates UI.

A separation that mirrors how teams already think:

  • Designers and engineers define Vessels. Every component has a clear contract — variants, actions, sizes — documented in code.
  • Product decides which Vessels exist in the app. The set of Vessels is the product surface.
  • The LLM decides which ones to show, how, and where. Not what to build. What to use.

And a separation that matches how systems should be built:

  • Protocol — transport-agnostic types both sides speak.
  • Engine — decision core, swappable; holds no layout state, and conversation history is opt-in.
  • Client — framework-agnostic runtime that applies decisions.
  • Framework bindings — @maelstrom-co/react today, others later.

Each layer is independently testable, independently replaceable, and small enough to read in an afternoon. See the Architecture Overview for how the pieces fit together.

A Maelstrom interface moves from “find the right screen” to “say the goal.” The user still sees real product components, but the arrangement reflects the current task instead of a fixed navigation path. A planning assistant keeps tasks compact while expanding the calendar. A support dashboard puts customer history next to live diagnostics. A workspace rearranges visible tools without asking the user to drag panels by hand.

Maelstrom adds orchestration machinery, so it is not the right default for every app.

Use a conventional UI when:

  • the product has one primary flow and little need for runtime layout changes
  • every screen must be fully deterministic for compliance or training reasons
  • the application has too few components for orchestration to add value
  • you are not prepared to write clear Vessel descriptions and action boundaries

Use Maelstrom when the interface has multiple meaningful components and the best combination depends on what the user asks next.