Skip to content

Use a custom LLM provider

Provider choice changes the three things users feel first: structured-output quality, latency, and cost. Stronger models usually need less healing, faster models reduce turn time, and cheaper models matter when your Vessel catalogue makes prompts large.

The reference engine talks to the model through a TanStack AI text adapter. Swapping providers is a single field on llmPlanner({ adapter }) from @maelstrom-co/planner-llm: the prompt builder, response parser, heal layer, and retry policy are all provider-agnostic.

  • An installed @maelstrom-co/engine package and an installed @maelstrom-co/planner-llm package.
  • An API key for the provider you want to use, exposed as the environment variable the adapter expects.
  1. Install the adapter package for the provider you want.

    Terminal window
    bun add @tanstack/ai-gemini
  2. Expose the provider’s API key. Each TanStack AI adapter looks up its key from an env var owned by the adapter package itself, not by @maelstrom-co/engine — check the provider package’s README if a name changes.

    .env
    GOOGLE_API_KEY=...
    # Or GEMINI_API_KEY — the adapter tries both.
  3. Pass the adapter to llmPlanner, and the planner to createEngine. The adapter factory takes a model id; the engine carries the planner through every chart_request.

    src/server/engine.ts
    import { createEngine } from '@maelstrom-co/engine';
    import { llmPlanner } from '@maelstrom-co/planner-llm';
    import { geminiText } from '@tanstack/ai-gemini';
    export const engine = createEngine({
    planner: llmPlanner({ adapter: geminiText('gemini-2.0-flash') }),
    });

That’s the entire swap. Send a chart_request and the engine will route through the new provider.

Provider catalogues change faster than these docs do — the model ids above are illustrative, not recommendations. Check the provider and TanStack AI adapter documentation before pinning a production model. Pick from the provider’s own catalogue, then weight the choice along three axes:

  • Structured output quality. The Engine asks for JSON validated against the protocol schema. Its heal layer is a coherence pass that drops invalid layout or Action entries from an otherwise valid response. Stronger models need that recovery less often, which trims retries and latency. The bundled retry policy (default retry: { max: 1 }) papers over most slips but is not free.
  • Latency. The LLM call is the dominant cost of a turn. “Flash” / “mini” tier models trade some accuracy for first-token latency; the heal layer can recover from some invalid entries, but it cannot guarantee a useful response.
  • Cost. Maelstrom prompts include the vessel/variant catalogue, so input tokens scale with how many vessels you register, not how chatty the user is.

The adapter is set at llmPlanner time and is the same for every turn. If you need per-request routing (A/B test, role-based model selection, fallback on quota errors), stand up multiple engines:

src/server/route.ts
import { createEngine } from '@maelstrom-co/engine';
import { llmPlanner } from '@maelstrom-co/planner-llm';
import { geminiText } from '@tanstack/ai-gemini';
import { openaiText } from '@tanstack/ai-openai';
const flashEngine = createEngine({
planner: llmPlanner({ adapter: geminiText('<flash-model-id>') }),
});
const proEngine = createEngine({
planner: llmPlanner({ adapter: openaiText('<pro-model-id>') }),
});
function pickEngine(userTier: 'free' | 'pro') {
return userTier === 'pro' ? proEngine : flashEngine;
}

Each engine carries its own history manager and concurrency state. Don’t share a single HistoryStore between engines that target different model families unless your deriveLLMContext projection accounts for the mix.

LlmPlannerOptions.adapter is typed as AnyTextAdapter from @tanstack/ai. Any object satisfying that contract — including your own — is accepted.

Implementing a custom adapter is out of scope for this guide. See the TanStack AI docs for the AnyTextAdapter contract; neither llmPlanner nor the engine extends it.

Use this when you need to call a model the TanStack AI catalogue does not cover yet (an internal endpoint, a fine-tuned deployment, a gateway), or when you want to wrap an existing adapter to add logging, request signing, or fallback chains.

Send a chart_request through the engine and inspect the result. If the adapter errors, llmPlanner logs LLM call failed through the engine’s configured logger with the structured EngineError attached, and the engine returns Err(ProviderError):

import { isOk } from '@maelstrom-co/protocol';
const result = await engine.handle(chartRequestMessage);
if (!isOk(result) && result.error._tag === 'ProviderError') {
console.error('Adapter error:', result.error.message);
}

If the result is Ok with a chart_response whose response.userMessage is non-empty, the new adapter is live.