fabro/docs/public/core-concepts/models.mdx
Bryan Helmkamp 3e065b806b
Bump lithos-llm to 43a42ac and migrate catalogs to the codecs schema
Move the lithos-llm pin from 55add459 to 43a42ac28e9d9bcf40a91abc02be4f12ca274ebb,
and the three Pebble pins from a39f43e to 67c9f48, Pebble `main`, which pins
that same lithos-llm revision so Cargo holds one lithos-llm crate. lithos-llm
`main` (ca19fac) is one commit further; that commit touches only its nightly
workflow, so this pin stays on the revision Pebble unifies with.

The `openai`, `anthropic`, `gemini`, and `openai-compatible` features are
gone upstream; each expanded to `runtime`, which `bedrock` implies, so the
four names leave the fabro-llm feature list. Every other manifest already
names `runtime`.

The catalog schema now names one adapter and many codecs per provider.
`adapter` defaults to `http`, `codecs = [...]` replaces `codec` and defaults
to `["openai-chat"]`, and the loader rejects the old `codec` key and the four
protocol-named adapter ids. Every inline catalog in tests and docs moves to
the new shape: the `openai-compatible` + `openai-chat` pair is dropped as the
default, `adapter = "openai"` + `codec = "openai-responses"` becomes
`codecs = ["openai-responses"]`, and the one test that swaps in a custom
adapter id now adds the line instead of replacing one. The settings
reference, the API schema's `Provider.adapter` description, and the SDK page
describe the new fields; the three `docs/superpowers/plans/` files that show
the old shape are dated, unchecked historical plans and are left as they are.

The implied agent profile for an operator provider that declares none used
to read the removed protocol adapter ids; it now reads the provider's first
codec (Anthropic Messages and Gemini map to their harnesses, the `bedrock`
adapter to Anthropic, everything else to OpenAI), with a test for the codec
path.

Absorbing the rest of the range: OpenRouter and Fireworks now ship enabled,
so the two fabro-llm tests that used OpenRouter as the disabled fixture use
`bedrock-openai`, and the docs and comments that said the two ship disabled
are corrected. The built-in catalog grew past 100 enabled model rows
(Vercel, TypeSafe, and the enabled OpenRouter and Fireworks rosters), so the
pagination shape test walks `page[offset]` to the last page instead of
assuming one page fits.

`cargo update -p` on the four crates also re-resolved a few already-locked
edges to match the lithos-llm lockfile: `windows-sys` 0.61.2/0.60.2 ->
0.59.0 under dirs-sys, errno, nu-ansi-term, quinn-udp, rustix,
rustls-platform-verifier, tempfile, terminal_size, and winapi-util;
`windows-core` 0.61.2 -> 0.62.2 under iana-time-zone; `errno` 0.2.8 ->
0.3.14 under signal-hook-registry; and `indexmap` 2.13.0 as a new public
dependency of lithos-llm. No package version was added or removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:07:00 -04:00

316 lines
18 KiB
Text

---
title: "Models"
description: "How Fabro routes tasks to LLM models and providers"
---
No single model is best at everything. Fabro lets you assign the right model to each workflow step — cheap, fast models for boilerplate, frontier models for hard reasoning, and a different provider for cross-critique so the reviewer brings fresh eyes. When a provider goes down, Fabro can fail over automatically.
<Frame>
<img src="/images/ensemble-workflow.svg" alt="Ensemble workflow: fan out to Opus and Gemini Pro, merge, then synthesize" />
</Frame>
## Model catalog
Fabro keeps five related concepts separate:
- A **provider** serves requests, such as `openai` or `openrouter`.
- A **model slug** is Fabro's canonical, human-facing model ID, such as `gpt-5.6-sol`.
- An **alias** is another user-facing selector, such as `gpt-56-sol`.
- A **family** is display and matching metadata. It does not affect routing identity.
- An **API ID** is the opaque model string sent on the provider wire. Workflows should never reference it.
One provider's route to one model slug is an **offering**, identified by `(provider, model slug)`. The same slug and alias may appear on several providers. Within one provider, however, every slug or alias must identify exactly one offering.
For an unqualified selector, Fabro checks canonical slugs before aliases, filters the candidates to providers whose adapters are ready, and then chooses the highest provider priority. Equal priorities use canonical provider ID order. An explicit provider restricts selection to that provider and is a pin: if it is unavailable, Fabro reports the error instead of silently switching.
For example, the shared `gpt-56-sol` alias can be portable across direct OpenAI and OpenRouter offerings:
| Ready providers | Selector | Selected offering |
|---|---|---|
| OpenAI only | `gpt-56-sol` | `openai/gpt-5.6-sol` |
| OpenRouter only | `gpt-56-sol` | `openrouter/gpt-5.6-sol` |
| OpenAI and OpenRouter | `gpt-56-sol` | OpenAI, because it has higher priority |
| Both, with `provider = "openrouter"` | `gpt-56-sol` | OpenRouter, because the provider is pinned |
Fabro performs this selection once when creating a run and persists the chosen provider and canonical slug in the run settings and graph. Resuming that run does not reconsider provider priority when credentials change. Runtime [model fallbacks](/execution/failures#model-fallbacks) are the separate mechanism for handling a later provider failure.
| Model | Provider | Aliases | Context | Cost (in/out per Mtok) | Speed |
|---|---|---|---|---|---|
| `claude-fable-5` | anthropic | `fable`, `claude-fable` | 1M | $10.00 / $50.00 | n/a |
| `claude-opus-5` | anthropic | `opus`, `claude-opus` | 1M | $5.00 / $25.00 | n/a |
| `claude-opus-4-8` | anthropic | | 1M | $5.00 / $25.00 | 25 tok/s |
| `claude-opus-4-7` | anthropic | | 1M | $5.00 / $25.00 | 25 tok/s |
| `claude-opus-4-6` | anthropic | | 1M | $5.00 / $25.00 | 25 tok/s |
| `claude-sonnet-4-6` | anthropic | `sonnet`, `claude-sonnet` | 200K | $3.00 / $15.00 | 50 tok/s |
| `claude-sonnet-4-5` | anthropic | | 200K | $3.00 / $15.00 | 50 tok/s |
| `claude-haiku-4-5` | anthropic | `haiku`, `claude-haiku` | 200K | $0.80 / $4.00 | 100 tok/s |
| `gpt-5.6-sol` | openai | `sol`, `gpt-sol`, `gpt56-sol`, `gpt-56-sol`, `gpt-5.6`, `gpt56`, `gpt-56` | 272K | $5.00 / $30.00 | n/a |
| `gpt-5.6-terra` | openai | `terra`, `gpt-terra`, `gpt56-terra`, `gpt-56-terra` | 272K | $2.50 / $15.00 | n/a |
| `gpt-5.6-luna` | openai | `luna`, `gpt-luna`, `gpt56-luna`, `gpt-56-luna` | 272K | $1.00 / $6.00 | n/a |
| `gpt-5.4` | openai | `gpt54`, `gpt5`, `codex` | 272K | $2.50 / $15.00 | 70 tok/s |
| `gpt-5.5` | openai | `gpt55` | 272K | $5.00 / $30.00 | 70 tok/s |
| `gpt-5.5-pro` | openai | `gpt55-pro` | 1M | $30.00 / $180.00 | 20 tok/s |
| `gpt-5.4-mini` | openai | `gpt54-mini`, `codex-spark` | 272K | $0.75 / $4.50 | 140 tok/s |
| `gpt-5.4-pro` | openai | `gpt54-pro` | 1M | $30.00 / $180.00 | 20 tok/s |
| `gemini-3.1-pro-preview` | gemini | `gemini-pro` | 1M | $2.00 / $12.00 | 85 tok/s |
| `gemini-3.1-pro-preview-customtools` | gemini | `gemini-customtools` | 1M | $2.00 / $12.00 | 85 tok/s |
| `gemini-3.5-flash` | gemini | `gemini-35-flash` | 1M | $1.50 / $9.00 | 150 tok/s |
| `gemini-3-flash-preview` | gemini | `gemini-flash` | 1M | $0.50 / $3.00 | 150 tok/s |
| `gemini-3.1-flash-lite` | gemini | `gemini-flash-lite`, `gemini-3.1-flash-lite-preview` | 1M | $0.25 / $1.50 | 200 tok/s |
| `kimi-k2.5` | moonshot | | 262K | $0.60 / $3.00 | 50 tok/s |
| `kimi-k3` | moonshot | `kimi` | 1M | $3.00 / $15.00 | n/a |
| `kimi-k3-fast` | venice | `kimi-fast` | 1M | $4.50 / $22.50 | n/a |
| `deepseek-v4-flash` | deepseek | `deepseek`, `deepseek-v4`, `deepseek-flash` | 1,048,576 | $0.14 / $0.28 | n/a |
| `deepseek-v4-pro` | deepseek | | 1,048,576 | $0.435 / $0.87 | n/a |
| `grok-4.6` | venice | `grok`, `grok46`, `grok-46` | 500K | $2.27 / $6.80 | n/a |
| `laguna-s-2.1` | poolside | `laguna`, `laguna-s` | 1M | $0.10 / $0.20 | n/a |
| `laguna-xs-2.1` | poolside | `laguna-xs` | 262K | $0.10 / $0.20 | n/a |
| `glm-5.2` | zai | `glm`, `glm5`, `glm52`, `glm5.2` | 1M | $1.40 / $4.40 | n/a |
| `glm-5.3` | venice | `glm`, `glm5`, `glm53`, `glm5.3`, `glm-5-3` | 1M | $1.75 / $5.50 | n/a |
| `minimax-m2.5` | minimax | `minimax` | 197K | $0.30 / $1.20 | 45 tok/s |
| `mercury-2` | inception | `mercury` | 131K | $0.25 / $0.75 | 1000 tok/s |
| `qwen3.8-max` | venice | `qwen`, `qwen-max`, `qwen3.8`, `qwen-3.8`, `qwen38`, `qwen-3.8-max`, `qwen38-max` | 1M | $2.50 / $7.50 | n/a |
| `qwen3.8-27b` | venice | `qwen-27b`, `qwen-3.8-27b`, `qwen38-27b` | 262K | $0.45 / $3.20 | n/a |
Each provider requires its own API key. Server-backed workflows read provider credentials from the server vault (for example `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `DEEPSEEK_API_KEY`, or `POOLSIDE_API_KEY` set with `fabro secret set` or `fabro provider login`). Standalone SDK/CLI flows can opt into env-backed credential sources explicitly. See the [Quick Start](/getting-started/quick-start) for setup.
Claude Fable 5 is available as an explicit model but is not the default Anthropic model. If Fable refuses a request, Fabro reports the refusal as a content-filter LLM error and applies the configured `run.model.fallbacks` chain when one is present.
## Configuring providers and models
Fabro's catalog is the [lithos-llm](https://docs.rs/lithos-llm) built-in catalog. The `[llm]` table in settings is a second layer over it: a lithos catalog overlay that adds providers and models or changes existing entries. Later layers win. Tables merge key by key and every other value replaces. Models are nested under their provider, so two providers can expose the same model id without overwriting each other.
Provider and model facts use lithos field names: `adapter`, `codecs`, `base_url`, `auth`, `enabled`, `limits`, `capabilities`, `pricing`, `small_default`, `probe`, `family`, and the cutoffs. The coding harness a model expects lives under `metadata.agent`, a namespace lithos ships and other agents such as Pebble read too. See [Settings Configuration](/reference/user-configuration#llm) for every key.
```toml title="settings.toml"
[llm.providers.proxy]
display_name = "Acme Gateway"
base_url = "https://llm-gateway.example.com/v1"
auth = { type = "bearer" }
aliases = ["gateway"]
default_model = "team-code-large"
[llm.providers.proxy.default_headers]
x-portkey-api-key = "{{ secrets.PORTKEY_API_KEY }}"
x-portkey-config = "@bedrock-prod"
[llm.providers.proxy.metadata.agent]
profile = "anthropic"
[llm.providers.proxy.models."team-code-large"]
display_name = "Team Code Large"
aliases = ["team-code"]
api_model = "provider-wire-model-name"
limits = { context_tokens = 200000, max_output_tokens = 32000 }
capabilities = { text = true, tools = true, reasoning = true, caching = true, reasoning_effort = { low = true, medium = true, high = true } }
protocol_options = { reasoning_effort_levels = true }
pricing = { input_usd_micros_per_million = 1500000, output_usd_micros_per_million = 8000000, cached_input_usd_micros_per_million = 300000 }
family = "team-code"
small_default = true
estimated_output_tps = 80
```
The gateway's API key is `PROXY_API_KEY`: lithos derives the secret name from the provider id (upper case, `-` and `.` as `_`, then `_API_KEY`). Store it with `fabro secret set PROXY_API_KEY ...`.
For [LiteLLM](/integrations/litellm), Fabro ships a disabled provider entry. Enable it in settings and declare the models your proxy exposes:
```toml title="settings.toml"
[llm.providers.litellm]
base_url = "http://localhost:4000/v1"
default_model = "litellm-gpt-5"
enabled = true
[llm.providers.litellm.models."litellm-gpt-5"]
display_name = "LiteLLM GPT-5"
api_model = "gpt-5"
limits = { context_tokens = 128000, max_output_tokens = 8192 }
capabilities = { text = true, tools = true }
```
`api_model` is the model name sent to that provider's API. It defaults to the exact model id, so omit it when the two strings match. Fabro does not infer vendor prefixes or rewrite the value.
<Note>
A `provider/model` selector such as `openai/gpt-5.6-sol` pins the provider and names the model by id, alias, or wire id. A bare selector with no provider pin picks the highest-priority ready offering; a separate `provider = "openrouter"` pin selects the OpenRouter offering. Providers with `allow_passthrough = true` also accept `provider/model` selectors for models the catalog does not list.
</Note>
Model roles are separate: the provider's `default_model` controls normal model selection for workflow execution, while `small_default = true` on a model row marks the provider's small utility model for metadata tasks such as generated run titles. If a provider has no small default, Fabro falls back to that provider's default model.
Provider auth has two parts. The lithos `auth` scheme says how a credential is sent: `{ type = "bearer" }`, `{ type = "header", name = "x-api-key" }`, `{ type = "headers" }` for providers that take several secret headers, `{ type = "none" }`, or `{ type = "aws" }`. lithos also says which secret names a provider reads: `OPENAI_API_KEY` for `openai`, `GEMINI_API_KEY` then `GOOGLE_API_KEY` for `gemini`, `MODAL_TOKEN_ID` and `MODAL_TOKEN_SECRET` for `modal`, and `<PROVIDER>_API_KEY` for a provider you define. Fabro looks each name up in the process environment first and the server vault second. Custom headers for any provider go in `default_headers` as literal text or `{{ secrets.NAME }}` tokens; put credentials in secrets and reference them with `{{ secrets.NAME }}` instead of a bare literal.
Workflow runs also add `x-session-id: <run-id>` to every LLM request so compatible gateways can group requests from the same run. An explicitly configured `x-session-id` in provider `default_headers` takes precedence.
Provider `metadata.agent.profile` defaults from `adapter` and controls profile-specific behavior such as which tools the agent registers, project-memory filenames, CLI/ACP command selection, and native session routing. Valid values are `anthropic`, `claude-5`, `openai`, `gemini`, `kimi`, `gpt56`, and `gpt6`; model-level values override provider-level values.
Three profiles are selected per model rather than per provider, because they follow the model wherever it is served: `claude-5` for Claude 5 models, `kimi` for Kimi models, and `gpt56` for the GPT-5.6 models (Sol, Terra, Luna); `gpt6` for GPT-6 Astra runs on the same harness as `gpt56`. The `gpt56` profile uses Codex's narrow core surface — `shell_command`, `apply_patch`, and `update_plan`, plus optional credential-backed `web_search` — instead of fabro's dedicated file-read, discovery, and `web_fetch` tools. On OpenAI-compatible routes that cannot carry the freeform `apply_patch` grammar, it substitutes the JSON-schema `edit_file` tool. Session features may add their own question, skill, or subagent tools separately.
Costs come from the lithos `pricing` table on each model row. Each token bucket (input, output, reasoning, cache read, cache write) prices at its own rate, with optional long-context and speed tiers. Providers that return an authoritative charge, such as OpenRouter, override the catalog estimate; the billing record says which source it came from.
<Note>
Provider fields in configuration, APIs, and model routing are provider ID strings. Built-in names like `anthropic`, `openai`, and `gemini` still work, but custom IDs like `proxy` work anywhere a provider ID is accepted.
</Note>
### Venice
Fabro ships a built-in [Venice](/integrations/venice) provider with a curated catalog of Venice-hosted Kimi, Grok, GLM, DeepSeek, and Qwen models. Store its API key with `fabro provider login --provider venice`. Pin `provider = "venice"` when a shared model slug must use Venice instead of a higher-priority direct provider.
### Poolside
Fabro ships a built-in [Poolside](/integrations/poolside) provider for Laguna S 2.1 and Laguna XS 2.1 over Poolside's OpenAI-compatible API. Store a direct API key with `fabro provider login --provider poolside`. The same model slugs are also available through the opt-in OpenRouter provider; its vendor-namespaced strings remain provider-only `api_model` values.
### OpenRouter
Fabro ships an [OpenRouter](/integrations/openrouter) provider definition with a curated model catalog, disabled by default. Enable it in settings and store an API key with `fabro provider login --provider openrouter`:
```toml title="settings.toml"
[llm.providers.openrouter]
enabled = true
```
### Modal
Fabro ships a [Modal](/integrations/modal) provider definition for Kimi K3, disabled by default. Modal assigns the endpoint URL and authenticates requests with a two-part proxy token:
```toml title="settings.toml"
[llm.providers.modal]
base_url = "https://your-endpoint.modal.run/v1"
enabled = true
```
Store both token values in the Fabro server vault:
```bash
fabro secret set MODAL_TOKEN_ID wk-...
fabro secret set MODAL_TOKEN_SECRET ws-...
```
### Amazon Bedrock
Fabro ships an [Amazon Bedrock](/integrations/bedrock) provider definition with a curated multi-vendor catalog over Bedrock's Converse API, disabled by default. Enable it and authenticate with a Bedrock API key or AWS SigV4 credentials:
```toml title="settings.toml"
[llm.providers.bedrock]
base_url = "https://bedrock-runtime.us-east-1.amazonaws.com"
enabled = true
```
### Ollama
Fabro ships an Ollama provider definition that is disabled by default. Enable it in settings when you want Fabro to route through a local Ollama server:
```toml title="settings.toml"
[llm.providers.ollama]
enabled = true
```
Enabling the provider alone does not expose any models — until #267 adds auto-discovery, add explicit `[llm.providers.ollama.models."<model-id>"]` blocks for each Ollama model you have pulled locally. Ollama's OpenAI-compatible endpoint accepts any bearer token, so local users can set `OLLAMA_API_KEY=ollama`.
## Default models
When no model or provider is specified, Fabro chooses the default offering on the highest-priority ready provider. If no provider adapter is ready, run creation reports that no eligible offering is available. Each provider has its own default model:
| Provider | Default model |
|---|---|
| `anthropic` | `claude-sonnet-5` |
| `openai` | `gpt-5.6-sol` |
| `gemini` | `gemini-3.5-flash` |
| `moonshot` | `kimi-k3` |
| `poolside` | `laguna-s-2.1` |
| `zai` | `glm-5.2` |
| `venice` | `deepseek-v4-flash` |
| `minimax` | `minimax-m2.5` |
| `inception` | `mercury-2` |
## Using models in workflows
Assign models to workflow nodes using [model stylesheets](/workflows/stylesheets), which use a CSS-like syntax:
```dot title="example.fabro"
digraph Example {
graph [
model_stylesheet="
* { model: claude-haiku-4-5; }
.coding { model: claude-sonnet-4-5; reasoning_effort: high; }
#review { model: gemini-3.1-pro-preview; }
"
]
spec [label="Write Spec"]
implement [label="Implement", class="coding"]
review [label="Review"]
}
```
This routes the spec node to Haiku (the default), implementation to Sonnet, and review to Gemini Pro.
## Overriding the default model
Model stylesheets set per-node models inside the workflow graph, but you can also override the default model for an entire run. This is useful for quick experimentation or when you want to swap models without editing the Graphviz file.
### CLI flags
Pass `--model` and optionally `--provider` to `fabro run`:
```bash
fabro run docs/internal/demo/01-hello.fabro --model claude-opus-4-6
fabro run docs/internal/demo/04-pipeline.fabro --model gemini-3.1-pro-preview
```
These flags set the default model for all nodes that don't have an explicit model assigned via a stylesheet. Without `--provider`, Fabro selects among ready offerings by priority. Add `--provider` to pin an exact provider, including for an uncatalogued provider model string.
### Run config TOML
For repeatable runs, set the model in a run config file:
```toml title="run.toml"
_version = 1
[workflow]
graph = "implement.fabro"
[run]
goal = "Implement the feature"
[run.model]
name = "claude-sonnet-4-5"
[run.model.fallbacks]
"claude-sonnet-4-5" = ["gemini", "openai"]
```
Then launch with:
```bash
fabro run run.toml
```
The `[run.model.fallbacks]` table is optional. Each key names the originally requested model. Its value is the ordered list to try after that model's active provider fails. An entry may be a bare provider token (like `"gemini"`), a bare model ID or alias (like `"gpt-terra"`), or a qualified `"provider:selector"` reference. A qualified selector may be the provider's canonical model ID, alias, or API ID, including API IDs with slashes such as `"openrouter:moonshotai/kimi-k3"`.
Fabro selects one chain from the original request. It does not switch to the chain configured for a fallback target. When the requested reasoning level is unavailable on a fallback target, Fabro uses the nearest supported level. Equal-distance choices round up.
<Note>
The precedence order is: node-level stylesheet > run config TOML > CLI flags > server defaults. More specific settings always win.
</Note>
## CLI commands
### List models
View all available models, or filter by provider:
```bash
fabro model list
fabro model list --provider anthropic
fabro model list --query codex
```
### Test models
Verify that your API keys are working by sending a test prompt to each configured provider:
```bash
fabro model test
fabro model test --model claude-sonnet-4-5
fabro model test --provider openai
```
This is useful for confirming connectivity after setup or when adding a new provider key.