fabro/docs/public/integrations/fireworks.mdx
Bryan Helmkamp 3e065b806b
Bump lithos-llm to 43a42ac and migrate catalogs to the codecs schema
Move the lithos-llm pin from 55add459 to 43a42ac28e9d9bcf40a91abc02be4f12ca274ebb,
and the three Pebble pins from a39f43e to 67c9f48, Pebble `main`, which pins
that same lithos-llm revision so Cargo holds one lithos-llm crate. lithos-llm
`main` (ca19fac) is one commit further; that commit touches only its nightly
workflow, so this pin stays on the revision Pebble unifies with.

The `openai`, `anthropic`, `gemini`, and `openai-compatible` features are
gone upstream; each expanded to `runtime`, which `bedrock` implies, so the
four names leave the fabro-llm feature list. Every other manifest already
names `runtime`.

The catalog schema now names one adapter and many codecs per provider.
`adapter` defaults to `http`, `codecs = [...]` replaces `codec` and defaults
to `["openai-chat"]`, and the loader rejects the old `codec` key and the four
protocol-named adapter ids. Every inline catalog in tests and docs moves to
the new shape: the `openai-compatible` + `openai-chat` pair is dropped as the
default, `adapter = "openai"` + `codec = "openai-responses"` becomes
`codecs = ["openai-responses"]`, and the one test that swaps in a custom
adapter id now adds the line instead of replacing one. The settings
reference, the API schema's `Provider.adapter` description, and the SDK page
describe the new fields; the three `docs/superpowers/plans/` files that show
the old shape are dated, unchecked historical plans and are left as they are.

The implied agent profile for an operator provider that declares none used
to read the removed protocol adapter ids; it now reads the provider's first
codec (Anthropic Messages and Gemini map to their harnesses, the `bedrock`
adapter to Anthropic, everything else to OpenAI), with a test for the codec
path.

Absorbing the rest of the range: OpenRouter and Fireworks now ship enabled,
so the two fabro-llm tests that used OpenRouter as the disabled fixture use
`bedrock-openai`, and the docs and comments that said the two ship disabled
are corrected. The built-in catalog grew past 100 enabled model rows
(Vercel, TypeSafe, and the enabled OpenRouter and Fireworks rosters), so the
pagination shape test walks `page[offset]` to the last page instead of
assuming one page fits.

`cargo update -p` on the four crates also re-resolved a few already-locked
edges to match the lithos-llm lockfile: `windows-sys` 0.61.2/0.60.2 ->
0.59.0 under dirs-sys, errno, nu-ansi-term, quinn-udp, rustix,
rustls-platform-verifier, tempfile, terminal_size, and winapi-util;
`windows-core` 0.61.2 -> 0.62.2 under iana-time-zone; `errno` 0.2.8 ->
0.3.14 under signal-hook-registry; and `indexmap` 2.13.0 as a new public
dependency of lithos-llm. No package version was added or removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:07:00 -04:00

134 lines
6.3 KiB
Text

---
title: "Fireworks AI"
description: "Run open-weights models on Fireworks AI's serverless inference platform"
---
[Fireworks AI](https://fireworks.ai/) serves open-weights models (Kimi, DeepSeek, GLM, Qwen, GPT-OSS, and more) behind an OpenAI-compatible API. Fabro ships an enabled `fireworks` provider entry with a curated model catalog; it needs only an API key, and `settings.toml` can adjust it without changing Fabro code.
## Prerequisites
- A [Fireworks AI account](https://fireworks.ai/) with serverless credit
- An API key from [app.fireworks.ai/settings/users/api-keys](https://app.fireworks.ai/settings/users/api-keys)
## Enable the provider
Fabro runs execute through a Fabro server. Add the provider override to the settings file used by that server. For a local server, this is usually `~/.fabro/settings.toml`; for a remote deployment, update the server host's Fabro settings.
```toml title="settings.toml"
_version = 1
[llm.providers.fireworks]
enabled = true
```
## Configure credentials
Store the key in the target Fabro server vault:
```bash
fabro provider login --provider fireworks
# For a non-default remote server:
fabro provider login --server https://your-fabro.example --provider fireworks
# Or set the vault token directly:
fabro secret set FIREWORKS_API_KEY fw_...
fabro secret --server https://your-fabro.example set FIREWORKS_API_KEY fw_...
```
Direct SDK usage outside a Fabro server can use an env-backed credential source explicitly:
```bash
export FIREWORKS_API_KEY=fw_...
```
## Included models
The built-in catalog gives Fireworks offerings the same human-facing model slugs used by other providers. Fireworks account-scoped model paths remain opaque `api_model` values:
| Fabro model slug | Fireworks API ID / notes |
| --- | --- |
| `kimi-k3` | `accounts/fireworks/models/kimi-k3` |
| `kimi-k3-fast` | `accounts/fireworks/routers/kimi-k3-fast`; Fast tier at 50% above standard rates |
| `kimi-k2.7-code` | `accounts/fireworks/models/kimi-k2p7-code`; provider default |
| `kimi-k2.6` | `accounts/fireworks/models/kimi-k2p6` |
| `deepseek-v4-pro`, `deepseek-v4-flash` (`deepseek`, `deepseek-v4`, `deepseek-flash`) | `accounts/fireworks/models/deepseek-v4-...` |
| `glm-5.2` | `accounts/fireworks/models/glm-5p2` |
| `minimax-m2.7` | `accounts/fireworks/models/minimax-m2p7` |
| `qwen3.7-plus` | `accounts/fireworks/models/qwen3p7-plus` |
| `gpt-oss-120b` | `accounts/fireworks/models/gpt-oss-120b` |
| `gpt-oss-20b` | `accounts/fireworks/models/gpt-oss-20b`; provider small default |
Any other Fireworks serverless model can be added under the provider. Choose a stable Fabro model slug as the table key and put the Fireworks account-scoped path in `api_model` (dots in upstream model names become `p`, e.g. `glm-5.2` → `glm-5p2`):
```toml title="settings.toml"
[llm.providers.fireworks.models."llama-4-maverick"]
display_name = "Llama 4 Maverick"
api_model = "accounts/fireworks/models/llama4-maverick-instruct-basic"
limits = { context_tokens = 1000000, max_output_tokens = 16384 }
capabilities = { text = true, tools = true }
```
Note that Fireworks' `GET /v1/models` endpoint only returns a featured subset of serverless models; a model absent from that list may still be servable. Verify custom additions with `fabro model test`.
## Use Fireworks models
```bash
fabro model list --provider fireworks
fabro model test --provider fireworks --model kimi-k3-fast
fabro run workflow.fabro --provider fireworks --model kimi-k3-fast
```
When targeting a non-default remote server, pass the same `--server` value to verification commands:
```bash
fabro model list --server https://your-fabro.example --provider fireworks
fabro model test --server https://your-fabro.example --provider fireworks --model kimi-k3-fast
```
In workflow stylesheets:
```dot title="workflow.fabro"
digraph Example {
graph [
model_stylesheet="
* { model: fireworks/kimi-k3-fast; }
"
]
start [shape=Mdiamond, label="Start"]
work [label="Work", prompt="Use the configured Fireworks model."]
exit [shape=Msquare, label="Exit"]
start -> work -> exit
}
```
## Prompt caching
Fireworks caches prompt prefixes automatically — no cache breakpoints or request changes are needed. Serverless responses report cached tokens in the usage body, and cached input tokens are billed at a per-model discount (typically 50% or better). Fabro reads the cached-token counts and applies the catalog's `cache_input_cost_per_mtok` rates when estimating costs.
## Costs
Catalog prices mirror [Fireworks serverless pricing](https://docs.fireworks.ai/serverless/pricing). Fireworks does not return in-band billing, so Fabro reports the cost source as `catalog`. `kimi-k3-fast` uses the published 50% Fast tier premium. Other Fast model variants and the Priority service tier are not included in the built-in catalog.
## Troubleshooting
**"No API key configured"** — Set the key on the target server with `fabro provider login --provider fireworks` or `fabro secret set FIREWORKS_API_KEY ...`. For direct SDK usage outside a Fabro server, export `FIREWORKS_API_KEY` in the invoking shell.
**"provider 'fireworks' is not configured in the server model catalog"** — Confirm the server host's `settings.toml` has `[llm.providers.fireworks]` with `enabled = true`. Fabro live-reloads `settings.toml` within a few seconds; after that, `fabro model list --provider fireworks` against the same server should show the enabled catalog.
**402 / insufficient credits** — Serverless inference requires prepaid credit; check your balance in the [Fireworks billing dashboard](https://app.fireworks.ai/settings/billing).
**Unknown model** — Confirm the model's `api_model` matches a Fireworks account-scoped model or router path exactly (`accounts/fireworks/models/...` or `accounts/fireworks/routers/...`), then run `fabro model test --model <fabro-model-id>`. Remember that `GET /v1/models` only lists a featured subset, so absence from that list is not conclusive.
## Further reading
<Columns cols={2}>
<Card title="Models" icon="microchip" href="/core-concepts/models">
How Fabro routes model IDs, providers, and fallbacks.
</Card>
<Card title="Settings Configuration" icon="gear" href="/reference/user-configuration">
Full reference for provider settings and provider-scoped model offerings.
</Card>
</Columns>