fabro/docs/public/integrations/litellm.mdx
Bryan Helmkamp 671324a06f
docs(secrets): document settings-declared credentials, fix stale local-run guidance
server-secrets-strategy.md described only two credential mechanisms — bootstrap
ServerSecrets and vault-only optional integrations — and stated its most
restrictive rule in terms of "server runtime", which is ambiguous now that every
run is a server process plus a worker. It omitted the third mechanism actually
used by operator-configured integrations: settings-declared credentials in
InterpString fields, resolved at consumption time from {{ env.NAME }} or
{{ secrets.NAME }}, as LLM provider extra_headers already does.

Add a "Which process resolves what" table keyed on resolving process and timing,
a "Settings-declared credentials" section with the extra_headers precedent, and a
mechanism table at the head of "Adding A New Server Secret". Replace "server
runtime" with per-process statements, and describe where CredentialResolver's
process-env fallback is actually live.

Also correct six docs that told operators to export provider keys for "standalone
local runs". There is no CLI-local run execution: runs always execute in a worker
whose environment is cleared and repopulated from WORKER_ENV_ALLOWLIST, which
excludes provider API keys. Those instructions could not have worked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 09:04:59 -04:00

131 lines
4.3 KiB
Text

---
title: "LiteLLM"
description: "Route Fabro models through a LiteLLM proxy"
---
[LiteLLM](https://docs.litellm.ai/) can run as an OpenAI-compatible proxy in front of many model providers. Fabro includes a disabled `litellm` provider entry so you can opt in from `settings.toml` without changing Fabro code.
## Prerequisites
- A running LiteLLM proxy reachable from the Fabro process
- At least one LiteLLM model name you want Fabro to route to
- A LiteLLM key or placeholder key available to Fabro
Fabro's built-in LiteLLM provider points at `http://localhost:4000/v1`. Change `base_url` if your proxy is hosted elsewhere.
## Enable the provider
Add the provider override and one or more model entries to `~/.fabro/settings.toml`:
```toml title="settings.toml"
_version = 1
[llm.providers.litellm]
enabled = true
base_url = "http://localhost:4000/v1"
[llm.providers.litellm.models."litellm-gpt-5"]
api_id = "gpt-5"
display_name = "LiteLLM GPT-5"
family = "litellm"
default = true
[llm.providers.litellm.models."litellm-gpt-5".limits]
context_window = 128000
max_output = 8192
[llm.providers.litellm.models."litellm-gpt-5".features]
tools = true
vision = false
reasoning = false
```
`api_id` is the model name Fabro sends to LiteLLM. It should match a model name configured in your LiteLLM proxy.
## Configure credentials
For server-backed runs, store `LITELLM_API_KEY` in the Fabro server vault.
For a server-owned secret:
```bash
fabro secret set LITELLM_API_KEY sk-proxy-key
```
`fabro exec` and direct `fabro-llm` SDK usage can use an env-backed credential source explicitly:
```bash
export LITELLM_API_KEY=sk-proxy-key
```
If your local LiteLLM proxy does not enforce authentication, use a placeholder value such as `anything`; the OpenAI-compatible client still needs a credential value.
## Use LiteLLM models
Once the provider is enabled and at least one model is declared, use the Fabro model ID like any other catalog model:
```bash
fabro model list --provider litellm
fabro model test --model litellm-gpt-5
fabro run workflow.fabro --model litellm-gpt-5
```
In workflow stylesheets:
```dot title="workflow.fabro"
digraph Example {
graph [
model_stylesheet="
* { model: litellm-gpt-5; }
"
]
start [shape=Mdiamond, label="Start"]
work [label="Work", prompt="Use the configured LiteLLM model."]
exit [shape=Msquare, label="Exit"]
start -> work -> exit
}
```
## Declaring more models
Declare each LiteLLM-routed model explicitly so Fabro knows its provider, context window, tool support, and routing defaults:
```toml title="settings.toml"
[llm.providers.litellm.models."litellm-fast"]
api_id = "fast-model"
display_name = "LiteLLM Fast"
family = "litellm"
aliases = ["fast"]
[llm.providers.litellm.models."litellm-fast".limits]
context_window = 64000
max_output = 4096
[llm.providers.litellm.models."litellm-fast".features]
tools = true
vision = false
reasoning = false
```
Only one model for a provider should set `default = true`. You may also mark one small/cheap utility model with `small_default = true`; Fabro uses it for metadata tasks such as generated run titles and falls back to the provider default when it is omitted.
## Troubleshooting
**"No API key configured"** — For runs, set `vault:LITELLM_API_KEY` with `fabro secret set LITELLM_API_KEY ...`. Exporting it in the server's shell has no effect on runs: workers start from a cleared environment and provider keys are not inherited. For `fabro exec` or direct SDK usage, export `LITELLM_API_KEY` in the invoking shell and use an env-backed credential source.
**Connection refused** — Confirm the LiteLLM proxy is running and that `base_url` is reachable from the Fabro process. For Docker deployments, `localhost` means the Fabro container unless you point it at a host or service name.
**Unknown model from LiteLLM** — Check that the model's `api_id` matches the model name configured in LiteLLM, then run `fabro model test --model <fabro-model-id>`.
## Further reading
<Columns cols={2}>
<Card title="Models" icon="microchip" href="/core-concepts/models">
How Fabro routes model IDs, providers, and fallbacks.
</Card>
<Card title="Settings Configuration" icon="gear" href="/reference/user-configuration">
Full reference for provider settings and provider-scoped model offerings.
</Card>
</Columns>