Adds **Amazon Bedrock** as an opt-in built-in provider, over Bedrock's unified **Converse / ConverseStream** API. One codec serves every Converse-capable family — Claude, Amazon Nova, Meta Llama, Mistral, DeepSeek, Moonshot Kimi, Z.AI GLM, MiniMax, NVIDIA Nemotron, and OpenAI gpt-oss — because AWS translates the envelope to each model's native dialect server-side. Auth is either **AWS SigV4** (the default credential chain — env / profile / IMDS / IRSA / SSO, resolved per request so sessions refresh) or a **Bedrock API key** (`AWS_BEARER_TOKEN_BEDROCK`, bearer). Disabled by default (the Ollama / OpenRouter opt-in pattern). This is the redo of #459's original Claude-only `InvokeModel` adapter, rebuilt on the gateway-refactor seams (#481–#497). @depopry's SigV4 signer, AWS event-stream frame decoder, `BedrockAuth`, the `aws_sigv4` credential grammar, `AdapterKind::Bedrock`, region-from-base_url, and the lean-deps decision are preserved and authored by him on the first two commits; the per-family `BedrockCodec` trait he wrote turned out to be the crate-wide `Codec` seam in miniature, so the refactor promoted exactly that shape. The original Claude-only description is preserved in a comment below. ## What's here - **`AdapterKind::Bedrock` × `CodecKind::BedrockConverse`** on the route, plus the `aws_sigv4` credential source (no static secret — the adapter signs at request time; `fabro-auth` stays AWS-free). *(@depopry)* - **SigV4 signer + AWS event-stream `FrameDecoder`** on the lean AWS stack (no `aws-sdk-bedrockruntime`; transport stays on `fabro-http`). Re-targeted at Converse's direct-JSON stream frames; the signer resolves credentials per request. *(@depopry)* - **`bedrock_converse` codec** — Converse envelope (`system[]`, typed content blocks, `inferenceConfig`, `toolConfig`), prompt caching via `cachePoint`, thinking-signature round-trip through `reasoningContent`, usage mapped onto the disjoint `TokenCounts` buckets, `provider_options.bedrock` passthrough. Plus the adapter shell and an event-stream byte loop beside the transport's shared SSE loop. - **Catalog**: `bedrock.toml` (Claude incl. Fable 5, Nova 2, Llama 4, Mistral, DeepSeek, Kimi, GLM, MiniMax, Nemotron, gpt-oss — cross-region inference-profile ids, per-model `billing_policy` so Claude bills Anthropic-style) and a companion **`bedrock-openai`** provider for GPT-5.5/5.4 over the `bedrock-mantle` Responses endpoint (pure config over the existing `openai_responses` codec, zero new code). - Secrets registry (`AWS_BEARER_TOKEN_BEDROCK`), gitleaks rules for both Bedrock key formats, the `docs/integrations/bedrock` guide, and live e2e tests. ## Live verification (confirmed end-to-end against a real AWS account) Verified on a real Bedrock account (us-east-2, SigV4 + bearer): - **SigV4 + Converse** — multiple families (Claude, Nova, DeepSeek, …) via the full settings → catalog → route → adapter → codec path. - **ConverseStream** — streaming deltas through the workflow engine. - **Multi-turn tool use** — agent loop with tool calls round-tripping (no-arg tools included). - **Multi-model routing** — Claude + DeepSeek pinned in one run through the single Converse codec. - **mantle Responses** — `openai.gpt-5.5` answered via the `bedrock-openai` provider (bearer auth). The exercise caught and fixed several issues that unit tests (static creds, mocked transports) could not — see the follow-up commits below. ## Follow-up fixes from live testing (commits on top of the foundation) 1. **Worker AWS env** — the workflow worker scrubs its env to an allowlist, so SigV4 (which re-resolves from the ambient chain per request) couldn't work through `fabro run`. The AWS credential-chain inputs now cross into the worker. 2. **Vault bearer key** — Bedrock was the only key-based provider missing a `vault:` credential ref, so `fabro secret set AWS_BEARER_TOKEN_BEDROCK` silently didn't feed it. Now resolves env → vault → SigV4. 3. **Converse tool-encoding hardening** — a no-arg tool call's `toolUse.input` is now a `{}` object (Bedrock rejects null), and every tool `inputSchema` gets a top-level `type: "object"` (strict families like DeepSeek reject a typeless schema Claude tolerates). 4. **Nova output cap** — `amazon.nova-2-lite` max_output 65536 → 65535 (Bedrock's per-request limit). Earlier fixes already folded into the foundation commits: the `aws-config` sleep-impl (default chain panicked) and AWS error-body decoding (top-level `message`/`Message`/`__type` → proper messages instead of "Unknown error"). ## Manual testing & setup See `docs/integrations/bedrock` — now documents the non-obvious account setup that live testing surfaced: the per-Region Anthropic use-case approval, `aws-marketplace:Subscribe` for third-party models, the Fable 5 / Mythos-class data-sharing opt-in, and the bearer-vs-SigV4 precedence override for running Converse + mantle side by side. ## Open decision / discussion - **Model-id naming** — Bedrock rows use dotted ids mirroring Bedrock's native inference-profile ids (`us.anthropic.claude-sonnet-4-6`, `openai.gpt-5.5`), which also makes them the wire `api_id`. Third scheme alongside bare ids and OpenRouter's `vendor/model` slashes. No collision risk (enforced at catalog build). Open to a uniform scheme if preferred. - **`BEDROCK_API_KEY` alias** — see the comment thread; the AWS console hands some users `export BEDROCK_API_KEY=` while the SDK-standard var is `AWS_BEARER_TOKEN_BEDROCK`. Question of whether to accept both. ## Deferred (named follow-ups) - **`qwen.qwen3-coder-next`** — omitted pending a verified Bedrock model/inference-profile id (its fabro id isn't a valid Bedrock identifier; needs an explicit `api_id`). Re-add once confirmed via `aws bedrock list-inference-profiles`. - **Claude Mythos 5** — Anthropic-Messages-only on `bedrock-mantle` (limited preview). - **Converse structured output** (`response_format` rejected with a clear error). - **`reasoning_effort` on Converse rows** via `additionalModelRequestFields` (the `bedrock-openai` GPT rows already accept effort levels). - **CountTokens** route (`count_input_tokens` returns `None`). ## Verification `cargo nextest run --workspace`: green except the pre-existing environment-dependent fabro-workflow failures (identical on main). clippy `-D warnings` + pinned-nightly fmt clean. Codec unit tests + adapter httpmock tests + frame-decoder/signer locks. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Scott Werner <scott@sublayer.com> Co-authored-by: Scott Werner <stwerner@vt.edu> |
||
|---|---|---|
| .. | ||
| src | ||
| tests | ||
| Cargo.toml | ||
| README.md | ||
fabro-llm
A unified async Rust client library for multiple LLM providers. Write your LLM integration code once and switch between Anthropic, OpenAI, and Google Gemini without changing your application logic.
Key concepts
- Client -- Routes requests to registered provider adapters. Build it from a
CredentialSourceor explicit typed credentials. - ProviderAdapter -- The trait every provider implements (
completeandstream). Built-in adapters:AnthropicAdapter,OpenAiAdapter,GeminiAdapter,OpenAiCompatibleAdapter. - Middleware -- Intercepts requests/responses for logging, caching, or transformation. Supports both blocking and streaming paths.
- generate() -- High-level function that wraps
Client.complete()with automatic tool execution loops, retries, timeouts, and cancellation. - Tool -- Active tools (with an execute handler) run automatically in the tool loop. Passive tools (no handler) surface tool calls back to the caller.
- Model catalog -- Built-in metadata for common models. Advisory only; unknown model strings pass through.
Providers
| Provider | Adapter | API | Env var |
|---|---|---|---|
| Anthropic | AnthropicAdapter |
Messages API | ANTHROPIC_API_KEY |
| OpenAI | OpenAiAdapter |
Responses API | OPENAI_API_KEY |
| Google Gemini | GeminiAdapter |
generateContent | GEMINI_API_KEY or GOOGLE_API_KEY |
| OpenAI-compatible | OpenAiCompatibleAdapter |
Chat Completions | (custom) |
All adapters support streaming, tool calling, structured output (response_format), and provider-specific options via provider_options.
Usage
Create from an environment-backed credential source
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::types::{Message, Request};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let request = Request {
model: "claude-sonnet-4-5".to_string(),
messages: vec![Message::user("What is the capital of France?")],
provider: None,
tools: None,
tool_choice: None,
response_format: None,
temperature: Some(0.0),
top_p: None,
max_tokens: Some(100),
stop_sequences: None,
reasoning_effort: None,
metadata: None,
provider_options: None,
};
let response = client.complete(&request).await?;
println!("{}", response.text());
High-level generate()
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::generate::{generate, GenerateParams};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let result = generate(
GenerateParams::new("claude-sonnet-4-5", client.clone())
.prompt("Explain monads in one sentence")
.system("You are a concise programming tutor.")
.max_tokens(200)
).await?;
println!("{}", result.text());
Tool calling
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::generate::{generate, GenerateParams};
use fabro_llm::tools::Tool;
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let weather_tool = Tool::active(
"get_weather",
"Get the current weather for a city",
serde_json::json!({
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"]
}),
|args, _ctx| async move {
let city = args["city"].as_str().unwrap_or("unknown");
Ok(serde_json::json!({"temp": "72F", "city": city}))
},
);
let result = generate(
GenerateParams::new("claude-sonnet-4-5", client.clone())
.prompt("What's the weather in San Francisco?")
.tools(vec![weather_tool])
.max_tool_rounds(3)
).await?;
Streaming
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::types::{Message, Request, StreamEvent};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use futures::StreamExt;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let request = Request {
model: "claude-sonnet-4-5".to_string(),
messages: vec![Message::user("Tell me a joke")],
// ...other fields set to None/defaults
# provider: None, tools: None, tool_choice: None,
# response_format: None, temperature: None, top_p: None,
# max_tokens: None, stop_sequences: None, reasoning_effort: None,
# metadata: None, provider_options: None,
};
let mut stream = client.stream(&request).await?;
while let Some(event) = stream.next().await {
match event? {
StreamEvent::TextDelta { delta, .. } => print!("{delta}"),
StreamEvent::Finish { response, .. } => {
println!("\nTokens used: {}", response.usage.total_tokens);
}
_ => {}
}
}
Middleware
use fabro_llm::error::Error;
use fabro_llm::middleware::{Middleware, NextFn, NextStreamFn};
use fabro_llm::provider::StreamEventStream;
use fabro_llm::types::{Request, Response};
struct LoggingMiddleware;
#[async_trait::async_trait]
impl Middleware for LoggingMiddleware {
async fn handle_complete(
&self,
request: Request,
next: NextFn,
) -> Result<Response, Error> {
eprintln!("Request to model: {}", request.model);
let response = next(request).await?;
eprintln!("Response tokens: {}", response.usage.total_tokens);
Ok(response)
}
async fn handle_stream(
&self,
request: Request,
next: NextStreamFn,
) -> Result<StreamEventStream, Error> {
next(request).await
}
}
OpenAI-compatible providers
use fabro_llm::providers::OpenAiCompatibleAdapter;
use std::sync::Arc;
let adapter = OpenAiCompatibleAdapter::new("your-api-key", "https://api.groq.com/openai/v1")
.with_name("groq");
Model catalog
use fabro_llm::catalog::{get_latest_model, get_model_info, list_models};
let info = get_model_info("claude-opus-4-6");
let anthropic_models = list_models(Some("anthropic"));
let best_reasoner = get_latest_model("anthropic", Some("reasoning"));
Input token counting
Use count_input_tokens when you need the current model-visible context size
without creating a completion:
use fabro_llm::{InputTokenCountPreference, Client};
let count = client
.count_input_tokens(&request, InputTokenCountPreference::PreferProvider)
.await?;
InputTokenCountPreference controls precision and data exposure:
PreferProvidersends the provider-serialized request to the upstream token-count endpoint when supported, then falls back to a local estimate only for unsupported adapters, network/timeout failures, rate limits, and provider server errors.RequireProvidersends the provider-serialized request and returns either a provider count or an error. It never returns a local estimate.EstimateOnlyvalidates and resolves the provider locally, does not call the adapter count endpoint, and returns a deterministic local estimate.
Provider-native counting sends model-visible request content to the provider's
token-count endpoint. That can include messages, system/developer instructions,
tools, schemas, structured content, and media metadata/content after provider
serialization. Use EstimateOnly when that extra upstream exposure is not
acceptable.
InputTokenCount is for input/context sizing. It is not billing usage and does
not include output, reasoning-output, cache-read, or cache-write token buckets.
Key types
| Type | Description |
|---|---|
Request |
Unified request with model, messages, tools, temperature, etc. |
Response |
Unified response with message, finish reason, usage, rate limit info |
Message |
A message with role, content parts, and optional tool call ID |
ContentPart |
Text, Image, Audio, Document, ToolCall, ToolResult, Thinking |
StreamEvent |
Events for streaming: TextDelta, ToolCallStart/Delta/End, Finish, etc. |
SdkError |
Typed errors with retryability, status codes, and provider error kinds |
GenerateParams |
Builder for the high-level generate() function |
GenerateResult |
Result containing response, tool results, total usage, and step history |
ToolDefinition |
Tool name, description, and JSON Schema parameters |
ToolChoice |
Auto, None, Required, or Named tool selection |
InputTokenCount |
Input/context token count from a provider count API or local estimate |
TokenCounts |
Billing-oriented token counts including input, output, reasoning, and cache tokens |
RetryPolicy |
Configurable retry with exponential backoff, jitter, and max delay |
Model |
Metadata about a model (context window, capabilities, costs) |
Error handling
SdkError provides structured error variants with built-in retryability classification:
- Retryable:
RateLimit,Server,Network,Stream,RequestTimeout - Non-retryable:
Authentication,AccessDenied,InvalidRequest,ContextLength,Configuration
The retry() function and generate() respect Retry-After headers and use exponential backoff with jitter.
Provider-specific options
Pass provider-specific parameters via provider_options without losing portability:
use fabro_llm::types::Request;
let request = Request {
provider_options: Some(serde_json::json!({
"anthropic": {
"thinking": {"type": "enabled", "budget_tokens": 10000},
"auto_cache": true
},
"openai": {
"store": true,
"previous_response_id": "resp_abc123"
},
"gemini": {
"safetySettings": [
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"}
]
}
})),
// ...other fields
# model: String::new(), messages: vec![], provider: None, tools: None,
# tool_choice: None, response_format: None, temperature: None,
# top_p: None, max_tokens: None, stop_sequences: None,
# reasoning_effort: None, metadata: None,
};