mirror of
https://github.com/fabro-sh/fabro.git
synced 2026-10-05 02:41:45 +00:00
## Summary
This is the foundational step of the unified agent transcript
implementation: it establishes one canonical set of replay types in
`fabro-types` and threads them into the existing `agent.message`,
`agent.tool.started`, and `agent.tool.completed` event shapes — without
breaking any existing producers or consumers.
## What changed
**New `fabro-types::transcript` module** owns `ContentPart`,
`ImageData`, `AudioData`, `DocumentData`, `ThinkingData`, `ToolCall`,
`ToolResult`, `MessageKind`, `MessageSource`, `PairMessageRef`,
`TranscriptMessage`, and `MessageId`. These were previously defined in
`fabro-llm::types`; they now live at the canonical layer.
**`fabro-llm::types`** drops its local definitions and re-exports from
`fabro-types` so every existing `fabro_llm::types::*` import keeps
compiling without change.
**`AgentMessageProps`** gains an optional `message:
Option<TranscriptMessage>` field; `AgentToolStartedProps` gains
`tool_call`, `turn_id`, and `parent_message_id`;
`AgentToolCompletedProps` gains `tool_result` and `turn_id`. All new
fields use `#[serde(default, skip_serializing_if = "Option::is_none")]`
so existing stored events deserialize cleanly.
**All current event emitters** (`fabro-workflow/event/convert.rs`, demo
fixtures, test helpers) are updated to set the new fields to `None` —
this is a mechanical compatibility update; actual enrichment comes in
later tasks.
### Design decisions worth noting
- `MessageKind` captures LLM role semantics (system / user / reasoning /
agent); `MessageSource` captures audit provenance (steer, pair,
loop_detection, …). They are intentionally kept separate so a steering
message can be `kind=user, source=steer` without collapsing the
distinction.
- `TranscriptMessage` is named with the `Transcript` prefix specifically
to avoid import ambiguity with `fabro_agent::Message` and
`fabro_llm::types::Message`.
- `ProviderAnswer` and `ProviderReasoning` are included as
`MessageSource` variants so committed model outputs carry a first-class
audit label distinct from user-originated inputs.
- The new fields are additive-only; no narrow legacy fields were
removed. Consumer migration is a separate step.
### Fabro Details
<details>
<summary>Ran 9 stages in 47m 27s for $12.79</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 6s | – | 0 |
| preflight_lint | 2m 21s | – | 0 |
| implement | 20m 6s | $8.53 | 0 |
| simplify_opus | 13m 36s | $2.68 | 0 |
| simplify_gpt | 4m 17s | $1.58 | 0 |
| verify | 4m 7s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **47m 27s** | **$12.79** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
|
||
|---|---|---|
| .. | ||
| src | ||
| tests | ||
| Cargo.toml | ||
| README.md | ||
fabro-llm
A unified async Rust client library for multiple LLM providers. Write your LLM integration code once and switch between Anthropic, OpenAI, and Google Gemini without changing your application logic.
Key concepts
- Client -- Routes requests to registered provider adapters. Build it from a
CredentialSourceor explicit typed credentials. - ProviderAdapter -- The trait every provider implements (
completeandstream). Built-in adapters:AnthropicAdapter,OpenAiAdapter,GeminiAdapter,OpenAiCompatibleAdapter. - Middleware -- Intercepts requests/responses for logging, caching, or transformation. Supports both blocking and streaming paths.
- generate() -- High-level function that wraps
Client.complete()with automatic tool execution loops, retries, timeouts, and cancellation. - Tool -- Active tools (with an execute handler) run automatically in the tool loop. Passive tools (no handler) surface tool calls back to the caller.
- Model catalog -- Built-in metadata for common models. Advisory only; unknown model strings pass through.
Providers
| Provider | Adapter | API | Env var |
|---|---|---|---|
| Anthropic | AnthropicAdapter |
Messages API | ANTHROPIC_API_KEY |
| OpenAI | OpenAiAdapter |
Responses API | OPENAI_API_KEY |
| Google Gemini | GeminiAdapter |
generateContent | GEMINI_API_KEY or GOOGLE_API_KEY |
| OpenAI-compatible | OpenAiCompatibleAdapter |
Chat Completions | (custom) |
All adapters support streaming, tool calling, structured output (response_format), and provider-specific options via provider_options.
Usage
Create from an environment-backed credential source
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::types::{Message, Request};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let request = Request {
model: "claude-sonnet-4-5".to_string(),
messages: vec![Message::user("What is the capital of France?")],
provider: None,
tools: None,
tool_choice: None,
response_format: None,
temperature: Some(0.0),
top_p: None,
max_tokens: Some(100),
stop_sequences: None,
reasoning_effort: None,
metadata: None,
provider_options: None,
};
let response = client.complete(&request).await?;
println!("{}", response.text());
High-level generate()
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::generate::{generate, GenerateParams};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let result = generate(
GenerateParams::new("claude-sonnet-4-5", client.clone())
.prompt("Explain monads in one sentence")
.system("You are a concise programming tutor.")
.max_tokens(200)
).await?;
println!("{}", result.text());
Tool calling
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::generate::{generate, GenerateParams};
use fabro_llm::tools::Tool;
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let weather_tool = Tool::active(
"get_weather",
"Get the current weather for a city",
serde_json::json!({
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"]
}),
|args, _ctx| async move {
let city = args["city"].as_str().unwrap_or("unknown");
Ok(serde_json::json!({"temp": "72F", "city": city}))
},
);
let result = generate(
GenerateParams::new("claude-sonnet-4-5", client.clone())
.prompt("What's the weather in San Francisco?")
.tools(vec![weather_tool])
.max_tool_rounds(3)
).await?;
Streaming
use fabro_auth::EnvCredentialSource;
use fabro_llm::client::Client;
use fabro_llm::types::{Message, Request, StreamEvent};
use fabro_model::catalog::LlmCatalogSettings;
use fabro_model::Catalog;
use futures::StreamExt;
use std::sync::Arc;
let source = EnvCredentialSource::new();
let catalog = Arc::new(Catalog::from_builtin_with_overrides(&LlmCatalogSettings::default())?);
let client = Client::from_source(&source, Arc::clone(&catalog)).await?;
let request = Request {
model: "claude-sonnet-4-5".to_string(),
messages: vec![Message::user("Tell me a joke")],
// ...other fields set to None/defaults
# provider: None, tools: None, tool_choice: None,
# response_format: None, temperature: None, top_p: None,
# max_tokens: None, stop_sequences: None, reasoning_effort: None,
# metadata: None, provider_options: None,
};
let mut stream = client.stream(&request).await?;
while let Some(event) = stream.next().await {
match event? {
StreamEvent::TextDelta { delta, .. } => print!("{delta}"),
StreamEvent::Finish { response, .. } => {
println!("\nTokens used: {}", response.usage.total_tokens);
}
_ => {}
}
}
Middleware
use fabro_llm::error::Error;
use fabro_llm::middleware::{Middleware, NextFn, NextStreamFn};
use fabro_llm::provider::StreamEventStream;
use fabro_llm::types::{Request, Response};
struct LoggingMiddleware;
#[async_trait::async_trait]
impl Middleware for LoggingMiddleware {
async fn handle_complete(
&self,
request: Request,
next: NextFn,
) -> Result<Response, Error> {
eprintln!("Request to model: {}", request.model);
let response = next(request).await?;
eprintln!("Response tokens: {}", response.usage.total_tokens);
Ok(response)
}
async fn handle_stream(
&self,
request: Request,
next: NextStreamFn,
) -> Result<StreamEventStream, Error> {
next(request).await
}
}
OpenAI-compatible providers
use fabro_llm::providers::OpenAiCompatibleAdapter;
use std::sync::Arc;
let adapter = OpenAiCompatibleAdapter::new("your-api-key", "https://api.groq.com/openai/v1")
.with_name("groq");
Model catalog
use fabro_llm::catalog::{get_latest_model, get_model_info, list_models};
let info = get_model_info("claude-opus-4-6");
let anthropic_models = list_models(Some("anthropic"));
let best_reasoner = get_latest_model("anthropic", Some("reasoning"));
Key types
| Type | Description |
|---|---|
Request |
Unified request with model, messages, tools, temperature, etc. |
Response |
Unified response with message, finish reason, usage, rate limit info |
Message |
A message with role, content parts, and optional tool call ID |
ContentPart |
Text, Image, Audio, Document, ToolCall, ToolResult, Thinking |
StreamEvent |
Events for streaming: TextDelta, ToolCallStart/Delta/End, Finish, etc. |
SdkError |
Typed errors with retryability, status codes, and provider error kinds |
GenerateParams |
Builder for the high-level generate() function |
GenerateResult |
Result containing response, tool results, total usage, and step history |
ToolDefinition |
Tool name, description, and JSON Schema parameters |
ToolChoice |
Auto, None, Required, or Named tool selection |
Usage |
Token counts including input, output, reasoning, and cache tokens |
RetryPolicy |
Configurable retry with exponential backoff, jitter, and max delay |
Model |
Metadata about a model (context window, capabilities, costs) |
Error handling
SdkError provides structured error variants with built-in retryability classification:
- Retryable:
RateLimit,Server,Network,Stream,RequestTimeout - Non-retryable:
Authentication,AccessDenied,InvalidRequest,ContextLength,Configuration
The retry() function and generate() respect Retry-After headers and use exponential backoff with jitter.
Provider-specific options
Pass provider-specific parameters via provider_options without losing portability:
use fabro_llm::types::Request;
let request = Request {
provider_options: Some(serde_json::json!({
"anthropic": {
"thinking": {"type": "enabled", "budget_tokens": 10000},
"auto_cache": true
},
"openai": {
"store": true,
"previous_response_id": "resp_abc123"
},
"gemini": {
"safetySettings": [
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"}
]
}
})),
// ...other fields
# model: String::new(), messages: vec![], provider: None, tools: None,
# tool_choice: None, response_format: None, temperature: None,
# top_p: None, max_tokens: None, stop_sequences: None,
# reasoning_effort: None, metadata: None,
};