* refactor(rust): give Messages configs typed litellm params
* feat(rust): port _normalize_system_role_messages for Messages hosts
mid_conversation_system mirrors Python's module function for function: a
leading run of system turns is hoisted into the top-level system, a later
turn stays in place when the model map flags supports_mid_conversation_system
and otherwise becomes a user turn where it was, never between a tool_use and
its tool_result, and billing blocks are stripped from the hoisted system. The
capability joins MessagesModelCapabilities and the bridge projection. Azure AI
uses it in place of fold_system_role_messages, which hoisted every turn.
* fix(rust): keep the Bedrock runtime endpoint behind a blank api_base
* refactor(rust): import LitellmParams from its crate instead of a re-export
* refactor(rust): template the Messages request decode errors
Both decode failures go through ErrorDetail::invalid with the subject and the
serde error as source instead of a preformatted sentence.
* refactor(rust_bridge): read the module globals the param specs name
The Messages host no longer keeps its own list of litellm globals; a spec
whose spellings the call leaves out names the global to read, so the next
family that has one needs no host change.
* refactor(rust): give each route lookup Python's implicit None
ProviderConfigManager's per-route methods list only the providers the route
serves and fall through to None; the Rust lookups now do the same, so a new
LlmProviders variant touches only the route that serves it.
litellm-router-types mirrors litellm/types/router.py: LitellmParams holds the
model, the credentials, one flattened group per credential family from
auth-types (aws, vertex), the deployment settings a config spells, and every
other key in extra, as Python's extra="allow" keeps it. Config parses it
directly, so config::LiteLlmParams and its untyped additional_fields are gone
and a mistyped provider param is a parse error. fields() and specs() derive
from the families' param specs for hosts that project kwargs and fold module
globals. Spelled<T> is the one untagged shape for a typed value or the text a
config spells, organization takes the list Python's router expands, and the
callback shorthand keeps a value-level one-or-many. No credentials wrapper
group: serde's flatten only consumes keys for struct-shaped children.
VertexParams in auth-types holds the vertex_* fields of CredentialLiteLLMParams
in both spellings, with one ParamSpec per setting: the wire names, the litellm
module global VertexBase.get_vertex_ai_* consults, and the environment names.
A credential deserializes from text or the JSON object a config spells, an
empty object counts as absent, and both credential fields are redacted in
Debug. auth-gcp is split so lib.rs is the entrypoint (config, auth,
constants), its lookups resolve through the specs, secret_names() derives
from them, and get_vertex_ai_project_from_credentials reads a service
account's project_id for hosts that need it before a token exchange. The OCR
host folds the module globals in by spec at projection, so OcrSettings no
longer carries vertex_project and vertex_location.
AwsParams in auth-types holds the aws_* fields of CredentialLiteLLMParams
with one ParamSpec per field: its wire name and the environment names it
falls back to. fields() and secret_names() derive from the specs, the AWS
helpers take the struct instead of a map, aws_auth_config and the region
resolution read through the specs, and a missing region is
AwsParams::REGION.missing("AWS"), whose message is rendered from the spec.
Bedrock Converse, Bedrock transcription, Textract OCR and the Bedrock
Messages config build the struct at their one untyped boundary.
* refactor(rust): derive strum VariantArray and string conversions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): use strum conversions directly with explicit spellings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): use rstest values for key management systems
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): add typed LLM wire contracts to llms-types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(rust): drop blank lines left in new wire types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): complete additive wire type contracts
* feat(rust): complete additive nested wire contracts
* feat(rust): give Messages content blocks concrete per-block contracts
Replace the shared all-optional ContentBlockPayload with one struct per
documented block and nested tagged unions for server tool results, so
required fields from the official Messages reference are enforced.
Make CustomTool a plain struct with required name and input_schema, add
MessagesToolParam plus the toolset and undated tool-search tags, model
MiniMax media sources as a tagged union, require documented fields on
chat audio, image URLs, and logprobs, and rename the new block enum to
MessagesContentPart so it no longer shadows the streaming struct.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(rust): tighten new wire contracts against the official references
Give each Messages builtin tool a concrete contract with its documented
required fields, fixed names, and typed allowed callers in place of the
shared all-optional ToolDefinition bag. Make MCP tool result content and
fallback triggers optional as the request params document, and narrow
tool search references, web fetch content, and bash results to their
documented block types.
Keep unmodeled Responses output items and new usage iteration and stop
detail types as Recognized values instead of rejecting the whole payload,
accept find_in_page and optional open_page URLs, preserve JSON Schema
property order, and move prompt_cache_breakpoint to Chat Completions
content parts where OpenAI documents it.
Drop OCR table and key/value types that only matched Azure Document
Intelligence and would reject other providers' normalized passthrough,
along with unsourced bounding box and applied edit fields.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(rust): type every documented Responses output item and close wire gaps
Add the 18 remaining documented Responses output item variants with
concrete nested actions, outputs, and safety checks, and add the
documented namespace, async, caller, and image generation fields to the
existing items.
Restrict MCP tool result content to a string or text blocks, model the
Messages diagnostics cache-miss union and its request param with null
distinct from absent, and accept draft-07 tuple items and prefixItems in
JSON Schema.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(rust): keep Messages contract assertions in test bodies
Compare built-in tool decodes against fully built expected values and
split the required-field rejections per type, so no case carries its
assertions in a callback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(rust): cite the Mistral OCR reference beside the OCR wire types
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.
Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(rust): prune unused inference crate dependencies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): rewrite inference layering docs for the split format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): rename litellm-core to litellm-inference
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): expose inference base API for format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move shared inference test helpers behind test-support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(clickhouse): harden migration execution and reuse retention SQL
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): keep ClickHouse retention TTLs in the current policy list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): own trace reads behind a cached TraceStore port
Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.
Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.
Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): seed large profiles with server-side copies and long sessions
Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chores
* style(lens): float the investigation setup badge on the tab edge
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): restructure the trace drawer and polish its layout
Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.
Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* typesafety
* wip
* fmt
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* feat(tracing): add SQL queries and schema-aware query help
* test(tracing): verify help requests and sync API types
* refactor(tracing): render query help with Askama
* refactor(tracing): use jinja extension for query guide
* fix(tracing): preserve query help when discovery fails
* feat(tracing): enforce team SQL scope with managed ClickHouse readers
* test(tracing): verify reads with one ClickHouse URL
* fix(tracing): revoke rotated trace reader credentials
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): streamline query help catalog assembly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tracing): run query help discovery sequentially
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tracing): update reader setup request expectations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): embed migration folders with a shared migrate! macro
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): enable syn proc-macro feature for litellm-migrate-macros
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): reject signed versions and symlinks in migrate!
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* wip
* feat(traces): establish shared Rust storage foundation
* fix(traces): escape ClickHouse text parameters
* test(traces): exercise response cap with bounded strings
* fix(traces): remove unnecessary lint expectation
* fix(traces): encode ClickHouse timestamp units in Rust
* test(traces): mark exception match as a regex
* refactor(traces): execute schema setup in Rust
* refactor(traces): use shared logging execution wrapper
* docs(traces): replace foundation README with boundary rules
* fix(traces): use current bridge execution facade
* fix(traces): account for protocol cast in lint budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): prepare inference and auth foundations
* fix(rust): keep textract operations parsing from kebab-case model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* done
* refactor(types): derive Anthropic beta string conversions with Strum
* fix(anthropic): report missing max_tokens as a missing field
* refactor(rust): type Anthropic messages headers and auth after the Python layout
Delete anthropic/messages/headers.rs. Its OAuth handling, credential ladder
and beta merging move to anthropic/common_utils.rs where Python keeps them
(optionally_handle_anthropic_oauth, get_auth_header, _merge_beta_headers),
and the feature beta injection becomes update_headers_with_anthropic_beta on
the messages config, as in Python. The BaseAnthropicMessagesConfig impl is
unchanged apart from the bodies of validate_environment and request_headers
Beta values are now the AnthropicBeta enum and BetaSet, which sort, dedupe
and comma-join by construction. Request params gain typed speed, tools and
context_management through Recognized, so the beta logic matches on enums
instead of string-comparing JSON. OauthToken parses the sk-ant-oat token once
and the chat config shares that detection instead of its own copy
Case-insensitive header helpers move next to has_header in litellm-http.
One deliberate divergence: a Bearer-prefixed OAuth key configured through
api_key or ANTHROPIC_API_KEY is sent with a single Bearer scheme, where
Python would emit "Bearer Bearer"
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* done
* fix(rust): repair test compilation and clippy failures
resolve auth before building the outbound request in prepare tests, give the host hook tests their own error type, and drop the disallowed reqwest client and err().expect() from core tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>