Commit graph

199 commits

Author SHA1 Message Date
yujonglee
ee295aea68
refactor(rust): give Messages configs typed litellm params (#45611)
* refactor(rust): give Messages configs typed litellm params

* feat(rust): port _normalize_system_role_messages for Messages hosts

mid_conversation_system mirrors Python's module function for function: a
leading run of system turns is hoisted into the top-level system, a later
turn stays in place when the model map flags supports_mid_conversation_system
and otherwise becomes a user turn where it was, never between a tool_use and
its tool_result, and billing blocks are stripped from the hoisted system. The
capability joins MessagesModelCapabilities and the bridge projection. Azure AI
uses it in place of fold_system_role_messages, which hoisted every turn.

* fix(rust): keep the Bedrock runtime endpoint behind a blank api_base

* refactor(rust): import LitellmParams from its crate instead of a re-export

* refactor(rust): template the Messages request decode errors

Both decode failures go through ErrorDetail::invalid with the subject and the
serde error as source instead of a preformatted sentence.

* refactor(rust_bridge): read the module globals the param specs name

The Messages host no longer keeps its own list of litellm globals; a spec
whose spellings the call leaves out names the global to read, so the next
family that has one needs no host change.

* refactor(rust): give each route lookup Python's implicit None

ProviderConfigManager's per-route methods list only the providers the route
serves and fall through to None; the Rust lookups now do the same, so a new
LlmProviders variant touches only the route that serves it.
2026-10-09 16:29:45 -07:00
yujonglee
3d2abf9a44
feat(rust): one LitellmParams type for a config and a caller (#45642)
litellm-router-types mirrors litellm/types/router.py: LitellmParams holds the
model, the credentials, one flattened group per credential family from
auth-types (aws, vertex), the deployment settings a config spells, and every
other key in extra, as Python's extra="allow" keeps it. Config parses it
directly, so config::LiteLlmParams and its untyped additional_fields are gone
and a mistyped provider param is a parse error. fields() and specs() derive
from the families' param specs for hosts that project kwargs and fold module
globals. Spelled<T> is the one untagged shape for a typed value or the text a
config spells, organization takes the list Python's router expands, and the
callback shorthand keeps a value-level one-or-many. No credentials wrapper
group: serde's flatten only consumes keys for struct-shaped children.
2026-10-09 16:29:45 -07:00
yujonglee
04a95f60c1
feat(rust): type the Vertex AI connection params and where Python reads them (#45641)
VertexParams in auth-types holds the vertex_* fields of CredentialLiteLLMParams
in both spellings, with one ParamSpec per setting: the wire names, the litellm
module global VertexBase.get_vertex_ai_* consults, and the environment names.
A credential deserializes from text or the JSON object a config spells, an
empty object counts as absent, and both credential fields are redacted in
Debug. auth-gcp is split so lib.rs is the entrypoint (config, auth,
constants), its lookups resolve through the specs, secret_names() derives
from them, and get_vertex_ai_project_from_credentials reads a service
account's project_id for hosts that need it before a token exchange. The OCR
host folds the module globals in by spec at projection, so OcrSettings no
longer carries vertex_project and vertex_location.
2026-10-09 16:29:44 -07:00
yujonglee
0410abea8b
feat(rust): type the AWS connection params and where Python reads them (#45610)
AwsParams in auth-types holds the aws_* fields of CredentialLiteLLMParams
with one ParamSpec per field: its wire name and the environment names it
falls back to. fields() and secret_names() derive from the specs, the AWS
helpers take the struct instead of a map, aws_auth_config and the region
resolution read through the specs, and a missing region is
AwsParams::REGION.missing("AWS"), whose message is rendered from the spec.
Bedrock Converse, Bedrock transcription, Textract OCR and the Bedrock
Messages config build the struct at their one untyped boundary.
2026-10-09 16:29:44 -07:00
moe-berri
22ad3c5fff
refactor(lens): connect LiteLLM to the independent Lens service (#45529)
* feat(lens): consume the independent Lens service and shared UI

* fix(lens): keep embedded setup stable and vendor UI before builds

* chore(lens): pin embedded UI provenance to published Lens source

* fix(lens): complete embedded UI extraction across CI builds

* docs(lens): record latest main and removal-boundary validation

* fix(lens): wire external chart credentials and ingestion modes

* test(lens): cover adapter failures and fix extraction CI

* docs(lens): refresh bundled chart setup instructions

* test(lens): restore gateway adapter test package marker

* build(lens): pin clean-cut UI and supported storage chart

* test(ui): await guardrail scope menu before checking options

* perf(lens): omit unused trace-team lookup from product requests

* docs(lens): record final-image adapter permission checks

* test(lens): assert delegated investigation authorization

* refactor(lens): remove copied trace runtime and preserve spend logging

* test(lens): verify independent gateway image upgrades and outages

* test(lens): pin the chart qualification database image

* fix(lens): update embedded UI metadata filters

* test: wait for Lens chart pod identity convergence

* fix(e2e): select Lens chart pods by deployment ownership

* fix(e2e): select Helm migration ownership for Lens upgrades

* test(lens): retain migration Jobs through rollout assertions

* docs(helm): require existing database for migration hooks

* test(lens): qualify release boundaries with embedded navigation update

* docs(lens): record browser qualification for embedded tabs

* fix(lens): remove unrelated CI and guardrail test changes

* test(lens): give qualification tenants unique key aliases

* feat(lens): embed the latest canonical Lens interface

* fix(lens): qualify split inference through the gateway

* fix(lens): preserve scoped auth and isolate delegated credentials

* chore(lens): remove unrelated documentation and lint changes

* fix(lens): bound concurrent forwarding buffers through response delivery

* fix(lens): preserve OAuth2 dispatch without inference policies

* chore(lens): adopt current shared onboarding UI

* fix(lens): refresh shared setup UI and review context

* fix(lens): bound uploads before gateway authentication

* docs(lens): remove extraction evidence from the gateway repo

* docs(lens): refresh shared setup prompt for paired releases
2026-10-09 16:17:05 -07:00
moe-berri
3ebc71be2a
fix(rust): update serde_with for serialization advisory (#45498) 2026-10-09 13:31:41 -07:00
yujonglee
0a98c7be10
refactor(rust): derive strum VariantArray and string conversions (#45434)
* refactor(rust): derive strum VariantArray and string conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): use strum conversions directly with explicit spellings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use rstest values for key management systems

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 21:24:49 +00:00
yujonglee
35b0992533
feat(rust): add standalone typed LLM wire contracts (#45188)
* refactor(rust): add typed LLM wire contracts to llms-types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): drop blank lines left in new wire types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): complete additive wire type contracts

* feat(rust): complete additive nested wire contracts

* feat(rust): give Messages content blocks concrete per-block contracts

Replace the shared all-optional ContentBlockPayload with one struct per
documented block and nested tagged unions for server tool results, so
required fields from the official Messages reference are enforced.

Make CustomTool a plain struct with required name and input_schema, add
MessagesToolParam plus the toolset and undated tool-search tags, model
MiniMax media sources as a tagged union, require documented fields on
chat audio, image URLs, and logprobs, and rename the new block enum to
MessagesContentPart so it no longer shadows the streaming struct.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(rust): tighten new wire contracts against the official references

Give each Messages builtin tool a concrete contract with its documented
required fields, fixed names, and typed allowed callers in place of the
shared all-optional ToolDefinition bag. Make MCP tool result content and
fallback triggers optional as the request params document, and narrow
tool search references, web fetch content, and bash results to their
documented block types.

Keep unmodeled Responses output items and new usage iteration and stop
detail types as Recognized values instead of rejecting the whole payload,
accept find_in_page and optional open_page URLs, preserve JSON Schema
property order, and move prompt_cache_breakpoint to Chat Completions
content parts where OpenAI documents it.

Drop OCR table and key/value types that only matched Azure Document
Intelligence and would reject other providers' normalized passthrough,
along with unsourced bounding box and applied edit fields.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(rust): type every documented Responses output item and close wire gaps

Add the 18 remaining documented Responses output item variants with
concrete nested actions, outputs, and safety checks, and add the
documented namespace, async, caller, and image generation fields to the
existing items.

Restrict MCP tool result content to a string or text blocks, model the
Messages diagnostics cache-miss union and its request param with null
distinct from absent, and accept draft-07 tuple items and prefixItems in
JSON Schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(rust): keep Messages contract assertions in test bodies

Compare built-in tool decodes against fully built expected values and
split the required-field rejections per type, so no case carries its
assertions in a callback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): cite the Mistral OCR reference beside the OCR wire types

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 16:27:12 +00:00
moe-berri
e9cfba2c17
feat(lens): isolate ingestion and investigations in a Rust service (#45148)
* feat(lens): isolate trace storage and investigation in a Rust service

* fix(lens): include Rust sources in the image build context

* feat(lens): wire service setup, scoped delivery receipts and lease attempts

* fix(lens): complete service routing and reject stale investigation results

* fix(lens): retry key propagation and validate isolated Compose setup

* fix(lens): seed through isolated ingestion and preserve upstream queue fixes

* chore: sync schema.prisma copies from root

* fix(lens): bind nullable due timestamps as text for Prisma

* chore(ui): remove stale lint suppressions

* fix(lens): address CI failures and review findings

* refactor(lens): remove retired Python worker and run evaluations in Rust

* fix(lens): reuse control connections and satisfy review checks

* test(lens): install and upgrade both Helm charts on Kubernetes

* test(lens): run connection reuse coverage as an integration test

* fix(ui): upgrade Next.js to 16.3.8 security release

* fix(lens): fence stale attempts and preserve reviewed evidence

* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"

This reverts commit 2f79a51b25.

* fix(lens): stop failed investigations and stream history excerpts

* test(lens): cover model tool and result contracts

* test(lens): fix retired routes and reuse installation build artifacts

* test(lens): use portable grep in Helm installation smoke

* test(lens): wait for migrations before forwarding Helm services

* fix(lens): keep failed evidence reads retryable

* fix(lens): preserve sandbox output during process exit

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-10-07 17:18:58 -07:00
devin-ai-integration[bot]
878ba39e7c
refactor(rust): extract inference-testing crate (#44873)
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.

Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 16:33:06 +00:00
devin-ai-integration[bot]
de74f81c69
chore(rust): prune inference deps and rewrite layering docs (#44836)
* chore(rust): prune unused inference crate dependencies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): rewrite inference layering docs for the split format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:27 -07:00
devin-ai-integration[bot]
ae35c9d775
refactor(rust): extract inference-ocr crate (#44832)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
6a55e0a8aa
refactor(rust): extract inference-chat crate (#44827)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
d7f5b40ab4
refactor(rust): extract inference-messages crate (#44818)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
785cb12f37
refactor(rust): extract inference-responses crate (#44811)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:25 -07:00
devin-ai-integration[bot]
63babf23e6
refactor(rust): extract inference-transcription crate (#44809)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
1a9b533e71
refactor(rust): rename litellm-core to litellm-inference (#44802)
* refactor(rust): rename litellm-core to litellm-inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): expose inference base API for format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move shared inference test helpers behind test-support

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
5f9eff55f8
refactor(cache): remove dead Cache._native_cache runtime path (#44785)
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:58:19 -07:00
moe-berri
2190003bb8
fix(lens): reconstruct native coding agent conversations (#44711)
* fix(lens): reconstruct native coding agent conversations

* fix(lens): display timestamps used for conversation ordering

* fix(lens): preserve coding trace identity and message provenance

* fix(lens): complete capture checks after trace pagination

* fix(lens): show loaded replies during trace pagination
2026-10-05 17:51:05 -07:00
devin-ai-integration[bot]
b9251dafad
refactor(rust): derive string enum serde through strum and serde_with (#44675)
* docs(rust): document string enum serde conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): derive string enum serde through strum and serde_with

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:34:34 -07:00
devin-ai-integration[bot]
74ad633e53
refactor(rust): rename litellm-framing crate to litellm-framer (#44683)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:32:00 +00:00
devin-ai-integration[bot]
fdc0e7dc2b
refactor(rust): track ClickHouse migrations in a checksummed ledger (#44580)
* refactor(rust): track ClickHouse migrations in a checksummed ledger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(clickhouse): harden migration execution and reuse retention SQL

* refactor(rust): track ClickHouse migrations in a checksummed ledger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): keep ClickHouse retention TTLs in the current policy list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 16:04:44 +00:00
devin-ai-integration[bot]
461a58c40a
refactor(lens): storage-independent trace reads, shared keyset pager, typed read failures (#44422)
* feat(lens): own trace reads behind a cached TraceStore port

Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.

Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.

Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(lens): seed large profiles with server-side copies and long sessions

Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* style(lens): float the investigation setup badge on the tab edge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens): restructure the trace drawer and polish its layout

Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.

Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* typesafety

* wip

* fmt

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:39:47 -07:00
yujonglee
c7e60f03de
fix(tracing): preserve spend identity and gateway correlation (#44421)
* fix(tracing): preserve spend identity and gateway correlation

* test(tracing): refresh real SDK spend captures

* fix(tracing): resolve complete gateway attempt costs across SDKs

* docs: add trace cost screenshot for PR 44421

* update fixtures

* wip

* docs: remove trace cost screenshot from PR evidence
2026-10-03 16:00:49 -07:00
devin-ai-integration[bot]
f445e466b4
refactor(traces): extract snapshot cache (#44424)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 13:03:30 -07:00
moe-berri
50190134c3
fix(lens): batch run reads and reset trace pagination (#44398)
* fix(lens): batch run reads and reset trace pagination

* fix(lens): scope batched list spend to each run
2026-10-03 18:43:23 +00:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
yujonglee
8d28e8d776
feat(tracing): add scoped SQL queries and schema-aware help (#44085)
* feat(tracing): add SQL queries and schema-aware query help

* test(tracing): verify help requests and sync API types

* refactor(tracing): render query help with Askama

* refactor(tracing): use jinja extension for query guide

* fix(tracing): preserve query help when discovery fails

* feat(tracing): enforce team SQL scope with managed ClickHouse readers

* test(tracing): verify reads with one ClickHouse URL

* fix(tracing): revoke rotated trace reader credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): streamline query help catalog assembly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): run query help discovery sequentially

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): update reader setup request expectations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:06:14 -07:00
yujonglee
ae6b762190
refactor(tracing): normalize agent spans in Rust (#44071)
* refactor(tracing): normalize agent spans in Rust

* refactor(tracing): generate dashboard trace types from API

* test(tracing): use complete trace response fixtures

* fix(ui): expose generated span error response type

* fix(tracing): retain full tool call payloads

* fix(tracing): preserve decoded attribute tuple shape

* fix(tracing): type consumed attributes as tuple
2026-10-01 18:01:20 -07:00
devin-ai-integration[bot]
e3c15c9d22
feat(rust): embed migration folders with a shared migrate! macro (#44104)
* feat(rust): embed migration folders with a shared migrate! macro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): enable syn proc-macro feature for litellm-migrate-macros

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): reject signed versions and symlinks in migrate!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 17:52:54 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
yujonglee
be67fce26a
refactor(proxy): inject tracing receiver and access context (#44035)
* refactor(proxy): inject tracing receiver and access context

* refactor(proxy): own tracing resources through FastAPI lifespan

* test(proxy): pass tracing dependency in Lens lifecycle

* refactor(proxy): stop tracing logger cooperatively

* refactor(proxy): derive tracing permissions in one place

* refactor(proxy): compose application lifespan state

* refactor(proxy): give Lens tracing storage directly

* refactor(tracing): name shared ClickHouse storage explicitly

* refactor(tracing): extract shared ClickHouse storage crate

* test(proxy): isolate db push timeout from Lens safety check

* fix(tracing): drain spend retries during shutdown
2026-10-01 13:45:32 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
yujonglee
268eb4d6e6
feat(tracing): add OTLP trace ingestion and reads (#43915) 2026-09-30 21:12:29 +00:00
yujonglee
41df8cf4d0
feat(traces): add Rust storage foundation (#43819)
* wip

* feat(traces): establish shared Rust storage foundation

* fix(traces): escape ClickHouse text parameters

* test(traces): exercise response cap with bounded strings

* fix(traces): remove unnecessary lint expectation

* fix(traces): encode ClickHouse timestamp units in Rust

* test(traces): mark exception match as a regex

* refactor(traces): execute schema setup in Rust

* refactor(traces): use shared logging execution wrapper

* docs(traces): replace foundation README with boundary rules

* fix(traces): use current bridge execution facade

* fix(traces): account for protocol cast in lint budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 12:00:00 -07:00
devin-ai-integration[bot]
66db132627
refactor(rust): add shared llms wire type derives (#43730) 2026-09-29 09:19:33 -07:00
devin-ai-integration[bot]
5e38a08741
feat(cache): select Rust caching through explicit cache objects (#43601)
* refactor(cache): organize v2 cache as a package

* docs: clarify experimental v2 guidance

* fix(cache): verify cache-hit accounting and preserve logging metadata

* refactor(cache): separate execution facts from host accounting

* refactor(rust): build messages routes with named dependencies

* wip

* fix(cache): preserve facade policy and preflight fallback

* refactor(cache): defer shared Python logging changes

* test(gateway-inference): allow dead code in shared test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cache): key prepared requests and honor facade controls

* feat(cache): use Python caches from Rust Messages inference

* refactor(cache): separate native and Python cache adapters

* refactor(cache): enforce shared composition and adapter boundaries

* fix(cache): let Python key delegated Rust Messages entries

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 00:01:44 +00:00
devin-ai-integration[bot]
e4190d86a6
refactor(rust): centralize host execution and compose callbacks (#43515)
* refactor(rust): extract litellm-host-native as the shared Rust host driver

Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): interrupt the machine when the in-process stream consumer fails

Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): separate the machine contract from coroutine execution

* auth update

* refactor(rust): use standard flow control for host requests

* style(rust): keep host driver imports formatted

* chores

* mostly relocation

* refactor(rust): separate interceptors from queued observers

* refactor(rust): centralize legacy callback mappings and lifecycle

* docs: define Python host boundaries and migration plan

* refactor: enforce Python host and bridge boundaries

* refactor(rust): separate operations from callback composition

* refactor(rust): compose SDK policy through call hooks

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:20:27 +00:00
devin-ai-integration[bot]
f184ace25b
feat(rust): add the MCP gateway (#43470)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
6e0926edde
feat(rust): add gateway UI login and sessions (#43469)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
ed43556e92
feat(rust): add virtual key storage contracts (#43468)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
5f637a2b11
feat(rust): separate gateway authentication and authorization (#43467)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
876539e1b3
feat(rust): add structured route lifecycle tracing (#43466)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:18:58 -07:00
devin-ai-integration[bot]
9b08112ed2
feat(rust): add litellm-db and litellm-db-testing workspace scaffolding (#43504)
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 18:48:45 -07:00
devin-ai-integration[bot]
1ceeefbf84
refactor(rust): use shared execution in gateway inference (#43463)
* feat(rust): add the HTTP host driver

* refactor(rust): use shared execution in gateway inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:50:05 -07:00
devin-ai-integration[bot]
18933c8a21
feat(rust): add the HTTP host driver (#43462)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 16:18:10 -07:00
devin-ai-integration[bot]
36784e3b79
refactor(rust): share call lifecycle across route-owned inference (#43461)
* feat(rust): expand gateway configuration parsing

* refactor(rust): unify core calls and host lifecycle

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): unify core calls and host lifecycle

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): read environment secrets when litellm is not importable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(core): drop the duplicate rstest attribute

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): make shared route dispatch route-owned

* fix(rust): satisfy Clippy in messages regression test

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:05:30 -07:00
devin-ai-integration[bot]
268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
devin-ai-integration[bot]
90873c46de
refactor(rust): expand logging and test coverage across gateway and Anthropic messages (#43295)
* refactor(rust): prepare inference and auth foundations

* fix(rust): keep textract operations parsing from kebab-case model names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* done

* refactor(types): derive Anthropic beta string conversions with Strum

* fix(anthropic): report missing max_tokens as a missing field

* refactor(rust): type Anthropic messages headers and auth after the Python layout

Delete anthropic/messages/headers.rs. Its OAuth handling, credential ladder
and beta merging move to anthropic/common_utils.rs where Python keeps them
(optionally_handle_anthropic_oauth, get_auth_header, _merge_beta_headers),
and the feature beta injection becomes update_headers_with_anthropic_beta on
the messages config, as in Python. The BaseAnthropicMessagesConfig impl is
unchanged apart from the bodies of validate_environment and request_headers

Beta values are now the AnthropicBeta enum and BetaSet, which sort, dedupe
and comma-join by construction. Request params gain typed speed, tools and
context_management through Recognized, so the beta logic matches on enums
instead of string-comparing JSON. OauthToken parses the sk-ant-oat token once
and the chat config shares that detection instead of its own copy

Case-insensitive header helpers move next to has_header in litellm-http.
One deliberate divergence: a Bearer-prefixed OAuth key configured through
api_key or ANTHROPIC_API_KEY is sent with a single Bearer scheme, where
Python would emit "Bearer Bearer"

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* done

* fix(rust): repair test compilation and clippy failures

resolve auth before building the outbound request in prepare tests, give the host hook tests their own error type, and drop the disallowed reqwest client and err().expect() from core tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 08:04:19 +00:00