Commit graph

401 commits

Author SHA1 Message Date
devin-ai-integration[bot]
dd86ca5175
refactor(traces): type the ClickHouse query help response (#44285)
* refactor(traces): type the ClickHouse query help response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): include agent names and frameworks in named contract round trips

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): cover native query help validation in the storage adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 18:01:27 -07:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
ishaan-berri
481a403090
feat(tracing): support claude agent sdk traces with agent name, logo and chat content (#44248)
* feat(traces): add Framework column to otel_traces

* feat(traces): pass span events to normalizers and add framework field

* feat(traces): add Claude Code and Agent SDK span normalizer

* feat(traces): decode events before normalizing and apply tool span names

* feat(traces): list distinct frameworks per trace

* feat(traces): return span framework in trace spans query

* test(traces): add scrubbed Claude Agent SDK OTLP fixtures

* test(traces): cover Claude Agent SDK normalization from real exports

* test(traces): assert trace list frameworks stay scoped per trace

* feat(tracing): validate framework in native normalized spans

* feat(tracing): add framework to Span and frameworks to TraceSummary

* feat(tracing): store normalized framework on span rows

* feat(tracing): surface span framework and trace frameworks

* test(tracing): cover framework aggregation in trace summaries

* test(tracing): decode Claude Agent SDK rows with framework and tool args

* chore(ui): regenerate API types for trace frameworks

* feat(ui): add trace framework registry for Claude Agent SDK and Claude Code

* feat(ui): show SDK logo and label in the runs list Agent column

* feat(ui): show SDK logo and label in the run header

* test(ui): cover SDK label and logo in the runs list

* test(ui): cover SDK label and logo in the run header

* feat(tracing): show the agent's final answer as claude agent span output

* feat(tracing): name claude code agents after their otel service

* test(tracing): cover claude code agent naming from the service

* fix(tracing): mark the span row framework field read-only

* test(tracing): scrub host os details from the claude sdk fixture

* test(tracing): scrub host os details from the detailed claude sdk fixture

* fix(ui): hide the decorative sdk logo from screen readers

* feat(ui): show the agent name with the sdk logo in the runs list

* feat(ui): show the agent name with the sdk logo in the run header

* test(ui): cover agent names beside the sdk logo in the runs list

* test(ui): cover the agent name in the run header
2026-10-02 21:24:56 +00:00
moe-berri
2584721ca3
fix(lens): preserve framework agent names and GenAI message content (#44218)
* fix(lens): use recorded agent identities across framework traces

* fix(lens): tighten agent identity and bound trace lookups

* style(tracing): wrap framework agent identity test case
2026-10-02 11:58:32 -07:00
yujonglee
276fc9c63a
fix(tracing): unify ClickHouse storage configuration (#43941)
* fix(tracing): use ClickHouse URL for reads by default

* fix(tracing): unify ClickHouse storage configuration

* fix(tracing): update dashboard setup copy for one URL

* test(tracing): make tests/unit/tracing a package

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(config): drop legacy string tracing store variant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): own ClickHouse defaults in constants and reject unset env references

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): use raw regex patterns in config tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): read ClickHouse env defaults when tracing config resolves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): split audit log query guard to fit condition-chain budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 16:31:26 +00:00
yujonglee
8d28e8d776
feat(tracing): add scoped SQL queries and schema-aware help (#44085)
* feat(tracing): add SQL queries and schema-aware query help

* test(tracing): verify help requests and sync API types

* refactor(tracing): render query help with Askama

* refactor(tracing): use jinja extension for query guide

* fix(tracing): preserve query help when discovery fails

* feat(tracing): enforce team SQL scope with managed ClickHouse readers

* test(tracing): verify reads with one ClickHouse URL

* fix(tracing): revoke rotated trace reader credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): streamline query help catalog assembly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): run query help discovery sequentially

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): update reader setup request expectations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:06:14 -07:00
devin-ai-integration[bot]
3f39fef52f
perf(traces): recalculate ClickHouse TTL info only on retention changes (#44117)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 01:22:37 +00:00
yujonglee
ae6b762190
refactor(tracing): normalize agent spans in Rust (#44071)
* refactor(tracing): normalize agent spans in Rust

* refactor(tracing): generate dashboard trace types from API

* test(tracing): use complete trace response fixtures

* fix(ui): expose generated span error response type

* fix(tracing): retain full tool call payloads

* fix(tracing): preserve decoded attribute tuple shape

* fix(tracing): type consumed attributes as tuple
2026-10-01 18:01:20 -07:00
devin-ai-integration[bot]
e3c15c9d22
feat(rust): embed migration folders with a shared migrate! macro (#44104)
* feat(rust): embed migration folders with a shared migrate! macro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): enable syn proc-macro feature for litellm-migrate-macros

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): reject signed versions and symlinks in migrate!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 17:52:54 -07:00
moe-berri
cc42a352cb
feat(lens): simplify setup and investigation workflow (#44089)
* feat(lens): simplify investigation setup and results

* fix(lens): pin worker with actionable failure diagnostics

* fix(lens): focus worker success on starting an investigation

* feat(lens): simplify investigation setup and worker defaults

* fix(lens): remove setup repetition and label billing access

* fix(lens): finish agent selection and setup readiness

* fix(lens): handle unavailable setup dependencies and restore UI build
2026-10-01 17:21:41 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
yujonglee
be67fce26a
refactor(proxy): inject tracing receiver and access context (#44035)
* refactor(proxy): inject tracing receiver and access context

* refactor(proxy): own tracing resources through FastAPI lifespan

* test(proxy): pass tracing dependency in Lens lifecycle

* refactor(proxy): stop tracing logger cooperatively

* refactor(proxy): derive tracing permissions in one place

* refactor(proxy): compose application lifespan state

* refactor(proxy): give Lens tracing storage directly

* refactor(tracing): name shared ClickHouse storage explicitly

* refactor(tracing): extract shared ClickHouse storage crate

* test(proxy): isolate db push timeout from Lens safety check

* fix(tracing): drain spend retries during shutdown
2026-10-01 13:45:32 -07:00
moe-berri
6d7d183a80
feat(lens): investigate sampled traces and retain batch results (#43942)
* fix(lens): parallelize scan analysis with bounded concurrency

* feat(lens): investigate sampled activity and preserve scan results

* fix(lens): pin the compatible investigation worker image

* fix(lens): report incomplete reviews and simplify setup validation

* fix(lens): stabilize large investigations and preserve incomplete results

* fix(lens): preserve bounded readers and distinguish counterexamples

* fix(lens): pin compatible worker and verify batched grouping cost

* fix(lens): exclude counterexamples from finding recurrence

* feat(lens): show completed scan duration in results and history

* fix(lens): fold batch selection into results navigation
2026-09-30 22:30:04 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
yujonglee
629c2b5808
feat(tracing): store spend in ClickHouse automatically (#43928) 2026-09-30 22:17:43 +00:00
yuneng-jiang
a6f6c64b6e
feat(ui): adopt the new LiteLLM logo and monogram (#43913)
* feat(ui): adopt the new LiteLLM logo and monogram

Swap the bundled admin UI logos for the new brand assets: the primary
logo in blue for light mode and white for dark mode, and the monogram for
the collapsed sidebar, favicons, and the built-in guardrail cards.

/get_image gains a variant=monogram query parameter so the collapsed
sidebar can request the monogram while admin-configured UI_LOGO_PATH /
UI_LOGO_PATH_DARK logos still take precedence. The bundled light logo
moves from JPEG to a transparent PNG.

* test(ui): query collapsed sidebar logos by role to stay within the lint budget

* fix: point remaining logo consumers at the new bundled assets

The Rust gateway UI served /get_image from the removed litellm_logo.jpg,
the non-root get_image tests pinned logo.jpg, and a cookbook script read
litellm/proxy/logo.jpg. Point them at the monogram and logo.png.

* fix(mcp): serve the BYOK OAuth page logo from /get_image

The page pointed at /ui/assets/logos/litellm_logo.jpg, which the rebrand
removes from the dashboard sources, so the next UI build would drop it.
/get_image?variant=monogram is always served by the proxy and follows any
admin-configured logo.

* fix(ui): invert the LiteLLM monogram on dark guardrail cards

The blue monogram has a transparent train cut-out, so on a dark card it
read as a muddy blue block. Inverting it yields the brand's white mark,
which the logo guidelines prescribe for dark backgrounds.

* fix(gateway-ui): serve theme and variant aware logos from the dashboard export

The Rust gateway served one monogram for every /get_image request, and the
committed export lacked it, so /get_image returned 404 until the next UI
release build. Pick the full or monogram logo in light or dark from the
query, ship those assets in the dashboard's public dir and the committed
export, and drop two comments that restated asserted paths.
2026-09-30 14:51:06 -07:00
yujonglee
268eb4d6e6
feat(tracing): add OTLP trace ingestion and reads (#43915) 2026-09-30 21:12:29 +00:00
devin-ai-integration[bot]
1fa3cde6a2
fix(traces): correct ClickHouse rollup partitioning, dedupe keys, and retention changes (#43901)
* fix(traces): correct ClickHouse rollup partitioning, dedupe keys, and retention changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): pin spend dedupe timestamps within one second

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 19:45:11 +00:00
yujonglee
41df8cf4d0
feat(traces): add Rust storage foundation (#43819)
* wip

* feat(traces): establish shared Rust storage foundation

* fix(traces): escape ClickHouse text parameters

* test(traces): exercise response cap with bounded strings

* fix(traces): remove unnecessary lint expectation

* fix(traces): encode ClickHouse timestamp units in Rust

* test(traces): mark exception match as a regex

* refactor(traces): execute schema setup in Rust

* refactor(traces): use shared logging execution wrapper

* docs(traces): replace foundation README with boundary rules

* fix(traces): use current bridge execution facade

* fix(traces): account for protocol cast in lint budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 12:00:00 -07:00
devin-ai-integration[bot]
b80052839e
refactor(rust): centralize bridge execution wrappers (#43871)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 15:54:27 +00:00
devin-ai-integration[bot]
9dda4d895f
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764)
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-09-30 07:32:16 -07:00
devin-ai-integration[bot]
273489824a
refactor(rust): orchestrate Messages route execution (#43719)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 19:54:12 +00:00
berriai-litellm-provider-info-sync[bot]
f4a7c04d99
chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page (#43745)
* chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page

Price-Sync: litellm-providers

* feat(cost): support openai ultrafast tier fields in the model catalog

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry <kerry@berri.ai>
2026-09-29 19:05:57 +00:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
devin-ai-integration[bot]
66db132627
refactor(rust): add shared llms wire type derives (#43730) 2026-09-29 09:19:33 -07:00
devin-ai-integration[bot]
5e38a08741
feat(cache): select Rust caching through explicit cache objects (#43601)
* refactor(cache): organize v2 cache as a package

* docs: clarify experimental v2 guidance

* fix(cache): verify cache-hit accounting and preserve logging metadata

* refactor(cache): separate execution facts from host accounting

* refactor(rust): build messages routes with named dependencies

* wip

* fix(cache): preserve facade policy and preflight fallback

* refactor(cache): defer shared Python logging changes

* test(gateway-inference): allow dead code in shared test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cache): key prepared requests and honor facade controls

* feat(cache): use Python caches from Rust Messages inference

* refactor(cache): separate native and Python cache adapters

* refactor(cache): enforce shared composition and adapter boundaries

* fix(cache): let Python key delegated Rust Messages entries

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 00:01:44 +00:00
devin-ai-integration[bot]
e4190d86a6
refactor(rust): centralize host execution and compose callbacks (#43515)
* refactor(rust): extract litellm-host-native as the shared Rust host driver

Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): interrupt the machine when the in-process stream consumer fails

Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): separate the machine contract from coroutine execution

* auth update

* refactor(rust): use standard flow control for host requests

* style(rust): keep host driver imports formatted

* chores

* mostly relocation

* refactor(rust): separate interceptors from queued observers

* refactor(rust): centralize legacy callback mappings and lifecycle

* docs: define Python host boundaries and migration plan

* refactor: enforce Python host and bridge boundaries

* refactor(rust): separate operations from callback composition

* refactor(rust): compose SDK policy through call hooks

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:20:27 +00:00
devin-ai-integration[bot]
74cad08997
refactor(rust): remove delivery routing abstraction (#43514)
* wip

* refactor: finish removing delivery routing abstraction

* refactor(rust): remove chat completion decline admission

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 03:31:22 +00:00
devin-ai-integration[bot]
f184ace25b
feat(rust): add the MCP gateway (#43470)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
6e0926edde
feat(rust): add gateway UI login and sessions (#43469)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
ed43556e92
feat(rust): add virtual key storage contracts (#43468)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
5f637a2b11
feat(rust): separate gateway authentication and authorization (#43467)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
876539e1b3
feat(rust): add structured route lifecycle tracing (#43466)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:18:58 -07:00
devin-ai-integration[bot]
8ab124309f
feat(rust): connect Python inference bindings to shared routes (#43465)
* feat(rust): support the HTTP Responses API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): enable native Python inference opt-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust_bridge): cover only python-only routes in the native load guard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep Python inference rollout disabled

* fix(rust): preserve inference defaults and continuation IDs

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:05:39 -07:00
devin-ai-integration[bot]
9b08112ed2
feat(rust): add litellm-db and litellm-db-testing workspace scaffolding (#43504)
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 18:48:45 -07:00
devin-ai-integration[bot]
0a21f24d50
feat(rust): support the HTTP Responses API (#43464)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 17:38:15 -07:00
devin-ai-integration[bot]
1ceeefbf84
refactor(rust): use shared execution in gateway inference (#43463)
* feat(rust): add the HTTP host driver

* refactor(rust): use shared execution in gateway inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:50:05 -07:00
devin-ai-integration[bot]
438bffc26e
build(rust): package the gateway container (#43471)
* build(rust): package the gateway container

* ci: exempt the gateway Dockerfile from the CI coverage gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:49:39 -07:00
devin-ai-integration[bot]
18933c8a21
feat(rust): add the HTTP host driver (#43462)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 16:18:10 -07:00
devin-ai-integration[bot]
36784e3b79
refactor(rust): share call lifecycle across route-owned inference (#43461)
* feat(rust): expand gateway configuration parsing

* refactor(rust): unify core calls and host lifecycle

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): unify core calls and host lifecycle

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): read environment secrets when litellm is not importable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(core): drop the duplicate rstest attribute

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): make shared route dispatch route-owned

* fix(rust): satisfy Clippy in messages regression test

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:05:30 -07:00
devin-ai-integration[bot]
b94f5bdbed
feat(rust): expand gateway configuration parsing (#43460)
* feat(rust): expand gateway configuration parsing

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:54:04 -07:00
devin-ai-integration[bot]
268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
devin-ai-integration[bot]
501ef23f4a
feat(rust): add the openai_like chat config foundation (#43379)
* feat(rust): add the openai_like chat config foundation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): let max_completion_tokens outrank max_tokens and decline refusal responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 18:07:09 -07:00
devin-ai-integration[bot]
3a6744cd02
feat(sail): add Sail as a provider with service_tier mapped to its completion window (#42840)
Register Sail (providers.json, LlmProviders.SAIL, OpenAI-compatible lists,
ProviderConfigManager) for chat, Responses and /v1/messages, and add its 12
models to both cost maps with asap, balanced and flex price columns.

Sail picks speed and price with metadata.completion_window and rejects
service_tier, so the Sail chat and Responses configs translate the tier:
default and priority to asap, flex to flex, balanced to balanced, auto to no
window. Billing prices the window that was sent. A tier Sail has no window
for, or a window or tier set where billing cannot see it (request metadata,
extra_body), is a 400 unless drop_params is set.

Add balanced to ServiceTier and its _balanced price columns to the model
info types, the Rust catalog and the dashboard schema. A transform_extra_body
hook on the chat and Responses base configs, which returns extra_body
unchanged by default, lets Sail keep the window when a caller also sends
extra_body.metadata. Sail is listed in the Add Model form and model picker.

Co-authored-by: shrey kharbanda <shrey@berri.ai>
2026-09-26 12:57:48 -07:00
yuneng-jiang
14f4c34c61
fix(ci): stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts (#43294)
* fix(ci): stop five stale or flaky CI reds and retry CyberArk policy-load conflicts

The Langfuse redaction unit test exports to a local OTLP capture instead of
polling Langfuse Cloud through a recorded lookup. The passthrough worker-kill
test only requires spend rows for requests the surviving worker served. The
spend-routes sweep treats the intentional /spend/capture_rate 503 as expected.
CyberArk retries a 409 policy load in Python, Rust and the e2e Conjur helper
instead of reading it as "variable exists". The integration egress guard now
matches the script's own cgroup, so it no longer blocks the CircleCI agent,
which runs as the same user.

* fix(ci): keep the policy-load backoff typed as float

* fix(ci): retry CyberArk policy loads without blocking the event loop and tighten the worker-kill and Langfuse tests

* fix(secrets): load CyberArk policy one request at a time per manager

* test(secrets): pin that non-conflict CyberArk policy failures are not retried

* test(unit): run tests/unit with only an allowlisted host environment

CircleCI's unit job inherits every project env var, so real provider keys,
REDIS_HOST, DATABASE_URL and AWS or Azure credentials reached tests that
assume none are set. Locally, litellm's import-time load_dotenv did the same
from any .env up the tree. The unit conftest now drops every variable outside
a small allowlist and disables dotenv before litellm is imported.

* test(e2e): name a failed search and the stuck batch status instead of misattributing them

The websearch session test read an empty web_search_tool_result_error block as a
successful search, so a failing search tool surfaced as a session billing bug.
The batch cancellation timeout now reports the last status the proxy returned.

* fix(ci): scrub the host environment per unit test instead of for the whole pytest process

GHA shards run tests/unit next to other suites in one process, so the import-time
scrub deleted MCP_TEST_PEER_PYTHON before tests/mcp_tests read it and the MCP
upstream fell back to the SDK2 interpreter. The two websearch tests that called
OpenAI and Perplexity live are removed: tests/unit no longer sees their keys.

* fix(ci): scrub only the host variables present before litellm is imported

The per-test scrub also deleted TIKTOKEN_CACHE_DIR, which litellm sets at import to
its bundled encodings, so tokenizer paths tried to download them and hit the
socket guard. The prisma setup test now passes its own database URL instead of
reading one another test leaked into the process environment.

* fix(ci): stop the order-dependent unit reds and settle logging tasks on their own queue

LoggingWorker marked a task done on whichever queue was current when the callback
finished, so a callback that outlived an event-loop change raised "task_done()
called too many times" or undercounted the new loop's queue. It now settles the
queue the task came from.

The rest are test isolation fixes for failures that only appeared when another
file ran first on the same xdist worker: a replaced user_api_key_cache, breaker
metrics unregistered by prometheus tests, semantic_router's health-check filter on
uvicorn.access, logging tasks carried over from bedrock tests, a Router-written
model_cost entry, and a stray post captured by the langflow test. The token
counter check now asserts bounded chunking instead of wall-clock time.

* test(e2e/ui): wait for the logout redirect before visiting a protected page

Logout revokes the session server-side before clearing cookies and navigating, so an immediate page.goto either ran with the cookie still set or was aborted by the logout redirect (net::ERR_ABORTED).

* test(unit): restore the prometheus metrics config per test and settle logs carried from earlier tests in the a2a cost tests

* test(router): pin the router clock in the usage counter tests so a minute rollover cannot empty the read

* test(e2e/ui): wait for logout to clear the token cookie instead of for a login redirect

* test(integration/mcp): answer the model-info probe another test's proxy sends to the model double
2026-09-26 09:25:13 -07:00
devin-ai-integration[bot]
90873c46de
refactor(rust): expand logging and test coverage across gateway and Anthropic messages (#43295)
* refactor(rust): prepare inference and auth foundations

* fix(rust): keep textract operations parsing from kebab-case model names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* done

* refactor(types): derive Anthropic beta string conversions with Strum

* fix(anthropic): report missing max_tokens as a missing field

* refactor(rust): type Anthropic messages headers and auth after the Python layout

Delete anthropic/messages/headers.rs. Its OAuth handling, credential ladder
and beta merging move to anthropic/common_utils.rs where Python keeps them
(optionally_handle_anthropic_oauth, get_auth_header, _merge_beta_headers),
and the feature beta injection becomes update_headers_with_anthropic_beta on
the messages config, as in Python. The BaseAnthropicMessagesConfig impl is
unchanged apart from the bodies of validate_environment and request_headers

Beta values are now the AnthropicBeta enum and BetaSet, which sort, dedupe
and comma-join by construction. Request params gain typed speed, tools and
context_management through Recognized, so the beta logic matches on enums
instead of string-comparing JSON. OauthToken parses the sk-ant-oat token once
and the chat config shares that detection instead of its own copy

Case-insensitive header helpers move next to has_header in litellm-http.
One deliberate divergence: a Bearer-prefixed OAuth key configured through
api_key or ANTHROPIC_API_KEY is sent with a single Bearer scheme, where
Python would emit "Bearer Bearer"

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* done

* fix(rust): repair test compilation and clippy failures

resolve auth before building the outbound request in prepare tests, give the host hook tests their own error type, and drop the disallowed reqwest client and err().expect() from core tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 08:04:19 +00:00
devin-ai-integration[bot]
affb547525
feat(rust): add config router and gateway crates (#43289)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-25 23:12:48 -07:00
devin-ai-integration[bot]
7ae721bf79
refactor(rust): prepare inference and auth foundations for the gateway (#43287)
* refactor(rust): prepare inference and auth foundations

* fix(rust): keep textract operations parsing from kebab-case model names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 23:12:48 -07:00
devin-ai-integration[bot]
4eb340abe0
fix(callbacks-legacy-python): traverse and release the retained headers dict (#43274)
* fix(callbacks-legacy-python): traverse and release the retained headers dict

LegacyLogging keeps the headers dict it hands to pre_call and post_call, but
its traverse never reported that edge to the collector and close never dropped
it. A cycle a callback builds through that dict could not be collected, and a
closed call kept the dict alive until the driver dropped the whole adapter.
Visit and clear headers like body, with regression tests for both

* refactor(callbacks-legacy-python): move the test support module into its own file

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-25 20:24:58 -07:00
devin-ai-integration[bot]
e1a9378093
fix(rust): preserve nested optional import failures (#43265)
* fix(rust): preserve nested optional import failures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): restore Python modules after settings tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 19:41:53 -07:00