Commit graph

73 commits

Author SHA1 Message Date
tin-berri
db14931401
fix(auto-router): attribute router-day savings without spend logs or a session (#44226)
Per-router auto-router savings were missing on proxies with
disable_spend_logs and for requests without a session id under
missing_session_id: omit. The router-day row is written by the
auto-router turn, whose enqueue sat behind the spend-logs flag and whose
builder dropped sessionless requests, while LiteLLM_DailyUserSpend has
neither gate. That money then showed only as unattributed savings.

Enqueue the turn whether or not spend logs are kept, and keep a
sessionless turn with an empty session id that writes the router-day row
while the session upserts skip it. The day and session rows still commit
in one statement, so late baseline corrections keep their ordering.

Without spend logs, write only the router-day aggregate: the turn drops
its session id, so no per-session row is stored, and baseline capture is
skipped, since a baseline observation can only publish once its
request's spend log exists. The flag keeps its meaning of no per-request
or per-session data, while the daily per-router money matches
LiteLLM_DailyUserSpend.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 20:56:45 -07:00
devin-ai-integration[bot]
ca1994e403
fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice (#44295)
* fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): preserve deferred logging bridge coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover gpt-6 family bridge routing and merged tool calls in integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 20:17:55 -07:00
ishaan-berri
2ccb7ed06e
feat(lens): issue briefs with problem, user goal, outcome and test cases (#44311)
* refactor(lens): move analysis prompts into markdown files

* feat(lens): ask the investigator for a scoped agent fix brief with two options

* test(lens): cover the agent fix brief through investigation and merges

* chore(ui): regenerate api types for the lens fix brief

* feat(ui): build copyable lens fix prompts

* feat(ui): show the lens fix brief with copy buttons for claude code and codex

* test(ui): cover copying a lens fix option

* refactor(lens): replace the fix options with a plain issue brief

* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes

* test(lens): cover the issue brief through investigation and merges

* chore(ui): regenerate api types for the lens issue brief

* refactor(ui): drop the lens fix prompt builders

* feat(ui): add a lens issue brief panel

* feat(ui): show the lens issue brief in the finding drawer

* test(ui): cover the lens issue brief and the legacy fallback

* feat(ui): render a lens issue brief as a markdown document

* test(ui): pin the lens issue brief markdown layout

* feat(ui): show the issue brief as a copyable file with claude code and codex buttons

* feat(ui): pass the finding title into the issue brief

* test(ui): cover copying the issue brief for claude code and codex

* feat(ui): bold the input and expected labels in lens test cases

* test(ui): pin the bold test case labels in the issue brief

* feat(ui): render the issue brief as formatted markdown

* test(ui): cover the rendered issue brief sections and raw markdown copy
2026-10-03 03:00:29 +00:00
tin-berri
f63d989ff9
feat: add Bespoke Nimble gateway and OSS classifier support (#44246)
* feat: add Bespoke Nimble gateway and OSS classifier support

* feat: accept Ollama's nimble model name for the Bespoke provider

* test: exempt the POST-only bespoke decisions route from the all-methods check

test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 19:37:08 -07:00
yujonglee
677205b3f5
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy

* test(proxy): use named bindings for spend scope regression

* test(proxy): reuse existing spend log query harness

* test(proxy): cover spend log permission lookup adoption

* chore(proxy): relocate existing spend query baseline

* refactor(proxy): make scope query returns explicit

* refactor(proxy): inject deferred log permission lookup

* test(proxy): cover teamless management compatibility lookup

* refactor(proxy): compose user and team log grants

* refactor(proxy): share generic authorization composition

* refactor(proxy): compose trace read permissions

* refactor(proxy): centralize spend and trace authorization

* refactor(proxy): strengthen spend and trace scope types

* refactor(proxy): flatten log read scope into owned logs

Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.

A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(proxy): run spend scope tests through one SQLite emulator

Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.

load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): unify log and trace ownership permissions

* test(tracing): align fixtures with ownership read scopes

* refactor(tracing): align query scopes with row ownership

* refactor(spend): make ownership SQL predicates explicit

* test(spend): validate ownership SQL against PostgreSQL

* docs(traces): drop key-row visibility from query help guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): reach the empty-memberships branch in team lookup test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 01:38:54 +00:00
yuneng-jiang
b877a38e5f
fix(proxy): stop queued registry read-throughs spending the resync budget (#44277)
* fix(proxy): stop queued registry read-throughs spending the resync budget

RegistryReadThrough.attempt serializes misses behind one lock, but every
request that was queued behind the first one still spent a unit of the
20-per-5s resync budget and re-ran the DB resync, even though the first
request had already loaded the object. A burst of more than 20 requests
for a model created on another worker therefore exhausted the budget and
the rest got 400 "Invalid model name".

attempt now checks whether the key is already loaded once it holds the
lock and returns early without touching the budget. Models check the
router's model names and deployment ids, guardrails and agents reuse
their existing registry lookups.

* test(proxy): gate the queued read-through test on events and cover each registry's loaded check

The queued-requests test now holds the first resync on an asyncio.Event instead of
a timed sleep and records calls in a recorder with tuple and frozenset state. New
tests show the model, guardrail and agent read-throughs each answer an object that
is already loaded without reading the database, so rewiring any registry's loaded
check now fails a test

* test(proxy): keep the queued read-through recorder inside its test and type the agent registry fixture
2026-10-02 18:12:56 -07:00
devin-ai-integration[bot]
dd86ca5175
refactor(traces): type the ClickHouse query help response (#44285)
* refactor(traces): type the ClickHouse query help response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): include agent names and frameworks in named contract round trips

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): cover native query help validation in the storage adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 18:01:27 -07:00
moe-berri
f498176a27
feat(lens): add sample previews and improve setup and worker feedback (#44268)
* feat(lens): add interactive traces and investigations demo

* fix(lens): limit sample previews to setup screens

* fix(lens): simplify tracing setup and align sample previews

* fix(ui): unify Lens and ROI demo notices

* fix(lens): distinguish preparation from zero selected runs

* fix(lens): report incompatible workers and address setup review
2026-10-02 17:27:13 -07:00
yuneng-jiang
a292fd409f
test: fix three order-dependent and timing-flaky tests (#44271)
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer

The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test

* test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding

/ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin

* test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router

test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture
2026-10-02 16:29:57 -07:00
ishaan-berri
a2068820ee
perf(proxy): batch daily model usage writes instead of upserting per request (#44243)
* perf(proxy): aggregate daily model usage per flush instead of upserting per request

* perf(proxy): drain queued model usage in the spend log flush job

* perf(proxy): queue model usage at request time instead of writing to the db

* test(proxy): cover batched daily model usage aggregation and retries

* test(proxy): read back model insights written by the batched flush

* fix(proxy): drain the whole model usage queue each flush so it cannot grow unbounded

* test(proxy): give the mock prisma client a model usage queue

* test(proxy): cover draining a model usage queue larger than one spend log batch
2026-10-02 16:07:18 -07:00
joshua-berri
04bc354525
refactor(mcp): extract upstream preparation and support modern clients (#44232)
* refactor(mcp): extract upstream preparation and support modern clients

* fix(mcp): reject incompatible upstream transport before saving

* fix(mcp): serialize protocol validation with server updates

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-02 16:04:11 -07:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
yuneng-jiang
9b5562f89b
test: repair stale and polluting tests red on scheduled main CI (#44229)
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
2026-10-02 21:32:17 +00:00
moe-berri
9fb327e8c6
fix(lens): run investigations with configured wildcard models (#44233)
* fix(lens): run investigations with configured wildcard models

* fix(lens): validate worker model access and pricing before analysis

* fix(lens): bound worker validation and preserve unrelated edits
2026-10-02 14:20:59 -07:00
devin-ai-integration[bot]
0238ec9721
fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key (#43978)
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key

A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a
rename) carries a partial Group resource. Each of its attributes now applies as
if sent with that path, so displayName updates the team alias and externalId
and members get their usual handling, and the pushed attributes merge into the
scim_data snapshot the PUT path already writes. A path-less remove or a
path-less op without an object value is rejected with a 400. Any group PATCH
drops an empty metadata key an earlier push left behind, and the Admin UI
metadata form skips an empty key so an affected team can save its settings.

* fix(scim): let a later path op win over an earlier path-less value in the group snapshot

* fix(scim): type the stored team metadata before the JSON object check

* test(scim): run the real group transformation in the path-less replace test

* test(scim): assert the renamed group comes back from the path-less replace

* test(scim): audit the path-less group PATCH on the live proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 14:20:04 -07:00
devin-ai-integration[bot]
f932e292c5
fix(proxy): carry key, team and project tags into pass-through spend logs (#42662)
* fix(proxy): carry key, team and project tags into pass-through spend logs

Pass-through endpoints built their request metadata without the key, team and project controls that native routes apply, so spend rows for configured routes and provider pass-throughs like /anthropic dropped the key, team and project tags and the key and team spend_logs_metadata. The native team and project controls now live in a shared helper that both paths call, key spend_logs_metadata is copied instead of aliased from the cached key, client metadata cannot overwrite user_api_key_ fields, and header tags dedupe with the same merge used on native routes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop covers markers from pass-through tag tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): type the shared team and project control helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): validate pass-through endpoint list instead of suppressing pyright

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover body, streaming, hostile, forged and native cells for pass-through tags

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: mock pass-through request url as httpx.URL after rebase on main

* refactor(proxy): merge spend_logs_metadata sources without a stacked comprehension

* refactor(proxy): name the team and request spend_logs_metadata merge

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 19:38:44 +00:00
devin-ai-integration[bot]
b21e44cbf9
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(jwt): route existing-key lookup through VerificationTokenRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key

Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.

* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team

Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.

With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.

* fix(jwt): keep the early return when no master key is set

Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.

Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.

* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes

A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)

Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget

Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401

Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter

* test(e2e): create the reused key in the team the JWT resolves to

The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass

* test(integration): match the held statement across TCP reads

The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read

* fix(jwt): gate key reuse on the claim field, not on the claim value

Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path

* fix(jwt): let an issuer's own user field replace the global one when gating key reuse

An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key

* test(jwt): make the flag-off test fail when the flag no longer gates key reuse

The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings

* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:11:33 -07:00
tin-berri
8aaa766734
fix(auto-router): count usage savings by selected UTC request day (#44115)
The auto-router usage view summed the lifetime savings of every session
overlapping the date range, so it disagreed with the Overall savings view,
which sums daily rollups by request day.

Record auto-routed money per UTC request day and router in one new table,
written in the same statement as the session rollup and corrected in the
same transaction as late baseline estimates. The all-router headline reads
the same daily rows and filters as Overall; savings no router day row
accounts for are reported as unattributed and void the baseline comparison.
Session shape and caching stay whole-session and are labelled so; the
savings-per-session tile is removed.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 11:57:42 -07:00
tin-berri
0dc23406eb
feat: add Laya gateway and OSS classifier providers (#43626)
* feat: add Laya gateway and classifier backend

* test: cover the pass-through model_group pin and repair the shard fakes

MockRequest in tests/pass_through_unit_tests gains an httpx.URL and an ASGI
scope, which get_request_route now reads inside
_init_kwargs_for_pass_through_endpoint, and the POST-only /laya/v1/systemone
route joins the protocol-constrained exemptions. A built-in pass-through pins
metadata.model_group to the resolved model so a client cannot choose its own
per-model budget key; test_pass_through_endpoints now proves that on a
non-Laya route and drops a duplicated assertion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 11:29:36 -07:00
tin-berri
3ae491a06c
fix(proxy): preserve state through composed lifespans (#44214) 2026-10-02 10:56:52 -07:00
devin-ai-integration[bot]
2c9a971171
fix(proxy): exit when DATABASE_URL is set but the Prisma toolchain is missing (#44207)
With DATABASE_URL set and no way to run the Prisma CLI (not on PATH, not importable), run_server used to print a plain notice and keep booting. The server then crashed later inside the DB exception handler with an unrelated ModuleNotFoundError traceback, and the migration-only entrypoint (--skip_server_startup) exited 0 without migrating. It now exits 1 with a red one-line message naming the missing toolchain and how to install it. The no-DATABASE_URL path is unchanged.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 10:32:04 -07:00
yuneng-jiang
2b85808011
feat(mcp)!: disable stdio MCP servers by default (#44066)
* feat(mcp)!: disable stdio MCP servers by default

stdio MCP servers now only run when the proxy is started with
LITELLM_ENABLE_MCP_STDIO=true. While it is off, existing stdio servers stay
registered but never start: tool listings skip them quietly, direct tool
calls and health checks return a 403 naming the env var, and creating or
updating a stdio server is rejected. The flag is read from the process
environment only, so DB-stored environment_variables cannot turn it on.

The UI reads mcp_stdio_enabled from /.well-known/litellm-ui-config to grey
out the stdio transport, show a banner on stdio forms, and badge stdio
server cards.

BREAKING CHANGE: stdio MCP servers are off by default. Set
LITELLM_ENABLE_MCP_STDIO=true in the proxy environment and restart to keep
using them.

* fix(mcp): ignore stdio flag from config file and read UI flag from the selected worker

LITELLM_ENABLE_MCP_STDIO set under environment_variables in config.yaml is now skipped like the DB-stored value, so only the process environment can enable stdio. The dashboard reads mcp_stdio_enabled from the proxy it is managing, so a control plane shows each worker's own setting.

* test(mcp): cover non-mapping payloads in the shared transport validator

* fix(mcp): skip blocked stdio servers quietly in every listing and keep the UI unchanged until the flag loads

Prompt, resource and resource-template listings now skip a blocked stdio server at debug level like tool listing does, instead of logging a warning per server on every call. The dashboard only treats stdio as disabled once the proxy explicitly reports mcp_stdio_enabled false, so a proxy with the flag on, or an older one without the field, renders exactly as before with no flicker while loading.

* fix(mcp): route blocked stdio tool calls to the flag error and warn once per server

A gateway tools/call naming a blocked stdio server's tool now returns the
LITELLM_ENABLE_MCP_STDIO message instead of "Tool not found".

The "will not start" warning moves out of build_mcp_server_from_table, which
DB reload re-runs on every cycle for rows with a NULL updated_at and which
drafts and test-connection also call. It now fires when a row first enters
the registry or changes transport.

* fix(ui): explain on the server detail page why a stdio server is inert

The Overview and MCP Tools tabs showed "No tools available" with no reason
while stdio is disabled. The detail page now shows the same warning banner
as the edit form, and hands off to the form's banner once editing starts.

* refactor(ui): name the stdio banner conditions on the server detail page

Keeps local/no-long-condition-chain within its budget

* fix(proxy): log the ignored DB-stored LITELLM_ENABLE_MCP_STDIO warning once

The DB config sync re-reads environment_variables on every cycle, so a stored
flag logged the warning on each sync per worker
2026-10-02 10:28:04 -07:00
Jim Aldon D'Souza
19da81579b
fix(ui): register tencent in the Add Model provider dropdown (#40924)
* fix(ui): register tencent in the Add Model provider dropdown

The Add Model provider dropdown is driven by the proxy's
/public/providers/fields endpoint, which serves
provider_create_fields.json. Tencent was frozen in the test's
ADD_MODEL_UNLISTED_PROVIDERS set, so it never appeared in the dropdown.

Add a Tencent entry (optional api_base + required api_key, matching
TENCENT_API_BASE/TENCENT_API_KEY) and unfreeze it in the backend test.
Register Tencent in the UI Providers enum, provider_map, and placeholder
map so the dropdown resolves the display name and model placeholder.

* fix(tencent): drop test docstring to satisfy comment policy

* fix(ui): bundle the Tencent Cloud logo for the provider dropdown
2026-10-02 10:15:41 -07:00
yujonglee
276fc9c63a
fix(tracing): unify ClickHouse storage configuration (#43941)
* fix(tracing): use ClickHouse URL for reads by default

* fix(tracing): unify ClickHouse storage configuration

* fix(tracing): update dashboard setup copy for one URL

* test(tracing): make tests/unit/tracing a package

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(config): drop legacy string tracing store variant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): own ClickHouse defaults in constants and reject unset env references

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): use raw regex patterns in config tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): read ClickHouse env defaults when tracing config resolves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): split audit log query guard to fit condition-chain budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 16:31:26 +00:00
devin-ai-integration[bot]
e32b25817f
fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs (#44075)
* fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): audit cells for stored spend-log tool payloads and logprobs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert cache-hit spend-log rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): handle model-list requests in spend-log audit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert unauthenticated requests never reach upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:57:11 -07:00
devin-ai-integration[bot]
d131c43782
fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging (#44067)
* fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): use transport-level doubles in Azure dispatch regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add Azure dispatch audit matrix integration cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): restore global callback lists after Router reset in SDK cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): make proxy restart cell independent of shutdown timing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:30:32 -07:00
PhimmStraiker
826b21aab6
fix(guardrails): straiker v3 routes sk_agt_ keys to v3 and fails closed on a missing verdict (#44011)
* fix(guardrails): straiker v3 routes sk_agt_ keys to v3 and stops reading a missing verdict as allow

An sk_agt_ key always calls /api/v3/detect, even when the guardrail was saved with
api_version 'v1' by the old shared default: the v1 webhook rejects that key with 401.
A 200 with no decision, or permissionDecision 'ask', now takes the failure policy
instead of allowing the request. A detect-mode action is still not a block.
A response-phase block is no longer remembered under the request, so asking the same
question again is scored instead of refused from memory. The replay memory is scoped
by principal and session together, so two principals on one session id never share a
block. A text-only /guardrails/apply_guardrail call relays the text as a user turn.
The agent_ref description now matches the code: the configured value wins.

* fix(guardrails): straiker v3 relays text beside an empty messages list and keys the memory on the key

/guardrails/apply_guardrail sends `messages: []` beside `text`; an empty list is no
conversation, so the text is relayed as the user turn. A verdict whose blocked_by is not
a list states no decision and takes the failure policy, the same way a missing decision
does. A key that names no user is still the caller, so the replay memory is keyed on the
key when no user is known.

* test(guardrails): build the straiker replay-scope request data without mutating it

---------

Co-authored-by: PhimmStraiker <PhimmStraiker@users.noreply.github.com>
2026-10-01 22:38:04 -07:00
yujonglee
8d28e8d776
feat(tracing): add scoped SQL queries and schema-aware help (#44085)
* feat(tracing): add SQL queries and schema-aware query help

* test(tracing): verify help requests and sync API types

* refactor(tracing): render query help with Askama

* refactor(tracing): use jinja extension for query guide

* fix(tracing): preserve query help when discovery fails

* feat(tracing): enforce team SQL scope with managed ClickHouse readers

* test(tracing): verify reads with one ClickHouse URL

* fix(tracing): revoke rotated trace reader credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): streamline query help catalog assembly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): run query help discovery sequentially

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): update reader setup request expectations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:06:14 -07:00
devin-ai-integration[bot]
e6639d15b5
fix(proxy): always exit when database setup fails at boot (#44141)
Remove the ENFORCE_PRISMA_MIGRATION_CHECK opt-in. When PrismaManager.setup_database returns False (database unreachable, connection retries exhausted, or prisma migrate deploy failing after retries) the proxy now always prints the red failure message and exits 1 instead of serving requests against a database whose schema may be behind the code. The --enforce_prisma_migration_check flag stays as a hidden no-op that prints a one-line deprecation warning so existing container args keep parsing; the env var is no longer read anywhere. The standalone migration entrypoint always runs run_server(("--skip_server_startup",)), and the integration launchers drop the flag.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:51:06 -07:00
devin-ai-integration[bot]
5ccb1a143b
fix(proxy): reject non-canonical daily activity dates (#44143)
The daily spend tables store date as text, so a request date that strptime accepts but that is not spelled YYYY-MM-DD (2026-9-24, 2026-09-4, full-width digits) was compared as raw text against canonical rows and matched nothing, and the export route copied it into Content-Disposition, which fails latin-1 encoding and returned 500. A shared parse_canonical_date_range now rejects any spelling whose round trip differs from the input, so every bounded daily activity route (user, team, tag, organization, customer, agent: aggregated, aggregated/keys, search, model top keys, export, cache leakage) and the paginated get_daily_activity path answer 400 before touching the repository, and the export filename is built from the validated dates.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:19:51 -07:00
devin-ai-integration[bot]
b9e71e990a
feat(mcp): add Microsoft 365 (Graph) server to the MCP catalog (#43099)
* feat(mcp): add Microsoft 365 (Graph) server to the MCP catalog

* fix(mcp): signpost the self-hosted Microsoft 365 URL and pin the catalog entry in tests

* test(mcp): pin both shipped copies of a catalog icon to the same bytes

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 02:50:29 +00:00
yujonglee
ae6b762190
refactor(tracing): normalize agent spans in Rust (#44071)
* refactor(tracing): normalize agent spans in Rust

* refactor(tracing): generate dashboard trace types from API

* test(tracing): use complete trace response fixtures

* fix(ui): expose generated span error response type

* fix(tracing): retain full tool call payloads

* fix(tracing): preserve decoded attribute tuple shape

* fix(tracing): type consumed attributes as tuple
2026-10-01 18:01:20 -07:00
moe-berri
cc42a352cb
feat(lens): simplify setup and investigation workflow (#44089)
* feat(lens): simplify investigation setup and results

* fix(lens): pin worker with actionable failure diagnostics

* fix(lens): focus worker success on starting an investigation

* feat(lens): simplify investigation setup and worker defaults

* fix(lens): remove setup repetition and label billing access

* fix(lens): finish agent selection and setup readiness

* fix(lens): handle unavailable setup dependencies and restore UI build
2026-10-01 17:21:41 -07:00
devin-ai-integration[bot]
9351463755
fix(proxy): persist SSO display name as user_alias on login (#44065)
* fix(proxy): persist SSO display name as user_alias on login

Generic/Microsoft SSO already parsed the IdP display_name, first_name and last_name into the SSO result, but the user upsert only wrote user_email and user_role, so the Users table never showed a name. Store the display name (first + last as fallback) as user_alias on first login and on later logins of users whose alias is still empty; never overwrite an alias already set. Whitespace-only names are treated as missing.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep stored user_email when SSO login carries no email claim

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert: keep stored user_email change, login writes the IdP email as before

Reverts 47c65cecff. A stored email staying eligible for email-based account linking after the IdP stops sending it is not wanted; the PR goes back to the user_alias fix only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 17:14:31 -07:00
devin-ai-integration[bot]
54a51c80df
feat(proxy): bounded daily activity routes for all usage entities (#43408)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:57:23 -07:00
devin-ai-integration[bot]
aa601ce4e8
refactor(repositories): daily activity repository with centralized bounded usage queries (#43398)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:02:01 -07:00
devin-ai-integration[bot]
03743ae020
feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup (#43911)
* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): share one optional-feature install path between lazy and eager registration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): say eager lazy routes register at worker startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): name the import callable passed to _install

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prove a startup hook can drop eager lazy routes for good

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): strip the lazy-routes flag from the lazy-mode control proxies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover the lazy warmup route registering a feature

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES

Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart

* fix(proxy): keep config pass-through routes ahead of eagerly registered features

With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:50:41 +00:00
devin-ai-integration[bot]
d7e6154863
fix(mcp): resolve team-granted toolsets for non-admin keys and dashboard sessions (#43908)
* fix(mcp): expand team and dashboard grants when listing and serving toolsets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): honor team toolset grants on the responses gateway path and expose a public team permission lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(mcp): drop the toolset route docstring tweak so the OpenAPI snapshot stays unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): scope inherited toolset grants by the key's own MCP ceiling

A key that declares any MCP grant of its own keeps only its own toolsets, one
that declares none inherits its team's, and require_key_mcp_access_defined
stops a virtual key inheriting while dashboard sessions and admitted users
still do. Adds the direct, no-grant, admin and key-ceiling integration cases
and makes the LLM gateway toolset case discriminate a scoped toolset from the
aggregate grant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): resolve toolset grants per admitted source and enforce the live team roster

Dashboard sessions and gateway-admitted users now expand into the admitted
subject's per-team sources when resolving toolset grants, so a team-granted
toolset is not capped by the user's own MCP row and is reachable on the
namespaced route. A cached team id no longer grants a toolset unless the live
roster still lists the user, a team lookup fault only drops that team's
inherited grant, and /team/member_add evicts the cached team object so the new
member is authoritative immediately

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): read a key's named object permission before letting it inherit team toolsets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): inject the toolset grant resolver into scope helpers so tests stop patching MCPRequestHandler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): drop the duplicate admitted_subject_sources wrapper after merging main and follow its renamed resolvers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): honour fresh policy and the session resource scope on pinned toolsets

A pinned toolset scope now reads the toolset through the writer when the admitted session
requires fresh policy, so a tool revoked from the toolset is gone on the next request. A
gateway bearer scoped to one server can only open a toolset that names that server

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep operator-open servers out of toolset gateway urls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): audit cells for team-granted toolsets across every surface

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): bound toolset edit convergence by both cache layers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): drop toolset integration cells that test behavior this PR does not change

Repeat-read byte identity, 20 concurrent calls, a stopped peer and a killed worker are covered
generically by test_mcp_resilience.py and test_mcp_user_env_vars.py. The toolset edit cell asserts
pre-existing cache propagation and flaked locally with connection resets while polling

* chore(mcp): drop mutable-ok suppressions that main's LIT013 now flags as unused

* chore(mcp): keep the require_key_mcp_access_defined read from adding an unknown-argument type error

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 14:15:06 -07:00
devin-ai-integration[bot]
c66c8288c3
fix(proxy-extras): build the SpendLogs indexes in the migration job instead of in migrations (#43948)
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 14:12:22 -07:00
devin-ai-integration[bot]
a38fff6560
fix(proxy): enforce key/team vector_stores allowlist on /v1/rag/query (#43953)
* add test case for /rag/query and stronger auth check

* style(proxy): ruff format auth_checks.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): build rag query vector store ids immutably and test the no-registry path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): integration coverage for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): audit cells for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): consolidate vector store allowlist audit coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type RAG vector store request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 13:48:41 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
yujonglee
be67fce26a
refactor(proxy): inject tracing receiver and access context (#44035)
* refactor(proxy): inject tracing receiver and access context

* refactor(proxy): own tracing resources through FastAPI lifespan

* test(proxy): pass tracing dependency in Lens lifecycle

* refactor(proxy): stop tracing logger cooperatively

* refactor(proxy): derive tracing permissions in one place

* refactor(proxy): compose application lifespan state

* refactor(proxy): give Lens tracing storage directly

* refactor(tracing): name shared ClickHouse storage explicitly

* refactor(tracing): extract shared ClickHouse storage crate

* test(proxy): isolate db push timeout from Lens safety check

* fix(tracing): drain spend retries during shutdown
2026-10-01 13:45:32 -07:00
tin-berri
163ebccad5
fix(auto-router): show actual and baseline spend for historical savings (#44057)
The usage card hid actual and baseline spend unless every older session
could be rebuilt from SpendLogs within two seconds, which on a real
gateway it never was. Each complexity router's actual spend is now its
rollup spend and its baseline is spend plus recorded savings, for old
and new requests alike, so the benchmarks and session endpoints never
scan SpendLogs. Adaptive and quality routers record no savings baseline
and stay out of the compared totals; savings_estimated_classifier_cost
is kept and covers the same compared requests
2026-10-01 13:44:32 -07:00
tin-berri
eb103334ee
feat(proxy): gzip buffered responses for clients that accept it (#44052)
* feat(proxy): gzip buffered responses for clients that accept it

Large JSON reads like /user/daily/activity/aggregated shipped tens of MB
uncompressed. Compress single-message bodies of 500B or more when the
client's Accept-Encoding allows gzip (q-values and the wildcard honored).
Streamed and etagged responses pass through untouched, every negotiable
response carries Vary: Accept-Encoding, and bodies of 1MB or more are
compressed in a worker thread

* fix(proxy): skip partial and no-transform responses in gzip and always release the held start

The gzip gate now also skips 206 Partial Content and Cache-Control: no-transform,
since compressing either breaks byte ranges or ignores an explicit ban on transforms.
A response start without a headers key no longer raises, and a start the app never
follows with a body message is forwarded when the app returns instead of being dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:43:18 -07:00
devin-ai-integration[bot]
008fcb4fe3
feat(tool-policies): show the user who owns the key that discovered a tool (#43892)
* feat(tool-policies): show the user who owns the key that discovered a tool

GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tool-policies): bound the owner lookup with chunked membership queries

The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tool-policies): cover the owner column across the tool routes and the dashboard

Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 20:34:57 +00:00
yuneng-jiang
bba85f0b6c
chore(lint): remove the LIT002 mutable-construction rule (#43971)
* chore(lint): remove the LIT002 mutable-construction rule

Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.

Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.

* chore(lint): keep the leftover mutable-ok markers for a follow-up

Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.

* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"

This reverts commit c35bc0b84e.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-01 12:24:02 -07:00
ishaan-berri
fcf87972fd
feat(ui): show daily token totals on the model leaderboard (#44044)
* feat(ui): bucket model leaderboard usage by day or week

* test(ui): cover daily buckets in model leaderboard series

* feat(ui): add daily/weekly toggle and per-bucket total to model leaderboard chart

* test(ui): cover the daily/weekly toggle on the model leaderboard

* feat(model-insights): add gateway-wide daily totals to the response type

* feat(model-insights): return per-day totals across every model, not just the top ranked ones

* test(model-insights): daily totals include models outside the top ranking

* chore(model-insights): regenerate lazy openapi snapshot for daily totals

* chore(ui): regenerate api types for model insights daily totals

* fix(ui): compute leaderboard bucket totals from gateway-wide daily totals

* fix(ui): show the gateway total, not the top-ten subtotal, in the leaderboard tooltip

* test(ui): cover gateway-wide bucket totals on the model leaderboard

* fix(model-insights): type daily totals as an immutable tuple

* fix(model-insights): build daily totals without new mutable collections
2026-10-01 12:12:30 -07:00
devin-ai-integration[bot]
2cfa5ec126
test(proxy): delete the legacy proxy test tree and shard tests/unit/proxy by glob (#44018)
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exercise the shard check directly for unit_selection-owned children

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): serve the redirect test from respx instead of a socket

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): credit shard ownership only to unit flags wired in gha

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): implement the wired-flag shard crediting the tests assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): split the root proxy test files into their own unit shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:40:45 -07:00
devin-ai-integration[bot]
a76b59db9f
test(proxy): move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy (#44015)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:23:31 +00:00
devin-ai-integration[bot]
24584d3d3d
test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy (#44012)
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): keep tuple identity in proxy state restore and fix misc target paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:14:24 +00:00