Commit graph

19818 commits

Author SHA1 Message Date
yuneng-jiang
2b85808011
feat(mcp)!: disable stdio MCP servers by default (#44066)
* feat(mcp)!: disable stdio MCP servers by default

stdio MCP servers now only run when the proxy is started with
LITELLM_ENABLE_MCP_STDIO=true. While it is off, existing stdio servers stay
registered but never start: tool listings skip them quietly, direct tool
calls and health checks return a 403 naming the env var, and creating or
updating a stdio server is rejected. The flag is read from the process
environment only, so DB-stored environment_variables cannot turn it on.

The UI reads mcp_stdio_enabled from /.well-known/litellm-ui-config to grey
out the stdio transport, show a banner on stdio forms, and badge stdio
server cards.

BREAKING CHANGE: stdio MCP servers are off by default. Set
LITELLM_ENABLE_MCP_STDIO=true in the proxy environment and restart to keep
using them.

* fix(mcp): ignore stdio flag from config file and read UI flag from the selected worker

LITELLM_ENABLE_MCP_STDIO set under environment_variables in config.yaml is now skipped like the DB-stored value, so only the process environment can enable stdio. The dashboard reads mcp_stdio_enabled from the proxy it is managing, so a control plane shows each worker's own setting.

* test(mcp): cover non-mapping payloads in the shared transport validator

* fix(mcp): skip blocked stdio servers quietly in every listing and keep the UI unchanged until the flag loads

Prompt, resource and resource-template listings now skip a blocked stdio server at debug level like tool listing does, instead of logging a warning per server on every call. The dashboard only treats stdio as disabled once the proxy explicitly reports mcp_stdio_enabled false, so a proxy with the flag on, or an older one without the field, renders exactly as before with no flicker while loading.

* fix(mcp): route blocked stdio tool calls to the flag error and warn once per server

A gateway tools/call naming a blocked stdio server's tool now returns the
LITELLM_ENABLE_MCP_STDIO message instead of "Tool not found".

The "will not start" warning moves out of build_mcp_server_from_table, which
DB reload re-runs on every cycle for rows with a NULL updated_at and which
drafts and test-connection also call. It now fires when a row first enters
the registry or changes transport.

* fix(ui): explain on the server detail page why a stdio server is inert

The Overview and MCP Tools tabs showed "No tools available" with no reason
while stdio is disabled. The detail page now shows the same warning banner
as the edit form, and hands off to the form's banner once editing starts.

* refactor(ui): name the stdio banner conditions on the server detail page

Keeps local/no-long-condition-chain within its budget

* fix(proxy): log the ignored DB-stored LITELLM_ENABLE_MCP_STDIO warning once

The DB config sync re-reads environment_variables on every cycle, so a stored
flag logged the warning on each sync per worker
2026-10-02 10:28:04 -07:00
devin-ai-integration[bot]
6c2ede00ac
test: remove 130 legacy tests owned by stronger unit proofs (#44157)
* test: remove 130 legacy tests owned by stronger unit proofs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: list router _embedding and _aembedding as covered via public embedding calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 10:18:50 -07:00
Jim Aldon D'Souza
19da81579b
fix(ui): register tencent in the Add Model provider dropdown (#40924)
* fix(ui): register tencent in the Add Model provider dropdown

The Add Model provider dropdown is driven by the proxy's
/public/providers/fields endpoint, which serves
provider_create_fields.json. Tencent was frozen in the test's
ADD_MODEL_UNLISTED_PROVIDERS set, so it never appeared in the dropdown.

Add a Tencent entry (optional api_base + required api_key, matching
TENCENT_API_BASE/TENCENT_API_KEY) and unfreeze it in the backend test.
Register Tencent in the UI Providers enum, provider_map, and placeholder
map so the dropdown resolves the display name and model placeholder.

* fix(tencent): drop test docstring to satisfy comment policy

* fix(ui): bundle the Tencent Cloud logo for the provider dropdown
2026-10-02 10:15:41 -07:00
yujonglee
276fc9c63a
fix(tracing): unify ClickHouse storage configuration (#43941)
* fix(tracing): use ClickHouse URL for reads by default

* fix(tracing): unify ClickHouse storage configuration

* fix(tracing): update dashboard setup copy for one URL

* test(tracing): make tests/unit/tracing a package

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(config): drop legacy string tracing store variant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): own ClickHouse defaults in constants and reject unset env references

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): use raw regex patterns in config tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): read ClickHouse env defaults when tracing config resolves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): split audit log query guard to fit condition-chain budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 16:31:26 +00:00
devin-ai-integration[bot]
615ed7900f
test(mcp): fire MCP client test deadlines on conditions instead of wall-clock time (#44166)
* test(mcp): fire MCP client test deadlines on conditions instead of wall-clock time

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): bound MCP client test failure paths independently of the triggered deadline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): measure outer-deadline cleanup budget from cancellation, not setup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 05:44:21 -07:00
devin-ai-integration[bot]
fe9b6fd603
refactor(types): replace Any with proven types in 7 files (#43844)
* refactor(types): replace Any with proven types in 9 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): pin xecguard-adjacent guardrail retry, mcp mixed tools, jwt routing and complexity router paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep only live-provable Any removals

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): type cache_hit as bool | None on the sync success path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 02:47:47 -07:00
devin-ai-integration[bot]
7be2983f11
test(straiker): assert a saved api_version v1 with an sk_agt_ key routes to v3 (#44153)
* test(straiker): assert a saved api_version v1 with an sk_agt_ key routes to v3 (#44106)

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(straiker): ignore zombie workers when counting uvicorn children after a kill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(straiker): describe the saved v1 cell by the v3 route it asserts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-02 01:31:02 -07:00
yuneng-jiang
7d50a31eb5
test(e2e): move live-provider legacy tests into tests/e2e (#44120)
* test(e2e): move live-provider legacy tests into tests/e2e

Port legacy tests that exercise real providers into the tests/e2e suites that own them, using the harness (/model/new plus deferred cleanup) and asserting on what the caller receives. Delete legacy tests already covered at equal or stronger strength by e2e, integration or unit tests, and drop the now empty ocr_testing CircleCI job

* test(e2e): address review on the live-provider test move

Assert the SSE error frame a client actually receives when a post_call guardrail blocks a stream, and require a tool call for every requested city before checking the answer. Restore the OCR matrix and its CircleCI job, the Claude Agent SDK streaming test, and test_async_create_batch, since their SDK-level and callback assertions have no equivalent in tests/e2e

* test(e2e): accept both guardrail block shapes on a blocked stream

A post_call block before the first chunk reaches the client as HTTP 400 with either a JSON error body or a single SSE error frame, depending on whether the block surfaced as an exception or an error chunk. Assert the policy message is present and the blocked output is absent in both

* test(realtime): restore direct SDK realtime tests against OpenAI

The e2e realtime tests go through the proxy and the remaining SDK tests either mock the upstream or assert less, so keep the direct litellm._arealtime tests with and without intent, and TestOpenAIRealtime::test_realtime_connection, in place

* test: make realtime and Nova stream checks deterministic

The direct SDK realtime tests now fail on a refused connection instead of skipping. The with-intent test asserts OpenAI rejects the exact intent value sent, which only happens when the intent is forwarded. The Nova /v1/messages stream test asserts stream structure, stop reason and usage instead of model wording

* test(realtime): own intent forwarding with a unit test instead of a live rejection

Assert litellm._arealtime passes the intent query param into the OpenAI realtime websocket URL, which is the behavior LiteLLM owns, and drop the live test that depended on OpenAI's rejection wording
2026-10-02 00:02:18 -07:00
devin-ai-integration[bot]
e32b25817f
fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs (#44075)
* fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): audit cells for stored spend-log tool payloads and logprobs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert cache-hit spend-log rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): handle model-list requests in spend-log audit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert unauthenticated requests never reach upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:57:11 -07:00
devin-ai-integration[bot]
d131c43782
fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging (#44067)
* fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): use transport-level doubles in Azure dispatch regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add Azure dispatch audit matrix integration cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): restore global callback lists after Router reset in SDK cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): make proxy restart cell independent of shutdown timing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:30:32 -07:00
devin-ai-integration[bot]
5a3a31ea9a
fix(azure_storage): keep client call ids from sharing one Data Lake file (#44099)
* fix(azure_storage): keep client call ids from sharing one Data Lake file

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(azure_storage): tighten adls_safe_file_name docstring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure_storage): give ids with dot or empty path segments their own Data Lake file

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): cover failure, cache-hit and edge call ids in Data Lake file names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): run the Data Lake file name cells on two proxy workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:23:37 -07:00
yuneng-jiang
a93c396a5c
test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128)
* test(integration): move legacy proxy, router and Redis tests into tests/integration

Port 39 legacy tests to the integration tier that owns them, running against
the scripted upstream, local Postgres and Redis, test-owned wire peers and
owned proxies. Delete 5 legacy tests whose contract is already owned by an
existing integration test, and remove the legacy functions, files and helpers
left unused.

* test(integration): cover recovery of a spent key after its budget is raised
2026-10-01 23:00:54 -07:00
PhimmStraiker
826b21aab6
fix(guardrails): straiker v3 routes sk_agt_ keys to v3 and fails closed on a missing verdict (#44011)
* fix(guardrails): straiker v3 routes sk_agt_ keys to v3 and stops reading a missing verdict as allow

An sk_agt_ key always calls /api/v3/detect, even when the guardrail was saved with
api_version 'v1' by the old shared default: the v1 webhook rejects that key with 401.
A 200 with no decision, or permissionDecision 'ask', now takes the failure policy
instead of allowing the request. A detect-mode action is still not a block.
A response-phase block is no longer remembered under the request, so asking the same
question again is scored instead of refused from memory. The replay memory is scoped
by principal and session together, so two principals on one session id never share a
block. A text-only /guardrails/apply_guardrail call relays the text as a user turn.
The agent_ref description now matches the code: the configured value wins.

* fix(guardrails): straiker v3 relays text beside an empty messages list and keys the memory on the key

/guardrails/apply_guardrail sends `messages: []` beside `text`; an empty list is no
conversation, so the text is relayed as the user turn. A verdict whose blocked_by is not
a list states no decision and takes the failure policy, the same way a missing decision
does. A key that names no user is still the caller, so the replay memory is keyed on the
key when no user is known.

* test(guardrails): build the straiker replay-scope request data without mutating it

---------

Co-authored-by: PhimmStraiker <PhimmStraiker@users.noreply.github.com>
2026-10-01 22:38:04 -07:00
yujonglee
8d28e8d776
feat(tracing): add scoped SQL queries and schema-aware help (#44085)
* feat(tracing): add SQL queries and schema-aware query help

* test(tracing): verify help requests and sync API types

* refactor(tracing): render query help with Askama

* refactor(tracing): use jinja extension for query guide

* fix(tracing): preserve query help when discovery fails

* feat(tracing): enforce team SQL scope with managed ClickHouse readers

* test(tracing): verify reads with one ClickHouse URL

* fix(tracing): revoke rotated trace reader credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): streamline query help catalog assembly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tracing): run query help discovery sequentially

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tracing): update reader setup request expectations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:06:14 -07:00
devin-ai-integration[bot]
e6639d15b5
fix(proxy): always exit when database setup fails at boot (#44141)
Remove the ENFORCE_PRISMA_MIGRATION_CHECK opt-in. When PrismaManager.setup_database returns False (database unreachable, connection retries exhausted, or prisma migrate deploy failing after retries) the proxy now always prints the red failure message and exits 1 instead of serving requests against a database whose schema may be behind the code. The --enforce_prisma_migration_check flag stays as a hidden no-op that prints a one-line deprecation warning so existing container args keep parsing; the env var is no longer read anywhere. The standalone migration entrypoint always runs run_server(("--skip_server_startup",)), and the integration launchers drop the flag.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:51:06 -07:00
devin-ai-integration[bot]
5ccb1a143b
fix(proxy): reject non-canonical daily activity dates (#44143)
The daily spend tables store date as text, so a request date that strptime accepts but that is not spelled YYYY-MM-DD (2026-9-24, 2026-09-4, full-width digits) was compared as raw text against canonical rows and matched nothing, and the export route copied it into Content-Disposition, which fails latin-1 encoding and returned 500. A shared parse_canonical_date_range now rejects any spelling whose round trip differs from the input, so every bounded daily activity route (user, team, tag, organization, customer, agent: aggregated, aggregated/keys, search, model top keys, export, cache leakage) and the paginated get_daily_activity path answer 400 before touching the repository, and the export filename is built from the validated dates.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:19:51 -07:00
devin-ai-integration[bot]
c208306f60
fix(daily_activity): keep NULL entity ids when excluding entity ids (#44139)
The shared exclusion predicate negated a PostgreSQL ANY comparison without handling NULL, so NULL entity ids evaluated to UNKNOWN and dropped out of the Unassigned bucket whenever exclude_*_ids was set. The paginated daily rows Prisma filter had the same NOT IN shape and gets the same IS NULL OR NOT IN treatment

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:19:48 -07:00
devin-ai-integration[bot]
b9e71e990a
feat(mcp): add Microsoft 365 (Graph) server to the MCP catalog (#43099)
* feat(mcp): add Microsoft 365 (Graph) server to the MCP catalog

* fix(mcp): signpost the self-hosted Microsoft 365 URL and pin the catalog entry in tests

* test(mcp): pin both shipped copies of a catalog icon to the same bytes

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 02:50:29 +00:00
devin-ai-integration[bot]
123237c8a8
fix(proxy-extras): bound the lock waits of the partitioned SpendLogs index build (#44109)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:28:50 -07:00
yujonglee
ae6b762190
refactor(tracing): normalize agent spans in Rust (#44071)
* refactor(tracing): normalize agent spans in Rust

* refactor(tracing): generate dashboard trace types from API

* test(tracing): use complete trace response fixtures

* fix(ui): expose generated span error response type

* fix(tracing): retain full tool call payloads

* fix(tracing): preserve decoded attribute tuple shape

* fix(tracing): type consumed attributes as tuple
2026-10-01 18:01:20 -07:00
moe-berri
cc42a352cb
feat(lens): simplify setup and investigation workflow (#44089)
* feat(lens): simplify investigation setup and results

* fix(lens): pin worker with actionable failure diagnostics

* fix(lens): focus worker success on starting an investigation

* feat(lens): simplify investigation setup and worker defaults

* fix(lens): remove setup repetition and label billing access

* fix(lens): finish agent selection and setup readiness

* fix(lens): handle unavailable setup dependencies and restore UI build
2026-10-01 17:21:41 -07:00
devin-ai-integration[bot]
9351463755
fix(proxy): persist SSO display name as user_alias on login (#44065)
* fix(proxy): persist SSO display name as user_alias on login

Generic/Microsoft SSO already parsed the IdP display_name, first_name and last_name into the SSO result, but the user upsert only wrote user_email and user_role, so the Users table never showed a name. Store the display name (first + last as fallback) as user_alias on first login and on later logins of users whose alias is still empty; never overwrite an alias already set. Whitespace-only names are treated as missing.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep stored user_email when SSO login carries no email claim

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert: keep stored user_email change, login writes the IdP email as before

Reverts 47c65cecff. A stored email staying eligible for email-based account linking after the IdP stops sending it is not wanted; the PR goes back to the user_alias fix only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 17:14:31 -07:00
devin-ai-integration[bot]
54a51c80df
feat(proxy): bounded daily activity routes for all usage entities (#43408)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:57:23 -07:00
shrey-berri
339f5b6a4d
fix(bedrock): add beta for mid-conversation tool changes (#43833) 2026-10-01 16:42:46 -07:00
devin-ai-integration[bot]
aa601ce4e8
refactor(repositories): daily activity repository with centralized bounded usage queries (#43398)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:02:01 -07:00
yuneng-jiang
4b1d9bf148
test(e2e): keep 1ms-timeout deployments off the provider cache (#44082)
The timeout reliability tests rely on a 1ms deadline the real backend always
misses. With E2E_PROVIDER_CACHE on, the deployment pointed at the cache edge,
and its healthy sibling in the same test had already recorded a response for
the same canonical request, so the edge answered from Redis inside the 1ms read
window. Build 342 of litellm-e2e saw test_timeout_trips_cooldown_then_recovers
get a 200 from the timing-out deployment itself, with a recording made about
12 hours earlier. Both timeout helpers now register on the live provider path,
which PROVIDER_CACHE.md reserves for tests that need real provider timing
2026-10-01 15:54:02 -07:00
devin-ai-integration[bot]
03743ae020
feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup (#43911)
* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): share one optional-feature install path between lazy and eager registration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): say eager lazy routes register at worker startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): name the import callable passed to _install

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prove a startup hook can drop eager lazy routes for good

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): strip the lazy-routes flag from the lazy-mode control proxies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover the lazy warmup route registering a feature

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES

Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart

* fix(proxy): keep config pass-through routes ahead of eagerly registered features

With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 22:50:41 +00:00
ishaan-berri
0f6a06a6b9
feat: add litellm.agent() to run claude code, codex, opencode and deep agents through the ai gateway (#43885)
* feat(harness): add litellm/__init__.py

* feat(harness): add litellm/constants.py

* feat(harness): add litellm/harness/__init__.py

* feat(harness): add litellm/harness/adapters/__init__.py

* feat(harness): add litellm/harness/adapters/base.py

* feat(harness): add litellm/harness/adapters/claude_code.py

* feat(harness): add litellm/harness/adapters/codex.py

* feat(harness): add litellm/harness/adapters/opencode.py

* feat(harness): add litellm/harness/endpoint.py

* feat(harness): add litellm/harness/errors.py

* feat(harness): add litellm/harness/options.py

* feat(harness): add litellm/harness/runtime.py

* feat(harness): add litellm/harness/sandbox/__init__.py

* feat(harness): add litellm/harness/sandbox/base.py

* feat(harness): add litellm/harness/sandbox/docker.py

* feat(harness): add litellm/harness/sandbox/local.py

* feat(harness): add litellm/harness/sandbox/snapshot.py

* feat(harness): add litellm/harness/sync.py

* feat(harness): add litellm/harness/types.py

* feat(harness): add litellm/sandbox/__init__.py

* feat(harness): add README.md

* feat(harness): add tests/harness_e2e/__init__.py

* feat(harness): add tests/harness_e2e/conftest.py

* feat(harness): add tests/harness_e2e/test_harness_e2e.py

* feat(harness): add tests/test_litellm/harness/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/test_claude_code.py

* feat(harness): add tests/test_litellm/harness/adapters/test_codex.py

* feat(harness): add tests/test_litellm/harness/adapters/test_opencode.py

* feat(harness): add tests/test_litellm/harness/core_fakes.py

* feat(harness): add tests/test_litellm/harness/sandbox/__init__.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_docker.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_local.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_snapshot.py

* feat(harness): add tests/test_litellm/harness/test_endpoint.py

* feat(harness): add tests/test_litellm/harness/test_init.py

* feat(harness): add tests/test_litellm/harness/test_runtime.py

* feat(harness): add tests/test_litellm/harness/test_sync.py

* feat(harness): add tests/test_litellm/harness/test_types.py

* test(harness): use word recall in stream e2e test

* refactor(harness): update litellm/__init__.py

* refactor(harness): update litellm/constants.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): remove litellm/harness/adapters/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/__init__.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/__init__.py

* refactor(harness): update litellm/llms/codex/harness/__init__.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/utils.py

* refactor(harness): update README.md

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): remove tests/test_litellm/harness/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/test_claude_code.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_codex.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_opencode.py

* refactor(harness): remove tests/test_litellm/harness/core_fakes.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/__init__.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_docker.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_local.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_snapshot.py

* refactor(harness): remove tests/test_litellm/harness/test_endpoint.py

* refactor(harness): remove tests/test_litellm/harness/test_init.py

* refactor(harness): remove tests/test_litellm/harness/test_runtime.py

* refactor(harness): remove tests/test_litellm/harness/test_sync.py

* refactor(harness): remove tests/test_litellm/harness/test_types.py

* refactor(harness): update tests/unit/harness/__init__.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/__init__.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/sandbox/__init__.py

* refactor(harness): update tests/unit/harness/sandbox/test_docker.py

* refactor(harness): update tests/unit/harness/sandbox/test_local.py

* refactor(harness): update tests/unit/harness/sandbox/test_snapshot.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/harness/test_types.py

* refactor(harness): update tests/unit/llms/claude_code/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/max_turns.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/resume_turn.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/success_tools.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/codex/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/reasoning.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn1_bash.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn2_resume_apply_patch.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn_failed.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/endpoint_requests.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/readonly_denied_bash.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn1_write_read.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn2_session_skill.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* ci: allowlist tests/harness_e2e, which needs live runtimes and a gateway

* fix(harness): bound the turn event queue

* fix(harness): bound the turn event queue with backpressure

* style: sort imports in utils

* ci: exclude agent-harness config folders from provider docs check

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/options.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/sandbox/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update tests/code_coverage_tests/recursive_detector.py

* refactor(harness): update tests/unit/llms/base_llm/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/__init__.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/__init__.py

* refactor(harness): update litellm/harness/sandbox/base.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/types/llms/custom_http.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py
2026-10-01 22:27:49 +00:00
devin-ai-integration[bot]
610576d247
fix(bedrock): accept Converse messages with no content key (#43936)
* fix(bedrock): accept Converse messages with no content key

A user or tool message whose content key is missing (or null, which the
message cleanup strips) made every Bedrock Converse request fail with
APIConnectionError 'content' before reaching Bedrock. The Converse
transform now reads content with .get for those messages, as it already
did for assistant messages: a content-less user message adds no block
and a content-less tool message becomes a toolResult with empty content.
The str branch also sends the continue message text instead of the
original whitespace-only text.

* fix(bedrock): send the continue message for a content-less user turn

* test(bedrock): type the content-less Converse message test parameters

* fix(bedrock): accept a Converse system message with no content key

* refactor(bedrock): read the system message content with get

* test(integration): audit Bedrock Converse messages without content

Adds the /audit cells for a chat message whose content key is missing or
null on a Converse-routed Bedrock deployment: happy, sad, edge, and chaos
rows through the OpenAI SDK, the Anthropic SDK, and raw httpx against the
scripted upstream, asserting the caller's response, the body the peer
received, and the spend row. The owned-proxy readiness deadline in the
integration harness is now INTEGRATION_PROXY_READY_SECONDS (default 70).

* test(integration): bound stray spend rows in the mid-burst restart cell

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-01 15:14:34 -07:00
devin-ai-integration[bot]
d7e6154863
fix(mcp): resolve team-granted toolsets for non-admin keys and dashboard sessions (#43908)
* fix(mcp): expand team and dashboard grants when listing and serving toolsets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): honor team toolset grants on the responses gateway path and expose a public team permission lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(mcp): drop the toolset route docstring tweak so the OpenAPI snapshot stays unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): scope inherited toolset grants by the key's own MCP ceiling

A key that declares any MCP grant of its own keeps only its own toolsets, one
that declares none inherits its team's, and require_key_mcp_access_defined
stops a virtual key inheriting while dashboard sessions and admitted users
still do. Adds the direct, no-grant, admin and key-ceiling integration cases
and makes the LLM gateway toolset case discriminate a scoped toolset from the
aggregate grant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): resolve toolset grants per admitted source and enforce the live team roster

Dashboard sessions and gateway-admitted users now expand into the admitted
subject's per-team sources when resolving toolset grants, so a team-granted
toolset is not capped by the user's own MCP row and is reachable on the
namespaced route. A cached team id no longer grants a toolset unless the live
roster still lists the user, a team lookup fault only drops that team's
inherited grant, and /team/member_add evicts the cached team object so the new
member is authoritative immediately

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): read a key's named object permission before letting it inherit team toolsets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): inject the toolset grant resolver into scope helpers so tests stop patching MCPRequestHandler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): drop the duplicate admitted_subject_sources wrapper after merging main and follow its renamed resolvers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): honour fresh policy and the session resource scope on pinned toolsets

A pinned toolset scope now reads the toolset through the writer when the admitted session
requires fresh policy, so a tool revoked from the toolset is gone on the next request. A
gateway bearer scoped to one server can only open a toolset that names that server

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep operator-open servers out of toolset gateway urls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): audit cells for team-granted toolsets across every surface

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): bound toolset edit convergence by both cache layers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): drop toolset integration cells that test behavior this PR does not change

Repeat-read byte identity, 20 concurrent calls, a stopped peer and a killed worker are covered
generically by test_mcp_resilience.py and test_mcp_user_env_vars.py. The toolset edit cell asserts
pre-existing cache propagation and flaked locally with connection resets while polling

* chore(mcp): drop mutable-ok suppressions that main's LIT013 now flags as unused

* chore(mcp): keep the require_key_mcp_access_defined read from adding an unknown-argument type error

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 14:15:06 -07:00
devin-ai-integration[bot]
c66c8288c3
fix(proxy-extras): build the SpendLogs indexes in the migration job instead of in migrations (#43948)
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 14:12:22 -07:00
devin-ai-integration[bot]
a38fff6560
fix(proxy): enforce key/team vector_stores allowlist on /v1/rag/query (#43953)
* add test case for /rag/query and stronger auth check

* style(proxy): ruff format auth_checks.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): build rag query vector store ids immutably and test the no-registry path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): integration coverage for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): audit cells for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): consolidate vector store allowlist audit coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type RAG vector store request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 13:48:41 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
yujonglee
be67fce26a
refactor(proxy): inject tracing receiver and access context (#44035)
* refactor(proxy): inject tracing receiver and access context

* refactor(proxy): own tracing resources through FastAPI lifespan

* test(proxy): pass tracing dependency in Lens lifecycle

* refactor(proxy): stop tracing logger cooperatively

* refactor(proxy): derive tracing permissions in one place

* refactor(proxy): compose application lifespan state

* refactor(proxy): give Lens tracing storage directly

* refactor(tracing): name shared ClickHouse storage explicitly

* refactor(tracing): extract shared ClickHouse storage crate

* test(proxy): isolate db push timeout from Lens safety check

* fix(tracing): drain spend retries during shutdown
2026-10-01 13:45:32 -07:00
tin-berri
163ebccad5
fix(auto-router): show actual and baseline spend for historical savings (#44057)
The usage card hid actual and baseline spend unless every older session
could be rebuilt from SpendLogs within two seconds, which on a real
gateway it never was. Each complexity router's actual spend is now its
rollup spend and its baseline is spend plus recorded savings, for old
and new requests alike, so the benchmarks and session endpoints never
scan SpendLogs. Adaptive and quality routers record no savings baseline
and stay out of the compared totals; savings_estimated_classifier_cost
is kept and covers the same compared requests
2026-10-01 13:44:32 -07:00
tin-berri
eb103334ee
feat(proxy): gzip buffered responses for clients that accept it (#44052)
* feat(proxy): gzip buffered responses for clients that accept it

Large JSON reads like /user/daily/activity/aggregated shipped tens of MB
uncompressed. Compress single-message bodies of 500B or more when the
client's Accept-Encoding allows gzip (q-values and the wildcard honored).
Streamed and etagged responses pass through untouched, every negotiable
response carries Vary: Accept-Encoding, and bodies of 1MB or more are
compressed in a worker thread

* fix(proxy): skip partial and no-transform responses in gzip and always release the held start

The gzip gate now also skips 206 Partial Content and Cache-Control: no-transform,
since compressing either breaks byte ranges or ignores an explicit ban on transforms.
A response start without a headers key no longer raises, and a start the app never
follows with a body message is forwarded when the app returns instead of being dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:43:18 -07:00
devin-ai-integration[bot]
008fcb4fe3
feat(tool-policies): show the user who owns the key that discovered a tool (#43892)
* feat(tool-policies): show the user who owns the key that discovered a tool

GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tool-policies): bound the owner lookup with chunked membership queries

The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tool-policies): cover the owner column across the tool routes and the dashboard

Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 20:34:57 +00:00
devin-ai-integration[bot]
ba8cd1ee31
fix(router): keep silent_model out of embedding provider requests (#44064)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 13:06:29 -07:00
yuneng-jiang
65a316fc92
ci(circleci): test Redis behavior against local Redis and print short tracebacks (#44062)
* ci(circleci): test Redis behavior against local Redis and print short tracebacks

The redis_caching_unit_tests job ran three legacy files against the shared remote Redis. The DualCache and
batch-read logic that never needed a server now lives in tests/unit/caching/test_dual_cache.py with a mocked
RedisCache, and the behavior that does need one (the increment-with-floor Lua script, read-through, deletes,
batch reads) moved to tests/integration, which starts a local Redis. test_returned_settings only read
REDIS_PORT and is replaced by a unit test of Router.get_settings

CircleCI pytest runs now use --tb=short so failure output stays readable in the test results tab

* ci(integration): print short tracebacks from run.py and allow Redis in the sdk shard
2026-10-01 13:00:54 -07:00
yuneng-jiang
0da00d4b2e
test(e2e): bill Sail windows that synchronous calls can still use (#44058)
Sail now rejects completion_window "flex" on synchronous requests with a
400 saying flex is only for background responses or Batch work. The chat
flex case and the responses flex case have failed on every scheduled
litellm-e2e run in builds 337, 340 and 341. The chat cases keep balanced and
auto, and the responses case sends a caller metadata.completion_window of
balanced, so both still prove the window reaches Sail and the bill uses
that window's distinct rates
2026-10-01 12:54:26 -07:00
yuneng-jiang
bba85f0b6c
chore(lint): remove the LIT002 mutable-construction rule (#43971)
* chore(lint): remove the LIT002 mutable-construction rule

Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.

Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.

* chore(lint): keep the leftover mutable-ok markers for a follow-up

Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.

* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"

This reverts commit c35bc0b84e.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-01 12:24:02 -07:00
ishaan-berri
fcf87972fd
feat(ui): show daily token totals on the model leaderboard (#44044)
* feat(ui): bucket model leaderboard usage by day or week

* test(ui): cover daily buckets in model leaderboard series

* feat(ui): add daily/weekly toggle and per-bucket total to model leaderboard chart

* test(ui): cover the daily/weekly toggle on the model leaderboard

* feat(model-insights): add gateway-wide daily totals to the response type

* feat(model-insights): return per-day totals across every model, not just the top ranked ones

* test(model-insights): daily totals include models outside the top ranking

* chore(model-insights): regenerate lazy openapi snapshot for daily totals

* chore(ui): regenerate api types for model insights daily totals

* fix(ui): compute leaderboard bucket totals from gateway-wide daily totals

* fix(ui): show the gateway total, not the top-ten subtotal, in the leaderboard tooltip

* test(ui): cover gateway-wide bucket totals on the model leaderboard

* fix(model-insights): type daily totals as an immutable tuple

* fix(model-insights): build daily totals without new mutable collections
2026-10-01 12:12:30 -07:00
yuneng-jiang
e725832cd8
chore(deps): drop unused pytest-postgresql dev dependency (#44056)
The pytest-postgresql based proxy tests moved to tests/integration on the real
Postgres harness in #43996, so nothing loads the plugin anymore. uv.lock is
edited by hand to drop the package and its now orphaned mirakuru and port-for
deps; uv lock --check passes and a full relock resolves the same package set
2026-10-01 18:53:54 +00:00
devin-ai-integration[bot]
2cfa5ec126
test(proxy): delete the legacy proxy test tree and shard tests/unit/proxy by glob (#44018)
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exercise the shard check directly for unit_selection-owned children

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): serve the redirect test from respx instead of a socket

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): credit shard ownership only to unit flags wired in gha

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): implement the wired-flag shard crediting the tests assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): split the root proxy test files into their own unit shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:40:45 -07:00
devin-ai-integration[bot]
a76b59db9f
test(proxy): move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy (#44015)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:23:31 +00:00
devin-ai-integration[bot]
24584d3d3d
test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy (#44012)
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): keep tuple identity in proxy state restore and fix misc target paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:14:24 +00:00
devin-ai-integration[bot]
25109a523b
test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy (#44006)
* test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved unit test directories

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exclude proxy-db-owned files from the misc target

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): drop the redundant fixture docstrings in the proxy conftest

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:06:42 -07:00
moe-berri
259c166ef6
refactor(lens)!: rename internal engine code and API (#44034)
* chore(lens): remove deployment screenshots

* refactor(lens)!: rename internal engine package and API

* fix(lens): pin worker image for renamed API

* test(lens): cover fresh and populated rename migrations

* fix(lens): protect db-push upgrades and restore routing and CI

* fix(lens): resolve migration tables across schemas and include database driver
2026-10-01 11:02:14 -07:00
devin-ai-integration[bot]
bfd3f39dca
feat(s3_v2): add s3_partition_granularity option for hourly S3 folders (#43748)
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack

The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the sink outage burst models in config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(s3_v2): read cold storage metadata without an empty dict default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the proxy to reconnect before the postgres outage recovery request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-01 11:00:21 -07:00
ishaan-berri
d24240c014
feat(ui): agent traces open in a side drawer with a chat-style run view (#43972)
* feat(ui): redesign agent trace run view with chat-style detail pane

Tree with connector lines, typed icon tiles and provider logos, hover
cards with timing, and Input/Output sections rendered as message cards.

* fix(tracing): show text for block-list message content and split normalizers per convention

OpenAI responses-style content (reasoning + text blocks) rendered as raw
JSON in the trace view. Keep the text blocks and drop opaque reasoning.

Move each convention into litellm/tracing/normalizers with an ordered
registry so new frameworks plug in without touching OTLP decoding.

* fix(tracing): keep long message histories as valid JSON and parse function_call blocks

* feat(ui): open agent traces in a resizable side drawer with a devtool-style tree

Clicking a run opens it in a drawer over the list instead of a full page.
j/k and the header arrows switch runs, Esc closes. The tree gets dashed
connectors, per-span waterfall bars, mono tool names and real provider
logos. AI messages with reasoning/function_call blocks render as text.

* fix(trace-ui): address review: valid JSON trimming, drawer keys, reduced motion, narrow screens

* feat(tracing): serve span content in a standard LiteLLM UI format

GET /v1/traces/{trace_id}/spans/{span_id} now also returns input_ui and
output_ui, a tagged union of messages, fields or text built server side by
litellm/tracing/ui_format.py. The trace UI renders from those fields and only
falls back to client-side parsing when talking to an older proxy. The raw
input and output strings are unchanged, and so is storage

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tracing): fall back to an elision marker when shortened messages still exceed the size limit

* fix(tracing): keep both messages when tool_calls are oversized and keep failed-tool styling

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 17:55:51 +00:00