* fix(proxy): persist SSO display name as user_alias on login
Generic/Microsoft SSO already parsed the IdP display_name, first_name and last_name into the SSO result, but the user upsert only wrote user_email and user_role, so the Users table never showed a name. Store the display name (first + last as fallback) as user_alias on first login and on later logins of users whose alias is still empty; never overwrite an alias already set. Whitespace-only names are treated as missing.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep stored user_email when SSO login carries no email claim
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert: keep stored user_email change, login writes the IdP email as before
Reverts 47c65cecff. A stored email staying eligible for email-based account linking after the IdP stops sending it is not wanted; the PR goes back to the user_alias fix only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
uv sync --frozen ignores --no-sources-package, so the "published" branch
already installed litellm-proxy-extras from the workspace (the shipped
main-stable non-root image records file:///app/litellm-proxy-extras in its
direct_url.json). Both branches ran the same install. Keep the single
uv sync that the main and database images use.
Resolves LIT-8899
Usage, cost optimization, user and team pages read the aggregate, paginated key, search, model top key, cache leakage and export routes added by the lower layers instead of downloading every key's daily rows into the browser. Key detail, model top key and search failures render explicit errors with retry controls, Overall Usage shows a loader while the aggregate is in flight, Retry on a failed first key page shows the loading state while it refetches, a short query keeps the loader or first page error visible instead of No keys match, key paging advances by the server offset so a page of already loaded keys moves on and only an empty page ends paging, and search results are stored as rows so a new teams array from the parent does not restart an in-flight search, and the global Cost tab shows a loader instead of zero totals while the aggregate reloads.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The timeout reliability tests rely on a 1ms deadline the real backend always
misses. With E2E_PROVIDER_CACHE on, the deployment pointed at the cache edge,
and its healthy sibling in the same test had already recorded a response for
the same canonical request, so the edge answered from Redis inside the 1ms read
window. Build 342 of litellm-e2e saw test_timeout_trips_cooldown_then_recovers
get a 200 from the timing-out deployment itself, with a recording made about
12 hours earlier. Both timeout helpers now register on the live provider path,
which PROVIDER_CACHE.md reserves for tests that need real provider timing
* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): share one optional-feature install path between lazy and eager registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(proxy): say eager lazy routes register at worker startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): name the import callable passed to _install
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove a startup hook can drop eager lazy routes for good
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): strip the lazy-routes flag from the lazy-mode control proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the lazy warmup route registering a feature
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES
Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart
* fix(proxy): keep config pass-through routes ahead of eagerly registered features
With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): add a call that posts an OTLP export to the proxy
* feat(ui): build a sample agent run as an OTLP export
* feat(ui): add a pulsing dot for active tracing
* feat(ui): render a crisp sample run preview from real trace components
* chore(ui): remove the pixelated agent traces preview image
* feat(ui): add test trace, tracing key, otel endpoints and more frameworks to tracing setup
* test(ui): cover tracing setup test trace, key masking, endpoints and frameworks
* feat(ui): open the received test trace and mark tracing as active
* test(ui): mock the new tracing setup network calls
* fix(ui): let tracing setup use the full page width
* fix(ui): send OTLP/JSON sample trace ids as hex per the spec
* test(ui): cover the sample trace export shape and hex ids
* fix(bedrock): accept Converse messages with no content key
A user or tool message whose content key is missing (or null, which the
message cleanup strips) made every Bedrock Converse request fail with
APIConnectionError 'content' before reaching Bedrock. The Converse
transform now reads content with .get for those messages, as it already
did for assistant messages: a content-less user message adds no block
and a content-less tool message becomes a toolResult with empty content.
The str branch also sends the continue message text instead of the
original whitespace-only text.
* fix(bedrock): send the continue message for a content-less user turn
* test(bedrock): type the content-less Converse message test parameters
* fix(bedrock): accept a Converse system message with no content key
* refactor(bedrock): read the system message content with get
* test(integration): audit Bedrock Converse messages without content
Adds the /audit cells for a chat message whose content key is missing or
null on a Converse-routed Bedrock deployment: happy, sad, edge, and chaos
rows through the OpenAI SDK, the Anthropic SDK, and raw httpx against the
scripted upstream, asserting the caller's response, the body the peer
received, and the spend row. The owned-proxy readiness deadline in the
integration harness is now INTEGRATION_PROXY_READY_SECONDS (default 70).
* test(integration): bound stray spend rows in the mid-burst restart cell
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(mcp): expand team and dashboard grants when listing and serving toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honor team toolset grants on the responses gateway path and expose a public team permission lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(mcp): drop the toolset route docstring tweak so the OpenAPI snapshot stays unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): scope inherited toolset grants by the key's own MCP ceiling
A key that declares any MCP grant of its own keeps only its own toolsets, one
that declares none inherits its team's, and require_key_mcp_access_defined
stops a virtual key inheriting while dashboard sessions and admitted users
still do. Adds the direct, no-grant, admin and key-ceiling integration cases
and makes the LLM gateway toolset case discriminate a scoped toolset from the
aggregate grant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): resolve toolset grants per admitted source and enforce the live team roster
Dashboard sessions and gateway-admitted users now expand into the admitted
subject's per-team sources when resolving toolset grants, so a team-granted
toolset is not capped by the user's own MCP row and is reachable on the
namespaced route. A cached team id no longer grants a toolset unless the live
roster still lists the user, a team lookup fault only drops that team's
inherited grant, and /team/member_add evicts the cached team object so the new
member is authoritative immediately
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): read a key's named object permission before letting it inherit team toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): inject the toolset grant resolver into scope helpers so tests stop patching MCPRequestHandler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): drop the duplicate admitted_subject_sources wrapper after merging main and follow its renamed resolvers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honour fresh policy and the session resource scope on pinned toolsets
A pinned toolset scope now reads the toolset through the writer when the admitted session
requires fresh policy, so a tool revoked from the toolset is gone on the next request. A
gateway bearer scoped to one server can only open a toolset that names that server
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep operator-open servers out of toolset gateway urls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): audit cells for team-granted toolsets across every surface
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): bound toolset edit convergence by both cache layers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): drop toolset integration cells that test behavior this PR does not change
Repeat-read byte identity, 20 concurrent calls, a stopped peer and a killed worker are covered
generically by test_mcp_resilience.py and test_mcp_user_env_vars.py. The toolset edit cell asserts
pre-existing cache propagation and flaked locally with connection resets while polling
* chore(mcp): drop mutable-ok suppressions that main's LIT013 now flags as unused
* chore(mcp): keep the require_key_mcp_access_defined read from adding an unknown-argument type error
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Lens rename check added in #44034 refuses prisma db push when it cannot
reach the database. This test pointed DATABASE_URL at a dead port to keep the
real CI database out of the run, so it now trips that check before the push
timeout it exists to cover. Unsetting DATABASE_URL skips every database probe
and keeps the test focused on the timeout hint.
* add test case for /rag/query and stronger auth check
* style(proxy): ruff format auth_checks.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): build rag query vector store ids immutably and test the no-registry path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): integration coverage for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): audit cells for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): consolidate vector store allowlist audit coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type RAG vector store request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The usage card hid actual and baseline spend unless every older session
could be rebuilt from SpendLogs within two seconds, which on a real
gateway it never was. Each complexity router's actual spend is now its
rollup spend and its baseline is spend plus recorded savings, for old
and new requests alike, so the benchmarks and session endpoints never
scan SpendLogs. Adaptive and quality routers record no savings baseline
and stay out of the compared totals; savings_estimated_classifier_cost
is kept and covers the same compared requests
* feat(proxy): gzip buffered responses for clients that accept it
Large JSON reads like /user/daily/activity/aggregated shipped tens of MB
uncompressed. Compress single-message bodies of 500B or more when the
client's Accept-Encoding allows gzip (q-values and the wildcard honored).
Streamed and etagged responses pass through untouched, every negotiable
response carries Vary: Accept-Encoding, and bodies of 1MB or more are
compressed in a worker thread
* fix(proxy): skip partial and no-transform responses in gzip and always release the held start
The gzip gate now also skips 206 Partial Content and Cache-Control: no-transform,
since compressing either breaks byte ranges or ignores an explicit ban on transforms.
A response start without a headers key no longer raises, and a start the app never
follows with a body message is forwarded when the app returns instead of being dropped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(tool-policies): show the user who owns the key that discovered a tool
GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tool-policies): bound the owner lookup with chunked membership queries
The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tool-policies): cover the owner column across the tool routes and the dashboard
Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(circleci): test Redis behavior against local Redis and print short tracebacks
The redis_caching_unit_tests job ran three legacy files against the shared remote Redis. The DualCache and
batch-read logic that never needed a server now lives in tests/unit/caching/test_dual_cache.py with a mocked
RedisCache, and the behavior that does need one (the increment-with-floor Lua script, read-through, deletes,
batch reads) moved to tests/integration, which starts a local Redis. test_returned_settings only read
REDIS_PORT and is replaced by a unit test of Router.get_settings
CircleCI pytest runs now use --tb=short so failure output stays readable in the test results tab
* ci(integration): print short tracebacks from run.py and allow Redis in the sdk shard
Sail now rejects completion_window "flex" on synchronous requests with a
400 saying flex is only for background responses or Batch work. The chat
flex case and the responses flex case have failed on every scheduled
litellm-e2e run in builds 337, 340 and 341. The chat cases keep balanced and
auto, and the responses case sends a caller metadata.completion_window of
balanced, so both still prove the window reaches Sail and the bill uses
that window's distinct rates
* chore(lint): remove the LIT002 mutable-construction rule
Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.
Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.
* chore(lint): keep the leftover mutable-ok markers for a follow-up
Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.
* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"
This reverts commit c35bc0b84e.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(ui): bucket model leaderboard usage by day or week
* test(ui): cover daily buckets in model leaderboard series
* feat(ui): add daily/weekly toggle and per-bucket total to model leaderboard chart
* test(ui): cover the daily/weekly toggle on the model leaderboard
* feat(model-insights): add gateway-wide daily totals to the response type
* feat(model-insights): return per-day totals across every model, not just the top ranked ones
* test(model-insights): daily totals include models outside the top ranking
* chore(model-insights): regenerate lazy openapi snapshot for daily totals
* chore(ui): regenerate api types for model insights daily totals
* fix(ui): compute leaderboard bucket totals from gateway-wide daily totals
* fix(ui): show the gateway total, not the top-ten subtotal, in the leaderboard tooltip
* test(ui): cover gateway-wide bucket totals on the model leaderboard
* fix(model-insights): type daily totals as an immutable tuple
* fix(model-insights): build daily totals without new mutable collections
* docs(proxy): point mcp_server test references at tests/unit/proxy
The legacy tests/test_litellm/proxy tree was removed in #44018. Repoint the
mcp_server AGENTS.md mirror path, swap its auth example for a module that still
exists, and drop the utils.py comment block that named the old test path
* docs(proxy): fix remaining mcp_server legacy test path and note import-time env reads
Repoint the second tests/test_litellm reference in the mcp_server AGENTS.md
Tests section and move the import-time env guidance there from the removed
utils.py comment
The pytest-postgresql based proxy tests moved to tests/integration on the real
Postgres harness in #43996, so nothing loads the plugin anymore. uv.lock is
edited by hand to drop the package and its now orphaned mirakuru and port-for
deps; uv lock --check passes and a full relock resolves the same package set
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exercise the shard check directly for unit_selection-owned children
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): serve the redirect test from respx instead of a socket
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): credit shard ownership only to unit flags wired in gha
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): implement the wired-flag shard crediting the tests assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): split the root proxy test files into their own unit shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): keep tuple identity in proxy state restore and fix misc target paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved unit test directories
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exclude proxy-db-owned files from the misc target
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop the redundant fixture docstrings in the proxy conftest
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(lens): remove deployment screenshots
* refactor(lens)!: rename internal engine package and API
* fix(lens): pin worker image for renamed API
* test(lens): cover fresh and populated rename migrations
* fix(lens): protect db-push upgrades and restore routing and CI
* fix(lens): resolve migration tables across schemas and include database driver
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack
The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the sink outage burst models in config
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(s3_v2): read cold storage metadata without an empty dict default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the proxy to reconnect before the postgres outage recovery request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* feat(ui): redesign agent trace run view with chat-style detail pane
Tree with connector lines, typed icon tiles and provider logos, hover
cards with timing, and Input/Output sections rendered as message cards.
* fix(tracing): show text for block-list message content and split normalizers per convention
OpenAI responses-style content (reasoning + text blocks) rendered as raw
JSON in the trace view. Keep the text blocks and drop opaque reasoning.
Move each convention into litellm/tracing/normalizers with an ordered
registry so new frameworks plug in without touching OTLP decoding.
* fix(tracing): keep long message histories as valid JSON and parse function_call blocks
* feat(ui): open agent traces in a resizable side drawer with a devtool-style tree
Clicking a run opens it in a drawer over the list instead of a full page.
j/k and the header arrows switch runs, Esc closes. The tree gets dashed
connectors, per-span waterfall bars, mono tool names and real provider
logos. AI messages with reasoning/function_call blocks render as text.
* fix(trace-ui): address review: valid JSON trimming, drawer keys, reduced motion, narrow screens
* feat(tracing): serve span content in a standard LiteLLM UI format
GET /v1/traces/{trace_id}/spans/{span_id} now also returns input_ui and
output_ui, a tagged union of messages, fields or text built server side by
litellm/tracing/ui_format.py. The trace UI renders from those fields and only
falls back to client-side parsing when talking to an older proxy. The raw
input and output strings are unchanged, and so is storage
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(tracing): fall back to an elision marker when shortened messages still exceed the size limit
* fix(tracing): keep both messages when tool_calls are oversized and keep failed-tool styling
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): reuse the shared httpx transport fixture in moved proxy tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub outbound HTTP and package moved test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore the config server hostname in the mcp resolution test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): pin the completion tokenizer model in the straiker screening test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden
#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key
* test(ui): give the auto-router threshold save wait room for the availability debounce
#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed
* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default
#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against
* test(e2e-ui): wait for the call-id search before hovering the logs row
The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time
* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off
The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support
* test(integration): read the agent 365 guardrail status by its own name in spend logs
The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict
* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check
gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler
* test(e2e-ui): fill the create-tag fields inside the dialog
#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step
* test(integration): run integration proxies with the CI license
Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments
* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5
Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving
codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI
* test(unit): join the session-minting thread before collecting the handler
asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference
* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early
owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.
uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds
* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build
The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there
* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections
The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry
* ci: move the unit-test uv cache split into a composite action
check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes
* test(unit): point tiktoken at the bundled cache for every unit test
The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.
* test(integration): answer model discovery probes in the hosted_vllm wire tests
The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): migrate DB and Redis backed proxy tests into tests/integration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop a type suppression comment from the key metadata integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): scope integration test cleanup to owned rows and wait for backend stats flush
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): seed NULL cache_hit and bound recovery reads from below
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): group /v1/messages contracts under tests/integration/messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): make ci coverage census collect nested test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): nest /v1/messages contracts under messages_endpoint/providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): replay a real Claude Code /v1/messages request through the native wire
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop legacy covers marker from claude code wire test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): inline the Claude Code request instead of a json fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop the legacy covers marker from the new contract
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge) (#43386)
* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): strengthen bot-flagged assertions in the Claude Code matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): type the usage mapping parameter in the shared builders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop low-priority Claude Code error and count_tokens tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): send the full 24-tool Claude Code request and pin upstream headers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop responses bridge Claude Code tests to keep this PR Anthropic direct only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): fix duplicate WebSearch tool, drop mutation in stream builders, ignore pings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): assert upstream request order in multi-turn Claude Code tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): name Claude Code wire tests by behavior and move provider-agnostic ones to routing and streaming
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): sort imports after moving the Claude Code fixture into _support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): group anthropic messages tests into feature subfolders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): drop the pre-move anthropic test paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): cover native reasoning translation, response and pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): narrow PR to new reasoning tests, restore moved files and drop non-reasoning tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): scope reasoning tests to reasoning and cover betas, thinking usage and streamed pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): write reasoning cases as literal sent and received fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): require stopped stream blocks and check upstream model on switch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>