* test(ui-e2e): update usage page selectors after the #45221 redesign
* test(ui-e2e): drop the top keys locator comment
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(responses): keep tool_result next to tool_use on Anthropic previous_response_id continuations
Replayed history no longer re-adds each stored turn's instructions, and the
current request's instructions go before the replayed history instead of
between the last tool_use and its tool_result
* test(responses): fake the Anthropic transport instead of acompletion in the continuation regression test
* fix(responses): keep the previous turn's instructions when a continuation sends none
* fix(responses): carry only the latest stored instructions when a continuation sends none
Replaying each spend log's own instructions put a system message between a
replayed tool_use and its tool_result, which Anthropic rejects with a 400
* test: make legacy live tests in spend, batches, openai endpoints and audio dirs offline (partial)
* test: migrate wave-1b live tests offline (guardrails, images, ocr, search, openai endpoints)
* test: fix wave-1b review items, add responses/ocr integration tests and firecrawl unit test
* test: anthropic messages router/bedrock/openai-bridge unit tests for wave-1b nodes
* test: anthropic messages logging, prompt-caching and tool-search unit tests; drop migrated base nodes
* test: finish pass_through_unit_tests nodes, logging drain fix and mutations
* test: migrate anthropic passthrough tests to integration wire tests
* test: fix passthrough migration wire spend row lookup and wildcard config
* test: migrate hosted vllm and openai file passthrough tests offline
* test: move assemblyai and vertex passthrough nodes to in-process unit tests
* test: restore unlisted router node and fix logging worker drain in passthrough unit tests
* test: drop spend-row BUG skip and sharpen non-streaming skip reason for anthropic messages
* test: use public presidio alias after merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: restore the batch and file unit tests the migration rewrote away
* test: restore full legacy intent in anthropic messages router unit tests
Drop the false BUG skip on non-streaming aanthropic_messages logging (success
callbacks do fire; the skipped body filtered on the wrong model), assert the
logged model_group, messages, cost and usage, cover streaming logging for both
Anthropic and Bedrock invoke, assert dict content blocks for Anthropic, Bedrock
invoke and the OpenAI bridge, add the Bedrock invoke leg of the router test,
fall back from a real 401, test system-prompt caching and streaming
message_start cache fields on converse and invoke, send the legacy tool-search
tools and beta header, remove the type: ignore and bare dict helper, and add
in-process native /anthropic passthrough spend logging tests
* test: fix passthrough migration wire tests and drop false native spend BUG skip
The native spend row test read rows by the shared master-key digest, so it
matched other tests' rows; read each row by its own message id instead and
assert tokens, total, spend, tags, provider, api_base and end_user on both the
non-streaming and streaming native routes. Merge the streaming test that never
checked spend, assert exact tags from litellm_metadata, stop rebinding a Final
in a loop, require cost > 0, and add the chat-completions bridge cost case the
legacy test covered
* test: harden the spend, OCR, image, search and guardrail migration replacements
The OCR spend tests now build fresh kwargs per case instead of mutating a
shared fixture, and every payload case asserts the exact logged spend. The
OCR wire test reads its spend row by request id and checks exact page
pricing. Image edit, Nova Canvas, DuckDuckGo, Firecrawl, Bedrock guardrail
and Presidio replacements now fake only the provider HTTP boundary (respx or
an in-process aiohttp server) and assert the outbound request, so the
DuckDuckGo limit, Azure base_model pricing and guardrail masking are proven
rather than assumed
* test: cover Exa and Perplexity search structure and max_results offline and retire the two base search methods
* test: drive batch and file replacements through the provider HTTP boundary and real logging callback
Replace monkeypatched litellm.afile_content and AsyncHTTPHandler doubles with respx routes,
read batch logging metadata from a registered success callback instead of get_logging_payload,
use the real managed-files hook for the GEN-2166 regression, assert outbound request bodies,
pin poller ownership explicitly in the migrated DB-sync tests, and require the scripted
upstream to be hit in the responses error-status wire tests
* test: assert file content download headers pass through the proxy
* test: point spend coverage references at the tests that replaced the retired spend job
* test: assert passthrough identity, spend and route dispatch from the code under test
The AssemblyAI non-admin test asserted metadata it wrote itself and leaked a
background poll to the real AssemblyAI host. It now drives assemblyai_proxy_route
with a real Request and waits for the success callback for its own transcript id.
The Vertex spend test matches its log by call id instead of taking the first event.
The OpenAI files wire test hit the native /{provider}/v1/files route; it now calls
/openai/files so the passthrough is what forwards the upload and delete.
* test: mock only the HTTP boundary in the migrated audio tests
Vertex TTS tests no longer replace _ensure_access_token or AsyncHTTPHandler.post;
the token comes from a mocked Google OAuth endpoint and the synthesize call from
respx. Speech tests assert the outbound body, the transcription cache test polls
for the cache write instead of relying on test ordering, and the model pass-through
test checks the multipart model field per model.
* test: wait on a logger event instead of polling the clock in anthropic messages unit tests
Recorders keep payloads in a rebound tuple and set an asyncio.Event; tests
await it with asyncio.wait_for instead of a sleep-and-deadline poll loop
* test: freeze module-level batch and file response fixtures as Final MappingProxyType
* test: fake the presidio analyzer with an in-memory aiohttp connector
The blocked-entity tests started an aiohttp TestServer, which binds a local
socket. They now hand the guardrail a ClientSession whose connector answers
/analyze and /anonymize in process, so no socket is opened and the outbound
analyze text and entities are still asserted
* test: type anthropic messages router test helpers with LiteLLM's Anthropic TypedDicts
Messages, cached system blocks and tool-search tools now use
AnthropicMessagesUserMessageParam, AnthropicMessagesTextParam,
AnthropicToolSearchToolRegex and AnthropicMessagesTool instead of bare
dict shapes; tools are converted to plain dicts only at the acreate call,
whose tools parameter is list[dict]
* test: type batch limiter helpers with TypedDicts and wait on the logging callback event instead of polling
* test: give the migrated OCR, image and presidio helpers precise types
OCR spend helpers take ReadOnly TypedDicts for kwargs and responses and use
LiteLLM's OCRResponse/OCRUsageInfo instead of local pydantic stand-ins; spend
metadata is validated with a TypeAdapter. The presidio fake uses LiteLLM's
PresidioAnalyzeRequest/ResponseItem types, and the image-edit logger validates
the logged payload instead of storing an untyped dict
* test: signal callback and cache events instead of polling
Recorders keep tuples and set an asyncio.Event, thread-safely, when the payload
for this test's transcript id or upstream URL arrives. The transcription cache
test waits on a Cache subclass that signals after async_add_cache. No clock
polling or sleeps remain in these tests.
* test: assert the batch limiter hook updates the caller's request in place
* test: tolerate model-list probes and read native passthrough rows by owned key
The router's OpenAI-compatible model-info refresh (litellm/router.py:10710)
sends GET /v1/models to configured openai api_bases, so the wire answers it
with an empty list and excludes it from the provider-call assertions. Native
/anthropic spend rows are now read by a per-request virtual key digest and
call_type, then the row's request_id is checked against the message id
* test: expect the OCR alias in the proxy response model
The proxy restamps every OpenAI-compatible response model to the name the
client requested (_override_openai_response_model), so /v1/ocr returns the
scenario alias. The upstream model is now checked on the drained request body
instead of inside the peer, where a failed assert never reached the test
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Add @step labels to the HttpTransport methods, the poll and wait helpers and the boot helpers that did real IO without recording a step, so a test that reaches the proxy through them no longer reports an empty or gappy step timeline in the JUnit report.
* test(e2e): move the harness self-tests out of tests/e2e
The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead
* fix(ci): point the providers replay controls at tests/e2e_harness
The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension
* style(tests): mark the raw-HTTP check's bindings Final
* feat(ui): configure Anthropic workload identity federation from the dashboard
Add Credential and Edit Credential offer workload identity federation for Anthropic, Add Model creates a federated credential and attaches it, and the credentials table marks federated credentials. Editing a credential now sends only the values the admin changed, and switching provider no longer leaves the previous provider's default base URL on screen
* fix(ui): keep a credential edit to what the admin set in the federation form
* fix(ui): lock the provider in the federation dialog opened from Add Model
* fix(ui): require one federation id when the identity source is the proxy environment
* test(credentials): cover the Anthropic federation dashboard and credential routes
Integration cells for the credential routes every dashboard shape writes (round trips, PATCH set and delete, malformed bodies, non-admin refusals, the token-file allowlist and exchange-host checks, every identity source through chat and messages against a scripted exchange, concurrent writes across two workers and a worker kill mid burst), plus Playwright specs for the Add Credential, Edit Credential and Add Model federation flows and the team-admin view. The owned proxies boot with a 2 s config reload so both workers serve a stored credential inside the fixture budget.
* test(e2e): type the federation spec's captured bodies and clean up the Add Model deployment by its created id
captureRequestBody and postAsMaster take a type parameter instead of returning Record<string, any>, the spec names the credential and model write shapes it captures, and the Add Model cell reads the deployment id from the /model/new response right after the click so a later failing check no longer leaves the deployment behind.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(e2e): cover ollama and ollama_chat on chat completions, responses and messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): declare Subject metadata and check streamed tool call ids in the ollama suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag quota_management tests with Subject metadata and record budget client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps
* test(e2e): leave the batches cleanup harness unit tests untagged
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): name every driven model on the vllm batch, prompt caching and complexity router subjects
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps
* test(e2e): leave the guardrails and logging harness unit tests untagged
* test(e2e): let the inner create_model step name the guardrail backend deployment
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the default guardrail backend model on the tests that drive it
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag management tests with Subject metadata and record management client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): keep the prompt out of the chat_status step so polled retries collapse
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag claude_code tests with Subject metadata and record CLI driver steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): decorate run_claude directly so the label gate discovers its step
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag llm_translation tests with Subject metadata and record harness steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the realtime param tuples Final
Main already retired LIT002 (#43971). This narrows LIT001 to mutable sequences and sets, so dict, Dict,
DefaultDict, OrderedDict, Counter, ChainMap, defaultdict and MutableMapping annotations are allowed, nested
list/set inside a mapping still trips, and the 945 mutable-ok suppressions that only covered mapping
annotations are deleted (LIT013 now flags them). AGENTS.md guidance updated to match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): live e2e asserting nova sonic realtime delivers each assistant sentence once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): drop ticket reference from nova sonic e2e docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): forward each Nova Sonic assistant sentence once over the realtime API
Nova 2 Sonic sends every assistant text block twice, a SPECULATIVE preview
next to the audio and a FINAL transcript once the audio turn has ended. The
realtime bridge forwarded both, so voice clients rendered each sentence twice
and the FINAL copies opened extra responses after response.done, the last of
which never closed. FINAL assistant text blocks are now dropped whole, so a
turn carries each sentence once inside the one response with its audio
* test(bedrock): tag the live Nova Sonic test and type its helpers
* test(bedrock): pin that the Nova Sonic barge-in marker is dropped with its FINAL block
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): assert the sibling-replica cooldown through the router
* test(e2e): warm the cooldown reads concurrently so every pod's read lands just before the trip
* test(e2e): send the trip right behind the warm so every pod's cooldown read is pinned to it
* test(e2e): warm every pod with a canned-answer group and trip only after every warm call answered
* test(e2e): trim the sibling cell's module docstring to what the design needs
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(mcp): advertise the SDK's latest spec revision and validate the RFC 9207 iss
MCPSpecVersion stopped at 2025-06-18 while the pinned SDK negotiates 2025-11-25, and the version LiteLLM puts on its own outbound initialize was a hardcoded historical member. Add the missing revision, name the highest revision we speak once, and pin it to the SDK's LATEST_PROTOCOL_VERSION with a test so the two cannot drift apart silently.
/authorize now seals the issuer it sent the user to into the OAuth state, and /callback holds the authorization response's RFC 9207 iss against it, refusing to forward a code that came back from an authorization server we never sent the user to. An absent iss, an unanchored server row and a state minted before the seal all keep their current behavior.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep params, query and fragment significant in issuer comparison
The shared canonicalizer drops all three, so two issuers differing only outside the path compared equal and a response from another tenant's authorization server would have continued through the flow.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): refresh generated API snapshots
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): cover OAuth client isolation and lint checks
* fix(mcp): preserve registered clients in the existing save payload
* test(mcp): cover optional OAuth registration metadata
* fix(mcp): preserve compatible OAuth registrations across edits
* fix(mcp): retain OAuth state through pending authorization
* fix(mcp): guard pending OAuth at form submission
* fix(mcp): discard canceled OAuth edit snapshots
* test(mcp): preserve complete OAuth registration assertions
* refactor(mcp): construct OAuth credential updates without mutation
* fix(mcp): simplify issuer binding and reject unverifiable callbacks
* fix(mcp): preserve replacement clients and pending redirect bindings
* fix(mcp): retain clients with replacement authentication methods
* fix(mcp): preserve cached clients and pin manual OAuth issuers
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* ci: move Postgres, MCP and Redis suites to CircleCI integration
* ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails
* ci: revert throwaway assert-ci-coverage check
* ci: keep the e2e helpers the gate tests still use
* ci: move the roi-database Postgres shard to CircleCI integration
* ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path
* ci: match the GitHub env for the moved Postgres and Redis jobs
* ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* update logic that marks a logging callback as complete
* test(logging): cover streaming failure dedupe in mark_logging_complete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): cover streaming failure dedupe in S3 and DataDog
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): keep has_run_logging as a deprecated alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): audit streaming failure dedupe across surfaces, fallbacks and sink outage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): assert anthropic upstream path in streaming failure audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): count only provider posts in streaming audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): assert the sink outage rejects uploads in burst audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): configure datadog retries with router override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): fix four CircleCI regressions on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): stop reloading auth_checks in unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(constants): cover CLI JWT expiry env parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(gateway): restore proxy lifespan after importing gateway.main in launch tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): document working preview setup before coordinated releases
* docs(lens): keep current setup guidance factual and preserve Helm options
* fix(lens): support explicit worker images on Compose 2
* fix(caching): count tool_call cache_control marks in the injection census
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): remove cache census casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): only skip injection on message or content marks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip injection on messages whose tool calls carry marks
Reverts 9f08d8aef8. A default 5m mark injected on assistant text lands
before the client's 1h tool_use mark, which Anthropic rejects with a 400
because a 1h breakpoint must not follow a 5m one. Keeping the full census
in the skip check leaves the client's tool_call breakpoint as the only
one on that message.
* fix(caching): count every client tool_call cache mark in the breakpoint census
The census gated tool_call marks on type function and dict shape, so a client mark on a call without a type or with a string cache_control slipped past the count and injection overflowed the 4 breakpoint cap. Count any non-None tool_call mark, keep the server tool exclusion, and add integration cells for the capped surfaces, the yaml stand-down, Bedrock and Gemini, and router affinity
* test(integration): hold the upstream so the worker kill lands mid-burst
* refactor(caching): reuse the transform's server tool lookup in the breakpoint census
The census now calls the same helper the Anthropic transform uses to decide
whether a marked tool call becomes a server tool block, so the two cannot
drift apart. The owned-proxy burst test waits up to 90 seconds for the burst
to reach the wire before it kills a worker
* refactor(anthropic): move the server tool rebuild check under llms/anthropic
* test(integration): audit the tool call mark census across chat, messages, responses, and chaos
Adds the /audit cells for the breakpoint census on assistant tool_calls marks: the
Responses stream bridge, the OpenAI and Anthropic SDK clients, in-process Pydantic
messages, response cache twins, request-level points, a provider 401 on a capped request,
malformed provider_specific_fields and tool_call ids, null or empty points, a points
update mid-burst, a proxy restart mid-burst, and a worker SIGKILL that picks the worker
holding the burst's upstream connections
* test(integration): close the SDK clients the cache census cells open
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(ui): compose logs tabs directly in the route
* feat(ui): share composable dashboard page layouts
* refactor(ui): compose page header and logs toolbar from parts
PageHeader drops its icon/title/subtitle/primaryAction/tabs/utilities props and the
leadingControls render prop in favor of PageHeaderTitle, PageHeaderDescription and
PageHeaderControls that each wrap one element and forward native props.
LogsTableToolbar's 15 props collapse into one LogsTimeRange value plus composable
LogsToolbar, LogsTimeRangePicker and LogsToolbarSwitch parts assembled in the panel.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): express DataTable layout classes as cva variants
Replaces the hand-rolled class-pair constants with boolean cva variants,
which also brings DataTable back under the complexity budget.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer
The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test
* test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding
/ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin
* test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router
test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG
The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context
* test(secret-detection): give the hand-built redaction request an ASGI path
Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path
* test(integration): isolate litellm callback lists per sdk test
usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards
* test(integration): keep the owner-lookup fault proxy off the shared read replica
The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies
* test(integration): request every seeded key in the team owner breakdown
The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys
* test(integration): give every owned Redis its own port in the redis-cache container
On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"
Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause
* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach
The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail
* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves
vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to
* test(e2e): check only stored message content for a leaked card number
The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there
* test(e2e): assert the proxy decodes token-array embeddings for titan
The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer
* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking
us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5
* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests
Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed
* test(e2e): cite the tokenizer and date behind the titan token-array fixture
* test(e2e): let migration seed replicas finish their request-log indexes before cloning
Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema
* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope
#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake
* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list
test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key
A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a
rename) carries a partial Group resource. Each of its attributes now applies as
if sent with that path, so displayName updates the team alias and externalId
and members get their usual handling, and the pushed attributes merge into the
scim_data snapshot the PUT path already writes. A path-less remove or a
path-less op without an object value is rejected with a 400. Any group PATCH
drops an empty metadata key an earlier push left behind, and the Admin UI
metadata form skips an empty key so an affected team can save its settings.
* fix(scim): let a later path op win over an earlier path-less value in the group snapshot
* fix(scim): type the stored team metadata before the JSON object check
* test(scim): run the real group transformation in the path-less replace test
* test(scim): assert the renamed group comes back from the path-less replace
* test(scim): audit the path-less group PATCH on the live proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(e2e): jwt auto_register map-existing-key repro
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(jwt): route existing-key lookup through VerificationTokenRepository
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key
Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.
* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team
Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.
With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.
* fix(jwt): keep the early return when no master key is set
Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.
Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.
* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes
A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)
Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget
Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401
Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter
* test(e2e): create the reused key in the team the JWT resolves to
The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass
* test(integration): match the held statement across TCP reads
The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read
* fix(jwt): gate key reuse on the claim field, not on the claim value
Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path
* fix(jwt): let an issuer's own user field replace the global one when gating key reuse
An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key
* test(jwt): make the flag-off test fail when the flag no longer gates key reuse
The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings
* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(e2e): move live-provider legacy tests into tests/e2e
Port legacy tests that exercise real providers into the tests/e2e suites that own them, using the harness (/model/new plus deferred cleanup) and asserting on what the caller receives. Delete legacy tests already covered at equal or stronger strength by e2e, integration or unit tests, and drop the now empty ocr_testing CircleCI job
* test(e2e): address review on the live-provider test move
Assert the SSE error frame a client actually receives when a post_call guardrail blocks a stream, and require a tool call for every requested city before checking the answer. Restore the OCR matrix and its CircleCI job, the Claude Agent SDK streaming test, and test_async_create_batch, since their SDK-level and callback assertions have no equivalent in tests/e2e
* test(e2e): accept both guardrail block shapes on a blocked stream
A post_call block before the first chunk reaches the client as HTTP 400 with either a JSON error body or a single SSE error frame, depending on whether the block surfaced as an exception or an error chunk. Assert the policy message is present and the blocked output is absent in both
* test(realtime): restore direct SDK realtime tests against OpenAI
The e2e realtime tests go through the proxy and the remaining SDK tests either mock the upstream or assert less, so keep the direct litellm._arealtime tests with and without intent, and TestOpenAIRealtime::test_realtime_connection, in place
* test: make realtime and Nova stream checks deterministic
The direct SDK realtime tests now fail on a refused connection instead of skipping. The with-intent test asserts OpenAI rejects the exact intent value sent, which only happens when the intent is forwarded. The Nova /v1/messages stream test asserts stream structure, stop reason and usage instead of model wording
* test(realtime): own intent forwarding with a unit test instead of a live rejection
Assert litellm._arealtime passes the intent query param into the OpenAI realtime websocket URL, which is the behavior LiteLLM owns, and drop the live test that depended on OpenAI's rejection wording
The timeout reliability tests rely on a 1ms deadline the real backend always
misses. With E2E_PROVIDER_CACHE on, the deployment pointed at the cache edge,
and its healthy sibling in the same test had already recorded a response for
the same canonical request, so the edge answered from Redis inside the 1ms read
window. Build 342 of litellm-e2e saw test_timeout_trips_cooldown_then_recovers
get a 200 from the timing-out deployment itself, with a recording made about
12 hours earlier. Both timeout helpers now register on the live provider path,
which PROVIDER_CACHE.md reserves for tests that need real provider timing
* feat(tool-policies): show the user who owns the key that discovered a tool
GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tool-policies): bound the owner lookup with chunked membership queries
The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tool-policies): cover the owner column across the tool routes and the dashboard
Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sail now rejects completion_window "flex" on synchronous requests with a
400 saying flex is only for background responses or Batch work. The chat
flex case and the responses flex case have failed on every scheduled
litellm-e2e run in builds 337, 340 and 341. The chat cases keep balanced and
auto, and the responses case sends a caller metadata.completion_window of
balanced, so both still prove the window reaches Sail and the bill uses
that window's distinct rates
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exercise the shard check directly for unit_selection-owned children
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): serve the redirect test from respx instead of a socket
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): credit shard ownership only to unit flags wired in gha
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): implement the wired-flag shard crediting the tests assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): split the root proxy test files into their own unit shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>