* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* feat(traces): add Framework column to otel_traces
* feat(traces): pass span events to normalizers and add framework field
* feat(traces): add Claude Code and Agent SDK span normalizer
* feat(traces): decode events before normalizing and apply tool span names
* feat(traces): list distinct frameworks per trace
* feat(traces): return span framework in trace spans query
* test(traces): add scrubbed Claude Agent SDK OTLP fixtures
* test(traces): cover Claude Agent SDK normalization from real exports
* test(traces): assert trace list frameworks stay scoped per trace
* feat(tracing): validate framework in native normalized spans
* feat(tracing): add framework to Span and frameworks to TraceSummary
* feat(tracing): store normalized framework on span rows
* feat(tracing): surface span framework and trace frameworks
* test(tracing): cover framework aggregation in trace summaries
* test(tracing): decode Claude Agent SDK rows with framework and tool args
* chore(ui): regenerate API types for trace frameworks
* feat(ui): add trace framework registry for Claude Agent SDK and Claude Code
* feat(ui): show SDK logo and label in the runs list Agent column
* feat(ui): show SDK logo and label in the run header
* test(ui): cover SDK label and logo in the runs list
* test(ui): cover SDK label and logo in the run header
* feat(tracing): show the agent's final answer as claude agent span output
* feat(tracing): name claude code agents after their otel service
* test(tracing): cover claude code agent naming from the service
* fix(tracing): mark the span row framework field read-only
* test(tracing): scrub host os details from the claude sdk fixture
* test(tracing): scrub host os details from the detailed claude sdk fixture
* fix(ui): hide the decorative sdk logo from screen readers
* feat(ui): show the agent name with the sdk logo in the runs list
* feat(ui): show the agent name with the sdk logo in the run header
* test(ui): cover agent names beside the sdk logo in the runs list
* test(ui): cover the agent name in the run header
* fix(lens): use recorded agent identities across framework traces
* fix(lens): tighten agent identity and bound trace lookups
* style(tracing): wrap framework agent identity test case
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exercise the shard check directly for unit_selection-owned children
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): serve the redirect test from respx instead of a socket
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): credit shard ownership only to unit flags wired in gha
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): implement the wired-flag shard crediting the tests assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): split the root proxy test files into their own unit shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): keep tuple identity in proxy state restore and fix misc target paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved unit test directories
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exclude proxy-db-owned files from the misc target
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop the redundant fixture docstrings in the proxy conftest
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): redesign agent trace run view with chat-style detail pane
Tree with connector lines, typed icon tiles and provider logos, hover
cards with timing, and Input/Output sections rendered as message cards.
* fix(tracing): show text for block-list message content and split normalizers per convention
OpenAI responses-style content (reasoning + text blocks) rendered as raw
JSON in the trace view. Keep the text blocks and drop opaque reasoning.
Move each convention into litellm/tracing/normalizers with an ordered
registry so new frameworks plug in without touching OTLP decoding.
* fix(tracing): keep long message histories as valid JSON and parse function_call blocks
* feat(ui): open agent traces in a resizable side drawer with a devtool-style tree
Clicking a run opens it in a drawer over the list instead of a full page.
j/k and the header arrows switch runs, Esc closes. The tree gets dashed
connectors, per-span waterfall bars, mono tool names and real provider
logos. AI messages with reasoning/function_call blocks render as text.
* fix(trace-ui): address review: valid JSON trimming, drawer keys, reduced motion, narrow screens
* feat(tracing): serve span content in a standard LiteLLM UI format
GET /v1/traces/{trace_id}/spans/{span_id} now also returns input_ui and
output_ui, a tagged union of messages, fields or text built server side by
litellm/tracing/ui_format.py. The trace UI renders from those fields and only
falls back to client-side parsing when talking to an older proxy. The raw
input and output strings are unchanged, and so is storage
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(tracing): fall back to an elision marker when shortened messages still exceed the size limit
* fix(tracing): keep both messages when tool_calls are oversized and keep failed-tool styling
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): reuse the shared httpx transport fixture in moved proxy tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub outbound HTTP and package moved test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore the config server hostname in the mcp resolution test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): pin the completion tokenizer model in the straiker screening test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): migrate DB and Redis backed proxy tests into tests/integration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop a type suppression comment from the key metadata integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): scope integration test cleanup to owned rows and wait for backend stats flush
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): seed NULL cache_hit and bound recovery reads from below
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject the HIBP client into the breached-password update test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): drive optional-discovery deadlines with an injected loop clock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: bound MCP deadline checks, support Python 3.10, pin HIBP URL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan Responses API input in Azure Text Moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): log Azure Text Moderation prompts at debug and cover streamed Responses blocking
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan Responses API input in Azure Prompt Shield and Text Moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate unmodeled Responses input items in Azure prompt extraction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): pick Azure prompt source by call type so a messages stub cannot hide Responses input
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): tighten Azure Content Safety endpoint test types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): suppress Azure cast lint violations with cast-ok reasons
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): audit Azure content safety across endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): isolate worker-kill audit rig and cover during_call on chat
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(guardrails): shorten Azure cast-ok reasons to fit the line limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): inline spend row count in the concurrency audit cell
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert caller-observed outcomes in Azure call type unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): reuse the existing text moderation response helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert no duplicate rows instead of exact row count after worker kill
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): poll worker-kill spend rows to settle before the duplicate check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep Azure Text Moderation on messages only so this PR stays Prompt Shield scoped
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(proxy): restore pre-config-wins handling of pass-through endpoints
Config-wins (#41779) made general_settings.pass_through_endpoints a config-owned key. The DB reader then got the config list back as if it were DB rows, re-registered each entry without forward_headers on every DB sync, and the stripped copy won the route lookup, so a config pass-through with forward_headers: true stopped forwarding Authorization. UI create, update and delete of pass-throughs were also rejected while the config declared any.
This puts pass-throughs back on their pre-#41779 path: the settings store no longer lets the config own the key, the config list is captured env-resolved at load_config, each DB sync merges DB entries with config entries on paths the DB does not declare, and /config/field/info reads the stored rows only. A UI pass-through write re-applies that merge immediately so the config entries stay served until the next sync.
* fix(proxy): keep config pass-throughs in every reload of the merged list
get_config now returns DB pass-throughs plus config ones on other paths,
each DB sync republishes that merged list, and /config/field/info reads
pass_through_endpoints from the DB row so a UI write never drops stored
entries when models are not stored in the DB
* fix(proxy): keep serving pass-throughs while the config file reloads
load_yaml cleared the runtime pass-through list, so auth: false routes
answered 401 while get_config awaited the database
* fix(proxy): read stored pass-throughs from the writer before a UI write
A lagging read replica could return an older list, and the UI create and
edit flows write the whole field back
* fix(proxy): apply config file pass-through auth changes on reload
The kept runtime list was merged as if it were DB entries, so an edited
config entry on the same path was dropped. Merge the stored DB row with
the fresh config instead, and give the field-info test mock a writer
* fix(proxy): keep pass-throughs served while a DB sync reads the database
get_config resets the stored DB rows before reading them again, which
cleared the served pass-through list and made auth: false routes answer
401 for the length of the read
* refactor(proxy): move the settings store reload out of the loop
basedpyright rejects a Final variable assigned inside a loop
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): narrow hiddenlayer startup auth timeout without cast
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer jwt refresh by configured timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover model_armor and run timeout probes concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): match sink calls to the exact guardrail name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups
Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names
* test(ci): move retired OpenAI text-completion fixtures to live vehicles
OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed
* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases
gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless
* test(ci): skip the ROI calculator repository listing in the security route sweep
GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
* feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token
Stamp metadata.used_client_oauth_token where the proxy decides to forward a
client's Anthropic OAuth token, carry it through StandardLoggingMetadata into
the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and
surface it on the Logs page as a Credential filter and drawer field. The token
itself never reaches the log
* fix(proxy): carry used_client_oauth_token onto failure spend rows for litellm_metadata routes
* fix(proxy): resolve used_client_oauth_token against the provider the call was sent to
* fix(proxy): keep the proxy's used_client_oauth_token stamp on failure rows and move the resolver under llms/anthropic
* fix(logging): read used_client_oauth_token from the proxy-stamped metadata slot
On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload
* fix(spend-logs): read used_client_oauth_token from the bucket the route stamped
A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): patch the shared proxy logger directly in the straiker api_version test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): keep the straiker stray-version block marker separate from the shared block marker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): drop redundant comments on the straiker api_version integration tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): send request conversation and tool calls to post-call monitor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(grayswan): tighten post-call context typing and wire test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): resolve post-call surface from request route before call_type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): omit tools from post-call monitor when request context is empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(grayswan): apply ruff format to post-call context changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): merge response text and tool calls into one assistant monitor message
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): only merge tool calls into the response text for single-choice responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): audit post-call context across endpoints, modes and outages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): share the upstream model probe reply across audit responders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): assert the full generic guardrail body and kill a real serving worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize the client user agent in the generic body assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize accept-encoding in generic body assertion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): keep volatile header placeholders only when the header is present
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): capture monitor calls immutably in the unit test client
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): type the test helper parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
TestPriceDataReloadAPI, TestPriceDataReloadIntegration and TestInvitationEndpoints set
app.dependency_overrides[user_api_key_auth] and never removed it. Under xdist the shared
proxy app kept the override, disabling auth for later tests on the same worker and failing
test_harness_smoke.py::test_auth_as_cleans_up_on_exit. Set it through monkeypatch.setitem
so it is undone at teardown
A per-test leak check over every tests/test_litellm/proxy file that touches
dependency_overrides found these 20 tests as the only leakers; it reports none after this change
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* refactor(proxy): route every team-admin decision through auth/team_access.py
Move the six team-admin helpers out of common_utils, team_endpoints and
key_management_endpoints into litellm/proxy/auth/team_access.py under public
names, and point every management route and helper at them. The key routes
keep checking team admin before org admin, so a team admin whose user row is
gone still passes as before. Status codes and bodies are unchanged, which the
223-case team-admin matrix confirms at the merge base and at the tip
common_utils keeps `_is_user_team_admin` as an alias because the published
litellm-enterprise 0.1.71 wheel still imports it from there
* refactor(proxy): answer every team access check with TeamAccess.allows
Replace the six helpers in auth/team_access.py with one resolver in
litellm/proxy/management/teams/access.py. Each route passes the roles it
accepts (TEAM_OR_ORG_ADMIN or TEAM_ADMIN_ONLY), and /team/update and
/team/info rank roles through strongest_role so org admin still outranks
team admin there
The org lookup moves behind an OrgRoles protocol, implemented by
PrismaOrgRoles in management/users/service.py, and get_team_access in
management/teams/dependencies.py is the only place that reads proxy_server
globals. _check_key_admin_access keeps its name and body from main
Routes that checked org admin first now read the roster first, so a team
admin whose org lookup errors now passes on /team/delete, /team/block,
/team/unblock, member reset_spend and reset_budget, and the team callback
routes. No allowed caller is denied
* fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python
pip install litellm fails on Microsoft Store Python because its user
site-packages is already 134 chars plus the profile name, and the content
filter guardrail ships YAML five directories deep under
litellm/proxy/guardrails/guardrail_hooks/litellm_content_filter/. The
existing wheel guard assumed a 100-char install prefix, so it never saw it.
Move categories/ and policy_templates/ to
litellm/proxy/guardrails/content_filter_data/ and drop the benchmark
fixtures from the wheel. Old category_file paths keep resolving because
the resolver only keys on the trailing categories/<file> or
policy_templates/<file> suffix.
Derive the guard's worst-case prefix from the Store Python site-packages
path with a 15-char profile name (149), fail files at 260 and directories
at 248 (CreateDirectoryW), and fix the off-by-one that let a 260-char
path through.
Fixes#43851
* ci: run the Windows wheel install guard on pull requests
The two Windows jobs live in CircleCI, which never runs on pull requests,
so nothing installs the wheel on Windows before merge. Add a GitHub Actions
job on windows-latest that builds the wheel and runs the guard.
Two things make the run deterministic instead of image dependent. The job
turns the LongPathsEnabled registry key off first, because runner images
ship with it on and python.exe is long-path aware, so a 300-char path would
install fine. The guard installs with pip instead of uv, because uv writes
files from Rust, which switches to extended-length paths on its own and can
never hit MAX_PATH.
* fix(guardrails): keep the old content filter package dir as a category search root
Deployments that copied their own category YAML into
guardrail_hooks/litellm_content_filter/ before the data move would have had
that file rejected by the new directory jail and missing from by-name loads,
inherit_from lookups, the UI category listing and the category YAML endpoint.
Every lookup now searches the bundled data dir first and the old package dir
second, with the bundled copy winning on a name clash.
* fix(guardrails): resolve category files through safe_join
By-name category lookups and the suffix search in the category_file resolver now go through safe_join, so a name or suffix that would escape its data root never reaches the filesystem. The LITELLM_CONTENT_FILTER_ALLOW_EXTERNAL_PATHS opt-out keeps its unjailed search. Clears the two CodeQL path-injection findings on the new lookup code.
* fix(guardrails): keep symlinked category files loadable by name
By-name category lookups resolved symlinks through safe_join, so a category file symlinked into the categories folder from elsewhere stopped loading. Those lookups now only reject names that leave the folder lexically and return the link untouched, matching how by-name loads behaved before the data move. The category_file resolver keeps its realpath jail as before.
* fix(guardrails): keep the category viewer inside the category folders
GET /guardrails/ui/category_yaml/{name} hands raw file contents to any valid key, and on main it refused a symlink whose target left the categories folder. The previous commit let by-name lookups follow symlinks again, which also let the viewer read whatever a symlink in a legacy categories folder pointed at. The viewer now checks the found file's real path against every categories folder it searches and answers 400 as before, while the guardrail's own by-name loads keep following symlinks
The roots come in through a FastAPI dependency so the check is testable against a temp folder, and the content filter's realpath containment moves to path_utils.is_within so both surfaces share it. The test that patched os.path.commonpath covered a branch that no longer exists and goes with it
* ci: drop the Windows wheel install job from pull requests
The job took about 13 minutes on every PR to guard an edge case. The
guard still runs its path-length check on Linux in base_sdk_install and
on Windows in the CircleCI windows_release_wheel job.
Callbacks saved from the Admin UI reach litellm through _add_custom_logger_callback_to_specific_event, which skipped the freshly built logger whenever a callback of the same exact class was already registered. With LITELLM_OTEL_V2 enabled every OTel preset (otel, arize, ...) is an OpenTelemetryV2, so a UI-added arize was treated as a duplicate of the yaml otel callback and never attached, and no trace ever reached Arize
The exists check now compares the exact class and the logger's callback_name, so presets that share a class register side by side while a true re-registration of the same preset is still skipped
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): adopt the new LiteLLM logo and monogram
Swap the bundled admin UI logos for the new brand assets: the primary
logo in blue for light mode and white for dark mode, and the monogram for
the collapsed sidebar, favicons, and the built-in guardrail cards.
/get_image gains a variant=monogram query parameter so the collapsed
sidebar can request the monogram while admin-configured UI_LOGO_PATH /
UI_LOGO_PATH_DARK logos still take precedence. The bundled light logo
moves from JPEG to a transparent PNG.
* test(ui): query collapsed sidebar logos by role to stay within the lint budget
* fix: point remaining logo consumers at the new bundled assets
The Rust gateway UI served /get_image from the removed litellm_logo.jpg,
the non-root get_image tests pinned logo.jpg, and a cookbook script read
litellm/proxy/logo.jpg. Point them at the monogram and logo.png.
* fix(mcp): serve the BYOK OAuth page logo from /get_image
The page pointed at /ui/assets/logos/litellm_logo.jpg, which the rebrand
removes from the dashboard sources, so the next UI build would drop it.
/get_image?variant=monogram is always served by the proxy and follows any
admin-configured logo.
* fix(ui): invert the LiteLLM monogram on dark guardrail cards
The blue monogram has a transparent train cut-out, so on a dark card it
read as a muddy blue block. Inverting it yields the brand's white mark,
which the logo guidelines prescribe for dark backgrounds.
* fix(gateway-ui): serve theme and variant aware logos from the dashboard export
The Rust gateway served one monogram for every /get_image request, and the
committed export lacked it, so /get_image returned 404 until the next UI
release build. Pick the full or monogram logo in light or dark from the
query, ship those assets in the dashboard's public dir and the committed
export, and drop two comments that restated asserted paths.
* fix(proxy): delete large teams without per-member transaction fan-out
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): evict email-only member caches and reset team members metric on delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep new delete-team literals within the LIT002 ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve deleted-team member ids before the locked delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup
`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.
* fix(repositories): slice find_by_emails into bounded IN statements
The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.
* fix(proxy): delete a team once when /team/delete repeats its id
The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.
* test(integration): audit cells for /team/delete on large, legacy and concurrent teams
Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.
Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.
Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.
* test(integration): pin each chaos outage to a live /team/delete
The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix(responses): scan and mask top-level instructions with guardrails
The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.
Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.
Resolves LIT-8931
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): reject empty guardrail rewrites instead of forwarding raw input
An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the texts-replacing guardrail helper explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): cover skip_system_message_in_guardrail on the live proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system
Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): annotate new guardrail tests with return types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the guardrail test doubles explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): strip caller credentials from websocket passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover configured x-api-key in websocket passthrough credential test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(agents): authoritative permissions
* fix: enforce authoritative managed agent permissions
* fix(agents): only consult the identity store for managed targets
is_agent_allowed entered the identity-store path whenever a prisma client
was configured, so an ordinary agent paired with an internal user returned
503 instead of 200. Classify the target from the registry first and fall
back to the store only when the registry has no entry, so an unmanaged
target never depends on the store being reachable.
* fix(agents): gate the managed path on an admitted policy object
Ten call sites branched on `managed_agent_policy is not None`, which any
MagicMock attribute satisfies, so the managed path fired on unmanaged
subjects and died in Pydantic validation as a 503. Route every check
through a shared helper that requires a real AgentResponse.
* test(mcp): stub the writer replica the fresh-policy reads use
reload_admitted_user now passes check_db_only through to get_user_object,
so the user row is read from writer_db. Point the mocks at the replica the
code actually reads and give each parametrized case its own user id.
* fix(agents): cap a managed agent at the invoking team's agents
resolve_agent_access returned the managed policy's grants before the
agent_caller ceiling was applied, so a managed agent acting on behalf of a
user reached agents that user's team was never granted. Intersect with the
caller ceiling the unmanaged path already honours.
* fix(agents): restore token narrowing and scope the private-access suppressions
The managed-model check lost its valid_token narrowing when it moved to the
shared helper. Make the caller-access resolver public rather than reaching
into it from module scope, and give each remaining private access a reason.
* docs(agents): drop the comment claiming admins skip the A2A permission check
The check has never had an admin bypass on this path, so the comment
described behaviour the code does not implement.
* test(proxy): stub the writer reads and restore the MCP manager singleton
Fresh-policy user lookups read writer_db, so the team and rest-endpoint
mocks stubbed a replica the code no longer reads, and the dashboard
session fake still had the pre-kwarg signature. The manager reload also
rebound global_mcp_server_manager in every MCP module without restoring
it, leaking an empty manager into later files.
* style: sort imports under the litellm package ruff config
* fix(mcp): cap a managed agent's servers and tools at the invoking caller
managed_agent_servers and managed_agent_tools returned the agent's own
grants without the agent_caller ceiling the unmanaged resolvers apply, so
a managed agent reached MCP servers and tools the echoed caller could not.
Call the existing ceiling helpers on both axes.
* refactor(mcp): return the caller-capped tools without an interim list
The ceiling helper already returns a sequence, so materializing it into a
list added a mutable collection for nothing. Sort at the return sites
instead, which also makes the tool order stable across both branches.
* fix(agents): preserve actor ceilings during managed target checks
* fix(agents): keep managed permission ceilings authoritative
* fix(mcp): fail closed on authoritative caller team outages
---------
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* fix(proxy): keep request-body aws credentials out of stored spend-log requests
* fix(proxy): redact every credential-named request-body field in stored spend-log requests
Replace the hard-coded AWS key check in the spend-log request-body sanitizer with
SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat,
Gemini and header credentials are redacted too. Proxy-stamped key identity metadata
is kept.
* fix(proxy): keep request identifiers named like keys in stored spend-log requests
* refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based
* refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list
* refactor(proxy): always redact credential-named fields in stored spend-log payloads
* fix(proxy): attribute completed batch cost rows to /batches in daily activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): wait for priced batch tokens before asserting team endpoint activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep MCP permissions visible after key, team and MCP server saves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type the object_permission include as a prisma TypedDict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): do not block key save confirmation on cache refetch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): look up hashed key names with two spend log rows per key
The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.
* fix(proxy): cap each spend log name probe at 100 rows per key
* fix(proxy): bound the newest-row probe at where the oldest probe stopped
The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.
* test(integration): add spend log alias probe cells for the daily activity routes
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate api types for spend logs search description
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop redundant comments from the call id logs helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(e2e): format logs call id helper and spec
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer
The Request ID cell shows only request_id again. When the row's litellm_call_id
differs, the cell tooltip lists it as x-litellm-call-id with its own copy button,
and the drawer header labels the second line x-litellm-call-id: instead of the
call id caption. Stacking two ids in every row made the column noisy for the
common case where the viewer only needs the row they searched for.
* test(e2e): cover the Request ID tooltip hover and copy path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: poll the clipboard after the tooltip copy and drop a jsdom aside
The e2e read navigator.clipboard right after the click, so a slow async write
could fail the check even though copy works. The unit test's fireEvent choice
(jsdom has no layout, so a real pointer move off the trigger closes the tooltip
before the click lands) is documented here instead of inline.
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context
UI and CLI session tokens are now always encrypted with AES-256-GCM and a
fixed session associated-data value, and the session-token check only accepts
AES-GCM values carrying that same value. Stored secrets keep their current
encryption and decrypt unchanged, so nothing needs migrating.
encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20
cannot bind associated data, so an AAD-bound value is always written as
AES-256-GCM, and an AAD-bound decrypt refuses the legacy format.
Session tokens issued before the upgrade stop validating, so UI and CLI users
sign in once more after upgrading.
* test(e2e): cover real SSO login through the dashboard and the lite CLI
Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts
against a live Keycloak stack. The dashboard spec checks that the SSO
session authorizes the Virtual Keys and Models data requests. The CLI
spec runs a real lite login in an isolated HOME with the keyring
disabled, then lists models and sends one chat completion with the
stored session. The main Playwright config now ignores oidc/.
* fix(auth): encode UI/CLI session tokens as unpadded base64url
Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them
Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token
* fix(auth): prefix UI/CLI session tokens with litellm_login_
A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs
The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.
Resolves LIT-9012
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates
Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys
An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.
The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.
Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep OAuth metadata generations only while a fetch is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.
The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)
Resolves LIT-8882
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): recover daily spend key owners
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): simplify daily spend owner recovery
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy): format daily activity metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover recovered owner metadata merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): bound the daily spend owner lookup with the statement timeout
* test(integration): audit the daily activity key owner fallback on every usage route
Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart
The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>