* feat(e2e): record each e2e test's steps, starting with ProxyClient
@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.
* docs(e2e): rewrite the recorded test steps guide in plain language
* fix(e2e): keep logging callback credentials out of recorded steps
* fix(e2e): mask the run's credentials in every recorded step
* fix(e2e): attach steps before the oauth failure snapshot
The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach
* fix(e2e): name the saved credential in its recorded step
The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups
Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names
* test(ci): move retired OpenAI text-completion fixtures to live vehicles
OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed
* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases
gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless
* test(ci): skip the ROI calculator repository listing in the security route sweep
GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
* feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token
Stamp metadata.used_client_oauth_token where the proxy decides to forward a
client's Anthropic OAuth token, carry it through StandardLoggingMetadata into
the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and
surface it on the Logs page as a Credential filter and drawer field. The token
itself never reaches the log
* fix(proxy): carry used_client_oauth_token onto failure spend rows for litellm_metadata routes
* fix(proxy): resolve used_client_oauth_token against the provider the call was sent to
* fix(proxy): keep the proxy's used_client_oauth_token stamp on failure rows and move the resolver under llms/anthropic
* fix(logging): read used_client_oauth_token from the proxy-stamped metadata slot
On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload
* fix(spend-logs): read used_client_oauth_token from the bucket the route stamped
A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
get_mcp_server_answering_to stops at the pass that finds an exact name or id and hides it from client_ip instead of falling through to the case-insensitive and prefix passes, and _scoped_server treats a name the registry knows for some caller but not this one as denied, so connect, discovery and scoped routing all refuse the hidden route instead of serving a public server whose alias only differs by case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Rework the Agent 365 sign-in provider for the guardrail shape 43189 landed on main: the OBO scope and
resource come from the fixed AGENT_365_PROD_* constants instead of the removed resource_app_id/api_base
fields, and the exchange runs through the shared TokenExchanger so the connect preflight and the tool
call reuse one cached token per caller assertion.
A present but non-JWS bearer is now rejected in preflight_caller_sign_in, so the connect answers 401
with the RFC 9728 challenge instead of letting the call reach tools/call and lose WWW-Authenticate in
the JSON-RPC error. The OBO-only tool-call challenge in operations.py stays narrowed to token_exchange
servers.
Immutable rewrites (tuple, MappingProxyType, explicit None checks) keep the LIT002 total within the
budget without a mutable-ok
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
LIT014 budgets one for clause per comprehension; chain.from_iterable keeps
the dedupe immutable
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
get_mcp_server_answering_to now falls back to get_mcp_server_by_name when
no published prefix form matches, preserving the exact-name lookup the
preemptive path had before the router-equivalent resolver. Applies ruff
format to caller_sign_in.py and agent_365.py per the lint gate.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Replace the parallel gateway sign-in provider registry with
CallerSignInProvider, merged into the existing oauth2_token_exchange
challenge: gates key off caller_sign_in_for() returning non-None, the
resolver answers by server_id, case-insensitive name, and short prefix
like the router, keeps_caller_authorization covers OBO servers, and the
Agent 365 guardrail exchanges through an injected TokenExchanger.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Listing used the resolved stored credential both to key the caller's catalog slot and as the
upstream transport header, so REST api_key and bearer_token listings sent the user's secret
instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client
and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly
as before the catalog existed, and the stored credential only names the slot tools/call reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun integrations shard after unrelated gitlab prompt manager timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit azure_storage client reuse against a local Data Lake sink
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): cover the exact TTL expiry boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): wait for a rejected write before flipping the sink back
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add azure_storage log delivery cells behind an opt-in lane
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): restart the proxy mid burst and bound the loss to the unflushed queue
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): drop the redundant stop after the owned proxy exits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): read azure_storage objects at the auth-mode-dependent layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): install the datalake sdk in the e2e lint environment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): drop the opt-in real Azure e2e cells and their e2e-dev dependency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
gitpython 3.1.62 (2026-09-07) and tornado 6.5.10 (2026-09-15) are past the
3-day uv cooldown. diskcache still has no fixed release, so its ignore moves
from 2026-10-01 to 2026-11-01
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): patch the shared proxy logger directly in the straiker api_version test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): keep the straiker stray-version block marker separate from the shared block marker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): drop redundant comments on the straiker api_version integration tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): send request conversation and tool calls to post-call monitor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(grayswan): tighten post-call context typing and wire test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): resolve post-call surface from request route before call_type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): omit tools from post-call monitor when request context is empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(grayswan): apply ruff format to post-call context changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): merge response text and tool calls into one assistant monitor message
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): only merge tool calls into the response text for single-choice responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): audit post-call context across endpoints, modes and outages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): share the upstream model probe reply across audit responders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): assert the full generic guardrail body and kill a real serving worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize the client user agent in the generic body assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize accept-encoding in generic body assertion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): keep volatile header placeholders only when the header is present
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): capture monitor calls immutably in the unit test client
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): type the test helper parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential,
one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the
pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
TestPriceDataReloadAPI, TestPriceDataReloadIntegration and TestInvitationEndpoints set
app.dependency_overrides[user_api_key_auth] and never removed it. Under xdist the shared
proxy app kept the override, disabling auth for later tests on the same worker and failing
test_harness_smoke.py::test_auth_as_cleans_up_on_exit. Set it through monkeypatch.setitem
so it is undone at teardown
A per-test leak check over every tests/test_litellm/proxy file that touches
dependency_overrides found these 20 tests as the only leakers; it reports none after this change
The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that
credential to the upstream client, which on an oauth2 server short-circuited the client_credentials
mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so
tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request
header, letting the M2M mint or signed JWT proceed as on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Map dangerous-tool-use-2026-09-03 for azure_ai in the beta header config so the
Claude Code auto mode beta reaches Foundry instead of being stripped, matching
the anthropic, bedrock, and vertex_ai entries
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(router): inherit catalog service-tier rates for custom-priced deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): apply tier-suffixed long-context rates when only tier thresholds are set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): cover canonical cost-map backend model resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): use descriptive names for service-tier pricing fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* test(bedrock): restore the AWS env after a failed live call in the auth tests
* test(bedrock): assert the regression test's failing call actually ran
* test(bedrock): drop the test that tests the auth tests
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(proxy): route every team-admin decision through auth/team_access.py
Move the six team-admin helpers out of common_utils, team_endpoints and
key_management_endpoints into litellm/proxy/auth/team_access.py under public
names, and point every management route and helper at them. The key routes
keep checking team admin before org admin, so a team admin whose user row is
gone still passes as before. Status codes and bodies are unchanged, which the
223-case team-admin matrix confirms at the merge base and at the tip
common_utils keeps `_is_user_team_admin` as an alias because the published
litellm-enterprise 0.1.71 wheel still imports it from there
* refactor(proxy): answer every team access check with TeamAccess.allows
Replace the six helpers in auth/team_access.py with one resolver in
litellm/proxy/management/teams/access.py. Each route passes the roles it
accepts (TEAM_OR_ORG_ADMIN or TEAM_ADMIN_ONLY), and /team/update and
/team/info rank roles through strongest_role so org admin still outranks
team admin there
The org lookup moves behind an OrgRoles protocol, implemented by
PrismaOrgRoles in management/users/service.py, and get_team_access in
management/teams/dependencies.py is the only place that reads proxy_server
globals. _check_key_admin_access keeps its name and body from main
Routes that checked org admin first now read the roster first, so a team
admin whose org lookup errors now passes on /team/delete, /team/block,
/team/unblock, member reset_spend and reset_budget, and the team callback
routes. No allowed caller is denied
* fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python
pip install litellm fails on Microsoft Store Python because its user
site-packages is already 134 chars plus the profile name, and the content
filter guardrail ships YAML five directories deep under
litellm/proxy/guardrails/guardrail_hooks/litellm_content_filter/. The
existing wheel guard assumed a 100-char install prefix, so it never saw it.
Move categories/ and policy_templates/ to
litellm/proxy/guardrails/content_filter_data/ and drop the benchmark
fixtures from the wheel. Old category_file paths keep resolving because
the resolver only keys on the trailing categories/<file> or
policy_templates/<file> suffix.
Derive the guard's worst-case prefix from the Store Python site-packages
path with a 15-char profile name (149), fail files at 260 and directories
at 248 (CreateDirectoryW), and fix the off-by-one that let a 260-char
path through.
Fixes#43851
* ci: run the Windows wheel install guard on pull requests
The two Windows jobs live in CircleCI, which never runs on pull requests,
so nothing installs the wheel on Windows before merge. Add a GitHub Actions
job on windows-latest that builds the wheel and runs the guard.
Two things make the run deterministic instead of image dependent. The job
turns the LongPathsEnabled registry key off first, because runner images
ship with it on and python.exe is long-path aware, so a 300-char path would
install fine. The guard installs with pip instead of uv, because uv writes
files from Rust, which switches to extended-length paths on its own and can
never hit MAX_PATH.
* fix(guardrails): keep the old content filter package dir as a category search root
Deployments that copied their own category YAML into
guardrail_hooks/litellm_content_filter/ before the data move would have had
that file rejected by the new directory jail and missing from by-name loads,
inherit_from lookups, the UI category listing and the category YAML endpoint.
Every lookup now searches the bundled data dir first and the old package dir
second, with the bundled copy winning on a name clash.
* fix(guardrails): resolve category files through safe_join
By-name category lookups and the suffix search in the category_file resolver now go through safe_join, so a name or suffix that would escape its data root never reaches the filesystem. The LITELLM_CONTENT_FILTER_ALLOW_EXTERNAL_PATHS opt-out keeps its unjailed search. Clears the two CodeQL path-injection findings on the new lookup code.
* fix(guardrails): keep symlinked category files loadable by name
By-name category lookups resolved symlinks through safe_join, so a category file symlinked into the categories folder from elsewhere stopped loading. Those lookups now only reject names that leave the folder lexically and return the link untouched, matching how by-name loads behaved before the data move. The category_file resolver keeps its realpath jail as before.
* fix(guardrails): keep the category viewer inside the category folders
GET /guardrails/ui/category_yaml/{name} hands raw file contents to any valid key, and on main it refused a symlink whose target left the categories folder. The previous commit let by-name lookups follow symlinks again, which also let the viewer read whatever a symlink in a legacy categories folder pointed at. The viewer now checks the found file's real path against every categories folder it searches and answers 400 as before, while the guardrail's own by-name loads keep following symlinks
The roots come in through a FastAPI dependency so the check is testable against a temp folder, and the content filter's realpath containment moves to path_utils.is_within so both surfaces share it. The test that patched os.path.commonpath covered a branch that no longer exists and goes with it
* ci: drop the Windows wheel install job from pull requests
The job took about 13 minutes on every PR to guard an edge case. The
guard still runs its path-length check on Linux in base_sdk_install and
on Windows in the CircleCI windows_release_wheel job.
* fix(transcription): honor base_url alias for Groq Whisper and report it as the api base
transcription() and speech() only accepted api_base, so a deployment configured with base_url leaked the alias into the provider params, which Groq rejected as an unknown param, and the request never reached the internal gateway. get_api_base() now reads the same alias so response headers and logs show the configured endpoint instead of the provider default
Resolves LIT-9071
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(speech): keep base_url after existing audio params, route Vertex speech to it, skip empty alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>