* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): share one optional-feature install path between lazy and eager registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(proxy): say eager lazy routes register at worker startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): name the import callable passed to _install
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove a startup hook can drop eager lazy routes for good
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): strip the lazy-routes flag from the lazy-mode control proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the lazy warmup route registering a feature
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES
Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart
* fix(proxy): keep config pass-through routes ahead of eagerly registered features
With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): accept Converse messages with no content key
A user or tool message whose content key is missing (or null, which the
message cleanup strips) made every Bedrock Converse request fail with
APIConnectionError 'content' before reaching Bedrock. The Converse
transform now reads content with .get for those messages, as it already
did for assistant messages: a content-less user message adds no block
and a content-less tool message becomes a toolResult with empty content.
The str branch also sends the continue message text instead of the
original whitespace-only text.
* fix(bedrock): send the continue message for a content-less user turn
* test(bedrock): type the content-less Converse message test parameters
* fix(bedrock): accept a Converse system message with no content key
* refactor(bedrock): read the system message content with get
* test(integration): audit Bedrock Converse messages without content
Adds the /audit cells for a chat message whose content key is missing or
null on a Converse-routed Bedrock deployment: happy, sad, edge, and chaos
rows through the OpenAI SDK, the Anthropic SDK, and raw httpx against the
scripted upstream, asserting the caller's response, the body the peer
received, and the spend row. The owned-proxy readiness deadline in the
integration harness is now INTEGRATION_PROXY_READY_SECONDS (default 70).
* test(integration): bound stray spend rows in the mid-burst restart cell
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(mcp): expand team and dashboard grants when listing and serving toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honor team toolset grants on the responses gateway path and expose a public team permission lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(mcp): drop the toolset route docstring tweak so the OpenAPI snapshot stays unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): scope inherited toolset grants by the key's own MCP ceiling
A key that declares any MCP grant of its own keeps only its own toolsets, one
that declares none inherits its team's, and require_key_mcp_access_defined
stops a virtual key inheriting while dashboard sessions and admitted users
still do. Adds the direct, no-grant, admin and key-ceiling integration cases
and makes the LLM gateway toolset case discriminate a scoped toolset from the
aggregate grant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): resolve toolset grants per admitted source and enforce the live team roster
Dashboard sessions and gateway-admitted users now expand into the admitted
subject's per-team sources when resolving toolset grants, so a team-granted
toolset is not capped by the user's own MCP row and is reachable on the
namespaced route. A cached team id no longer grants a toolset unless the live
roster still lists the user, a team lookup fault only drops that team's
inherited grant, and /team/member_add evicts the cached team object so the new
member is authoritative immediately
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): read a key's named object permission before letting it inherit team toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): inject the toolset grant resolver into scope helpers so tests stop patching MCPRequestHandler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): drop the duplicate admitted_subject_sources wrapper after merging main and follow its renamed resolvers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honour fresh policy and the session resource scope on pinned toolsets
A pinned toolset scope now reads the toolset through the writer when the admitted session
requires fresh policy, so a tool revoked from the toolset is gone on the next request. A
gateway bearer scoped to one server can only open a toolset that names that server
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep operator-open servers out of toolset gateway urls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): audit cells for team-granted toolsets across every surface
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): bound toolset edit convergence by both cache layers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): drop toolset integration cells that test behavior this PR does not change
Repeat-read byte identity, 20 concurrent calls, a stopped peer and a killed worker are covered
generically by test_mcp_resilience.py and test_mcp_user_env_vars.py. The toolset edit cell asserts
pre-existing cache propagation and flaked locally with connection resets while polling
* chore(mcp): drop mutable-ok suppressions that main's LIT013 now flags as unused
* chore(mcp): keep the require_key_mcp_access_defined read from adding an unknown-argument type error
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* add test case for /rag/query and stronger auth check
* style(proxy): ruff format auth_checks.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): build rag query vector store ids immutably and test the no-registry path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): integration coverage for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): audit cells for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): consolidate vector store allowlist audit coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type RAG vector store request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tool-policies): show the user who owns the key that discovered a tool
GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tool-policies): bound the owner lookup with chunked membership queries
The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tool-policies): cover the owner column across the tool routes and the dashboard
Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(circleci): test Redis behavior against local Redis and print short tracebacks
The redis_caching_unit_tests job ran three legacy files against the shared remote Redis. The DualCache and
batch-read logic that never needed a server now lives in tests/unit/caching/test_dual_cache.py with a mocked
RedisCache, and the behavior that does need one (the increment-with-floor Lua script, read-through, deletes,
batch reads) moved to tests/integration, which starts a local Redis. test_returned_settings only read
REDIS_PORT and is replaced by a unit test of Router.get_settings
CircleCI pytest runs now use --tb=short so failure output stays readable in the test results tab
* ci(integration): print short tracebacks from run.py and allow Redis in the sdk shard
* chore(lint): remove the LIT002 mutable-construction rule
Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.
Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.
* chore(lint): keep the leftover mutable-ok markers for a follow-up
Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.
* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"
This reverts commit c35bc0b84e.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* chore(lens): remove deployment screenshots
* refactor(lens)!: rename internal engine package and API
* fix(lens): pin worker image for renamed API
* test(lens): cover fresh and populated rename migrations
* fix(lens): protect db-push upgrades and restore routing and CI
* fix(lens): resolve migration tables across schemas and include database driver
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack
The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the sink outage burst models in config
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(s3_v2): read cold storage metadata without an empty dict default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the proxy to reconnect before the postgres outage recovery request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden
#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key
* test(ui): give the auto-router threshold save wait room for the availability debounce
#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed
* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default
#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against
* test(e2e-ui): wait for the call-id search before hovering the logs row
The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time
* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off
The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support
* test(integration): read the agent 365 guardrail status by its own name in spend logs
The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict
* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check
gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler
* test(e2e-ui): fill the create-tag fields inside the dialog
#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step
* test(integration): run integration proxies with the CI license
Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments
* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5
Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving
codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI
* test(unit): join the session-minting thread before collecting the handler
asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference
* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early
owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.
uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds
* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build
The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there
* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections
The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry
* ci: move the unit-test uv cache split into a composite action
check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes
* test(unit): point tiktoken at the bundled cache for every unit test
The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.
* test(integration): answer model discovery probes in the hosted_vllm wire tests
The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
* test(proxy): migrate DB and Redis backed proxy tests into tests/integration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop a type suppression comment from the key metadata integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): scope integration test cleanup to owned rows and wait for backend stats flush
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): seed NULL cache_hit and bound recovery reads from below
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): group /v1/messages contracts under tests/integration/messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): make ci coverage census collect nested test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): nest /v1/messages contracts under messages_endpoint/providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): replay a real Claude Code /v1/messages request through the native wire
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop legacy covers marker from claude code wire test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): inline the Claude Code request instead of a json fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop the legacy covers marker from the new contract
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge) (#43386)
* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): strengthen bot-flagged assertions in the Claude Code matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): type the usage mapping parameter in the shared builders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop low-priority Claude Code error and count_tokens tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): send the full 24-tool Claude Code request and pin upstream headers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop responses bridge Claude Code tests to keep this PR Anthropic direct only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): fix duplicate WebSearch tool, drop mutation in stream builders, ignore pings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): assert upstream request order in multi-turn Claude Code tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): name Claude Code wire tests by behavior and move provider-agnostic ones to routing and streaming
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): sort imports after moving the Claude Code fixture into _support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): group anthropic messages tests into feature subfolders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): drop the pre-move anthropic test paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): cover native reasoning translation, response and pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): narrow PR to new reasoning tests, restore moved files and drop non-reasoning tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): scope reasoning tests to reasoning and cover betas, thinking usage and streamed pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): write reasoning cases as literal sent and received fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): require stopped stream blocks and check upstream model on switch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan Responses API input in Azure Text Moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): log Azure Text Moderation prompts at debug and cover streamed Responses blocking
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan Responses API input in Azure Prompt Shield and Text Moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate unmodeled Responses input items in Azure prompt extraction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): pick Azure prompt source by call type so a messages stub cannot hide Responses input
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): tighten Azure Content Safety endpoint test types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): suppress Azure cast lint violations with cast-ok reasons
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): audit Azure content safety across endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): isolate worker-kill audit rig and cover during_call on chat
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(guardrails): shorten Azure cast-ok reasons to fit the line limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): inline spend row count in the concurrency audit cell
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert caller-observed outcomes in Azure call type unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): reuse the existing text moderation response helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert no duplicate rows instead of exact row count after worker kill
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): poll worker-kill spend rows to settle before the duplicate check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep Azure Text Moderation on messages only so this PR stays Prompt Shield scoped
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): narrow hiddenlayer startup auth timeout without cast
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer jwt refresh by configured timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover model_armor and run timeout probes concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): match sink calls to the exact guardrail name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups
Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names
* test(ci): move retired OpenAI text-completion fixtures to live vehicles
OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed
* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases
gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless
* test(ci): skip the ROI calculator repository listing in the security route sweep
GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
* fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun integrations shard after unrelated gitlab prompt manager timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit azure_storage client reuse against a local Data Lake sink
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): cover the exact TTL expiry boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): wait for a rejected write before flipping the sink back
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add azure_storage log delivery cells behind an opt-in lane
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): restart the proxy mid burst and bound the loss to the unflushed queue
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): drop the redundant stop after the owned proxy exits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): read azure_storage objects at the auth-mode-dependent layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): install the datalake sdk in the e2e lint environment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_storage): drop the opt-in real Azure e2e cells and their e2e-dev dependency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): patch the shared proxy logger directly in the straiker api_version test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): keep the straiker stray-version block marker separate from the shared block marker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): drop redundant comments on the straiker api_version integration tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): send request conversation and tool calls to post-call monitor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(grayswan): tighten post-call context typing and wire test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): resolve post-call surface from request route before call_type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): omit tools from post-call monitor when request context is empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(grayswan): apply ruff format to post-call context changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): merge response text and tool calls into one assistant monitor message
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(grayswan): only merge tool calls into the response text for single-choice responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): audit post-call context across endpoints, modes and outages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): share the upstream model probe reply across audit responders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): assert the full generic guardrail body and kill a real serving worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize the client user agent in the generic body assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): normalize accept-encoding in generic body assertion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): keep volatile header placeholders only when the header is present
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): capture monitor calls immutably in the unit test client
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(grayswan): type the test helper parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): inherit catalog service-tier rates for custom-priced deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): apply tier-suffixed long-context rates when only tier thresholds are set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): cover canonical cost-map backend model resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): use descriptive names for service-tier pricing fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* fix(proxy): delete large teams without per-member transaction fan-out
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): evict email-only member caches and reset team members metric on delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep new delete-team literals within the LIT002 ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve deleted-team member ids before the locked delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup
`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.
* fix(repositories): slice find_by_emails into bounded IN statements
The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.
* fix(proxy): delete a team once when /team/delete repeats its id
The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.
* test(integration): audit cells for /team/delete on large, legacy and concurrent teams
Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.
Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.
Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.
* test(integration): pin each chaos outage to a live /team/delete
The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages
vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(hosted_vllm): forward replayed reasoning_content only when it is a string
* test(integration): cover hosted_vllm reasoning_content replay across endpoints
* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): keep upstream support unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): excluded_services resolves from the otel callback config only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): name the otel callback logger so excluded_services owner lookup matches
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): assert no aux datastore traces reach the tenant sink
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): read bogus-start proxy log from the results dir
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): read only this invocation's bogus-start proxy log
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): assert operator kept db spans over the whole recorded window
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): split operator db-span asserts by trace scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): tolerate a bogus exclusion env when callback config wins
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): keep bogus exclusion env fatal when a preset parses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): hoist the preset check out of the callback loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): use a rule-scoped pyright suppression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): log and drop unknown excluded_services instead of failing boot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the operator spend-writer span before checking the tenant for postgres
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): leave callback init and boot untouched when excluded_services is unset
Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): log and ignore malformed excluded_services instead of failing startup
Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): cover failed upstream calls in the excluded_services matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
* fix(responses): scan and mask top-level instructions with guardrails
The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.
Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.
Resolves LIT-8931
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): reject empty guardrail rewrites instead of forwarding raw input
An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the texts-replacing guardrail helper explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): cover skip_system_message_in_guardrail on the live proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system
Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): annotate new guardrail tests with return types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the guardrail test doubles explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): attribute completed batch cost rows to /batches in daily activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): wait for priced batch tokens before asserting team endpoint activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor: clean up fresh tech debt from 2026-09-29
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(routing): pin usage-based routing Redis reads through the proxy and SDK
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): look up hashed key names with two spend log rows per key
The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.
* fix(proxy): cap each spend log name probe at 100 rows per key
* fix(proxy): bound the newest-row probe at where the oldest probe stopped
The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.
* test(integration): add spend log alias probe cells for the daily activity routes
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates
Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys
An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.
The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.
Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep OAuth metadata generations only while a fetch is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): recover daily spend key owners
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): simplify daily spend owner recovery
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy): format daily activity metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover recovered owner metadata merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): bound the daily spend owner lookup with the statement timeout
* test(integration): audit the daily activity key owner fallback on every usage route
Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart
The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): satisfy type-discipline and strict ruff budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stamp the served service_tier on every Responses bridge chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): cover anthropic and responses served-tier billing paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): bill disconnects through the router's anthropic stream wrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: apply ruff format to the anthropic stream changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(coverage): ignore delegating properties the ast scan cannot see
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: keep the cast-ok reasons on the cast call line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(streaming): parameterize delegated chunks and messages types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): follow the anthropic pass_through rename after merging main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drain the logging worker between response cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): keep the served service_tier on streamed chunks and bill it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): type the served service_tier chunk without a loose kwargs dict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill the served service_tier over the requested one
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): drop explanatory comment from the tier resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
* feat(cost_calculator): add cost_per_second for chat per-second pricing
Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins
Move Bedrock commitment rows to cost_per_second so they bill once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop legacy per-second fields from chat paths
Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost_calculator): recognize output-only per-second rates
Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>