* test(router): settle the shared logging worker before recording shadow callbacks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): always stop the shared logging worker after settling it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.
Resolves LIT-9012
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter
increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot
release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch
is sent once the success and failure callbacks have run, or on a deadline, and pending batches are
drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request
stay direct; numeric string TTLs keep the direct-path coercion.
Resolves LIT-8883
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.
The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)
Resolves LIT-8882
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Claude Code background sessions (claude --bg) stamp x-app: cli-bg on every
request, including main-loop turns. The session router binding only
accepted x-app: cli, so a background session never bound and its
subagents' concrete-model calls bypassed the router.
The binding write already requires the requested model to resolve to a
pre-routing strategy, so background side calls naming plain models still
never bind. Since f6eff1bde0 removed the clear path, the x-app check
guarded nothing else.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(responses): run aresponses through the async wrapper so the cache is read once
aresponses() sets kwargs["aresponses"] = True and runs the decorated sync
responses() on an executor, but _is_async_request() did not recognise that
flag, so the sync wrapper did a second cache lookup on the executor thread
with a differently ordered cache-key input. Every /v1/responses request paid
two cache GETs against two different keys. Recognising aresponses in
_is_async_request() leaves the async wrapper as the only cache reader and
writer for the async path, one GET per request, same key on read and write
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): let a responses cache entry cover aresponses so responses-only configs keep caching
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.
Resolves LIT-8881
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(router): fetch cooldown state and usage counters in one Redis round trip
The cooldown filter (CooldownCache) and usage-based-routing-v2 selection
(LowestTPMLoggingHandler_v2) each issued their own MGET on every request
because they live in different objects. RoutingReadBatch fetches both key
sets through DualCache.async_batch_get_cache_shared while the healthy
deployments are resolved and hands the usage slice to the strategy, so
selection does not read again. Each cache keeps its own memory tier,
throttling, reservation rollback and circuit-breaker handling, and the
strategy falls back to its own read when the prefetch does not cover its
keys. simple-shuffle keeps reading only cooldowns.
aresponses no longer issues a second, blocking response-cache read from
the worker thread that runs the sync wrapper.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep per-cache tier failures inside the shared batch read
Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared
so a failing tier degrades that cache's read to None the way async_batch_get_cache does,
instead of escaping into routing. Drop the aresponses sync-cache guard: for native
Responses models the worker-thread read is the one whose key matches the write, so
skipping it broke cached /v1/responses replays.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(router): rename usage key builder so the async cache-call check reads it as a key helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(alerting): narrow daily-report cache values before numeric comparison
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): type the shared batch-read helpers and merge Redis results without mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: fix import sort in test_dual_cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): flatten shared batch read keys without a stacked comprehension
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): satisfy type-discipline and strict ruff budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stamp the served service_tier on every Responses bridge chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): cover anthropic and responses served-tier billing paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): bill disconnects through the router's anthropic stream wrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: apply ruff format to the anthropic stream changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(coverage): ignore delegating properties the ast scan cannot see
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: keep the cast-ok reasons on the cast call line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(streaming): parameterize delegated chunks and messages types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): follow the anthropic pass_through rename after merging main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drain the logging worker between response cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): keep the served service_tier on streamed chunks and bill it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): type the served service_tier chunk without a loose kwargs dict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill the served service_tier over the requested one
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): drop explanatory comment from the tier resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
* fix(model-prices): correct azure/eu/gpt-6-astra to Data Zone rates
Co-authored-by: rain <1504569896@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): align groq, gemini and openai entries with official docs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): roll in verified Vertex, Gemini, OpenRouter and Azure AI registry fixes
Absorbs the fields from #43609, #43666, #43671 and #43644 that match the provider's own docs or price API today, and adds a cost test for the azure/eu/gpt-6-astra Data Zone tiers
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): add Copilot, Bedrock Kimi K3, Gemini Robotics and OpenRouter values from official sources
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: rain <1504569896@qq.com>
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
* feat(cost_calculator): add cost_per_second for chat per-second pricing
Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins
Move Bedrock commitment rows to cost_per_second so they bill once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop legacy per-second fields from chat paths
Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost_calculator): recognize output-only per-second rates
Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): classify every credential-bearing param for the canary suite
* test(proxy): classify gcs_path_service_account as secret, run registry in auth-checks shard, check slot ids at import
* test(proxy): use one generic slot id for callback and request-body credential params
* test(proxy): name a canary slot only for params an integration test plants
* test(proxy): classify the SigNoz callback params
* test(proxy): move the slot sync note into the module docstring
* test(proxy): model unplanted credential params as their own classification
* test(security): classify request-body api_key as unplanted until D1 exists; check registry slots against the harness
* test(security): classify request-body api_key under slot D1
* test(security): classify Langfuse and Datadog callback secrets under slots C1 and C3
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma
The noma_v2 guardrail accepts a gateway_name param, falling back to the
NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty
it goes out as a top-level gateway_name field on /litellm/guardrail. The
param works for both guardrail: noma_v2 and guardrail: noma with use_v2,
and it is appended after the existing constructor params so positional
callers keep their meaning
* chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name
The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so
the lazy snapshot and the generated dashboard types need regenerating
* Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over
The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.
Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): mirror every dispatcher fallback path in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): skip already-tried order levels in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(router): credit the #39566 branch this fix supersedes
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.
The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.
`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.
Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.
Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.
Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
* fix(otel): send cache and reasoning tokens in langfuse usage_details
The OTel V2 Langfuse mapper only sent input, output and total, so cache reads, cache writes and reasoning tokens never reached Langfuse. Emit them as input_cached_tokens, input_cache_creation and output_reasoning_tokens, and send input/output net of those buckets so Langfuse does not price the same tokens twice.
Fixes#43542
* fix(otel): drop redundant comments from the usage_details change
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): drive the router request test through an httpx MockTransport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): assert tool definitions reach the router upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* feat(providers): add Prism provider
* fix(providers): complete Prism registration
* feat(providers): expose Prism responses and messages
* feat(providers): add DeepSeek V4.1 Flash to Prism
* test(providers): exercise Prism endpoint requests
* fix(providers): align Prism pricing and limits with the live catalog
deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models
* test(prism): assert cost-map invariants instead of pinning catalog facts
* test(prism): derive the asserted model list from the cost map instead of pinning it
* test(prism): capture requests through respx instead of appending to a list and swapping the client transport
---------
Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
* feat(anthropic): add Claude Sonnet 5.5
Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys
Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit bc9b6f8a5c)
* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard
PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vertex_ai): reproduce traced Gemma Responses stream failure
* fix(vertex_ai): wrap Gemma fake streams for Responses tracing
* test(vertex_ai): cover Gemma traced streams and usage options
* test(vertex_ai): inject gemma test deps and assert hidden usage accounting
Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.
Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.
* test(vertex_ai): drop explanatory comment from usage-option assertions
* fix(gemini): forward seed to the Gemini API instead of rejecting it
The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough
* test(gemini): assert the forwarded seed without mutating shared state
* fix(tools): salvage concatenated JSON tool call arguments
* fix(tools): harden concatenated tool-call salvage for review findings
Skip non-dict JSON during split so salvage cannot emit empty tool calls.
Collapse srvtoolu_ expansions to the first object so server results stay paired.
Allocate __concat_n ids that cannot collide with sibling tool call ids.
Propagate cache_control onto every expanded Anthropic tool_use block.
Rename the XML invoke loop variable so the key-leak gate no longer flags {args}
* test(tools): cover concat id bump and srvtoolu array keep
Only collapse srvtoolu_ when concatenated salvage expanded; a valid JSON
array argument stays one server tool input
* revert(anthropic): drop concat expansion from pass-through adapter
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* revert(tools): keep concat salvage out of request-side tool converters
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): expand strictly salvaged concatenated tool arguments in normalized tool calls
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): retain at most the salvage cap while validating concatenated arguments
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(tools): assert concat sibling ids unique after sanitization
A sibling id that only collides after colon-to-underscore sanitization must force the next concat suffix
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* refactor(tools): drop unused strict mode from split_concatenated_json_objects
Strict mode had no production caller. Rejection cases now sit on salvage, and split matches upstream main
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
---------
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
_is_unsignable_thinking_block() only checked block["signature"], so a
thinking block with a valid-looking signature but empty (or
whitespace-only) thinking text sailed through _drop_unsignable_thinking_blocks
and into anthropic_messages_pt(). Anthropic rejects that with:
400 messages.N.content.M.thinking: each thinking block must contain thinking
This is reachable whenever a thinking_blocks history item gets replayed
through this Anthropic-shaped request path (e.g. a non-Anthropic reasoning
turn with no summary text), the same class of bug PR #36033 fixed on the
Responses adapter's own separate code path.
Now the signature check runs first (unsigned blocks are still dropped, same
as before), then an additional check drops the block if `thinking` is
missing, not a string, or strips to empty. redacted_thinking blocks are
untouched since they don't have type == "thinking".
* fix(vertex_ai): consider tools when validating context caching min tokens
Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.
Fixes#42804
* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
* feat(otel): add SigNoz preset for OpenTelemetry v2
Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite
Absorbs the work from https://github.com/BerriAI/litellm/pull/38206
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): drop explanatory comments from the SigNoz preset
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(types): keep signoz dynamic param lines within ruff format width
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(signoz): assert the missing-endpoint boot path directly instead of in an except block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for the signoz health service
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): stand default cache points down when extra_body hides a direct client mark
On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's.
* fix(caching): keep chat completions on the envelope-merged mark count for the default stand-down
Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides.