`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.
The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.
`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.
Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.
Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.
Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
* fix(otel): send cache and reasoning tokens in langfuse usage_details
The OTel V2 Langfuse mapper only sent input, output and total, so cache reads, cache writes and reasoning tokens never reached Langfuse. Emit them as input_cached_tokens, input_cache_creation and output_reasoning_tokens, and send input/output net of those buckets so Langfuse does not price the same tokens twice.
Fixes#43542
* fix(otel): drop redundant comments from the usage_details change
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): drive the router request test through an httpx MockTransport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): assert tool definitions reach the router upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* feat(providers): add Prism provider
* fix(providers): complete Prism registration
* feat(providers): expose Prism responses and messages
* feat(providers): add DeepSeek V4.1 Flash to Prism
* test(providers): exercise Prism endpoint requests
* fix(providers): align Prism pricing and limits with the live catalog
deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models
* test(prism): assert cost-map invariants instead of pinning catalog facts
* test(prism): derive the asserted model list from the cost map instead of pinning it
* test(prism): capture requests through respx instead of appending to a list and swapping the client transport
---------
Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
* feat(anthropic): add Claude Sonnet 5.5
Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys
Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit bc9b6f8a5c)
* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard
PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vertex_ai): reproduce traced Gemma Responses stream failure
* fix(vertex_ai): wrap Gemma fake streams for Responses tracing
* test(vertex_ai): cover Gemma traced streams and usage options
* test(vertex_ai): inject gemma test deps and assert hidden usage accounting
Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.
Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.
* test(vertex_ai): drop explanatory comment from usage-option assertions
* fix(gemini): forward seed to the Gemini API instead of rejecting it
The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough
* test(gemini): assert the forwarded seed without mutating shared state
* fix(tools): salvage concatenated JSON tool call arguments
* fix(tools): harden concatenated tool-call salvage for review findings
Skip non-dict JSON during split so salvage cannot emit empty tool calls.
Collapse srvtoolu_ expansions to the first object so server results stay paired.
Allocate __concat_n ids that cannot collide with sibling tool call ids.
Propagate cache_control onto every expanded Anthropic tool_use block.
Rename the XML invoke loop variable so the key-leak gate no longer flags {args}
* test(tools): cover concat id bump and srvtoolu array keep
Only collapse srvtoolu_ when concatenated salvage expanded; a valid JSON
array argument stays one server tool input
* revert(anthropic): drop concat expansion from pass-through adapter
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* revert(tools): keep concat salvage out of request-side tool converters
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): expand strictly salvaged concatenated tool arguments in normalized tool calls
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): retain at most the salvage cap while validating concatenated arguments
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(tools): assert concat sibling ids unique after sanitization
A sibling id that only collides after colon-to-underscore sanitization must force the next concat suffix
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* refactor(tools): drop unused strict mode from split_concatenated_json_objects
Strict mode had no production caller. Rejection cases now sit on salvage, and split matches upstream main
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
---------
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
_is_unsignable_thinking_block() only checked block["signature"], so a
thinking block with a valid-looking signature but empty (or
whitespace-only) thinking text sailed through _drop_unsignable_thinking_blocks
and into anthropic_messages_pt(). Anthropic rejects that with:
400 messages.N.content.M.thinking: each thinking block must contain thinking
This is reachable whenever a thinking_blocks history item gets replayed
through this Anthropic-shaped request path (e.g. a non-Anthropic reasoning
turn with no summary text), the same class of bug PR #36033 fixed on the
Responses adapter's own separate code path.
Now the signature check runs first (unsigned blocks are still dropped, same
as before), then an additional check drops the block if `thinking` is
missing, not a string, or strips to empty. redacted_thinking blocks are
untouched since they don't have type == "thinking".
* fix(vertex_ai): consider tools when validating context caching min tokens
Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.
Fixes#42804
* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
* feat(otel): add SigNoz preset for OpenTelemetry v2
Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite
Absorbs the work from https://github.com/BerriAI/litellm/pull/38206
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): drop explanatory comments from the SigNoz preset
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(types): keep signoz dynamic param lines within ruff format width
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(signoz): assert the missing-endpoint boot path directly instead of in an except block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for the signoz health service
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): stand default cache points down when extra_body hides a direct client mark
On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's.
* fix(caching): keep chat completions on the envelope-merged mark count for the default stand-down
Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides.
* test(langtrace): integration test for the built-in callback wire (path, x-api-key)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langtrace): deliver spans to app.langtrace.ai/api/trace with x-api-key
The built-in langtrace callback posted to the dead host langtrace.ai, sent
the key as api_key instead of x-api-key, and let the OTLP endpoint
normalizer append /v1/traces to the complete /api/trace path, so every
export returned 404. Default the host to https://app.langtrace.ai, honor
LANGTRACE_API_HOST for self-hosted servers, pass the key as an exporter
header instead of a process-wide env var, and keep the /api/trace path
unchanged for traces on the langtrace callback only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langtrace): keep LANGTRACE_API_HOST that already ends in /api/trace
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langtrace): deterministic audit inventory for the built-in callback (surfaces, failures, endpoints, chaos)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langtrace): build the repeated-request body once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langtrace): outage cell asserts at-most-once delivery, not span loss
The OTLP HTTP exporter reposts once on ConnectionError and the batch
processor may still be flushing the previous burst when the sink closes,
so whether the outage burst is lost or delivered after revival depends
on timing. The invariant is no duplicate and recovery on the same port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langtrace): delivered spans must carry the prompt in the gen_ai.content.prompt event
The upstream echoes the marker into the completion, so a whole-span
match alone would still pass if the prompt event disappeared
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langtrace): disable model info refresh so the scripted upstream only sees completion requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(openai): exclude fine-tuned and custom gpt-5-chat aliases from gpt-5 reasoning path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): keep gpt-5-chat alias regression test diff minimal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): cover temperature pass-through for gpt-5-chat aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): annotate locals and wrap long lines in gpt-5-chat alias test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(params): validate stream_chunk_size once and carry it as typed control options
Checks stream_chunk_size at the top of completion() and acompletion(), accepts
digit strings, returns a 400 naming the param unless drop_params is set, and
stores the checked value under _litellm_control. Bedrock Converse and Invoke
read it from litellm_params; the Bedrock-only checker and the dead Invoke pops
are gone. Owned-kwarg filtering now runs through one helper everywhere.
Refs LIT-8317
* test(bedrock): drop tests for the removed stream_chunk_size_from helper
Refs LIT-8317
* fix(params): check stream_chunk_size before the MCP gateway branch
Refs LIT-8317
* fix(params): return assert_never in the exhaustive control-options match
Refs LIT-8317
* fix(params): address council review of the control options change
Read all_litellm_params live so names registered after import stay
LiteLLM-owned, make litellm_params a required keyword on the stream
wrapper hooks, give digit strings and ints the same 18-digit range,
share the default-chunking test table, test the Responses bridge through
litellm.responses, and revert formatting-only churn in existing tests.
Refs LIT-8317
* fix(params): address the second council review of control options
Keep the Responses bridge on its original all_litellm_params forwarding,
narrow _int_from_decimal_string inline so it type-checks, bound nested
huge ints in the error message, store _litellm_control only when a value
is set, simplify the parser to its single field, drop the one-caller
wrapper, and tighten the tests.
Refs LIT-8317
* fix(params): keep the 18-digit length check on stream_chunk_size strings
A 19-character string with leading zeros such as 0000000000000000001 would
otherwise pass as 1, although the rule and the error message say at most
18 digits.
Refs LIT-8317
* test(params): tidy control options tests after council sign-off
Move the Responses bridge test into the existing bridge test file, drop the
rebind test that pinned an implementation detail, assert through
stored_control_options instead of the storage key, and cover
drop_params="true" through Bedrock streaming.
Refs LIT-8317
* test(params): wrap a chunking test row that went past 120 characters
Refs LIT-8317
* fix(tests): stop VCR recording and replaying a test's own localhost upstream
* test(vcr): prove a localhost response an earlier run stored is never replayed
* test(vcr): drive the localhost cassette checks in-process instead of through a loopback server
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* ci: cut CircleCI wall time without loosening test isolation
* fix(ci): parse integration split files that follow --results
The CircleCI machine image ships Python 3.12.2, whose argparse leaves the
files positional empty when it follows an option and another positional, so
every extensions node exited with 'unrecognized arguments'. Reproduced on
3.12.2; parse_intermixed_args selects the files on 3.12.2, 3.12.13 and 3.13
* test(ci): resolve command references in the Rust toolchain guard
The Windows rustup install moved into the install_windows_toolchain command,
which the guard only recognized for install_rust. It now accepts any command
that installs a pinned rustup and reads the Windows toolchain pin from it
* ci: cache the Windows release cargo build from main
windows_release_wheel rebuilt every dependency with fat LTO on each run. It now
restores the release target and cargo registry saved by main's scheduled run,
drops the workspace crates' fingerprints so they always rebuild from the
checked-out source, and still runs the full LTO link
* ci: run the Windows release wheel build on windows.xlarge
The fat-LTO release build is the slowest job in the pipeline; more cores
speed up the dependency compile ahead of the final link
* ci: skip the Windows fingerprint cleanup when the cargo cache missed
On a cold cache the release fingerprint directory does not exist, and the
CircleCI PowerShell wrapper failed the step on the suppressed not-found error
Kwargs LiteLLM code introduces for its own use were only kept out of provider
bodies if someone also listed them in all_litellm_params. Undeclared ones went
into extra_body or optional_params, reached the provider, and the provider
rejected the request. is_litellm_owned_kwarg in types/utils.py now defines
LiteLLM-owned once: a registered name, or any name starting with
INTERNAL_KWARG_PREFIX from litellm/constants.py. Every filter that builds
provider params from kwargs uses it: chat completion, transcription,
embedding, image generation and edit, search and video, ElevenLabs text to
speech, and the Bedrock batch mapper. The two untyped shared filters now take
Mapping[str, object]
The stream_chunk_size wire test becomes test_internal_params_wire.py. It also
sends an undeclared _litellm_ kwarg and asserts that no _litellm_ key reaches
any of the six provider bodies, while extra_body passthrough keeps working
Refs LIT-8318, LIT-8319
* fix(s3_v2): drop terminal upload failures, bound retries per flush and enforce the queue cap at enqueue
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): keep retrying credential-rotation 403s, only AccessDenied-style errors are terminal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(env_keys): exclude DEFAULT_S3_MAX_FLUSH_ATTEMPTS as an internal tuning var
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): retry every 5xx, warn on first queue overflow, validate the flush budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): read the queue cap defensively so un-initialized loggers still enqueue
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): drop the getattr in _enqueue and tighten the retry tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): keep the constructor flush budget when the callback override is invalid
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(s3_v2): adapt per-object upload concurrency to sink latency and throttling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(s3_v2): tidy adaptive limiter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(s3_v2): wake one waiter per released upload slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(s3_v2): make the enqueue queue cap configurable with s3_max_queue_size
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): audit cells for cache hits, coded 403 and callback modes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): retry bucket-wide failures by default, age-budget requeues and make terminal drops and adaptive concurrency opt-in
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(s3_v2): suppress the missing-waiter ValueError explicitly in the adaptive limiter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): count oldest events trimmed after a failed flush as callback failures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(s3_v2): move the mutable-ok marker onto the list literal it suppresses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): report post-flush overflow drops once and grow adaptive concurrency above the floor before asserting back-off
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): fail the SlowDown back-off test when the measured window sees no PUTs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): restore base retry defaults, opt-in age budget, no enqueue cap, back off outside the limiter slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): hoist the default no-op upload slot to a module constant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): drop unused mutable-ok suppressions on queue appends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): keep the retry queue oldest-first and prioritise fresh events at upload time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(s3_v2): keep the mutable-ok marker on the queue list literal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): build request bodies inside the upload slot and keep the sync retry set at base parity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): drop wall-clock sleeps from the unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): rebuild the request body inside the slot on every retry attempt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): anchor the backoff window on the first observed failure and tighten shard assertions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): default the upload slot to the logger limiter so monkeypatched doubles keep working
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): clear ambient AWS env credentials so the rotating profile signs the sync retry test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): shrink the linear send-batch perf test to 2k/8k elements
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): fix stale batch sizes in the perf test assert message
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): keep async in-call retries on the base 403/500/503 set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): hold the upload slot across retries, restore the bool upload contract, and fail safe on bool config
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): mark dropped uploads by element identity so a shared key cannot mask a retryable sibling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): make the per-flush drop lookup constant time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): take the upload slot in the caller like base, build the body once per attempt loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): match base retry, logging and hook behaviour unless the new options are opted in
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): assert the wire key in the init-bypassed sync upload test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): assert the signed headers and wire key in the init-bypassed sync upload test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): default the upload limiter at class level instead of reading it with getattr
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): drop the duplicate annotations that redeclare the class-level counters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(s3_v2): drop terminal-failed uploads by default and bound retry age to one hour
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): fall back to the configured retry age on invalid values
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): drive retry-age tests from a fixed clock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): audit cells for retry-age opt-out and 429 single-put parity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): scan retrieved vector store chunks with the request's pre-call guardrails
Vector store retrieval runs inside acompletion after the proxy's pre-call
guardrails have already seen the request, so a retrieved chunk carrying an
injection reached the prompt unscanned. Each retrieved context message now
goes through every pre-call guardrail the request is subject to before it
is injected: a block raises the same 400 the guardrail gives for user text,
a masking guardrail rewrites the context, and a guardrail that fails while
scanning fails the request instead of injecting the chunk unscanned
* fix(guardrails): return a guardrail block unmapped from exception_type so the Responses API surfaces the guardrail's own 400
* fix(guardrails): build the deployment hooks' identity from stamped metadata only
Top-level user_api_key_* fields in a request body are client controlled, so the pre-call, chunk scan, and post-call deployment hooks now take UserAPIKeyAuth from the metadata the proxy stamped, and the chunk scanner returns or raises on every branch.
* fix(guardrails): block route verdicts on retrieved chunks, keep guardrail verdicts out of router retries and fallbacks, and scan chunks against the client's request
* fix(guardrails): keep the merged guardrail list when scanning chunks against the client's request
The scan request laid the client's kept body over the deployment kwargs, so a client
that sent its own top-level guardrails list shadowed the merged metadata.guardrails
list and a key or team guardrail skipped the chunk scan. The kwargs now win and the
keys the proxy relocates into metadata are dropped from the body's contribution.
* fix(guardrails): strip the deployment's guardrail keys from the chunk scan request so merged team guardrails still run
* test(guardrails): type the vector store scan test doubles
* test(guardrails): type the scan double's request data
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* ci: warn on SQL IN lists with no written bound
Postgres caps a prepared statement at 32,767 bind parameters and a
membership filter binds one per value, so an IN list built from table
data breaks once the table outgrows the cap. That is how the budget reset
job froze every due budget (LIT-7535, #40564).
check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter
whose value has no fixed size and every raw SQL literal that splices a
list in after "IN (", unless the line carries "# bounded-ok: <reason>".
It only warns for now: the output is the inventory for RCA action item
AI-1, and it exits 0.
* ci: decide a constant IN list by its module binding, not its casing
An ALL_CAPS name imported or filled at runtime is as unbounded as any
other, so a name now passes only when the module binds it once to a
value of fixed size. Adds Final to the locals a loop does not forbid.
* ci: only a frozen module value makes an IN list constant
A module list bound once could still grow through append or extend, so
a name now counts as fixed only when it is bound to a tuple, frozenset
or constant. Trims the module docstring to what a reader needs.
* ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones
Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in
and delete_many_in split a deduplicated value list into 5,000-value chunks,
AND each chunk with the caller's where, run them in order (a transaction
handle works) and combine the results. Writes take a required atomicity
argument, and a where that already filters the chunked field is refused.
check_unbounded_in_lists.py now fails CI on any finding missing from
unbounded_in_baseline.txt and on any stale baseline entry, so the baseline
only shrinks. Entries are keyed by path, enclosing scope, kind, field and
occurrence, not line numbers. The helper module is exempt, a constant
spread into a frozen tuple counts as fixed, and messages point at the
helper for "in" and at an array parameter for "not_in" and raw SQL.
A real-Postgres integration test shows a raw 40,000-value filter rejected
for too many bind variables while the helpers handle it.
* refactor: rename bounded_in to chunked_in and let callers pick a chunk size
The helper module is litellm.repositories.chunked_in, and its unit and
integration tests, the checker's exemption path and its finding messages
follow the new name. The `# bounded-ok` marker is unchanged.
find_many_in, count_in, update_many_in and delete_many_in take a
keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value
below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before
any query, which leaves the rest of the filter headroom under Postgres's
32,767 bind-parameter cap.
* refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable
LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results.
* refactor: recover user details with find_many_in, sending chunks as lists
_details_for_user_ids reads users through find_many_in instead of a raw
"in" filter, so its lookup stays under the bind-parameter cap for any
number of recovered keys. Up to 5,000 ids it still sends one find_many
with the same where dict, and a PrismaError from any chunk is still
logged and treated as no details.
The helper now sends each chunk as a list, so a chunked filter equals
the dict a hand-written call would send and a migrated call site's
existing assertions keep passing.
The site's baseline entry is gone.
* ci: skip functional TypedDict field maps in the unbounded IN list check
The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156.
* fix: refuse an update_many_in whose data writes the chunked field
Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}.
* docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it
It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests.
* ci: key an unbounded IN list finding by its filtered expression too
A baseline key of path, scope, kind, field and occurrence let a PR delete
a baselined filter and add a different unbounded one on the same field in
the same function, and the new one took over the old key. The key now
also carries the filtered expression's source, whitespace-normalized
(the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as
one new and one stale entry and fails the run. The same expression
re-added in the same function is still the same finding.
Every baseline entry is rewritten in the new form; the 156 findings are
unchanged, and only occurrence indexes renumber where one field had
several different expressions.
Register Sail (providers.json, LlmProviders.SAIL, OpenAI-compatible lists,
ProviderConfigManager) for chat, Responses and /v1/messages, and add its 12
models to both cost maps with asap, balanced and flex price columns.
Sail picks speed and price with metadata.completion_window and rejects
service_tier, so the Sail chat and Responses configs translate the tier:
default and priority to asap, flex to flex, balanced to balanced, auto to no
window. Billing prices the window that was sent. A tier Sail has no window
for, or a window or tier set where billing cannot see it (request metadata,
extra_body), is a 400 unless drop_params is set.
Add balanced to ServiceTier and its _balanced price columns to the model
info types, the Rust catalog and the dashboard schema. A transform_extra_body
hook on the chat and Responses base configs, which returns extra_body
unchanged by default, lets Sail keep the window when a caller also sends
extra_body.metadata. Sail is listed in the Add Model form and model picker.
Co-authored-by: shrey kharbanda <shrey@berri.ai>
* fix(responses): run stream failure and success hooks on the iterating loop instead of blocking it
A dropped provider stream on native /v1/responses ran the failure logging through run_async_function from inside the async iterator, which parks the event loop thread on a helper-thread future until every failure callback returns, and never returns when a callback waits on state only that loop can advance. With a running loop the failure handlers (and the completed-stream success deployment hook) are now scheduled as tasks on it, the way chat streaming already does; the sync iterator keeps its blocking path
* fix(responses): await stream failure and success logging on the iterating loop before propagating
Keep the merge-base hook set for the native Responses stream: async_failure_handler plus the executor-thread failure_handler on failure, and the post-call success deployment hook on completion. Inside a running loop the async handler is scheduled as a task on that loop and the async iterator awaits it before re-raising, so the loop is never blocked on a foreign-loop future and the failure is attributed before the router's fallback wrapper re-enters the same logging object. The sync iterator inside a running loop keeps the task fire-and-forget with a strong reference.
* fix(responses): submit the sync failure handler only after the async one finishes
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(proxy): add fail_closed_rate_limit_enforcement to reject requests with 503 while Redis rate limit counters are unreachable
* fix(proxy): reject fail-closed rate limit checks before logging the in-memory fallback and pin the boot warning in the lifespan
* fix(proxy): coerce the fail-closed flag, fail closed on read-only checks, and refund partial cluster increments
* fix(proxy): window-guard rate limit refunds and catch the fail-closed rejection by type
* fix(proxy): read the compaction rate-limit gate's limiter from the proxy hook registry
* fix(proxy): count the pending request in read-only rate-limit checks and keep the compaction gate off the caller's parallel slot
The compaction polyfill's summary-model gate, once it ran against the real v3 limiter, showed two behaviors nobody had chosen. The read-only check compared the stored counter with the same `>` the increment path uses, but a read-only check decides a request that has not been counted yet, so a summary model exactly at its rpm limit still went out. The read-only path now adds the pending increment of 1 before comparing; the increment path is unchanged.
The gate also passed the key's max_parallel_requests gauge through, and the read-only gauge count includes the caller's own in-flight slot, so a key with max_parallel_requests: 1 never compacted. The gate now drops that gauge from its descriptors, since the summary call runs inside a request the limiter already admitted.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(otel): detach post-response service spans by request phase, name redis spans by operation
Service spans logged from the post-response phase (success callbacks, the response-cache write) now root their own trace linked to the request span even while the server span is still recording, instead of only when they happen to end after it. Redis service spans are named `redis <operation>`; the litellm call chain that issued them moves to the `litellm.service.caller` attribute via a typed `ServiceLoggerPayload.caller` field.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): keep the service caller on failure and legacy spans, test the production phase dispatch sites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): mark anthropic messages stream cache write as post-response phase
The /v1/messages streaming cache writer awaits async_add_cache inline
instead of going through create_cache_write_task, so its redis span
stayed parented under the request trace. Wrap the write in
post_response_phase so it detaches like the chat completions write.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic): write the Messages stream cache in a background task after handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): stop five stale or flaky CI reds and retry CyberArk policy-load conflicts
The Langfuse redaction unit test exports to a local OTLP capture instead of
polling Langfuse Cloud through a recorded lookup. The passthrough worker-kill
test only requires spend rows for requests the surviving worker served. The
spend-routes sweep treats the intentional /spend/capture_rate 503 as expected.
CyberArk retries a 409 policy load in Python, Rust and the e2e Conjur helper
instead of reading it as "variable exists". The integration egress guard now
matches the script's own cgroup, so it no longer blocks the CircleCI agent,
which runs as the same user.
* fix(ci): keep the policy-load backoff typed as float
* fix(ci): retry CyberArk policy loads without blocking the event loop and tighten the worker-kill and Langfuse tests
* fix(secrets): load CyberArk policy one request at a time per manager
* test(secrets): pin that non-conflict CyberArk policy failures are not retried
* test(unit): run tests/unit with only an allowlisted host environment
CircleCI's unit job inherits every project env var, so real provider keys,
REDIS_HOST, DATABASE_URL and AWS or Azure credentials reached tests that
assume none are set. Locally, litellm's import-time load_dotenv did the same
from any .env up the tree. The unit conftest now drops every variable outside
a small allowlist and disables dotenv before litellm is imported.
* test(e2e): name a failed search and the stuck batch status instead of misattributing them
The websearch session test read an empty web_search_tool_result_error block as a
successful search, so a failing search tool surfaced as a session billing bug.
The batch cancellation timeout now reports the last status the proxy returned.
* fix(ci): scrub the host environment per unit test instead of for the whole pytest process
GHA shards run tests/unit next to other suites in one process, so the import-time
scrub deleted MCP_TEST_PEER_PYTHON before tests/mcp_tests read it and the MCP
upstream fell back to the SDK2 interpreter. The two websearch tests that called
OpenAI and Perplexity live are removed: tests/unit no longer sees their keys.
* fix(ci): scrub only the host variables present before litellm is imported
The per-test scrub also deleted TIKTOKEN_CACHE_DIR, which litellm sets at import to
its bundled encodings, so tokenizer paths tried to download them and hit the
socket guard. The prisma setup test now passes its own database URL instead of
reading one another test leaked into the process environment.
* fix(ci): stop the order-dependent unit reds and settle logging tasks on their own queue
LoggingWorker marked a task done on whichever queue was current when the callback
finished, so a callback that outlived an event-loop change raised "task_done()
called too many times" or undercounted the new loop's queue. It now settles the
queue the task came from.
The rest are test isolation fixes for failures that only appeared when another
file ran first on the same xdist worker: a replaced user_api_key_cache, breaker
metrics unregistered by prometheus tests, semantic_router's health-check filter on
uvicorn.access, logging tasks carried over from bedrock tests, a Router-written
model_cost entry, and a stray post captured by the langflow test. The token
counter check now asserts bounded chunking instead of wall-clock time.
* test(e2e/ui): wait for the logout redirect before visiting a protected page
Logout revokes the session server-side before clearing cookies and navigating, so an immediate page.goto either ran with the cookie still set or was aborted by the logout redirect (net::ERR_ABORTED).
* test(unit): restore the prometheus metrics config per test and settle logs carried from earlier tests in the a2a cost tests
* test(router): pin the router clock in the usage counter tests so a minute rollover cannot empty the read
* test(e2e/ui): wait for logout to clear the token cookie instead of for a login redirect
* test(integration/mcp): answer the model-info probe another test's proxy sends to the model double
* refactor(rust): prepare inference and auth foundations
* fix(rust): keep textract operations parsing from kebab-case model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* done
* refactor(types): derive Anthropic beta string conversions with Strum
* fix(anthropic): report missing max_tokens as a missing field
* refactor(rust): type Anthropic messages headers and auth after the Python layout
Delete anthropic/messages/headers.rs. Its OAuth handling, credential ladder
and beta merging move to anthropic/common_utils.rs where Python keeps them
(optionally_handle_anthropic_oauth, get_auth_header, _merge_beta_headers),
and the feature beta injection becomes update_headers_with_anthropic_beta on
the messages config, as in Python. The BaseAnthropicMessagesConfig impl is
unchanged apart from the bodies of validate_environment and request_headers
Beta values are now the AnthropicBeta enum and BetaSet, which sort, dedupe
and comma-join by construction. Request params gain typed speed, tools and
context_management through Recognized, so the beta logic matches on enums
instead of string-comparing JSON. OauthToken parses the sk-ant-oat token once
and the chat config shares that detection instead of its own copy
Case-insensitive header helpers move next to has_header in litellm-http.
One deliberate divergence: a Bearer-prefixed OAuth key configured through
api_key or ANTHROPIC_API_KEY is sent with a single Bearer scheme, where
Python would emit "Bearer Bearer"
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* done
* fix(rust): repair test compilation and clippy failures
resolve auth before building the outbound request in prepare tests, give the host hook tests their own error type, and drop the disallowed reqwest client and err().expect() from core tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* test: move key-gated tests/test_litellm SDK tests into tests/llm_translation and drop empty folders
* test: make token counter and health check unit tests run offline
* ci: point unit shards, rust path filter, Makefile and docs at tests/unit
* docs: fix stale test_litellm run paths in moved llm_translation tests
* fix: correct databricks e2e sys.path depth and contributing example path
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): email alerts at configured percentages of a team member budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(alerting): label team member budget crossings as team member budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(auth): cover the team member alert dispatch from _check_team_member_budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(email): drop the emoji from the team member budget alert template
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): ignore team member alert thresholds outside 1 to 100 on both the backend and the dashboard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): bound team member alert threshold key length before int parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): drop the legacy covers marker from the team member alert test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(team): reject malformed team_member_max_budget_alert_emails on team writes
Thresholds outside 1-100, non-list recipients, and invalid emails now return 422 on
/team/new, /team/update and PATCH /team/{id} instead of being stored and silently
ignored. The value is stored canonically. Read-side LiteLLM_TeamTable is unchanged,
and the PATCH body stays a raw merge patch so a null threshold still deletes it.
* fix(auth): enforce and alert on team member budgets only in common_checks
The builder re-checked the team member budget inline before common_checks ran the same
check, so one request that crossed a team_member_max_budget_alert_emails threshold
dispatched two alerts. Drop the inline check; common_checks is the single authorization
point and already covers per-member rows, the team default member budget, zero-cost
skips and the cross-pod spend counter. Its 422 message now uses the TeamMember=user:team
form the builder and budget reservation already returned.
* Revert "fix(team): reject malformed team_member_max_budget_alert_emails on team writes"
This reverts commit 703e754b46.
* fix(alerting): keep BaseBudgetAlertType.get_event_message zero-arg
Requiring user_info broke existing callers and out-of-tree subclasses. The team member
label now comes from SlackAlerting.budget_alerts, so the interface and its Readme are
unchanged from main.
* fix(mcp): keep team member budget enforcement on the MCP OAuth auth dependency
The MCP OAuth dependency stops at _user_api_key_auth_builder and never reaches common_checks, so removing the builder's inline member budget check would have let over-budget members through there. Enforce it explicitly for that caller.
* fix(auth): keep main's team member budget enforcement, alert once per request
Restore the builder's team member budget check and 422 message exactly as on main and drop the MCP-only gate. The builder sends the member alert only on the request it rejects; common_checks sends it for requests that get past the builder, so no request alerts twice.
* test(integration): read team member alert deliveries without a shared accumulator
* test(integration): match team member alert deliveries by subject so other alerts cannot race the count
* refactor(proxy): build the team member alert threshold config without mutable collections
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): collapse the alert recipient isinstance checks into one call
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): read the SMTP sink through lock-guarded snapshots and assert the exact deliveries
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: load the embedding base image from a committed 100x100 PNG instead of downloading it
* test: move the volcengine embedding test into tests/unit
* test: check gpt2 and r50k_base tokenizer parity against committed tiktoken reference files
* test: check hub tokenizer selection against an in-memory Hugging Face hub
* test: serve image URLs from respx in the gemini tool-result and format-param tests
* ci: drop the emptied legacy core-utils test path
* test: cover the cohere and anthropic tokenizer paths in the hub tokenizer test
* test: fetch every format-param image through respx and check its bytes reach the request
* test: drop the gpt2 and r50k_base parity tests, which no litellm path uses
* test: drop comments that restate assertions in the format-param test