* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
aws_bedrock_project_id reached Bedrock Mantle as anthropic-workspace, a
header AWS ignores, so requests ran under account defaults and models
that require a project data-retention mode failed with a 400. Send the
header AWS documents, anthropic-workspace-id, on both Mantle routes.
Re-lands #31994 by @gunjanjaswal, merged into the since-deleted
litellm_oss_staging branch on 2026-08-22, which never reached main.
Fixes#31947
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Bedrock GPT-5.6 rejects reasoning effort minimal with 400 unsupported_value. The model map already
sets supports_minimal_reasoning_effort=false for these models, but neither the Converse path nor the
native Responses path read it, so the value was forwarded. Drop it under drop_params and raise
UnsupportedParamsError otherwise, matching how the OpenAI GPT-5 path treats explicitly disabled
effort levels
Co-authored-by: Aasif-Multani <20943280+Aasif-Multani@users.noreply.github.com>
* feat(bedrock): send grok chat completions through runtime openai path
Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse
* feat(bedrock): serve gpt-oss and gpt-5.6 chat completions on runtime's native openai path
* fix(bedrock): route gpt-oss response_format to Converse and decide the route once from the raw request
* fix(bedrock): serve region-path and GovCloud gpt-oss ids on native Chat Completions
The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects
* fix(bedrock): keep params AWS refuses natively off the chat completions route
Drop the params each family 400s or 503s on runtime Chat Completions from the native config's supported list (GPT-5.6 penalties, stop, and logprobs, Grok penalties, gpt-oss logit_bias) so drop_params drops them as Converse did, gate legacy functions on GPT-5.6 the same way as tools, and send an Anthropic-style thinking block to Converse, the only route that forwards it
* fix(bedrock): keep schema-less json_object on Converse for the chat completions models
* fix(bedrock): keep every json_object response_format on Converse for the chat completions models
* fix(rust): declare the bedrock runtime chat completions flags on ModelInfo
* fix(bedrock): opt into the native chat completions route through supported_endpoints
* docs(cost-map): describe the bedrock native chat completions capability flags
* revert: docs(cost-map): describe the bedrock native chat completions capability flags
This reverts commit 4101c0ceb2.
cost-map-guard runs main's schema generator under pull_request_target and compares
its output to the PR's committed schema, so a PR that changes the generator's output
cannot pass that required check until the generator change lands on main first. The
descriptions move to a follow-up that lands the generator change ahead of the schema
* test(bedrock): move the native chat completions tests under tests/unit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): drop reasoning_effort none for grok on the native chat completions route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): keep converse extension params on the converse route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): inline http image urls and keep stop on converse for native chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(bedrock): share the sync remote media inliner
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(image-handling): infer the image mime type when the server sends a generic content type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): stop sending aws_bedrock_project_id as OpenAI-Project on the runtime chat completions route
* feat(bedrock): make native chat completions an opt-in bedrock/chat_completions/ route
Bare Bedrock OpenAI and Grok model ids stay on Converse as on main. The
bedrock/chat_completions/<model> prefix opts a deployment into bedrock-runtime's
/openai/v1/chat/completions, and a request carrying a Converse-only param
still falls back to Converse. The cost map no longer decides the route.
* fix(bedrock): keep chat_completions/-prefixed deployments on the native Responses surface
* fix(bedrock): keep provider response headers on the runtime chat completions route
* feat(bedrock): serve gpt-5.6 and newer on runtime chat completions by default
Unprefixed bedrock/<gpt-5.6+> models whose cost-map row lists /v1/chat/completions
now route to the native OpenAI-compatible endpoint; converse/ pins Converse and
chat_completions/ still opts gpt-oss and Grok in. Guardrails, application inference
profile ARNs, and tools with reasoning keep falling back to Converse per request.
Hoist the remote-media url comprehension into a single-clause helper.
* fix(bedrock): refuse temperature and top_p natively on GPT 5.6 and newer like Converse does
AWS answers temperature and top_p with a 400 on the native Chat Completions endpoint for the GPT 5.6+ models, the same models whose Converse route already dropped both under drop_params via supports_sampling_params: false. The native config now honors that price-map flag, the gpt-6 and gpt-6.1 rows carry it, and the gpt-6 family joins gpt-5 in refusing frequency_penalty, presence_penalty, logprobs, and top_logprobs before the request reaches AWS.
* fix(bedrock): refuse GPT sampling and logprob params natively only while reasoning is on
On bedrock-runtime's native chat completions endpoint, GPT-5.x and GPT-6.x
accept temperature, top_p, frequency_penalty, presence_penalty, logprobs,
and top_logprobs once reasoning_effort is "none", and refuse them with any
other effort or when the effort is unset. The previous commit refused the
sampling params unconditionally from the cost map's supports_sampling_params
flag, which lost the reasoning-off case and never covered the penalties or
logprobs. The refusal now keys on the model being a GPT id and reasoning
being active, raises a 400 UnsupportedParamsError naming the params unless
drop_params drops them, and lets everything through under "none". Grok and
gpt-oss keep their unconditional family refusals.
* refactor(bedrock): keep the Converse route-prefix strip inside the bedrock llms module
* fix(bedrock): forward a non-string reasoning_effort on the native route instead of crashing
A list or dict reasoning_effort hit a frozenset membership test in
without_refused_reasoning_effort and raised TypeError, which the proxy
surfaced as a 500 APIConnectionError with no upstream call. The value is
now left alone unless it is a string Bedrock's native endpoint refuses,
so AWS answers the malformed value with its own 400 like it does for an int
* fix(bedrock): route overlong GPT version digits to Converse and send native chat completions to the runtime endpoint
A model id with more than 4300 version digits raised ValueError in the route check; the digits are now bounded so such ids fall back to Converse. The native chat completions URL now follows Converse's precedence: aws_bedrock_runtime_endpoint (or AWS_BEDROCK_RUNTIME_ENDPOINT) wins over api_base, so a deployment that sets both keeps sending to the same host
* fix(bedrock): route model_id overrides to Converse and never send an empty bearer natively
A deployment whose litellm_params carry model_id (an application inference profile or provisioned throughput ARN) went to the native Chat Completions route with the base model in the URL and model_id left in the body. It now takes Converse like the bedrock/arn:... model form, which encodes the override into the request URL
A blank api_key on a SigV4 deployment became an Authorization header reading Bearer with nothing after it on the native route, since the OpenAI-like header builder writes any non-None key and the signer keeps a non-AWS4 Authorization header. validate_environment now resolves the key through bedrock_bearer_token, so a blank key is signed with SigV4 the way Converse signs it
* test(bedrock): audit the native GPT chat completions route on the integration rig
Adds the /audit cells for the runtime chat completions route: the scripted Bedrock runtime peer, the happy and fallback wire tests, the sad-path and regex worst-case tests, the chaos burst tests, the Messages adapter tests, and the Responses native-route tests. Tests only, no product diff.
* test(bedrock): harden the runtime chat completions audit cells
The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row
* fix(bedrock): refuse or drop a non-string reasoning_effort before the native chat completions call
A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native
route serves now answers 400 from litellm before any wire request, naming the type and the
drop_params way out, and is dropped under drop_params so AWS applies its default effort, the
way Converse dropped it on main. The tip since a0cef91f0b forwarded it for AWS to refuse
* test(bedrock): pin router retries off and give the chaos bursts config deployments on an owned proxy
* test(bedrock): wait for the replacement worker before tearing down the sigkill chaos proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo <mateo@berri.ai>
* feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them
* fix(bedrock): rename the lookaround flag to supports_regex_lookaround and keep dropped patternProperties names allowed
The cost-map flag becomes a generic supports_regex_lookaround capability, which the
cost-map schema admits as a supports_* boolean, and the Converse transform now owns
the drop decision instead of the shared tools factory. A patternProperties key dropped
from an object closed by additionalProperties: false leaves its value schema as that
object's additionalProperties, so the names it allowed stay allowed, on the OpenAI
non-Python-regex drop too. tool_with_sanitized_parameters also cleans Anthropic-shape
tools (input_schema).
* fix(router): keep a deployment's supports_regex_lookaround off the shared cost-map entry
A deployment's model_info.supports_regex_lookaround was written to the shared
bedrock/<model> cost-map key, so every sibling deployment of that model id
inherited one deployment's choice. The flag now stays under the deployment's
own id, which is what the Converse lookaround check reads first, and the
shared entry keeps the cost map's value
* test(bedrock): audit the Converse lookaround drop on the wire across endpoints, SDKs, flags and chaos
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(responses): drop bridge-minted reasoning items from OpenAI replays
* fix(responses): send id-less stored reasoning items without a made-up id
The chat-to-Responses bridge gave a stored reasoning item with no id an rs_<n> id that OpenAI rejects (404 without encrypted content, 400 with it); an id-less item is accepted and verified by OpenAI itself. Decoding encrypted_content now keeps the verifiable thinking blocks of a mixed array instead of rejecting the whole array, and the verifiable-block rule lives in the shared module.
* test(integration): audit minted reasoning item replay across responses, chat bridge and messages
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(anthropic): stop repeating streamed thinking text in the signature chunk
The Anthropic stream handler emitted, on the signature_delta, a thinking block carrying every thinking delta seen so far plus the signature, after it had already streamed that text as per-delta blocks. Every additive consumer (the Agents SDK, stream_chunk_builder, the Responses bridge) stored the text twice under one signature and replayed the doubled block on the next turn
The signature chunk now carries a signature-only block, matching the Bedrock emitter, so accumulators rebuild the text once and the saved history replays to Anthropic exactly as it was streamed
* test(anthropic): audit the signature-only thinking chunk across wire providers
* test(anthropic): pin the cached replay's thinking shape in the wire audit
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat: add Bespoke Nimble gateway and OSS classifier support
* feat: accept Ollama's nimble model name for the Bespoke provider
* test: exempt the POST-only bespoke decisions route from the all-methods check
test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(exa): fall back to highlights/summary when text is missing (#36905)
* test: type exa search transformation test helpers
* test(exa): drop type: ignore in test helper and trim redundant comments
* test(exa): move snippet fallback test to tests/unit and parametrize it
Uses a real httpx.Response instead of a Mock and annotates snippet as Final.
* fix(exa): inline snippet fallback since Final is not allowed inside a loop
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* feat: add Laya gateway and classifier backend
* test: cover the pass-through model_group pin and repair the shard fakes
MockRequest in tests/pass_through_unit_tests gains an httpx.URL and an ASGI
scope, which get_request_route now reads inside
_init_kwargs_for_pass_through_endpoint, and the POST-only /laya/v1/systemone
route joins the protocol-constrained exemptions. A built-in pass-through pins
metadata.model_group to the resolved model so a client cannot choose its own
per-model budget key; test_pass_through_endpoints now proves that on a
non-Laya route and drops a duplicated assertion.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(chatgpt): preserve requested service tier in Responses calls
* fix(chatgpt): avoid extra mutable collections in tier filtering
* test(chatgpt): refresh service-tier cases after main rebase
Keep the tier-preservation regression on the current subscription model and match the moved unit-test suite's formatting. The adapter still forwards preferences without claiming backend priority entitlement.
* refactor(chatgpt): map service tier through a lookup table after the allowlist filter
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix(oci): resolve the GenAI endpoint realm from the region instead of hardcoding oraclecloud.com
Government (OC2/OC3/OC4) and other non-commercial realms live under a
different second-level domain, so a request for us-luke-1 was sent to
inference.generativeai.us-luke-1.oci.oraclecloud.com and failed DNS
Delegate the lookup to the OCI SDK's region registry when it is installed,
honour OCI_DEFAULT_REALM otherwise, and keep api_base as the explicit override
* fix(oci): read per-region realm metadata without the SDK and cover the registry path
Replace the global OCI_DEFAULT_REALM fallback, which would have redirected
commercial regions too in a mixed deployment, with the SDK's own per-region
sources: OCI_REGION_METADATA and ~/.oci/regions-config.json. Regions not
described anywhere keep their commercial endpoint
Exercise the SDK registry path with a fake oci.regions module so CI, which
has no SDK, still covers it, and skip the real-SDK test on oci.regions so a
namespace package named oci in the tests tree cannot masquerade as the SDK
* fix(oci): validate regions-config.json entries individually and tolerate undecodable files
One malformed entry no longer discards the valid ones, and the file is parsed
from bytes so an undecodable file is logged and ignored instead of failing
every OCI request built without the SDK
* fix(oci): resolve the realm from the compartment OCID so Government regions work without the SDK
The Docker image ships without the oci package and Government deployments rarely carry OCI_REGION_METADATA, so the reviewed fallback still sent us-luke-1 to oraclecloud.com. Every compartment OCID already names its realm (ocid1.compartment.oc2..), so map that key through the SDK's twenty realm domains first, then the metadata sources, then the SDK registry, then the commercial default.
* fix(oci): consult the SDK registry before hand-parsed region metadata and harden the fallback
Review follow-ups on the realm resolver. Read the compartment realm first, then the SDK registry when it is installed, and only then the hand-parsed metadata sources, so the same file is never parsed twice with different rules. Lowercase metadata values like the SDK does, accept single-label realm domains, and expand ~ with os.path so a container without a home directory cannot raise out of URL building. Type the compartment as str | None at the caller, drop populate_by_name, isolate the legacy region tests from the developer's ~/.oci, and keep the new tests on the immutable style.
* fix(bedrock): accept Converse messages with no content key
A user or tool message whose content key is missing (or null, which the
message cleanup strips) made every Bedrock Converse request fail with
APIConnectionError 'content' before reaching Bedrock. The Converse
transform now reads content with .get for those messages, as it already
did for assistant messages: a content-less user message adds no block
and a content-less tool message becomes a toolResult with empty content.
The str branch also sends the continue message text instead of the
original whitespace-only text.
* fix(bedrock): send the continue message for a content-less user turn
* test(bedrock): type the content-less Converse message test parameters
* fix(bedrock): accept a Converse system message with no content key
* refactor(bedrock): read the system message content with get
* test(integration): audit Bedrock Converse messages without content
Adds the /audit cells for a chat message whose content key is missing or
null on a Converse-routed Bedrock deployment: happy, sad, edge, and chaos
rows through the OpenAI SDK, the Anthropic SDK, and raw httpx against the
scripted upstream, asserting the caller's response, the body the peer
received, and the spend row. The owned-proxy readiness deadline in the
integration harness is now INTEGRATION_PROXY_READY_SECONDS (default 70).
* test(integration): bound stray spend rows in the mid-burst restart cell
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* chore(lint): remove the LIT002 mutable-construction rule
Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.
Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.
* chore(lint): keep the leftover mutable-ok markers for a follow-up
Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.
* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"
This reverts commit c35bc0b84e.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden
#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key
* test(ui): give the auto-router threshold save wait room for the availability debounce
#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed
* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default
#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against
* test(e2e-ui): wait for the call-id search before hovering the logs row
The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time
* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off
The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support
* test(integration): read the agent 365 guardrail status by its own name in spend logs
The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict
* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check
gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler
* test(e2e-ui): fill the create-tag fields inside the dialog
#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step
* test(integration): run integration proxies with the CI license
Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments
* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5
Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving
codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI
* test(unit): join the session-minting thread before collecting the handler
asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference
* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early
owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.
uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds
* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build
The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there
* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections
The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry
* ci: move the unit-test uv cache split into a composite action
check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes
* test(unit): point tiktoken at the bundled cache for every unit test
The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.
* test(integration): answer model discovery probes in the hosted_vllm wire tests
The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): narrow hiddenlayer startup auth timeout without cast
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): bound hiddenlayer jwt refresh by configured timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover model_armor and run timeout probes concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): match sink calls to the exact guardrail name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(transcription): honor base_url alias for Groq Whisper and report it as the api base
transcription() and speech() only accepted api_base, so a deployment configured with base_url leaked the alias into the provider params, which Groq rejected as an unknown param, and the request never reached the internal gateway. get_api_base() now reads the same alias so response headers and logs show the configured endpoint instead of the provider default
Resolves LIT-9071
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(speech): keep base_url after existing audio params, route Vertex speech to it, skip empty alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages
vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(hosted_vllm): forward replayed reasoning_content only when it is a string
* test(integration): cover hosted_vllm reasoning_content replay across endpoints
* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* fix(responses): scan and mask top-level instructions with guardrails
The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.
Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.
Resolves LIT-8931
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): reject empty guardrail rewrites instead of forwarding raw input
An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the texts-replacing guardrail helper explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): cover skip_system_message_in_guardrail on the live proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system
Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): annotate new guardrail tests with return types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the guardrail test doubles explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): satisfy type-discipline and strict ruff budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stamp the served service_tier on every Responses bridge chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): cover anthropic and responses served-tier billing paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): bill disconnects through the router's anthropic stream wrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: apply ruff format to the anthropic stream changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(coverage): ignore delegating properties the ast scan cannot see
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: keep the cast-ok reasons on the cast call line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(streaming): parameterize delegated chunks and messages types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): follow the anthropic pass_through rename after merging main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drain the logging worker between response cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): keep the served service_tier on streamed chunks and bill it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): type the served service_tier chunk without a loose kwargs dict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill the served service_tier over the requested one
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): drop explanatory comment from the tier resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over
The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.
Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): mirror every dispatcher fallback path in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): skip already-tried order levels in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(router): credit the #39566 branch this fix supersedes
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.
Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): drive the router request test through an httpx MockTransport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): assert tool definitions reach the router upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(providers): add Prism provider
* fix(providers): complete Prism registration
* feat(providers): expose Prism responses and messages
* feat(providers): add DeepSeek V4.1 Flash to Prism
* test(providers): exercise Prism endpoint requests
* fix(providers): align Prism pricing and limits with the live catalog
deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models
* test(prism): assert cost-map invariants instead of pinning catalog facts
* test(prism): derive the asserted model list from the cost map instead of pinning it
* test(prism): capture requests through respx instead of appending to a list and swapping the client transport
---------
Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit bc9b6f8a5c)
* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard
PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vertex_ai): reproduce traced Gemma Responses stream failure
* fix(vertex_ai): wrap Gemma fake streams for Responses tracing
* test(vertex_ai): cover Gemma traced streams and usage options
* test(vertex_ai): inject gemma test deps and assert hidden usage accounting
Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.
Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.
* test(vertex_ai): drop explanatory comment from usage-option assertions
* fix(gemini): forward seed to the Gemini API instead of rejecting it
The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough
* test(gemini): assert the forwarded seed without mutating shared state
* fix(vertex_ai): consider tools when validating context caching min tokens
Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.
Fixes#42804
* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
* fix(openai): exclude fine-tuned and custom gpt-5-chat aliases from gpt-5 reasoning path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): keep gpt-5-chat alias regression test diff minimal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): cover temperature pass-through for gpt-5-chat aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(openai): annotate locals and wrap long lines in gpt-5-chat alias test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(params): validate stream_chunk_size once and carry it as typed control options
Checks stream_chunk_size at the top of completion() and acompletion(), accepts
digit strings, returns a 400 naming the param unless drop_params is set, and
stores the checked value under _litellm_control. Bedrock Converse and Invoke
read it from litellm_params; the Bedrock-only checker and the dead Invoke pops
are gone. Owned-kwarg filtering now runs through one helper everywhere.
Refs LIT-8317
* test(bedrock): drop tests for the removed stream_chunk_size_from helper
Refs LIT-8317
* fix(params): check stream_chunk_size before the MCP gateway branch
Refs LIT-8317
* fix(params): return assert_never in the exhaustive control-options match
Refs LIT-8317
* fix(params): address council review of the control options change
Read all_litellm_params live so names registered after import stay
LiteLLM-owned, make litellm_params a required keyword on the stream
wrapper hooks, give digit strings and ints the same 18-digit range,
share the default-chunking test table, test the Responses bridge through
litellm.responses, and revert formatting-only churn in existing tests.
Refs LIT-8317
* fix(params): address the second council review of control options
Keep the Responses bridge on its original all_litellm_params forwarding,
narrow _int_from_decimal_string inline so it type-checks, bound nested
huge ints in the error message, store _litellm_control only when a value
is set, simplify the parser to its single field, drop the one-caller
wrapper, and tighten the tests.
Refs LIT-8317
* fix(params): keep the 18-digit length check on stream_chunk_size strings
A 19-character string with leading zeros such as 0000000000000000001 would
otherwise pass as 1, although the rule and the error message say at most
18 digits.
Refs LIT-8317
* test(params): tidy control options tests after council sign-off
Move the Responses bridge test into the existing bridge test file, drop the
rebind test that pinned an implementation detail, assert through
stored_control_options instead of the storage key, and cover
drop_params="true" through Bedrock streaming.
Refs LIT-8317
* test(params): wrap a chunking test row that went past 120 characters
Refs LIT-8317
Kwargs LiteLLM code introduces for its own use were only kept out of provider
bodies if someone also listed them in all_litellm_params. Undeclared ones went
into extra_body or optional_params, reached the provider, and the provider
rejected the request. is_litellm_owned_kwarg in types/utils.py now defines
LiteLLM-owned once: a registered name, or any name starting with
INTERNAL_KWARG_PREFIX from litellm/constants.py. Every filter that builds
provider params from kwargs uses it: chat completion, transcription,
embedding, image generation and edit, search and video, ElevenLabs text to
speech, and the Bedrock batch mapper. The two untyped shared filters now take
Mapping[str, object]
The stream_chunk_size wire test becomes test_internal_params_wire.py. It also
sends an undeclared _litellm_ kwarg and asserts that no _litellm_ key reaches
any of the six provider bodies, while extra_body passthrough keeps working
Refs LIT-8318, LIT-8319
Register Sail (providers.json, LlmProviders.SAIL, OpenAI-compatible lists,
ProviderConfigManager) for chat, Responses and /v1/messages, and add its 12
models to both cost maps with asap, balanced and flex price columns.
Sail picks speed and price with metadata.completion_window and rejects
service_tier, so the Sail chat and Responses configs translate the tier:
default and priority to asap, flex to flex, balanced to balanced, auto to no
window. Billing prices the window that was sent. A tier Sail has no window
for, or a window or tier set where billing cannot see it (request metadata,
extra_body), is a 400 unless drop_params is set.
Add balanced to ServiceTier and its _balanced price columns to the model
info types, the Rust catalog and the dashboard schema. A transform_extra_body
hook on the chat and Responses base configs, which returns extra_body
unchanged by default, lets Sail keep the window when a caller also sends
extra_body.metadata. Sail is listed in the Add Model form and model picker.
Co-authored-by: shrey kharbanda <shrey@berri.ai>
* feat(proxy): add fail_closed_rate_limit_enforcement to reject requests with 503 while Redis rate limit counters are unreachable
* fix(proxy): reject fail-closed rate limit checks before logging the in-memory fallback and pin the boot warning in the lifespan
* fix(proxy): coerce the fail-closed flag, fail closed on read-only checks, and refund partial cluster increments
* fix(proxy): window-guard rate limit refunds and catch the fail-closed rejection by type
* fix(proxy): read the compaction rate-limit gate's limiter from the proxy hook registry
* fix(proxy): count the pending request in read-only rate-limit checks and keep the compaction gate off the caller's parallel slot
The compaction polyfill's summary-model gate, once it ran against the real v3 limiter, showed two behaviors nobody had chosen. The read-only check compared the stored counter with the same `>` the increment path uses, but a read-only check decides a request that has not been counted yet, so a summary model exactly at its rpm limit still went out. The read-only path now adds the pending increment of 1 before comparing; the increment path is unchanged.
The gate also passed the key's max_parallel_requests gauge through, and the read-only gauge count includes the caller's own in-flight slot, so a key with max_parallel_requests: 1 never compacted. The gate now drops that gauge from its descriptors, since the summary call runs inside a request the limiter already admitted.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(otel): detach post-response service spans by request phase, name redis spans by operation
Service spans logged from the post-response phase (success callbacks, the response-cache write) now root their own trace linked to the request span even while the server span is still recording, instead of only when they happen to end after it. Redis service spans are named `redis <operation>`; the litellm call chain that issued them moves to the `litellm.service.caller` attribute via a typed `ServiceLoggerPayload.caller` field.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): keep the service caller on failure and legacy spans, test the production phase dispatch sites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): mark anthropic messages stream cache write as post-response phase
The /v1/messages streaming cache writer awaits async_add_cache inline
instead of going through create_cache_write_task, so its redis span
stayed parented under the request trace. Wrap the write in
post_response_phase so it detaches like the chat completions write.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic): write the Messages stream cache in a background task after handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): stop five stale or flaky CI reds and retry CyberArk policy-load conflicts
The Langfuse redaction unit test exports to a local OTLP capture instead of
polling Langfuse Cloud through a recorded lookup. The passthrough worker-kill
test only requires spend rows for requests the surviving worker served. The
spend-routes sweep treats the intentional /spend/capture_rate 503 as expected.
CyberArk retries a 409 policy load in Python, Rust and the e2e Conjur helper
instead of reading it as "variable exists". The integration egress guard now
matches the script's own cgroup, so it no longer blocks the CircleCI agent,
which runs as the same user.
* fix(ci): keep the policy-load backoff typed as float
* fix(ci): retry CyberArk policy loads without blocking the event loop and tighten the worker-kill and Langfuse tests
* fix(secrets): load CyberArk policy one request at a time per manager
* test(secrets): pin that non-conflict CyberArk policy failures are not retried
* test(unit): run tests/unit with only an allowlisted host environment
CircleCI's unit job inherits every project env var, so real provider keys,
REDIS_HOST, DATABASE_URL and AWS or Azure credentials reached tests that
assume none are set. Locally, litellm's import-time load_dotenv did the same
from any .env up the tree. The unit conftest now drops every variable outside
a small allowlist and disables dotenv before litellm is imported.
* test(e2e): name a failed search and the stuck batch status instead of misattributing them
The websearch session test read an empty web_search_tool_result_error block as a
successful search, so a failing search tool surfaced as a session billing bug.
The batch cancellation timeout now reports the last status the proxy returned.
* fix(ci): scrub the host environment per unit test instead of for the whole pytest process
GHA shards run tests/unit next to other suites in one process, so the import-time
scrub deleted MCP_TEST_PEER_PYTHON before tests/mcp_tests read it and the MCP
upstream fell back to the SDK2 interpreter. The two websearch tests that called
OpenAI and Perplexity live are removed: tests/unit no longer sees their keys.
* fix(ci): scrub only the host variables present before litellm is imported
The per-test scrub also deleted TIKTOKEN_CACHE_DIR, which litellm sets at import to
its bundled encodings, so tokenizer paths tried to download them and hit the
socket guard. The prisma setup test now passes its own database URL instead of
reading one another test leaked into the process environment.
* fix(ci): stop the order-dependent unit reds and settle logging tasks on their own queue
LoggingWorker marked a task done on whichever queue was current when the callback
finished, so a callback that outlived an event-loop change raised "task_done()
called too many times" or undercounted the new loop's queue. It now settles the
queue the task came from.
The rest are test isolation fixes for failures that only appeared when another
file ran first on the same xdist worker: a replaced user_api_key_cache, breaker
metrics unregistered by prometheus tests, semantic_router's health-check filter on
uvicorn.access, logging tasks carried over from bedrock tests, a Router-written
model_cost entry, and a stray post captured by the langflow test. The token
counter check now asserts bounded chunking instead of wall-clock time.
* test(e2e/ui): wait for the logout redirect before visiting a protected page
Logout revokes the session server-side before clearing cookies and navigating, so an immediate page.goto either ran with the cookie still set or was aborted by the logout redirect (net::ERR_ABORTED).
* test(unit): restore the prometheus metrics config per test and settle logs carried from earlier tests in the a2a cost tests
* test(router): pin the router clock in the usage counter tests so a minute rollover cannot empty the read
* test(e2e/ui): wait for logout to clear the token cookie instead of for a login redirect
* test(integration/mcp): answer the model-info probe another test's proxy sends to the model double