* fix(cli): pin pi compat flags for gateway models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): annotate pi compatibility field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): clarify pi compatibility suppression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): drop the stale mutable-ok suppression on the pi compat block
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(router): drop the encrypted reasoning a fallback hop's target cannot decrypt
An order-based or configured fallback hop replayed the failed provider's
encrypted reasoning items to the next deployment, which answered 400
(Bedrock Mantle: invalid encrypted reasoning; OpenAI:
invalid_encrypted_content), so every multi-turn Responses fallback for
Codex-style clients failed. The hop now drops the encrypted reasoning
its target cannot decrypt and keeps each item's readable summary. With
encrypted_content_affinity on, the pin narrows to the hop's target
order instead of emptying it, so the hop reaches the next order
instead of failing with no deployments available.
* fix(router): keep encrypted reasoning a same-boundary fallback hop can decrypt
Unmarked encrypted reasoning on a hop is attributed to the deployment that just
failed, read from the retry breadcrumb, so a hop to a deployment on the same
api_base and api_key keeps it and a cross-provider hop still drops it. The hop
tests script the upstream at the httpx boundary instead of doubling the handler,
and the router coverage script lists the two hop helpers with their tests
* fix(router): read the hop's failed deployment from its own metadata bucket and carry it into the Responses mid-stream snapshot
* test(router): use a real Router without the origin deployment in the hop strip test
* test(router): cover the hop strip when no failed deployment is known
* fix(proxy): drop the router's fallback hop state keys from the client body
* fix(proxy): keep a request's max_fallbacks cap, drop only the hop state keys
* test(integration): audit cells for the fallback hop encrypted reasoning strip
Forty-five checked-in cells under tests/integration/routing prove the hop strips the previous deployment's encrypted reasoning on /v1/responses, /v1/chat/completions and /v1/messages (httpx, OpenAI and Anthropic SDKs, sync and async, streaming and not), that the affinity pin yields to the hop, that a client-sent fallback_depth, _target_order and attempted_targets never move or strip a request, and that a concurrent burst, an order-1 outage and a killed worker keep every request stripped and logged once. Every call goes through a lane pinned to one worker that already lists the deployments it needs, because the peer worker learns a /model/new row through the config-sync resync up to sixteen seconds later
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(guardrails): support logging_only mode for the Akto guardrail
* test(guardrails): wait for the spend row before reading Akto calls and cover mixed modes
* test(guardrails): cover Akto logging_only on MCP tool calls
* test(guardrails): type the Akto logging_only unit tests and inject the HTTP handler
* test(guardrails): cover failing Akto replies, provider failures and a mixed outage burst under logging_only
* test(guardrails): cover an unreachable Akto with fail_open under logging_only
* ci(unit): add unit passed collector job and fold proxy-db shards into test-unit.yml
* test(ci): cover the unit passed gate's success, failure, cancelled and skipped results
* test(ci): run the unit passed gate test from the code quality workflow
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* refactor(python-bridge): ship a signature base and read the resolved call
NativeCall carries base (positionals by name plus signature defaults) instead
of the fully bound dict. resolved lays kwargs over base, which is what bound
held, so every pre-hook read and every route host keeps seeing the same values.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bedrock): build the transcription NativeCall with an empty base
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(dispatch): read the resolved call instead of bound
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(host-python): rename effective to effective_py_args and note the shallow copy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
#44800 re-raises the provider's own error from the chat adapter, so a bridged /v1/messages stream's error frame now reads litellm.RateLimitError without the litellm.MidStreamFallbackError: prefix that #44989's tests pinned shortly before #44800 merged. The three tests now expect the provider error text once and no sentinel, matching the chat_limited case and every other assertion in the two files
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(bedrock): split the <reasoning> tag for gpt-oss only on native Chat Completions
The native Chat Completions route moved a leading <reasoning>...</reasoning>
block into reasoning_content for every model, while only gpt-oss writes its
reasoning inline. A GPT 5.6 or Grok answer that starts with a literal
<reasoning> tag lost that text, streaming and non-streaming alike. Both paths
now split only when the model id is gpt-oss.
* fix(bedrock): drive the inline reasoning split from a cost-map flag
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test: replace 61 live logging and otel tests with offline unit and integration coverage
* test: restore datadog formatting, deliver datadog logs and redis failures over the wire, assert full router hook payloads, tighten stream usage and otel checks
Restores the pre-existing test_datadog.py lines the branch had reflowed. Datadog success, failure and redis-failure replacements now assert the gzip body posted to the intake, with redis failing through a real cache on a closed local port. Router hook sequence recorder checks the legacy field types and asserts concrete payloads, exact streaming and fallback sequences. Stream usage asserts the default include_usage request body and the redacted messages value. Otel asserts response id and token counts.
* test: assert budget envelope figures, guardrail inspection and exact prometheus samples per request
* test: isolate the prometheus latency test on its own deployment so counts do not depend on order
* test: cover timed slack delivery, keep the redis failure test off the network, and make new payload types read-only
* test: drain the redis test's logging and scope otel span checks to the test's own trace
* test: drive the periodic slack flush without a wall-clock interval
* test: scope router hook events to the test, script the db clock, split a nested comprehension
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* test(ui-e2e): update usage page selectors after the #45221 redesign
* test(ui-e2e): drop the top keys locator comment
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(otel): emit OpenInference tool calls and metadata on Arize OTel v2 spans
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): shed OpenInference output tool calls individually under the span attribute budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): avoid mutation in Arize OTel v2 integration helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): audit Arize OTel v2 OpenInference spans across endpoints, modes and chaos
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): wait for each exported span before the next request in Arize OTel v2 audit tests
* test(otel): make Arize OTel v2 audit absence and outage checks deterministic
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): collect Arize OTel v2 outage spans through an in-order sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): add skipped BUG cells for pre-existing Arize OTel v2 gaps
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): repair Arize OTel v2 regressions from #43698 (linear fit, metadata slot, repr tool args) (#44488)
* fix(otel): keep Arize OTel v2 regressions in check — O(n) fit, metadata slot, repr tool args
Three regressions from #43698's OpenInference tool-call/metadata emission:
1. Metadata evicted indexed message attributes: the new `metadata` key
competed for the 128-attribute span budget, and the fit sheds whole
message groups BEFORE `span.set_attribute`, so the SDK's dropped
counter stayed 0 — an invisible eviction (live A/B: input-message
attributes 86 -> 84). Two-part fix: the fit pins `metadata` behind
every message group (it sheds only once all indexed messages are
gone), and the span budget no longer charges pre-set attributes the
mappers overwrite in place — a boundary-opened LLM span already
carries keys like `gen_ai.request.model`, so the old accounting
reserved slots the fit could never spend. Live: 86 input-message
attributes with the metadata attribute riding alongside.
2. Quadratic shed on long prompts: `_message_shed_groups` rescanned the
full group map once per message (measured on a real acompletion:
0.032/0.128/0.478/1.910s at 1000/2000/4000/8000 messages vs
0.007/0.010/0.022/0.029s at base). Index the tool-call groups once by
(family, message index): the fit is linear again (0.008/0.007/0.014/
0.031s, same rig).
3. Malformed Python tool arguments lost the whole span: provider
adapters and `model_construct` responses hand over raw objects, and
`json.dumps` raises on tuple-keyed dicts (TypeError) and cycles
(ValueError) before the span is exported. Serialize with a repr
fallback; both cases now export with a readable arguments attribute.
The attribute budget change affects every boundary-opened LLM-call span
(strictly more attributes retained, never fewer); the mapper changes
only touch the OpenInference vocabulary.
* refactor(otel): build the tool-call group index in one shot
Review follow-up: the dict.setdefault/append seeding in
_tool_call_groups_by_message violated the no-mutation coding convention
(AGENTS.md: build values in one shot with comprehensions or generators
wrapped in tuple()/MappingProxyType()). Rebuild it as a sorted groupby
comprehension; randomized parity harness confirms the shed order is
byte-identical to the seeded version (400 trials).
Also pin the overflow corner Greptile asked about: a pre-set
indexed-message key the fit sheds keeps its earlier value in place, so
the span total can never exceed the SDK limit (new emitter test).
* fix(otel): key the groupby with an explicit tuple to keep basedpyright at budget
The slice-keyed groupby (group[:2]) widened the key to tuple[str | int],
adding one reportGeneralTypeIssues over the codebase ceiling. Key by the
explicit (family, message index) pair instead; shed order unchanged
(300-trial randomized parity harness).
* fix(otel): read pre-set span keys through a helper typed for both runtime shapes
The SDK annotates ReadableSpan.attributes as a Mapping, but an ended span
hands back a tuple of pairs, so the inline isinstance branch narrowed to
Never and pushed reportGeneralTypeIssues one over the codebase ceiling.
Extract _carried_keys with the runtime union declared on the parameter;
behavior unchanged.
* test(otel): drop the wall-clock bound from the long-prompt attribute fit test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): annotate to_openai_dict as Mapping after the rebase onto main
Main's type-discipline budget tightened since the branch point; the plain
dict return annotation is the one violation the rebased branch adds.
Callers only serialize the result, so the read-only view is accurate.
* chore(otel): drop the restating docstring from to_openai_dict
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): accept team-scoped models by their public name on POST /fallback
create_fallback validated the primary and fallback models against the
router's stored model names only, so a model created through POST
/model/new with model_info.team_id was rejected with a 404 unless the
caller used the generated model_name_<team_id>_<uuid> name. The
endpoint now also accepts the team public model names, which request
time fallback matching already keys on, and lists both kinds of names
in the 404's available_models.
* fix(proxy): read fallback rules fresh and clear the config cache after a fallback write
A second POST or DELETE /fallback within the 60 s config cache TTL started from
a cached copy of router_settings and dropped every rule stored since that copy
was taken, by any instance. Both endpoints now evict the cached row before the
read and invalidate it after the upsert
* test(proxy): type the fallback endpoint tests and prove a team request fails over by public name
* fix(proxy): keep fallback writes working when the Redis config cache is down and type the stored settings read
* fix(proxy): keep the non-standard fallback shapes the router accepts on writes and resolve the config cache at call time
* test(proxy): mark the config cache outage test's result as Final
* fix(proxy): replace a same-key fallback rule in place and read a null rule list as empty
* test(proxy): let the stored router settings fixture carry a null rule list
* test(proxy): audit cells for fallback rules by team public name
* test(proxy): delete the fallback rules the audit cells save
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): link model access group chips to the access group filter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format access group chip link changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): mock access group hook in affected suites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep access group chips unlinked until the group lookup resolves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep cached access group names when a refetch fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): show the Add Model picker once the model catalog loads after a provider is picked
* fix(ui): show the Add Model picker when a name typed before the catalog loaded was cleared
* feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs
Adds general_settings.spend_logs_metadata_fields with mutually exclusive include and exclude lists. The filter runs on a copy of the row right before it is queued for Postgres, so daily spend rollups, budgets and callbacks still see every key.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): keep excluded auto-router savings keys out of published spend log metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): cover metadata retention across endpoints, failures, cache hits, batches, runtime updates and auto-router publication
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for spend_logs_metadata_fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): filter metadata at DB write so guardrail usage sees full rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): poll guardrail daily metrics instead of reading once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): drop timeout comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): read spend_logs_metadata_fields through typed general settings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
#45214 routed both decision formats through the shared decisions IR in
litellm/llms/base_llm/decisions/transformation.py and litellm/types/decisions.py.
That left litellm/llms/base_llm/decisions/systemone.py (to_system_one_request,
question_keys, to_decisions_response, the SystemOne* models and their adapter)
and litellm/types/openai_decisions.py with no importer outside their own unit
tests, so both modules and both test files go.
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(rust): derive strum VariantArray and string conversions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): use strum conversions directly with explicit spellings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): use rstest values for key management systems
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(python-bridge): take NativeCall directly and fold routes into per-route folders
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(python-bridge): use rstest for updated tests and keep embedding's stub parameter name
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(guardrails): run each mode only on streaming, non-streaming, or both
Add stream_scope so a rail can target streaming inference, non-streaming inference, or both per pre, during, and post mode
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): honor stream_scope in pipelines and dashboard types
Pipeline steps skipped the stream_scope filter, direct construction ignored mixed-case maps, and schema.d.ts was stale.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): skip unmatched stream_scope steps instead of allowing
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ui): keep stored stream_scope keys for modes not on screen
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ci): format stream_scope helpers and update fork MCP unit tests
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(mcp): keep hang cancellation tests from timing out during setup
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): honor path-defined streaming for stream_scope
Passthrough routes like Gemini streamGenerateContent decide streaming from the URL, so stamp that onto hook data before guardrails run.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(rust): copy AnthropicModelCapabilities instead of cloning
Clippy treats clone-on-Copy as an error, which failed rust-lint on the messages request tests.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): trust only server stream classification
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): keep streaming marker through deepcopy
scan_raw_request snapshots copy each field, so a plain object() marker would lose identity and skip streaming-only rails.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(cost-map): drop duplicate perceptron-mk1.5 row
Two main cost-map PRs both added the OpenRouter model, so the merge left a second key that CI rejects.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(rust): expect native transcription 429 as RustUpstreamError
HTTP status errors from native routes map through route_error_to_pyerr, so the wheel SIGINT child was dying on an outdated RuntimeError check and never reached the hang probe.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(tests): follow Google Interactions OpenAPI without hardcoded names
The live spec dropped CreateModelInteractionParams and renamed the item path to {interactionsId}. Misc CI failed because the compliance tests still looked those names up as literals.
* test(guardrails): reproduce stream_scope bedrock and passthrough field gaps
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover invalid stored stream_scope reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): classify bedrock stream actions and keep caller is_streaming_request
Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): make stream_scope_allows public and drop mutable builds
Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): classify pass-through and Bedrock stream scopes
Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): harden stream classification and validation
Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): add stream scope integration audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): avoid mutating passthrough custom body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): run stream scope audit without enterprise license
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): keep stored scope restart cell on one worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): scope stream scope audit sink assertions to the rail under test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert logging_only scope absence behind an ordered barrier rail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): assert one logging_only scan per phase after the barrier rail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): set request scope on pass-through stream fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate invalid YAML stream_scope in v1 guardrails list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: merge main into litellm_guardrail_stream_scope_fixes
Update pass-through pre-call test callbacks for main's endpoint_type argument
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): add request paths to pass-through fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover websocket pass-through stream scope and LIT-9050 outage spend rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(lint): remove unused type discipline suppressions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): script the vertex live upstream in the websocket stream scope test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep stream marker json-serializable and restore pass-through helper names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lint): allow required Bedrock action re-export
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): strip the stream marker from pass-through payloads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): drop stream marker by value in scans and snapshots, plain-tuple stream scope state
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): render tag-scoped guardrail modes read-only in the custom code editor
A guardrail whose litellm_params.mode is the tag-scoped dict {tags, default}
crashed the Custom Code editor on open: normalizeMode wrapped the dict into
the mode array and StreamScopeFields rendered it as a React child (error #31,
whole dashboard unmounted). Treat a non-string non-array mode as no editable
modes, show formatGuardrailMode(mode) in a disabled input (read-only, matching
the guardrail info view), and keep mode/stream_scope out of the update payload
for such guardrails.
* fix(ui): resolve merge fallout in guardrails components
Deduplicate toModeArray import after the merge, and move the read-only
guardrail details block back into GuardrailReadOnlyDetails (now rendering
the shared mode/logging-only rows plus the stream-scope detail) so
guardrail_info.tsx stays under the 800-line lint budget. Guardrails UI
suite: 298 passed.
* fix(ui): drop duplicate toModeArray import reintroduced by merge
* chore(pass-through): document the deliberate in-place marker strip as mutable-ok
The clear/update on _parsed_body is load-bearing: rebinding to a fresh
mapping instead breaks 75 pass-through tests because the marker-free body
must propagate through the caller's request dict so downstream guardrail
scans and snapshots never observe the server streaming marker.
* fix(guardrails): move stream_scope after timeout in CustomGuardrail init
Inserting stream_scope before the existing timeout parameter shifted the
positional slot of timeout, so positional callers constructed with their
timeout bound to stream_scope (ValueError) and timeout silently None.
Restores the base parameter order; keyword callers are unaffected.
* fix(guardrails): typing pass for the lint gates
stream_scope leaves the declared constructor parameters (restoring the
base positional surface; keyword construction unchanged), unknown config
values crossing the new stream-scope code paths get typed locals or
cast-ok boundaries, and the passthrough payload literals are annotated.
All three lint gates pass against current main; the guardrail suites are
unchanged (236+5019 passing; one known anyio-driver failure pre-existing
on base).
* style: sort cast imports for the ruff gate
* chore(pass-through): drop unused BEDROCK_STREAMING_ACTIONS re-export
The streaming check now uses is_bedrock_streaming_endpoint; nothing in
the repo imports the name from this module.
---------
Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: gabriele <gabriele@berri.ai>
* feat(rust): add the Anthropic beta header policy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): resolve known betas held as Other in AnthropicBeta::on
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): use named rstest cases for Other beta resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): make the rstest named-case rule explicit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move BetaPolicy tests to the public API test crate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(gemini): accept file content blocks with video_metadata on multimodal embeddings
* test(gemini): annotate the embedding mock helper and keep the type comment short
* fix(gemini-embeddings): keep file blocks through a partial embedding cache hit and validate video_metadata strictly
* test(caching): type the partial-hit cache test and pin its clock
* fix(gemini-embeddings): accept the Files API URI /v1/files returns as a file block's file_id
* fix(gemini-embeddings): 400 on bad input shapes, drop unknown block keys under drop_params, cache multi-block inputs
A bare object `input` answered 500 from both the caching handler and the
transformation; both now answer 400. An unresolved `files/` reference on
vertex_ai/ answered a ValueError 500; it now answers 400 naming the gemini/
provider. An empty `format` passed through to the provider; it now answers 400
naming file.format. Unknown block keys (`detail`, an unknown video_metadata
key, a top-level block key) are dropped under drop_params, global or
per-request, and still answer 400 without it. The embedding cache counted
file blocks against `max_messages`, so a request with 5 or more blocks was
never cached; file blocks no longer count.
* fix(caching): answer 400 for an object embedding input on the cache lookup
* test(gemini-embeddings): annotate the new embedding tests with return types and Final
* test(gemini-embeddings): annotate the transformation and files batch tests with return types and Final
* test(gemini-embeddings): audit file content blocks on the wire, in the cache, in batch uploads and under chaos
Integration cells for gemini/ and vertex_ai/ embedding file blocks with video_metadata: the batchEmbedContents and embedContent wire shapes, format overrides, nested and repeated inputs, every malformed block answering 400 before any provider call, drop_params at the request, deployment and YAML levels, the OpenAI SDK sync and async clients, Redis cache fills, hits, partial hits and the metadata in the key with zero-spend hit rows, chat, responses and messages controls, the vertex_ai/ batch JSONL upload, a mixed fast/slow/malformed/dropped burst and a worker kill. The wire peer now reads chunked request bodies, which the streamed GCS media upload sends
* test(integration): treat an upload the client abandons before its terminating chunk as a disconnect, never a stored request
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(voyage): rebrand to VoyageAI by MongoDB, refresh model list, support both contextual input shapes
* fix(voyage): send auto-chunking params for flat contextual inputs
A flat list[str] or bare str for the contextualized embeddings endpoint is
only valid as documents with enable_auto_chunking=True and input_type=document,
or as queries with input_type=query. Default those params for non-query flat
inputs so the request matches the live API contract, while letting caller-set
values win. Nested list[list[str]] still passes through unchanged.
Tests now assert the auto-chunking params instead of only echoing inputs.
* feat(voyage): route MongoDB-issued keys to ai.mongodb.com
Mirror voyageai.util.get_default_base_url from the official SDK: a key with
the `al-` prefix is issued by MongoDB and is only valid on ai.mongodb.com,
every other key on api.voyageai.com. The choice now lives in one shared
helper that the embedding, contextual, multimodal, and rerank configs all
use for both the default base URL and the Authorization header, so the host
and the key always agree.
Also drops the voyage-4-nano and voyage-multilingual-2 model map entries:
voyage-4-nano is not served on the Voyage API and voyage-multilingual-2 is
an older model, so neither belongs in this change.
* fix(voyage): restore MongoDB routing on contextual endpoint and re-add voyage-4-nano after upstream merge
The upstream merge replaced the contextual config's get_complete_url with a
Voyage-only host and dropped voyage-4-nano from the model map. Reapply the
al- key routing via get_default_base_url and add voyage-4-nano (open-weight,
per docs.voyageai.com) so the branch keeps its task changes on top of upstream.
* fix(voyage): drop voyage-4-nano and the contextual tests upstream already covers
voyage-4-nano is not served on the Voyage API, so it does not belong in the
model map. The contextual input tests duplicate
tests/test_litellm/llms/voyage/test_voyage_contextual_embedding.py, which
landed upstream with the auto-chunking fix this branch was carrying.
* test(voyage): drop sys.path.insert from the voyage common utils test
* fix(voyage): route rerank on the key it authenticates with
get_complete_url picked the rerank host from the environment while the auth
header carried the request key, so an explicit MongoDB-issued key was posted
to api.voyageai.com. validate_environment now hands the resolved key to the
config instance and get_complete_url reads it back, so the host and the
credential always come from one key. The config is built per request, so
nothing carries over between them.
* test: allow ultrafast pricing keys in model prices schema
* test: drop ultrafast schema keys now added upstream
* test(integration): prove voyage key-based host routing on the wire
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(model_prices): add Mistral Large 4 and its OpenRouter route
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): sync Gemini API shutdown dates for TTS, Live, image and Omni previews
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): correct Claude Haiku 5.5 flags, add Anthropic web search flags and OpenRouter Claude Haiku 5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): add Bedrock flex tier prices and clear lapsed Together DeepSeek V4 dates
Absorbs #45320 (Bedrock flex prices, verified against the AWS Price List API)
and the DeepSeek V4 date removals from #45307 (verified against Together docs and API)
Co-authored-by: Roi <roi@roitev.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): Mistral Large 4 context, OpenRouter Claude latest aliases and batch pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mistral): expect 1048576 context for Mistral Large 4
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: Roi <roi@roitev.com>
* feat(fireworks_ai): forward the LiteLLM user id as user behind fireworks_forward_user_id
* fix(fireworks_ai): keep the forwarded user id when the caller sends extra_body.user
* fix(vertex_ai): forward the inline-tools-2026-09-15 beta to Vertex and Anthropic
* test(vertex_ai): trim the inline-tools beta test docstrings and type their locals
* test(vertex_ai): add audit cells for inline-tools beta forwarding on Vertex Anthropic
* test(vertex_ai): send legal whitespace in the hostile beta header audit cell
* test(vertex_ai): close the SDK clients in the inline-tools audit cells
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): retry lock-timed-out daily spend batches in place so the shutdown flush keeps them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): annotate the lock-timeout retry test and carry its context in assert messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): mark the remaining lock-timeout test local Final
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): audit the websocket rejection log on every live route
* test(proxy): scan the stopped proxy log for the websocket crash and drop the uvicorn wording pin
* fix(proxy): use budget reset window for projected spend alerts
* style(proxy): use PEP 585 tuple annotations in projection helpers
* fix(proxy): derive projection date from budget reset timezone
* fix(proxy): project spend in real time within the budget reset window
* fix(proxy): derive the projection window start from the reset schedule
30d and monthly keys reset on the 1st of the month and unrecognized spellings like 1hr reset at the next midnight, but the window start stepped back the literal duration, so now could land before the window and the elapsed floor inflated the projection into a false alert. The window start now mirrors get_next_standardized_reset_time, and a reset more than one window away no longer projects at all
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
rust-analyzer cannot resolve the bare alias names emitted by
macro_rules_attribute::apply, so every aliased type was invisible to it
(no go-to-definition or find-references). Referencing the aliases through
crate:: resolves them via the re-export that attribute_alias! generates.
Co-authored-by: Nate Armstrong <narmstrong@Nates-MBP.localdomain>
* fix(proxy): mark background responses stale_expired when provider returns 404
Responses deleted upstream (store=false / ZDR rows dropped after provider retention) now move to stale_expired on the poll instead of being retried every cycle until the 7 day stale cleanup.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): drop inline comment in check_responses_cost
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drive _get_response seam in 404 stale_expired tests
Assigning AsyncMock on the instance keeps the new tests inside the TQ002/TQ008 test-quality ceilings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only expire 404 responses polled through a resolved deployment
A NotFoundError from the bare SDK fallback can mean a missing or misconfigured deployment, which a config fix inside the staleness window can still recover. Rows whose poll went through their router deployment are the ones the provider actually dropped, so only those move to stale_expired.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve the deployment before fetching so via_router binds once
Drops the reinitialised loop flag flagged by review and modernizes the moved annotation.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drive the real response fetch in the poll 404 regression tests
* fix(enterprise): expire background response rows on any provider 404 status
The GET path maps the provider's 404 body to litellm.BadRequestError that still carries status_code 404, so an except on NotFoundError never fired and the poll job kept retrying the row every cycle. Key the terminal decision on the status code instead of the class
* fix(enterprise): expire a background response only on a 404 naming it and skip a bad deployment entry per job
* chore(enterprise): record the poll-cycle bound on the managed-object IN lists
* test(integration): cover background response poll retirement on the real proxy
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test: move 76 live top-level tests to offline unit and integration coverage
* test: keep the mock-param opt-in guard valid once the top-level senders are gone
* test: address review on offline replacements
* docs(test): drop the deleted harness test file from the harness check command
* test: make the key rebind, team member delete and routes integration tests exercise the legacy paths
* test: make fallback, rpm and spend integration contracts deterministic and clean up their rows, assert the rpm limit in usage-based routing
* test: expect the no-deployments error at the rpm limit and cover the strategy check without pre-call checks
* test: hand member cleanups to the scenario instead of growing a budget list, flatten callback kinds
* test: scope the admin health check to the test's own deployment
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* docs(rust): add ADR folder with ADR 000 and ADR 001 scaffold
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(rust): add ADR 000 and ADR 001 on typing requests inside Rust
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(rust): list four ADR sections in ADR 000
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(rust): move ADR process doc from adr_000 to AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>