Commit graph

53984 commits

Author SHA1 Message Date
devin-ai-integration[bot]
577d74c1ce
fix(bedrock): count tokens on bedrock-mantle when bedrock-runtime cannot count a Claude model (#45317) 2026-10-08 15:26:49 -07:00
devin-ai-integration[bot]
3822947b0d
fix(rust): add the inline-tools-2026-09-15 beta to AnthropicBeta (#45439)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:25:17 -07:00
yuneng-jiang
b1eae8855e
fix(ui): show the Add Model picker once the model catalog loads after a provider is picked (#45426)
* fix(ui): show the Add Model picker once the model catalog loads after a provider is picked

* fix(ui): show the Add Model picker when a name typed before the catalog loaded was cleared
2026-10-08 15:12:25 -07:00
devin-ai-integration[bot]
406514fcaf
feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs (#44659)
* feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs

Adds general_settings.spend_logs_metadata_fields with mutually exclusive include and exclude lists. The filter runs on a copy of the row right before it is queued for Postgres, so daily spend rollups, budgets and callbacks still see every key.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): keep excluded auto-router savings keys out of published spend log metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): cover metadata retention across endpoints, failures, cache hits, batches, runtime updates and auto-router publication

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for spend_logs_metadata_fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): filter metadata at DB write so guardrail usage sees full rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): poll guardrail daily metrics instead of reading once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): drop timeout comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): read spend_logs_metadata_fields through typed general settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 14:51:44 -07:00
devin-ai-integration[bot]
0e48048bd5
chore(decisions): remove the System One converters and OpenAI spec types left dead by #45214 (#45442)
#45214 routed both decision formats through the shared decisions IR in
litellm/llms/base_llm/decisions/transformation.py and litellm/types/decisions.py.
That left litellm/llms/base_llm/decisions/systemone.py (to_system_one_request,
question_keys, to_decisions_response, the SystemOne* models and their adapter)
and litellm/types/openai_decisions.py with no importer outside their own unit
tests, so both modules and both test files go.

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 14:34:06 -07:00
yujonglee
0a98c7be10
refactor(rust): derive strum VariantArray and string conversions (#45434)
* refactor(rust): derive strum VariantArray and string conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): use strum conversions directly with explicit spellings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use rstest values for key management systems

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 21:24:49 +00:00
moe-berri
4ed0267b5a
fix(lens): prevent progress updates from starving analysis budget reservations (#45432)
* fix(lens): serialize analysis budget reservations with progress updates

* test(lens): verify budget reservations block competing progress writes
2026-10-08 21:08:55 +00:00
yujonglee
0721cffab2
refactor(python-bridge): take NativeCall directly and fold routes into per-route folders (#45413)
* refactor(python-bridge): take NativeCall directly and fold routes into per-route folders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(python-bridge): use rstest for updated tests and keep embedding's stub parameter name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 13:48:28 -07:00
devin-ai-integration[bot]
85a3869dfd
feat(guardrails): per-mode stream_scope with bedrock stream and pass-through fixes (#43801)
* feat(guardrails): run each mode only on streaming, non-streaming, or both

Add stream_scope so a rail can target streaming inference, non-streaming inference, or both per pre, during, and post mode

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): honor stream_scope in pipelines and dashboard types

Pipeline steps skipped the stream_scope filter, direct construction ignored mixed-case maps, and schema.d.ts was stale.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): skip unmatched stream_scope steps instead of allowing

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): keep stored stream_scope keys for modes not on screen

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): format stream_scope helpers and update fork MCP unit tests

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(mcp): keep hang cancellation tests from timing out during setup

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): honor path-defined streaming for stream_scope

Passthrough routes like Gemini streamGenerateContent decide streaming from the URL, so stamp that onto hook data before guardrails run.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(rust): copy AnthropicModelCapabilities instead of cloning

Clippy treats clone-on-Copy as an error, which failed rust-lint on the messages request tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): trust only server stream classification

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): keep streaming marker through deepcopy

scan_raw_request snapshots copy each field, so a plain object() marker would lose identity and skip streaming-only rails.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cost-map): drop duplicate perceptron-mk1.5 row

Two main cost-map PRs both added the OpenRouter model, so the merge left a second key that CI rejects.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(rust): expect native transcription 429 as RustUpstreamError

HTTP status errors from native routes map through route_error_to_pyerr, so the wheel SIGINT child was dying on an outdated RuntimeError check and never reached the hang probe.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tests): follow Google Interactions OpenAPI without hardcoded names

The live spec dropped CreateModelInteractionParams and renamed the item path to {interactionsId}. Misc CI failed because the compliance tests still looked those names up as literals.

* test(guardrails): reproduce stream_scope bedrock and passthrough field gaps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover invalid stored stream_scope reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): classify bedrock stream actions and keep caller is_streaming_request

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): make stream_scope_allows public and drop mutable builds

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): classify pass-through and Bedrock stream scopes

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): harden stream classification and validation

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add stream scope integration audit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): avoid mutating passthrough custom body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): run stream scope audit without enterprise license

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): keep stored scope restart cell on one worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): scope stream scope audit sink assertions to the rail under test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert logging_only scope absence behind an ordered barrier rail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert one logging_only scan per phase after the barrier rail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): set request scope on pass-through stream fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): tolerate invalid YAML stream_scope in v1 guardrails list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: merge main into litellm_guardrail_stream_scope_fixes

Update pass-through pre-call test callbacks for main's endpoint_type argument

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add request paths to pass-through fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover websocket pass-through stream scope and LIT-9050 outage spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(lint): remove unused type discipline suppressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): script the vertex live upstream in the websocket stream scope test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep stream marker json-serializable and restore pass-through helper names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): allow required Bedrock action re-export

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): strip the stream marker from pass-through payloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): drop stream marker by value in scans and snapshots, plain-tuple stream scope state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): render tag-scoped guardrail modes read-only in the custom code editor

A guardrail whose litellm_params.mode is the tag-scoped dict {tags, default}
crashed the Custom Code editor on open: normalizeMode wrapped the dict into
the mode array and StreamScopeFields rendered it as a React child (error #31,
whole dashboard unmounted). Treat a non-string non-array mode as no editable
modes, show formatGuardrailMode(mode) in a disabled input (read-only, matching
the guardrail info view), and keep mode/stream_scope out of the update payload
for such guardrails.

* fix(ui): resolve merge fallout in guardrails components

Deduplicate toModeArray import after the merge, and move the read-only
guardrail details block back into GuardrailReadOnlyDetails (now rendering
the shared mode/logging-only rows plus the stream-scope detail) so
guardrail_info.tsx stays under the 800-line lint budget. Guardrails UI
suite: 298 passed.

* fix(ui): drop duplicate toModeArray import reintroduced by merge

* chore(pass-through): document the deliberate in-place marker strip as mutable-ok

The clear/update on _parsed_body is load-bearing: rebinding to a fresh
mapping instead breaks 75 pass-through tests because the marker-free body
must propagate through the caller's request dict so downstream guardrail
scans and snapshots never observe the server streaming marker.

* fix(guardrails): move stream_scope after timeout in CustomGuardrail init

Inserting stream_scope before the existing timeout parameter shifted the
positional slot of timeout, so positional callers constructed with their
timeout bound to stream_scope (ValueError) and timeout silently None.
Restores the base parameter order; keyword callers are unaffected.

* fix(guardrails): typing pass for the lint gates

stream_scope leaves the declared constructor parameters (restoring the
base positional surface; keyword construction unchanged), unknown config
values crossing the new stream-scope code paths get typed locals or
cast-ok boundaries, and the passthrough payload literals are annotated.
All three lint gates pass against current main; the guardrail suites are
unchanged (236+5019 passing; one known anyio-driver failure pre-existing
on base).

* style: sort cast imports for the ruff gate

* chore(pass-through): drop unused BEDROCK_STREAMING_ACTIONS re-export

The streaming check now uses is_bedrock_streaming_endpoint; nothing in
the repo imports the name from this module.

---------

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: gabriele <gabriele@berri.ai>
2026-10-08 13:46:53 -07:00
devin-ai-integration[bot]
7ec2a94cd8
fix(ci): drop the publicly known master key prefix from the voyage routing test (#45431)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 13:41:03 -07:00
yuneng-jiang
a7b04a8028
test(decisions): post System One bodies to /v1/systemone in the translation bases (#45327) 2026-10-08 13:23:43 -07:00
yujonglee
9505af0639
feat(rust): add the Anthropic beta header policy (#45419)
* feat(rust): add the Anthropic beta header policy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): resolve known betas held as Other in AnthropicBeta::on

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use named rstest cases for Other beta resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): make the rstest named-case rule explicit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move BetaPolicy tests to the public API test crate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 20:22:58 +00:00
devin-ai-integration[bot]
56bafc29cf
feat(gemini): accept file content blocks with video_metadata on multimodal embeddings (#45305)
* feat(gemini): accept file content blocks with video_metadata on multimodal embeddings

* test(gemini): annotate the embedding mock helper and keep the type comment short

* fix(gemini-embeddings): keep file blocks through a partial embedding cache hit and validate video_metadata strictly

* test(caching): type the partial-hit cache test and pin its clock

* fix(gemini-embeddings): accept the Files API URI /v1/files returns as a file block's file_id

* fix(gemini-embeddings): 400 on bad input shapes, drop unknown block keys under drop_params, cache multi-block inputs

A bare object `input` answered 500 from both the caching handler and the
transformation; both now answer 400. An unresolved `files/` reference on
vertex_ai/ answered a ValueError 500; it now answers 400 naming the gemini/
provider. An empty `format` passed through to the provider; it now answers 400
naming file.format. Unknown block keys (`detail`, an unknown video_metadata
key, a top-level block key) are dropped under drop_params, global or
per-request, and still answer 400 without it. The embedding cache counted
file blocks against `max_messages`, so a request with 5 or more blocks was
never cached; file blocks no longer count.

* fix(caching): answer 400 for an object embedding input on the cache lookup

* test(gemini-embeddings): annotate the new embedding tests with return types and Final

* test(gemini-embeddings): annotate the transformation and files batch tests with return types and Final

* test(gemini-embeddings): audit file content blocks on the wire, in the cache, in batch uploads and under chaos

Integration cells for gemini/ and vertex_ai/ embedding file blocks with video_metadata: the batchEmbedContents and embedContent wire shapes, format overrides, nested and repeated inputs, every malformed block answering 400 before any provider call, drop_params at the request, deployment and YAML levels, the OpenAI SDK sync and async clients, Redis cache fills, hits, partial hits and the metadata in the key with zero-spend hit rows, chat, responses and messages controls, the vertex_ai/ batch JSONL upload, a mixed fast/slow/malformed/dropped burst and a worker kill. The wire peer now reads chunked request bodies, which the streamed GCS media upload sends

* test(integration): treat an upload the client abandons before its terminating chunk as a disconnect, never a stored request

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 13:20:03 -07:00
fzowl
f663794fcd
feat(voyage): rebrand to VoyageAI by MongoDB and route MongoDB keys to ai.mongodb.com (#41812)
* feat(voyage): rebrand to VoyageAI by MongoDB, refresh model list, support both contextual input shapes

* fix(voyage): send auto-chunking params for flat contextual inputs

A flat list[str] or bare str for the contextualized embeddings endpoint is
only valid as documents with enable_auto_chunking=True and input_type=document,
or as queries with input_type=query. Default those params for non-query flat
inputs so the request matches the live API contract, while letting caller-set
values win. Nested list[list[str]] still passes through unchanged.

Tests now assert the auto-chunking params instead of only echoing inputs.

* feat(voyage): route MongoDB-issued keys to ai.mongodb.com

Mirror voyageai.util.get_default_base_url from the official SDK: a key with
the `al-` prefix is issued by MongoDB and is only valid on ai.mongodb.com,
every other key on api.voyageai.com. The choice now lives in one shared
helper that the embedding, contextual, multimodal, and rerank configs all
use for both the default base URL and the Authorization header, so the host
and the key always agree.

Also drops the voyage-4-nano and voyage-multilingual-2 model map entries:
voyage-4-nano is not served on the Voyage API and voyage-multilingual-2 is
an older model, so neither belongs in this change.

* fix(voyage): restore MongoDB routing on contextual endpoint and re-add voyage-4-nano after upstream merge

The upstream merge replaced the contextual config's get_complete_url with a
Voyage-only host and dropped voyage-4-nano from the model map. Reapply the
al- key routing via get_default_base_url and add voyage-4-nano (open-weight,
per docs.voyageai.com) so the branch keeps its task changes on top of upstream.

* fix(voyage): drop voyage-4-nano and the contextual tests upstream already covers

voyage-4-nano is not served on the Voyage API, so it does not belong in the
model map. The contextual input tests duplicate
tests/test_litellm/llms/voyage/test_voyage_contextual_embedding.py, which
landed upstream with the auto-chunking fix this branch was carrying.

* test(voyage): drop sys.path.insert from the voyage common utils test

* fix(voyage): route rerank on the key it authenticates with

get_complete_url picked the rerank host from the environment while the auth
header carried the request key, so an explicit MongoDB-issued key was posted
to api.voyageai.com. validate_environment now hands the resolved key to the
config instance and get_complete_url reads it back, so the host and the
credential always come from one key. The config is built per request, so
nothing carries over between them.

* test: allow ultrafast pricing keys in model prices schema

* test: drop ultrafast schema keys now added upstream

* test(integration): prove voyage key-based host routing on the wire

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 13:12:47 -07:00
devin-ai-integration[bot]
95dd90f637
fix(model_prices): add Bedrock flex prices, Anthropic web search flags, Gemini shutdown dates, OpenRouter alias drift, Mistral Large 4 context (#44910)
* feat(model_prices): add Mistral Large 4 and its OpenRouter route

Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): sync Gemini API shutdown dates for TTS, Live, image and Omni previews

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): correct Claude Haiku 5.5 flags, add Anthropic web search flags and OpenRouter Claude Haiku 5.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add Bedrock flex tier prices and clear lapsed Together DeepSeek V4 dates

Absorbs #45320 (Bedrock flex prices, verified against the AWS Price List API)
and the DeepSeek V4 date removals from #45307 (verified against Together docs and API)

Co-authored-by: Roi <roi@roitev.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): Mistral Large 4 context, OpenRouter Claude latest aliases and batch pricing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mistral): expect 1048576 context for Mistral Large 4

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: Roi <roi@roitev.com>
2026-10-08 13:12:42 -07:00
yucheng-berri
29f48e62bf
feat(fireworks_ai): forward the LiteLLM user id as user behind fireworks_forward_user_id (#45265)
* feat(fireworks_ai): forward the LiteLLM user id as user behind fireworks_forward_user_id

* fix(fireworks_ai): keep the forwarded user id when the caller sends extra_body.user
2026-10-08 12:48:16 -07:00
devin-ai-integration[bot]
3571b5dc72
fix(vertex_ai): forward the inline-tools-2026-09-15 beta to Vertex and Anthropic (#45297)
* fix(vertex_ai): forward the inline-tools-2026-09-15 beta to Vertex and Anthropic

* test(vertex_ai): trim the inline-tools beta test docstrings and type their locals

* test(vertex_ai): add audit cells for inline-tools beta forwarding on Vertex Anthropic

* test(vertex_ai): send legal whitespace in the hostile beta header audit cell

* test(vertex_ai): close the SDK clients in the inline-tools audit cells

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 12:27:52 -07:00
Mateo Wang
b9ff2d66a4
fix(proxy): retry lock-timed-out daily spend batches in place so the shutdown flush keeps them (#45273)
* fix(proxy): retry lock-timed-out daily spend batches in place so the shutdown flush keeps them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): annotate the lock-timeout retry test and carry its context in assert messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): mark the remaining lock-timeout test local Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 12:05:47 -07:00
Mateo Wang
487164f762
test(proxy): audit the websocket rejection log on every live route (#45293)
* test(proxy): audit the websocket rejection log on every live route

* test(proxy): scan the stopped proxy log for the websocket crash and drop the uvicorn wording pin
2026-10-08 12:04:23 -07:00
Solomon Mithra
0e455ae7a9
fix(proxy): use budget reset window for projected spend alerts (#31942)
* fix(proxy): use budget reset window for projected spend alerts

* style(proxy): use PEP 585 tuple annotations in projection helpers

* fix(proxy): derive projection date from budget reset timezone

* fix(proxy): project spend in real time within the budget reset window

* fix(proxy): derive the projection window start from the reset schedule

30d and monthly keys reset on the 1st of the month and unrecognized spellings like 1hr reset at the next midnight, but the window start stepped back the literal duration, so now could land before the window and the elapsed floor inflated the projection into a false alert. The window start now mirrors get_next_standardized_reset_time, and a reset more than one window away no longer projects at all

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-10-08 12:02:40 -07:00
nate-berri
e843cb7aca
fix(rust): path-qualify attribute aliases so rust-analyzer resolves them (#45416)
rust-analyzer cannot resolve the bare alias names emitted by
macro_rules_attribute::apply, so every aliased type was invisible to it
(no go-to-definition or find-references). Referencing the aliases through
crate:: resolves them via the re-export that attribute_alias! generates.

Co-authored-by: Nate Armstrong <narmstrong@Nates-MBP.localdomain>
2026-10-08 11:51:58 -07:00
berriai-litellm-provider-info-sync[bot]
b0e0a305f4
fix(bedrock_mantle): take gpt-oss output limits and EOL dates from the Bedrock model cards (#45417)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 11:33:28 -07:00
berriai-litellm-provider-info-sync[bot]
d8d8364ec7
chore(cost-map): add openai gpt-6.1-sol ultrafast tier prices from the pricing page (#45408)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 11:27:23 -07:00
berriai-litellm-provider-info-sync[bot]
26b9bf2e6c
fix(bedrock): take context, output limits and EOL dates from the Bedrock model cards (#45409)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 11:23:29 -07:00
devin-ai-integration[bot]
b7ecb05340
fix(proxy): mark background responses stale_expired when provider returns 404 (#41724)
* fix(proxy): mark background responses stale_expired when provider returns 404

Responses deleted upstream (store=false / ZDR rows dropped after provider retention) now move to stale_expired on the poll instead of being retried every cycle until the 7 day stale cleanup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop inline comment in check_responses_cost

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): drive _get_response seam in 404 stale_expired tests

Assigning AsyncMock on the instance keeps the new tests inside the TQ002/TQ008 test-quality ceilings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only expire 404 responses polled through a resolved deployment

A NotFoundError from the bare SDK fallback can mean a missing or misconfigured deployment, which a config fix inside the staleness window can still recover. Rows whose poll went through their router deployment are the ones the provider actually dropped, so only those move to stale_expired.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve the deployment before fetching so via_router binds once

Drops the reinitialised loop flag flagged by review and modernizes the moved annotation.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): drive the real response fetch in the poll 404 regression tests

* fix(enterprise): expire background response rows on any provider 404 status

The GET path maps the provider's 404 body to litellm.BadRequestError that still carries status_code 404, so an except on NotFoundError never fired and the poll job kept retrying the row every cycle. Key the terminal decision on the status code instead of the class

* fix(enterprise): expire a background response only on a 404 naming it and skip a bad deployment entry per job

* chore(enterprise): record the poll-cycle bound on the managed-object IN lists

* test(integration): cover background response poll retirement on the real proxy

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 10:35:52 -07:00
devin-ai-integration[bot]
d5173e3d70
test: move 76 live top-level tests to offline unit and integration coverage (#45352)
* test: move 76 live top-level tests to offline unit and integration coverage

* test: keep the mock-param opt-in guard valid once the top-level senders are gone

* test: address review on offline replacements

* docs(test): drop the deleted harness test file from the harness check command

* test: make the key rebind, team member delete and routes integration tests exercise the legacy paths

* test: make fallback, rpm and spend integration contracts deterministic and clean up their rows, assert the rpm limit in usage-based routing

* test: expect the no-deployments error at the rpm limit and cover the strategy check without pre-call checks

* test: hand member cleanups to the scenario instead of growing a budget list, flatten callback kinds

* test: scope the admin health check to the test's own deployment

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 10:33:35 -07:00
nate-berri
a81f507bca
docs(rust): add ADRs for the Rust core (#45205)
* docs(rust): add ADR folder with ADR 000 and ADR 001 scaffold

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): add ADR 000 and ADR 001 on typing requests inside Rust

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): list four ADR sections in ADR 000

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): move ADR process doc from adr_000 to AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 10:21:56 -07:00
joshua-berri
49266e3b17
fix(mcp): preserve elicitation context and report relay failures (#45255)
* fix(mcp): correlate elicitation relays and report failures

* fix(mcp): preserve elicitation timeout errors on Python 3.10

* fix(mcp): keep request ID alias in type-checking imports

* fix(mcp): preserve elicitation type narrowing and clarify HTTP limits

* ci(mcp): publish elicitation coverage from GitHub Actions

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 10:18:25 -07:00
devin-ai-integration[bot]
9f150f74ed
feat(prometheus): expose per-project per-model rate limit allowed and used gauges (#43561)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 12:03:29 -05:00
devin-ai-integration[bot]
6e54dced8c
test: realign stale decisions and lens tests with merged behaviour (#45346)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 09:44:18 -07:00
Mateo Wang
ce6889bbf7
fix(responses): keep tool_result next to tool_use on Anthropic previous_response_id continuations (#45322)
* fix(responses): keep tool_result next to tool_use on Anthropic previous_response_id continuations

Replayed history no longer re-adds each stored turn's instructions, and the
current request's instructions go before the replayed history instead of
between the last tool_use and its tool_result

* test(responses): fake the Anthropic transport instead of acompletion in the continuation regression test

* fix(responses): keep the previous turn's instructions when a continuation sends none

* fix(responses): carry only the latest stored instructions when a continuation sends none

Replaying each spend log's own instructions put a system message between a
replayed tool_use and its tool_result, which Anthropic rejects with a 400
2026-10-08 09:44:11 -07:00
devin-ai-integration[bot]
92654b69a9
test: delete 34 legacy tests already covered by e2e (#45340)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 09:43:56 -07:00
yujonglee
35b0992533
feat(rust): add standalone typed LLM wire contracts (#45188)
* refactor(rust): add typed LLM wire contracts to llms-types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): drop blank lines left in new wire types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): complete additive wire type contracts

* feat(rust): complete additive nested wire contracts

* feat(rust): give Messages content blocks concrete per-block contracts

Replace the shared all-optional ContentBlockPayload with one struct per
documented block and nested tagged unions for server tool results, so
required fields from the official Messages reference are enforced.

Make CustomTool a plain struct with required name and input_schema, add
MessagesToolParam plus the toolset and undated tool-search tags, model
MiniMax media sources as a tagged union, require documented fields on
chat audio, image URLs, and logprobs, and rename the new block enum to
MessagesContentPart so it no longer shadows the streaming struct.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(rust): tighten new wire contracts against the official references

Give each Messages builtin tool a concrete contract with its documented
required fields, fixed names, and typed allowed callers in place of the
shared all-optional ToolDefinition bag. Make MCP tool result content and
fallback triggers optional as the request params document, and narrow
tool search references, web fetch content, and bash results to their
documented block types.

Keep unmodeled Responses output items and new usage iteration and stop
detail types as Recognized values instead of rejecting the whole payload,
accept find_in_page and optional open_page URLs, preserve JSON Schema
property order, and move prompt_cache_breakpoint to Chat Completions
content parts where OpenAI documents it.

Drop OCR table and key/value types that only matched Azure Document
Intelligence and would reject other providers' normalized passthrough,
along with unsourced bounding box and applied edit fields.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(rust): type every documented Responses output item and close wire gaps

Add the 18 remaining documented Responses output item variants with
concrete nested actions, outputs, and safety checks, and add the
documented namespace, async, caller, and image generation fields to the
existing items.

Restrict MCP tool result content to a string or text blocks, model the
Messages diagnostics cache-miss union and its request param with null
distinct from absent, and accept draft-07 tuple items and prefixItems in
JSON Schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(rust): keep Messages contract assertions in test bodies

Compare built-in tool decodes against fully built expected values and
split the required-field rejections per type, so no case carries its
assertions in a callback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): cite the Mistral OCR reference beside the OCR wire types

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 16:27:12 +00:00
devin-ai-integration[bot]
82a5d9842a
fix(ui): list only catalog models in the Add Model picker (#44877)
The Add Model selector read the merged runtime cost map, so deployment
ids and provider-prefixed backend keys registered for proxy deployments
showed up as duplicate model choices. The public cost map endpoint now
accepts catalog_only=true to return the catalog as loaded, and the Add
Model panel requests that view. The default response is unchanged.

Resolves LIT-9263

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 11:18:29 -05:00
devin-ai-integration[bot]
98465fe2ba
fix(ci): trim rust-test debuginfo so the job fits on the runner disk (#45318)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:34:20 +00:00
ishaan-berri
f824d11a22
fix(lens): block teamless feedback writes and fix retention test (#45266)
* fix(lens): reject feedback writes from callers with no team or key

* test(lens): teamless callers cannot overwrite feedback

* test(lens): expect lens_feedback retention during schema setup

* refactor(lens): flatten feedback summaries without a stacked comprehension

* chore(lens-ui): drop routine comment on the feedback panel

* chore(lens-ui): drop routine comment on the feedback view

* chore(lens-ui): drop routine comment on the low feedback check

* chore(lens-ui): drop routine comment on the post stub
2026-10-08 07:54:11 -07:00
berriai-litellm-provider-info-sync[bot]
d8c0e2c715
fix(azure): move sora-2 retirement date to the later Models API date (#45355)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 04:53:13 -07:00
devin-ai-integration[bot]
df13d59306
fix(logging): skip sync success callbacks for internal sub-calls, deflake RAG and Langfuse tests (#45345)
* fix(logging): skip sync success callbacks for internal sub-calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langfuse): assert no retry sleep instead of a wall-clock budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: add types to deflake regression coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): avoid unrelated formatting changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 04:45:21 -07:00
devin-ai-integration[bot]
01166906fb
test(integration): run the response cache tool-call tests on a proxy without the message cap (#45350)
* test(integration): run the response cache tool-call tests on a proxy without the message cap

Since #43878 the response cache skips any request past 4 messages by default, so the three providers-lane tests that post a seven-message tool conversation twice never saw a cache hit. They now boot one owned proxy from the lane config with max_messages set to null and keep their assertions; the shared proxy and the caching-lane cap test are unchanged

* test(integration): give the uncapped-proxy cache tests the owned-proxy timeout budget

The providers lane runs pytest with a 90 second per-test timeout that includes fixture setup, and the first of the three tests to run pays the owned proxy boot before the spend-log test can poll for up to 70 seconds. The sibling owned-proxy tests already carry pytest.mark.timeout(240), so these three take the same budget

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:24:13 -07:00
devin-ai-integration[bot]
9a2c9e6099
fix(proxy-extras): name the database error when the Lens rename check cannot run (#45341)
* fix(proxy-extras): name the database error when the Lens rename check cannot run

The Lens rename check runs before prisma db push and raised a generic
"Cannot verify Lens data safety" RuntimeError on any psycopg error, with
the cause only in the chained exception. proxy_cli prints the message and
exits, so a Postgres at its connection limit read as a Lens problem.

The raised message now ends with the redacted database error text, and
the connection opener is an injected parameter so the unit test drives
the check without a database.

* test(proxy-extras): drop the Lens check test class docstring

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:18:25 -07:00
joshua-berri
d0202ac364
feat(mcp): support upstream OAuth client metadata identities (#45231)
* feat(mcp): support upstream OAuth client metadata identities

* fix(mcp): load client metadata on cold OAuth requests

* fix(mcp): discover client metadata with configured OAuth endpoints

* fix(mcp): preserve configured endpoints when discovery fails

* fix(mcp): retain CIMD identities across refresh and catalog reloads

* fix(mcp): reuse saved CIMD identity for token endpoint refresh

* fix(mcp): keep dynamic client registration ahead of CIMD unless the deployment opts in

A provider that advertises both a registration endpoint and client ID
metadata documents now gets the registration flow the gateway used
before, so a gateway on a private network keeps working against it;
the metadata document identity applies when the provider offers no
registration, or when general_settings.mcp_prefer_client_id_metadata_document
is true. The saved CIMD refresh identity compares tokens as bytes so a
non-ASCII refresh token cannot crash the token route.

* chore(ui): regenerate dashboard API types for the new MCP general setting

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:10:34 -07:00
devin-ai-integration[bot]
3a3b8cad80
feat(ui): add TypeSafe and Strands Decider to the Add Model provider list (#45335)
* feat(ui): add TypeSafe and Strands Decider to the Add Model provider list

* test(public_endpoints): type the provider fields helper against the route's model

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:04:07 -07:00
devin-ai-integration[bot]
0e26edfdb9
test: move 81 legacy live tests in pass-through, spend, batches, audio, search, guardrails, image and ocr dirs offline (#45288)
* test: make legacy live tests in spend, batches, openai endpoints and audio dirs offline (partial)

* test: migrate wave-1b live tests offline (guardrails, images, ocr, search, openai endpoints)

* test: fix wave-1b review items, add responses/ocr integration tests and firecrawl unit test

* test: anthropic messages router/bedrock/openai-bridge unit tests for wave-1b nodes

* test: anthropic messages logging, prompt-caching and tool-search unit tests; drop migrated base nodes

* test: finish pass_through_unit_tests nodes, logging drain fix and mutations

* test: migrate anthropic passthrough tests to integration wire tests

* test: fix passthrough migration wire spend row lookup and wildcard config

* test: migrate hosted vllm and openai file passthrough tests offline

* test: move assemblyai and vertex passthrough nodes to in-process unit tests

* test: restore unlisted router node and fix logging worker drain in passthrough unit tests

* test: drop spend-row BUG skip and sharpen non-streaming skip reason for anthropic messages

* test: use public presidio alias after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: restore the batch and file unit tests the migration rewrote away

* test: restore full legacy intent in anthropic messages router unit tests

Drop the false BUG skip on non-streaming aanthropic_messages logging (success
callbacks do fire; the skipped body filtered on the wrong model), assert the
logged model_group, messages, cost and usage, cover streaming logging for both
Anthropic and Bedrock invoke, assert dict content blocks for Anthropic, Bedrock
invoke and the OpenAI bridge, add the Bedrock invoke leg of the router test,
fall back from a real 401, test system-prompt caching and streaming
message_start cache fields on converse and invoke, send the legacy tool-search
tools and beta header, remove the type: ignore and bare dict helper, and add
in-process native /anthropic passthrough spend logging tests

* test: fix passthrough migration wire tests and drop false native spend BUG skip

The native spend row test read rows by the shared master-key digest, so it
matched other tests' rows; read each row by its own message id instead and
assert tokens, total, spend, tags, provider, api_base and end_user on both the
non-streaming and streaming native routes. Merge the streaming test that never
checked spend, assert exact tags from litellm_metadata, stop rebinding a Final
in a loop, require cost > 0, and add the chat-completions bridge cost case the
legacy test covered

* test: harden the spend, OCR, image, search and guardrail migration replacements

The OCR spend tests now build fresh kwargs per case instead of mutating a
shared fixture, and every payload case asserts the exact logged spend. The
OCR wire test reads its spend row by request id and checks exact page
pricing. Image edit, Nova Canvas, DuckDuckGo, Firecrawl, Bedrock guardrail
and Presidio replacements now fake only the provider HTTP boundary (respx or
an in-process aiohttp server) and assert the outbound request, so the
DuckDuckGo limit, Azure base_model pricing and guardrail masking are proven
rather than assumed

* test: cover Exa and Perplexity search structure and max_results offline and retire the two base search methods

* test: drive batch and file replacements through the provider HTTP boundary and real logging callback

Replace monkeypatched litellm.afile_content and AsyncHTTPHandler doubles with respx routes,
read batch logging metadata from a registered success callback instead of get_logging_payload,
use the real managed-files hook for the GEN-2166 regression, assert outbound request bodies,
pin poller ownership explicitly in the migrated DB-sync tests, and require the scripted
upstream to be hit in the responses error-status wire tests

* test: assert file content download headers pass through the proxy

* test: point spend coverage references at the tests that replaced the retired spend job

* test: assert passthrough identity, spend and route dispatch from the code under test

The AssemblyAI non-admin test asserted metadata it wrote itself and leaked a
background poll to the real AssemblyAI host. It now drives assemblyai_proxy_route
with a real Request and waits for the success callback for its own transcript id.
The Vertex spend test matches its log by call id instead of taking the first event.
The OpenAI files wire test hit the native /{provider}/v1/files route; it now calls
/openai/files so the passthrough is what forwards the upload and delete.

* test: mock only the HTTP boundary in the migrated audio tests

Vertex TTS tests no longer replace _ensure_access_token or AsyncHTTPHandler.post;
the token comes from a mocked Google OAuth endpoint and the synthesize call from
respx. Speech tests assert the outbound body, the transcription cache test polls
for the cache write instead of relying on test ordering, and the model pass-through
test checks the multipart model field per model.

* test: wait on a logger event instead of polling the clock in anthropic messages unit tests

Recorders keep payloads in a rebound tuple and set an asyncio.Event; tests
await it with asyncio.wait_for instead of a sleep-and-deadline poll loop

* test: freeze module-level batch and file response fixtures as Final MappingProxyType

* test: fake the presidio analyzer with an in-memory aiohttp connector

The blocked-entity tests started an aiohttp TestServer, which binds a local
socket. They now hand the guardrail a ClientSession whose connector answers
/analyze and /anonymize in process, so no socket is opened and the outbound
analyze text and entities are still asserted

* test: type anthropic messages router test helpers with LiteLLM's Anthropic TypedDicts

Messages, cached system blocks and tool-search tools now use
AnthropicMessagesUserMessageParam, AnthropicMessagesTextParam,
AnthropicToolSearchToolRegex and AnthropicMessagesTool instead of bare
dict shapes; tools are converted to plain dicts only at the acreate call,
whose tools parameter is list[dict]

* test: type batch limiter helpers with TypedDicts and wait on the logging callback event instead of polling

* test: give the migrated OCR, image and presidio helpers precise types

OCR spend helpers take ReadOnly TypedDicts for kwargs and responses and use
LiteLLM's OCRResponse/OCRUsageInfo instead of local pydantic stand-ins; spend
metadata is validated with a TypeAdapter. The presidio fake uses LiteLLM's
PresidioAnalyzeRequest/ResponseItem types, and the image-edit logger validates
the logged payload instead of storing an untyped dict

* test: signal callback and cache events instead of polling

Recorders keep tuples and set an asyncio.Event, thread-safely, when the payload
for this test's transcript id or upstream URL arrives. The transcription cache
test waits on a Cache subclass that signals after async_add_cache. No clock
polling or sleeps remain in these tests.

* test: assert the batch limiter hook updates the caller's request in place

* test: tolerate model-list probes and read native passthrough rows by owned key

The router's OpenAI-compatible model-info refresh (litellm/router.py:10710)
sends GET /v1/models to configured openai api_bases, so the wire answers it
with an empty list and excludes it from the provider-call assertions. Native
/anthropic spend rows are now read by a per-request virtual key digest and
call_type, then the row's request_id is checked against the message id

* test: expect the OCR alias in the proxy response model

The proxy restamps every OpenAI-compatible response model to the name the
client requested (_override_openai_response_model), so /v1/ocr returns the
scenario alias. The upstream model is now checked on the drained request body
instead of inside the peer, where a failed assert never reached the test

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 02:38:12 -07:00
devin-ai-integration[bot]
04d97abffb
test: make 77 legacy live tests offline in litellm_utils, router_unit and responses dirs (#45298)
* test: make 77 legacy live tests offline in litellm_utils, router_unit and responses dirs

* test: restore lost coverage in litellm_utils offline replacements

Bound live health-check tasks to the concurrency limit so eager task
creation behind a semaphore fails, route the Langfuse trace-id check
through completion -> Logging.get_trace_id with all four metadata
combinations, run the callback dedup checks through acompletion, pass
the dynamic key via litellm_params, cover the env-inferred default
model list, and add the per-provider audio transcription config lookup

* test: restore router coverage lost in the offline move

Add the missing non-stream LIT-3058 header/count-once test, pin the UTC minute on every
usage-counter test, assert router-level client reuse for transcription, and replace the
private selector-attribute checks with routing outcomes per strategy. Assert outbound bodies
for speech, rerank, image, assistants and moderation, add timeouts to event waits, and drop
the router_unit_tests husk files that no longer hold tests

* test: restore sync stream, router sync stream, error event and field-type coverage for migrated responses tests

Add the sync streaming logging and sync Router.responses streaming cases the
offline replacements dropped, port the legacy per-event and response field-type
validation, assert raw headers, the float created_at conversion, MCP tool headers,
the search_context_size input, and add an offline replacement for the in-stream
context-window error event. Remove helpers left dead in the legacy file.

* test: run numpydoc-backed unit tests in GHA and drop a test that pinned a crash

* test: restore sequence number, item id and content part checks in the responses stream validator

* test: match router logging events to the test's own deployment so late events from other tests are ignored

* test: restore the s3 cold storage history test to its original assertions

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 02:37:55 -07:00
Priyansh Nandwana
7c7b0ea85b
fix(bedrock_mantle): send OpenAI explicit prompt cache breakpoints for GPT-5.6 and newer (#38729)
* fix(cache_control): let a hosted deployment opt into the OpenAI cache dialect

_targets_openai_prompt_cache_breakpoint gated on custom_llm_provider == "openai"
unconditionally, so an OpenAI-shaped model served by another provider could never
qualify, even with supports_prompt_cache_breakpoint set explicitly on its own
cost-map entry. bedrock_mantle honours prompt_cache_breakpoint end to end, and
had no way to say so.

model_cost is keyed per exact deployment string, so a flag on the deployment's own
entry states the dialect more precisely than a provider name can. Consult it
before the provider check, and require the entry's litellm_provider to match the
serving provider so an openai entry cannot license another provider. Entries for
the openai provider keep their api_base check, so an OpenAI-compatible third-party
host is still not assumed to speak the dialect.

Fixes #38666

* test: drop a laziness assertion the existing suite already makes

test_provider_lookup_skipped_for_models_below_gpt_5_6 already pins that
_resolve_provider is not called for an unflagged model, and it covers the new
flag lookup unchanged. The duplicate patched a litellm internal for no added
coverage, which the test-quality gate counts against TQ008.

* fix(bedrock_mantle): flag GPT-5.6 and newer rows for OpenAI explicit prompt cache breakpoints

* fix(responses): predict the chat-completions bridge with the model name the dispatch resolves

* fix(cache_control): read a region-prefixed Mantle GPT deployment through its region-free price-map row

* fix(cache_control): let a bare hosted deployment name read its own provider's price-map row

* fix(responses): hand the hook the provider the router resolved for a Foundry GPT deployment

* fix(cache_control): read the deployment breakpoint flag strictly and default a null prompt_cache_options

A cost-map or deployment `supports_prompt_cache_breakpoint` that is not the boolean `true` (the string "true",
the integer 1, "false", 0, a long string) no longer opts a deployment into the OpenAI prompt cache dialect; only
`true` does, the same reading Bedrock Converse applies to its own flag.

A client that sends `"prompt_cache_options": null` on `/v1/chat/completions` or `/v1/messages` now gets the
implicit default the hook already applied on `/v1/responses` when it placed a breakpoint; before, the null
suppressed the default and the request left with a breakpoint and no options.

* test(integration): audit cells for the Mantle GPT prompt cache breakpoint dialect

Deterministic cells for the configured-breakpoint path on Bedrock Mantle GPT and Azure AI Foundry GPT-6
deployments: every endpoint, streaming and not, sync and async SDKs and raw httpx, the hostile option shapes,
flag precedence, the response cache twin, a two-worker burst, a worker kill and a graceful restart.

* fix(cache_control): ignore a malformed cache_control_injection_points value instead of failing the request

A deployment or client `cache_control_injection_points` that is not a list of points (a string, an
integer, a bare point dict, a list of strings) raised inside the prompt hook (`'str' object has no
attribute 'get'`, `'int' object is not iterable`) and turned every request to that deployment into a
500 on `/v1/chat/completions`, `/v1/messages` and `/v1/responses`. Every entry point now reads such a
value as no configured points, the way a `null` already read, and the request leaves without a
breakpoint; entries of a list that are not points are dropped and the point entries kept.

* test(integration): cover int, dict and string-list injection point shapes in the Mantle audit cells

The D9 cell now also sends an integer, a bare point dict and a list of strings as the deployment's
`cache_control_injection_points`, which the merge base answered with a 500 on every request.

* fix(cache_control): stamp the OpenAI dialect for a bare deployment name served by a flagged provider

A deployment written as `model: openai.gpt-5.6-sol` with `custom_llm_provider: bedrock_mantle` has no cost-map row
of its own and no openai row of the same name, so the dialect stamp's cheap gate (`supports_openai_prompt_cache_breakpoint`)
returned early and the configured point was never stamped. On `/v1/responses` and on the chat seeding path the hook
sees no provider, so it fell back to the Anthropic `cache_control` marker, which Bedrock Mantle strips.

The gate now also passes a model whose serving provider is already known and whose provider-keyed row carries the
flag, which is the same row the dialect resolution reads and costs no provider lookup. A bare name without a provider
is still left alone, so models below GPT-5.6 keep their points untouched and resolve nothing.

* test(integration): cover a bare Mantle deployment name with its provider in the audit cells

A deployment configured as `model: openai.gpt-5.6-sol` plus `custom_llm_provider: bedrock_mantle` sends the
breakpoint and the implicit options on all three endpoints; the merge base leaves it on the Anthropic dialect.

* fix(cache_control): keep a configured point that sits beside junk entries on every chat seed path

`configured_injection_points` kept the point dicts of a mixed `cache_control_injection_points` list as a tuple,
but only read a `list` back. The chat seed and the `/v1/responses` dialect stamp write the normalized value back
onto the request, and on a deployment the OpenAI dialect does not stamp (an Anthropic model, Claude on Bedrock
Mantle) that value is the tuple itself, so the hook then read it as no configured points and the valid point was
silently dropped. The stamped OpenAI dialect and the `/v1/messages` path kept it only because they build a new list.

The normalizer now reads back the tuple it wrote, so the point reaches the wire on every path; an all-dict list
still passes through as the same object.

* test(integration): cover a configured point beside junk entries on a Claude Mantle deployment

A deployment configured with `cache_control_injection_points: ["system", {system point}, 3]` marks the system
block on the Anthropic dialect on all three endpoints; the merge base answers 500 on the junk entry.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 02:07:22 -07:00
devin-ai-integration[bot]
7a659973e3
revert(lint-gates): accept an empty base scan again, since zero violations is a legitimate count (#45319)
This reverts commit 7c7efd4d55 (#45314).

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 08:40:08 +00:00
devin-ai-integration[bot]
7dd9aff244
refactor(proxy): type fresh Moyai connect helpers and drop MCP rpm getattr (#45315)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 01:31:59 -07:00
devin-ai-integration[bot]
7c7efd4d55
fix(lint-gates): fail on an empty base scan instead of blaming the change for every violation (#45314)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 08:06:18 +00:00
tin-berri
057034d0d3
fix(lens): refresh runs until gateway costs are complete (#45105) 2026-10-08 00:24:41 -07:00
devin-ai-integration[bot]
8683f1418e
fix(utils): make supports_audio_output read the supports_audio_output cost-map key (#45294)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 23:54:38 -07:00