Commit graph

53949 commits

Author SHA1 Message Date
ishaan-berri
f824d11a22
fix(lens): block teamless feedback writes and fix retention test (#45266)
* fix(lens): reject feedback writes from callers with no team or key

* test(lens): teamless callers cannot overwrite feedback

* test(lens): expect lens_feedback retention during schema setup

* refactor(lens): flatten feedback summaries without a stacked comprehension

* chore(lens-ui): drop routine comment on the feedback panel

* chore(lens-ui): drop routine comment on the feedback view

* chore(lens-ui): drop routine comment on the low feedback check

* chore(lens-ui): drop routine comment on the post stub
2026-10-08 07:54:11 -07:00
berriai-litellm-provider-info-sync[bot]
d8c0e2c715
fix(azure): move sora-2 retirement date to the later Models API date (#45355)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 04:53:13 -07:00
devin-ai-integration[bot]
df13d59306
fix(logging): skip sync success callbacks for internal sub-calls, deflake RAG and Langfuse tests (#45345)
* fix(logging): skip sync success callbacks for internal sub-calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langfuse): assert no retry sleep instead of a wall-clock budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: add types to deflake regression coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): avoid unrelated formatting changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 04:45:21 -07:00
devin-ai-integration[bot]
01166906fb
test(integration): run the response cache tool-call tests on a proxy without the message cap (#45350)
* test(integration): run the response cache tool-call tests on a proxy without the message cap

Since #43878 the response cache skips any request past 4 messages by default, so the three providers-lane tests that post a seven-message tool conversation twice never saw a cache hit. They now boot one owned proxy from the lane config with max_messages set to null and keep their assertions; the shared proxy and the caching-lane cap test are unchanged

* test(integration): give the uncapped-proxy cache tests the owned-proxy timeout budget

The providers lane runs pytest with a 90 second per-test timeout that includes fixture setup, and the first of the three tests to run pays the owned proxy boot before the spend-log test can poll for up to 70 seconds. The sibling owned-proxy tests already carry pytest.mark.timeout(240), so these three take the same budget

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:24:13 -07:00
devin-ai-integration[bot]
9a2c9e6099
fix(proxy-extras): name the database error when the Lens rename check cannot run (#45341)
* fix(proxy-extras): name the database error when the Lens rename check cannot run

The Lens rename check runs before prisma db push and raised a generic
"Cannot verify Lens data safety" RuntimeError on any psycopg error, with
the cause only in the chained exception. proxy_cli prints the message and
exits, so a Postgres at its connection limit read as a Lens problem.

The raised message now ends with the redacted database error text, and
the connection opener is an injected parameter so the unit test drives
the check without a database.

* test(proxy-extras): drop the Lens check test class docstring

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:18:25 -07:00
joshua-berri
d0202ac364
feat(mcp): support upstream OAuth client metadata identities (#45231)
* feat(mcp): support upstream OAuth client metadata identities

* fix(mcp): load client metadata on cold OAuth requests

* fix(mcp): discover client metadata with configured OAuth endpoints

* fix(mcp): preserve configured endpoints when discovery fails

* fix(mcp): retain CIMD identities across refresh and catalog reloads

* fix(mcp): reuse saved CIMD identity for token endpoint refresh

* fix(mcp): keep dynamic client registration ahead of CIMD unless the deployment opts in

A provider that advertises both a registration endpoint and client ID
metadata documents now gets the registration flow the gateway used
before, so a gateway on a private network keeps working against it;
the metadata document identity applies when the provider offers no
registration, or when general_settings.mcp_prefer_client_id_metadata_document
is true. The saved CIMD refresh identity compares tokens as bytes so a
non-ASCII refresh token cannot crash the token route.

* chore(ui): regenerate dashboard API types for the new MCP general setting

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:10:34 -07:00
devin-ai-integration[bot]
3a3b8cad80
feat(ui): add TypeSafe and Strands Decider to the Add Model provider list (#45335)
* feat(ui): add TypeSafe and Strands Decider to the Add Model provider list

* test(public_endpoints): type the provider fields helper against the route's model

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 04:04:07 -07:00
devin-ai-integration[bot]
0e26edfdb9
test: move 81 legacy live tests in pass-through, spend, batches, audio, search, guardrails, image and ocr dirs offline (#45288)
* test: make legacy live tests in spend, batches, openai endpoints and audio dirs offline (partial)

* test: migrate wave-1b live tests offline (guardrails, images, ocr, search, openai endpoints)

* test: fix wave-1b review items, add responses/ocr integration tests and firecrawl unit test

* test: anthropic messages router/bedrock/openai-bridge unit tests for wave-1b nodes

* test: anthropic messages logging, prompt-caching and tool-search unit tests; drop migrated base nodes

* test: finish pass_through_unit_tests nodes, logging drain fix and mutations

* test: migrate anthropic passthrough tests to integration wire tests

* test: fix passthrough migration wire spend row lookup and wildcard config

* test: migrate hosted vllm and openai file passthrough tests offline

* test: move assemblyai and vertex passthrough nodes to in-process unit tests

* test: restore unlisted router node and fix logging worker drain in passthrough unit tests

* test: drop spend-row BUG skip and sharpen non-streaming skip reason for anthropic messages

* test: use public presidio alias after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: restore the batch and file unit tests the migration rewrote away

* test: restore full legacy intent in anthropic messages router unit tests

Drop the false BUG skip on non-streaming aanthropic_messages logging (success
callbacks do fire; the skipped body filtered on the wrong model), assert the
logged model_group, messages, cost and usage, cover streaming logging for both
Anthropic and Bedrock invoke, assert dict content blocks for Anthropic, Bedrock
invoke and the OpenAI bridge, add the Bedrock invoke leg of the router test,
fall back from a real 401, test system-prompt caching and streaming
message_start cache fields on converse and invoke, send the legacy tool-search
tools and beta header, remove the type: ignore and bare dict helper, and add
in-process native /anthropic passthrough spend logging tests

* test: fix passthrough migration wire tests and drop false native spend BUG skip

The native spend row test read rows by the shared master-key digest, so it
matched other tests' rows; read each row by its own message id instead and
assert tokens, total, spend, tags, provider, api_base and end_user on both the
non-streaming and streaming native routes. Merge the streaming test that never
checked spend, assert exact tags from litellm_metadata, stop rebinding a Final
in a loop, require cost > 0, and add the chat-completions bridge cost case the
legacy test covered

* test: harden the spend, OCR, image, search and guardrail migration replacements

The OCR spend tests now build fresh kwargs per case instead of mutating a
shared fixture, and every payload case asserts the exact logged spend. The
OCR wire test reads its spend row by request id and checks exact page
pricing. Image edit, Nova Canvas, DuckDuckGo, Firecrawl, Bedrock guardrail
and Presidio replacements now fake only the provider HTTP boundary (respx or
an in-process aiohttp server) and assert the outbound request, so the
DuckDuckGo limit, Azure base_model pricing and guardrail masking are proven
rather than assumed

* test: cover Exa and Perplexity search structure and max_results offline and retire the two base search methods

* test: drive batch and file replacements through the provider HTTP boundary and real logging callback

Replace monkeypatched litellm.afile_content and AsyncHTTPHandler doubles with respx routes,
read batch logging metadata from a registered success callback instead of get_logging_payload,
use the real managed-files hook for the GEN-2166 regression, assert outbound request bodies,
pin poller ownership explicitly in the migrated DB-sync tests, and require the scripted
upstream to be hit in the responses error-status wire tests

* test: assert file content download headers pass through the proxy

* test: point spend coverage references at the tests that replaced the retired spend job

* test: assert passthrough identity, spend and route dispatch from the code under test

The AssemblyAI non-admin test asserted metadata it wrote itself and leaked a
background poll to the real AssemblyAI host. It now drives assemblyai_proxy_route
with a real Request and waits for the success callback for its own transcript id.
The Vertex spend test matches its log by call id instead of taking the first event.
The OpenAI files wire test hit the native /{provider}/v1/files route; it now calls
/openai/files so the passthrough is what forwards the upload and delete.

* test: mock only the HTTP boundary in the migrated audio tests

Vertex TTS tests no longer replace _ensure_access_token or AsyncHTTPHandler.post;
the token comes from a mocked Google OAuth endpoint and the synthesize call from
respx. Speech tests assert the outbound body, the transcription cache test polls
for the cache write instead of relying on test ordering, and the model pass-through
test checks the multipart model field per model.

* test: wait on a logger event instead of polling the clock in anthropic messages unit tests

Recorders keep payloads in a rebound tuple and set an asyncio.Event; tests
await it with asyncio.wait_for instead of a sleep-and-deadline poll loop

* test: freeze module-level batch and file response fixtures as Final MappingProxyType

* test: fake the presidio analyzer with an in-memory aiohttp connector

The blocked-entity tests started an aiohttp TestServer, which binds a local
socket. They now hand the guardrail a ClientSession whose connector answers
/analyze and /anonymize in process, so no socket is opened and the outbound
analyze text and entities are still asserted

* test: type anthropic messages router test helpers with LiteLLM's Anthropic TypedDicts

Messages, cached system blocks and tool-search tools now use
AnthropicMessagesUserMessageParam, AnthropicMessagesTextParam,
AnthropicToolSearchToolRegex and AnthropicMessagesTool instead of bare
dict shapes; tools are converted to plain dicts only at the acreate call,
whose tools parameter is list[dict]

* test: type batch limiter helpers with TypedDicts and wait on the logging callback event instead of polling

* test: give the migrated OCR, image and presidio helpers precise types

OCR spend helpers take ReadOnly TypedDicts for kwargs and responses and use
LiteLLM's OCRResponse/OCRUsageInfo instead of local pydantic stand-ins; spend
metadata is validated with a TypeAdapter. The presidio fake uses LiteLLM's
PresidioAnalyzeRequest/ResponseItem types, and the image-edit logger validates
the logged payload instead of storing an untyped dict

* test: signal callback and cache events instead of polling

Recorders keep tuples and set an asyncio.Event, thread-safely, when the payload
for this test's transcript id or upstream URL arrives. The transcription cache
test waits on a Cache subclass that signals after async_add_cache. No clock
polling or sleeps remain in these tests.

* test: assert the batch limiter hook updates the caller's request in place

* test: tolerate model-list probes and read native passthrough rows by owned key

The router's OpenAI-compatible model-info refresh (litellm/router.py:10710)
sends GET /v1/models to configured openai api_bases, so the wire answers it
with an empty list and excludes it from the provider-call assertions. Native
/anthropic spend rows are now read by a per-request virtual key digest and
call_type, then the row's request_id is checked against the message id

* test: expect the OCR alias in the proxy response model

The proxy restamps every OpenAI-compatible response model to the name the
client requested (_override_openai_response_model), so /v1/ocr returns the
scenario alias. The upstream model is now checked on the drained request body
instead of inside the peer, where a failed assert never reached the test

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 02:38:12 -07:00
devin-ai-integration[bot]
04d97abffb
test: make 77 legacy live tests offline in litellm_utils, router_unit and responses dirs (#45298)
* test: make 77 legacy live tests offline in litellm_utils, router_unit and responses dirs

* test: restore lost coverage in litellm_utils offline replacements

Bound live health-check tasks to the concurrency limit so eager task
creation behind a semaphore fails, route the Langfuse trace-id check
through completion -> Logging.get_trace_id with all four metadata
combinations, run the callback dedup checks through acompletion, pass
the dynamic key via litellm_params, cover the env-inferred default
model list, and add the per-provider audio transcription config lookup

* test: restore router coverage lost in the offline move

Add the missing non-stream LIT-3058 header/count-once test, pin the UTC minute on every
usage-counter test, assert router-level client reuse for transcription, and replace the
private selector-attribute checks with routing outcomes per strategy. Assert outbound bodies
for speech, rerank, image, assistants and moderation, add timeouts to event waits, and drop
the router_unit_tests husk files that no longer hold tests

* test: restore sync stream, router sync stream, error event and field-type coverage for migrated responses tests

Add the sync streaming logging and sync Router.responses streaming cases the
offline replacements dropped, port the legacy per-event and response field-type
validation, assert raw headers, the float created_at conversion, MCP tool headers,
the search_context_size input, and add an offline replacement for the in-stream
context-window error event. Remove helpers left dead in the legacy file.

* test: run numpydoc-backed unit tests in GHA and drop a test that pinned a crash

* test: restore sequence number, item id and content part checks in the responses stream validator

* test: match router logging events to the test's own deployment so late events from other tests are ignored

* test: restore the s3 cold storage history test to its original assertions

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 02:37:55 -07:00
Priyansh Nandwana
7c7b0ea85b
fix(bedrock_mantle): send OpenAI explicit prompt cache breakpoints for GPT-5.6 and newer (#38729)
* fix(cache_control): let a hosted deployment opt into the OpenAI cache dialect

_targets_openai_prompt_cache_breakpoint gated on custom_llm_provider == "openai"
unconditionally, so an OpenAI-shaped model served by another provider could never
qualify, even with supports_prompt_cache_breakpoint set explicitly on its own
cost-map entry. bedrock_mantle honours prompt_cache_breakpoint end to end, and
had no way to say so.

model_cost is keyed per exact deployment string, so a flag on the deployment's own
entry states the dialect more precisely than a provider name can. Consult it
before the provider check, and require the entry's litellm_provider to match the
serving provider so an openai entry cannot license another provider. Entries for
the openai provider keep their api_base check, so an OpenAI-compatible third-party
host is still not assumed to speak the dialect.

Fixes #38666

* test: drop a laziness assertion the existing suite already makes

test_provider_lookup_skipped_for_models_below_gpt_5_6 already pins that
_resolve_provider is not called for an unflagged model, and it covers the new
flag lookup unchanged. The duplicate patched a litellm internal for no added
coverage, which the test-quality gate counts against TQ008.

* fix(bedrock_mantle): flag GPT-5.6 and newer rows for OpenAI explicit prompt cache breakpoints

* fix(responses): predict the chat-completions bridge with the model name the dispatch resolves

* fix(cache_control): read a region-prefixed Mantle GPT deployment through its region-free price-map row

* fix(cache_control): let a bare hosted deployment name read its own provider's price-map row

* fix(responses): hand the hook the provider the router resolved for a Foundry GPT deployment

* fix(cache_control): read the deployment breakpoint flag strictly and default a null prompt_cache_options

A cost-map or deployment `supports_prompt_cache_breakpoint` that is not the boolean `true` (the string "true",
the integer 1, "false", 0, a long string) no longer opts a deployment into the OpenAI prompt cache dialect; only
`true` does, the same reading Bedrock Converse applies to its own flag.

A client that sends `"prompt_cache_options": null` on `/v1/chat/completions` or `/v1/messages` now gets the
implicit default the hook already applied on `/v1/responses` when it placed a breakpoint; before, the null
suppressed the default and the request left with a breakpoint and no options.

* test(integration): audit cells for the Mantle GPT prompt cache breakpoint dialect

Deterministic cells for the configured-breakpoint path on Bedrock Mantle GPT and Azure AI Foundry GPT-6
deployments: every endpoint, streaming and not, sync and async SDKs and raw httpx, the hostile option shapes,
flag precedence, the response cache twin, a two-worker burst, a worker kill and a graceful restart.

* fix(cache_control): ignore a malformed cache_control_injection_points value instead of failing the request

A deployment or client `cache_control_injection_points` that is not a list of points (a string, an
integer, a bare point dict, a list of strings) raised inside the prompt hook (`'str' object has no
attribute 'get'`, `'int' object is not iterable`) and turned every request to that deployment into a
500 on `/v1/chat/completions`, `/v1/messages` and `/v1/responses`. Every entry point now reads such a
value as no configured points, the way a `null` already read, and the request leaves without a
breakpoint; entries of a list that are not points are dropped and the point entries kept.

* test(integration): cover int, dict and string-list injection point shapes in the Mantle audit cells

The D9 cell now also sends an integer, a bare point dict and a list of strings as the deployment's
`cache_control_injection_points`, which the merge base answered with a 500 on every request.

* fix(cache_control): stamp the OpenAI dialect for a bare deployment name served by a flagged provider

A deployment written as `model: openai.gpt-5.6-sol` with `custom_llm_provider: bedrock_mantle` has no cost-map row
of its own and no openai row of the same name, so the dialect stamp's cheap gate (`supports_openai_prompt_cache_breakpoint`)
returned early and the configured point was never stamped. On `/v1/responses` and on the chat seeding path the hook
sees no provider, so it fell back to the Anthropic `cache_control` marker, which Bedrock Mantle strips.

The gate now also passes a model whose serving provider is already known and whose provider-keyed row carries the
flag, which is the same row the dialect resolution reads and costs no provider lookup. A bare name without a provider
is still left alone, so models below GPT-5.6 keep their points untouched and resolve nothing.

* test(integration): cover a bare Mantle deployment name with its provider in the audit cells

A deployment configured as `model: openai.gpt-5.6-sol` plus `custom_llm_provider: bedrock_mantle` sends the
breakpoint and the implicit options on all three endpoints; the merge base leaves it on the Anthropic dialect.

* fix(cache_control): keep a configured point that sits beside junk entries on every chat seed path

`configured_injection_points` kept the point dicts of a mixed `cache_control_injection_points` list as a tuple,
but only read a `list` back. The chat seed and the `/v1/responses` dialect stamp write the normalized value back
onto the request, and on a deployment the OpenAI dialect does not stamp (an Anthropic model, Claude on Bedrock
Mantle) that value is the tuple itself, so the hook then read it as no configured points and the valid point was
silently dropped. The stamped OpenAI dialect and the `/v1/messages` path kept it only because they build a new list.

The normalizer now reads back the tuple it wrote, so the point reaches the wire on every path; an all-dict list
still passes through as the same object.

* test(integration): cover a configured point beside junk entries on a Claude Mantle deployment

A deployment configured with `cache_control_injection_points: ["system", {system point}, 3]` marks the system
block on the Anthropic dialect on all three endpoints; the merge base answers 500 on the junk entry.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 02:07:22 -07:00
devin-ai-integration[bot]
7a659973e3
revert(lint-gates): accept an empty base scan again, since zero violations is a legitimate count (#45319)
This reverts commit 7c7efd4d55 (#45314).

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 08:40:08 +00:00
devin-ai-integration[bot]
7dd9aff244
refactor(proxy): type fresh Moyai connect helpers and drop MCP rpm getattr (#45315)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 01:31:59 -07:00
devin-ai-integration[bot]
7c7efd4d55
fix(lint-gates): fail on an empty base scan instead of blaming the change for every violation (#45314)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 08:06:18 +00:00
tin-berri
057034d0d3
fix(lens): refresh runs until gateway costs are complete (#45105) 2026-10-08 00:24:41 -07:00
devin-ai-integration[bot]
8683f1418e
fix(utils): make supports_audio_output read the supports_audio_output cost-map key (#45294)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 23:54:38 -07:00
devin-ai-integration[bot]
9df3d4c96b
test(integration): boot owned proxies with a readiness-sized worker healthcheck budget (#45292)
* test(integration): boot owned proxies with a readiness-sized worker healthcheck budget

* test(integration): keep owned_proxy_process's extra_arguments for cells that pass their own flags

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 23:40:26 -07:00
devin-ai-integration[bot]
91f36f1774
chore: bump litellm-proxy-extras to 0.4.107 and litellm-enterprise to 0.1.75 (#45291)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 23:03:01 -07:00
devin-ai-integration[bot]
da2bb5a6b1
test(e2e): cover responses API gaps on gemini, azure, compact and context management (#45248)
* test(e2e): cover responses API gaps on gemini, azure, compact and context management

* test(e2e): assert streamed usage cost and split anthropic strict schema test

* test(e2e): assert compaction items and final stream event in responses e2e

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 22:48:29 -07:00
devin-ai-integration[bot]
d6e453f74b
ci(lint): gate every rule on its merge-base count and cap Anys at a fixed total (#40193)
* ci(lint): derive gate ceilings from merge-base counts and drop the budget files

The four lint gates (ruff strict, type discipline, basedpyright, test quality)
now fail a branch only when a rule's codebase count grows past its count at the
merge-base with litellm_internal_staging plus a fixed per-rule headroom, which
is zero everywhere except the LIT010/LIT011 and reportAny/reportExplicitAny
seeds. Base counts come from a disk cache, then the CI artifact the renamed
publish-lint-base-counts workflow uploads for every staging push (all four
checkers, one artifact per checker and sha), then a scan of the base worktree.

The four *-budget.json files, make lint-budget-update, budget_ratchet_check.py,
and the unratcheted check are gone, so no PR carries a budget edit again.

* test(lint-gates): assert checker identities by behavior and type the new gate tests

* test(lint-gates): resolve base points over an injected git so unit tests stay in-process

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 05:39:29 +00:00
Mateo Wang
fbbb54c841
fix(proxy): answer a rejected websocket handshake without crashing the HTTP exception handler (#45251)
* fix(proxy): answer a rejected websocket handshake without crashing the HTTP exception handler

* fix(proxy): read the OTLP route check's scope keys with .get and drive the websocket auth tests through the real auth path

* test(proxy): build the websocket auth test keys without reading the clock
2026-10-07 22:27:46 -07:00
Mateo Wang
42d2158894
fix(responses): map input_audio blocks in the chat-to-Responses bridge (#45224)
* fix(responses): map input_audio blocks in the chat-to-Responses bridge

* fix(responses): honor base_model when deciding to drop input_audio under drop_params

* fix(responses): carry a dropped audio part's cache breakpoint to the preceding part

Dropping an input_audio part under drop_params lost the prompt_cache_breakpoint
riding on it, so an injected system marker landing on a trailing audio block left
prompt_cache_options explicit with no breakpoint on the wire. The marker now moves
to the nearest preceding kept part when that part has none.

Update the CircleCI-only Responses bridge wire tests that pinned input_audio to
its stringified input_text shape: video_url becomes the stringified kind so the
marker-on-stringified-block coverage survives, and dedicated audio tests assert
the forwarded part and the drop_params carry-over.

* refactor(responses): carry dropped audio markers in one pass

Replace the per-block slice that re-read the rest of the content list for
every retained part with a single reverse scan, and skip the copy for a
message that carries no audio part. Type the unit test helper's messages
as a sequence of mappings instead of bare lists.

* test(responses): add integration cells for input_audio parts through the bridge

Cover the chat-to-Responses bridge's input_audio handling on the real proxy: the part is forwarded
as input_audio without drop_params and dropped under drop_params unless the model or its base_model
supports audio input, across httpx, the OpenAI SDK (sync stream and async), the Anthropic SDK, native
/v1/responses, tool and assistant messages, hostile input_audio values, duplicate parts, an audio-only
message, marker carry-over, a null drop_params, a global litellm_settings.drop_params, a mixed burst,
and a burst with dropped upstream connections

* test(responses): check every upstream attempt of a dropped connection

The dropped-connection cell required exactly one upstream POST per call,
which tied it to the retry policy instead of the fix. It now groups the
recorded POSTs by marker, requires one attempt per answered call, and
checks that every attempt of a doomed call carries the dropped audio shape.
2026-10-07 22:25:35 -07:00
Mateo Wang
328f5a720c
fix(router): honor a per-request fallbacks list on a mid-stream fallback (#45227)
* fix(router): honor a per-request fallbacks list on a mid-stream fallback

The completion entrypoints record the per-request fallbacks, context_window_fallbacks, and content_policy_fallbacks in the mid-stream controls carrier the Responses and Messages paths already use, and each completion attempt restores them into the kwargs its stream re-enters the fallback chain with, so a streaming request that dies before its first chunk fails over to the list the request named instead of the router-level list only. The sync vs async parity cell now asserts the backup's text on both twins.

* test(router): annotate the new fallback test locals as Final and flatten the leak check

* fix(proxy): skip null key and team router settings when merging per-request overrides

* test(proxy): annotate the null router settings test locals and drop a redundant comment
2026-10-07 22:25:22 -07:00
Mateo Wang
a7ee038592
fix(router): keep include_fallback_errors off the provider call on the sync Router path (#45218)
* fix(router): keep include_fallback_errors off the provider call on the sync Router path

Router.completion spread include_fallback_errors into litellm.completion, so OpenAI
answered 400 Unknown parameter and a fallback walk lost the backup's reply. Only
Router._acompletion popped the router-only flags. Both paths now build the provider
call through one shared helper that drops them, and a wire-level respx test covers
each path.

* test(router): pin include_fallback_errors off the wire across the sync, async, and proxy paths
2026-10-07 22:25:12 -07:00
devin-ai-integration[bot]
a3b280f69d
feat(lint): add LIT015 requiring pydantic models to be frozen (#42348)
* feat(lint): add LIT013 requiring pydantic models to be frozen

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): let an explicit frozen=False override earlier or inherited frozen=True in LIT013

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(lint): annotate LIT013 helper locals with Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(lint): document the frozen pydantic model rule as LIT015

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(lint): honor the frozen class keyword in LIT015

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(lint): annotate main and scan_paths locals as Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): cover LiteLLMBaseModel, BaseSettings, and SettingsConfigDict in LIT015

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): resolve LIT015 model bases per class node

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): treat the OpenAIObject alias as a LIT015 model base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): resolve shared LIT015 config constants and list frozen-ok in the gate hint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): resolve LIT015 shared config at class definition

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 22:18:00 -07:00
Mateo Wang
33d908e0ae
feat(decisions): add OpenAI as a Decisions provider behind a shared decisions format (#45214)
* feat(decisions): serve System One format at /v1/systemone and OpenAI format at /v1/decisions

The System One request format moves to /v1/systemone and /systemone. /v1/decisions and
/decisions now accept the OpenAI Decisions API format, translate it into a System One
request, route it through the same pipeline, and translate the answers back.

* fix(decisions): import assert_never from typing_extensions for Python 3.10

* fix(decisions): cap OpenAI-format questions at the System One limit

/v1/decisions accepted up to 200 questions, but the System One request it translates to takes at most 128, so 129 to 200 questions failed with an internal validation error. Both now share MAX_DECISION_QUESTIONS.

* fix(decisions): return 400 for bodies that are not JSON

* test(decisions): send System One bodies to /v1/systemone in integration tests

* feat(decisions): add OpenAI as a Decisions provider

System One requests to openai/ models are translated to OpenAI's Decisions shape on the way out and OpenAI's answers are translated back, so both /v1/systemone and /v1/decisions can route to gpt-6-luna.

* refactor(decisions): move the OpenAI wire translation into llms/openai

* feat(decisions): route both request formats through a shared decisions IR

System One and OpenAI-format bodies now convert to one internal representation, and each provider translates it to its own wire format. OpenAI deployments receive the caller's messages, images, names, typed choice values, level descriptions and safety_identifier unchanged. System One providers return a 400 for image input. Cached and cache-write tokens are priced on both response shapes, and provider response extras survive the round trip.

* fix(decisions): keep provider fields inside System One answers

* fix(decisions): load on Python 3.10 and 3.11 and bill cached tokens once under custom pricing

Decisions IR dataclasses used a mappingproxy default, which is only hashable from Python 3.12, so import litellm failed on 3.10 and 3.11. Decisions usage now reaches cost_per_token as a Usage with prompt_tokens_details, so custom pricing no longer adds cached and cache-write tokens on top of input_tokens that already include them

* fix(decisions): honor litellm OpenAI key and base settings for the openai provider

OpenAI decisions now resolve the key from litellm.api_key, litellm.openai_key, then OPENAI_API_KEY, and the base from litellm.api_base, OPENAI_BASE_URL, then OPENAI_API_BASE, matching other OpenAI calls. Providers own their configured key and base lookup through the endpoint config.

* test(decisions): clear every OpenAI base setting in the proxy OpenAI deployment test

* fix(decisions): send OpenAI instructions for System One questions without them and reject single-option questions

* fix(decisions): return guardrail_information on OpenAI-format decisions when requested

* refactor(decisions): validate proxy request data before reading guardrail settings
2026-10-08 04:57:41 +00:00
Mateo Wang
8fc31b1e4e
feat(decisions): serve System One format at /v1/systemone and OpenAI format at /v1/decisions (#45184)
* feat(decisions): serve System One format at /v1/systemone and OpenAI format at /v1/decisions

The System One request format moves to /v1/systemone and /systemone. /v1/decisions and
/decisions now accept the OpenAI Decisions API format, translate it into a System One
request, route it through the same pipeline, and translate the answers back.

* fix(decisions): import assert_never from typing_extensions for Python 3.10

* fix(decisions): cap OpenAI-format questions at the System One limit

/v1/decisions accepted up to 200 questions, but the System One request it translates to takes at most 128, so 129 to 200 questions failed with an internal validation error. Both now share MAX_DECISION_QUESTIONS.

* fix(decisions): return 400 for bodies that are not JSON

* test(decisions): send System One bodies to /v1/systemone in integration tests

* fix(decisions): return guardrail_information on OpenAI-format decisions when requested

* refactor(decisions): validate proxy request data before reading guardrail settings
2026-10-07 21:41:03 -07:00
Mayuri
9506f9e58b
fix(proxy): admit litellm_proxy/hosted_vllm in provider-endpoint discovery (#38617)
* fix(proxy): admit litellm_proxy/hosted_vllm in provider-endpoint discovery

get_provider_models gated all endpoint discovery on membership in the
static litellm.models_by_provider dict, which litellm_proxy and
hosted_vllm are intentionally absent from since their model list only
exists behind the provider's own endpoint. GET /v1/models returned the
literal wildcard string instead of the expanded model list for these
providers. Now falls through to ProviderConfigManager, which already
knows about them, before giving up.

* fix(tests): stop patching an SDK internal in the discovery regression test

TQ008 flagged patching litellm.proxy.auth.model_checks.get_valid_models.
Assert the gate's own return value instead of mocking past it.

* chore: drop redundant comment per repo convention

* test(proxy): cover litellm_proxy wildcard model discovery end to end

---------

Signed-off-by: mayuriphad <163738104+mayuriphad@users.noreply.github.com>
2026-10-07 21:15:57 -07:00
ishaan-berri
03ba79d261
feat(lens): show end-user feedback on traces, stored in ClickHouse (#45171)
* feat(lens): add lens_feedback ClickHouse table for human trace scores

ReplacingMergeTree(UpdatedAt, IsDeleted) keyed like agent_traces_by_key so
feedback joins traces in sort order, with a bloom filter on TraceId, a score
CHECK, and the same retention as traces

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add scoped ClickHouse reads for trace feedback

feedback_target resolves a visible trace's team, key and trace_ref before a
write; feedback lists the latest live entry per author; feedback_summary
returns count, average and lowest score for a batch of traces

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens): pin the Claude session to trace id hash shared with feedback

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): expose feedback queries to the Python trace bridge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add typed feedback models and ClickHouse feedback store

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add /lens/feedback API for 0-10 human scores on traces

PUT saves or replaces the caller's score and comment, GET lists every
entry, DELETE removes the caller's, POST /summary feeds the trace list.
Callers can target a trace_id or a Claude Code session_id

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add typed feedback calls to the traces API

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): flag rated traces and filter the list by feedback

A Feedback column shows each run's average score and rating count, marked
red when any score is 4 or below, and a filter narrows to rated or
low-score runs through the URL

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show and edit human feedback inside a trace

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): let app keys post end-user feedback on their own traces

Any key can now PUT/DELETE feedback on traces its team or key sent, naming
the end user in an optional user field. Reading feedback stays with Lens
admins

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show end-user feedback first in a run and flag low-scored rows

Replace the edit popover and feedback filter with a read-only view: the
list shows each run's score in a fixed-width badge and tints the row red
when a user scored it 4 or below, and opening a run shows what users said
with their score right under the header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 20:14:30 -07:00
devin-ai-integration[bot]
c912236fa5
fix(proxy): declare the Moyai settings write's service target and allowlist its routes (#45220)
Some checks failed
Unit Tests / integrations (push) Blocked by required conditions
Unit Tests / OpenAI and Meta Providers (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Blocked by required conditions
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Blocked by required conditions
Unit Tests / proxy-extras (push) Blocked by required conditions
Unit Tests / Vertex AI (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-feature-endpoints (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests: Proxy DB Operations / Lens Python 3.10 (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Publish basedpyright base counts / publish (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
* fix(proxy): declare the config_params service target for the Moyai UI settings write

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exercise the Moyai settings write against a real DualCache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allowlist the Moyai quick-connect routes on the backend component

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): observe the Moyai settings write target through a composed DualCache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 20:03:44 -07:00
ishaan-berri
b85756104f
feat(lens): show agent, user and slack thread first in the run header (#45261)
* feat(lens): add who started the run to RunSource

* feat(lens): carry agent.source.user on trace span rows

* feat(lens): resolve the run source user from the span row

* feat(lens): read agent.source.user in trace_spans query

* feat(lens): read agent.source.user in trace_page_spans query

* feat(lens): read agent.source.user in trace_span_batch query

* feat(lens): read agent.source.user in trace_list_span_batch query

* feat(lens): decode the source user from clickhouse span rows

* test(lens): add source user to trace cache test rows

* test(lens): add source user to trace cache read fixtures

* test(lens): add source user to trace cache snapshot fixtures

* test(lens): add source user to capture fixture rows

* test(lens): round-trip source user on span row contract

* test(lens): cover run source user resolution

* chore(lens): regenerate python trace types with source user

* chore(lens): regenerate trace json schema with source user

* chore(lens): regenerate trace page json schema with source user

* chore(ui): add source user to api types

* feat(lens): add run user chip and slack thread chip

* feat(lens): put agent, user and thread on the first row of the run header

* test(lens): cover the run header identity row and compact totals

* feat(lens): start demo runs from a slack thread
2026-10-07 20:01:55 -07:00
devin-ai-integration[bot]
0e09e99146
refactor(decisions): dispatch /v1/decisions through provider configs and the shared HTTP handler (#45130)
* refactor(decisions): dispatch /v1/decisions through provider configs and the shared HTTP handler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): map OpenRouter connection failures to APIConnectionError and drop explanatory docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): keep APIConnectionError for unreachable deployments on every provider

`_map_upstream_exception` now maps a status `_handle_error` synthesized for a
non-HTTP failure to `APIConnectionError` before `exception_type` runs, so a
Perplexity (openai-compatible) deployment that cannot be reached no longer
answers `InternalServerError` where main answers `APIConnectionError`.

Audit cells: the unreachable-deployment integration cell is parametrized over
the five providers, OpenRouter connection failures are pinned on chat, stream,
embeddings, decisions and the SDK, a deployment whose provider has no decisions
config gets the 400 naming every supported provider, and a unit cell pins the
mapper alone. The first docstring sentence of the five translation bases is
dropped.

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 19:45:45 -07:00
ishaan-berri
e4d9a591dd
fix(lens-ui): open traces on the top-level step and stop showing model names as input (#45254)
* fix(lens-ui): open a trace on its top-level step instead of the first failed one

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): pin that failed runs open on the root step

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens-ui): drop the unused firstErrorSpan helper

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): stop testing the removed firstErrorSpan helper

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(lens-ui): stop showing the model name as a run's input when none was recorded

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover the trace list row with no recorded input

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 02:26:50 +00:00
devin-ai-integration[bot]
4f49f88172
fix(proxy): keep idle responses websockets open until a configurable session limit (#44433)
* add config-driven bound to /responses websocket connection

* test(proxy): cover responses websocket session limit and idle first frame

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ci): sync API schema and Ruff formatting

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): close responses websocket client before upstream cleanup at session limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert recorded websocket outcomes and simplify session reaping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): move responses websocket session limit timing coverage to integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): run provider-close checks off the event loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover handshake auth, subprotocol, DB override and worker kill for the responses websocket session limit

Adds the audit cells for the configurable session limit: a rejected bearer fails the handshake with 403, a requested Sec-WebSocket-Protocol is echoed back, a DB override through /config/field/update caps new sessions on every worker without a restart while sessions accepted before it keep their limit, deleting the override restores the default, out-of-range and wrong-type updates answer 400 and leave the stored row alone, and SIGKILLing the worker holding idle sockets drops only its sockets while the proxy keeps serving

* test(integration): open override sockets until both proxy workers hold one

uvicorn workers share one listening socket and a worker that wakes first accepts a whole simultaneous burst, so eight sockets opened at once can all land on one worker. The DB override cell now opens its sockets one at a time, alternating the two paths, until at least eight are open and every worker holds one, bounded at forty

---------

Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 19:26:30 -07:00
yucheng-berri
efe0715831
fix(proxy): accept router-wide default_litellm_params in required-param validation (#44483)
* fix(proxy): accept router-wide default_litellm_params in required-param validation

The required-body-param check added in #43787 only consulted request data
and deployment-level litellm_params, so a param supplied solely by
router_settings.default_litellm_params (e.g. max_tokens for
/v1/messages, documents for /rerank) was rejected with a 400 at route
entry even though the router would have injected it during dispatch.
Consult the router-wide defaults in the same place, treating None-valued
defaults as absent to mirror the setdefault merge.

* fix(proxy): consult only the defaults the dispatching router will merge

Review follow-up: router-wide defaults are now consulted only where
dispatch actually applies them. user_config requests dispatch on their
own throwaway Router, so they consult that config's defaults instead of
the global router's; search, managed-agent and eval routes, and
model-less direct dispatch never pass through the router's defaults
merge, so they keep the route-entry 400. Also adds rerank router-default
coverage (unit + e2e) and drops the redundant source comment.
2026-10-07 19:25:28 -07:00
devin-ai-integration[bot]
02b8b5cc80
fix(proxy): traceparent/baggage fallback must not override caller metadata (#43688)
* fix(proxy): traceparent/baggage fallback must not override caller metadata

The W3C traceparent/baggage fallback in
add_litellm_metadata_from_request_headers documents itself as last-resort:

    Lower priority than everything above - only fires when neither the
    explicit litellm headers nor the Anthropic-metadata path found anything

But it guards on the top-level `litellm_trace_id` / `litellm_session_id` body
keys and never checks `metadata`, which is the documented way callers set
trace_id / session_id (`metadata: {"trace_id": ...}` on /chat/completions).
A caller that explicitly sets metadata.trace_id has it silently replaced by the
header value, so the implementation contradicts its own stated precedence.

This is not a corner case on managed platforms: GCP's front end injects a
traceparent into every inbound request, so the fallback fires on traffic whose
caller never sent the header. The request still returns 200 and the trace still
reaches the logging backend, just under an id the caller never chose, so any
caller correlating by its own id silently fails to find its trace.

Guard both fallbacks on the caller's request-body metadata as well.
Deliberately narrow:
- x-litellm-trace-id still outranks the body (documented priority #1)
- a request that steers neither field still adopts traceparent/baggage exactly
  as before
- steering is per-field: setting only trace_id still lets session_id come from
  baggage
- litellm_metadata is checked too, for LITELLM_METADATA_ROUTES (/v1/responses,
  /v1/messages, batches, files)

Also corrects the comment, which understated the guard.

4 new tests; each fails without the source change. The existing traceparent and
baggage tests are unchanged and still pass.

* fix(proxy): check only the active metadata container and ignore empty values

Addresses review. The first version treated any value in either `metadata` or
`litellm_metadata` as caller steering. On LITELLM_METADATA_ROUTES the body
`metadata` is provider-facing and is only promoted later, so a session_id there
suppressed baggage before apply_missing_session_id_policy ran, and a request
with a usable baggage session id got a 400 under `missing_session_id: reject`.
An empty session_id did the same, since the policy treats "" as absent

Now only the active metadata container counts, and only a truthy value, which
matches how apply_missing_session_id_policy decides a session id is present.
The helper takes a typed `object` rather than a bare dict

Adds regression tests for both cases, including the end to end reject path,
and drops the test docstrings

* fix(proxy): count promoted caller trace ids on litellm_metadata routes

Addresses review. On LITELLM_METADATA_ROUTES the proxy promotes the caller's
trace control fields (trace_id, session_id, ...) from `metadata` into
`litellm_metadata`, but only after the header fallback runs. Checking only
`litellm_metadata` let traceparent and baggage fill those fields first, and the
promotion then skipped them because they were already set, so the trace was
recorded under the header ids

The check now also reads `metadata` for fields in
LITELLM_TRACE_CONTROL_METADATA_FIELDS on those routes, and
apply_missing_session_id_policy uses the same check, so `reject` no longer
refuses a request whose session id is about to be promoted

The earlier test asserting that a session_id in `metadata` must not block
baggage on /v1/responses had the premise wrong, since that value is promoted.
It is replaced by tests that assert the promoted caller ids win on
/v1/responses and /v1/messages, with no policy, reject, and generate

* fix(proxy): an empty trace field in litellm_metadata shadows the promoted one

Addresses review. Promotion copies a trace control field from `metadata` into
`litellm_metadata` only when the key is absent, so a key that is present but
empty in `litellm_metadata` wins and the `metadata` value is never used. The
check still looked at `metadata` in that case, which let
`missing_session_id: reject` accept a request that ends up with an empty
session id, and let an empty trace_id or session_id block the headers

The check now follows the same rule as promotion. If the active container has
the key, its value decides. Only when the key is absent does the promoted
`metadata` value count. This makes the conflicting-bucket cases behave exactly
as on main

* fix(proxy): stay within the LIT006 cast budget

The type-discipline gate failed: this branch added 4 unsuppressed `cast()`
calls and LIT006 was already at its limit. Each cast now carries a
`# cast-ok` reason. The two in the helper follow an `isinstance` check that
proves the Mapping. The two at the call sites stay because the method's
`data` parameter is a bare `dict`, and dropping those casts adds
basedpyright unknown-type errors instead. No behavior change

* test(proxy): cover caller metadata precedence over W3C trace headers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): generated session uses promoted caller trace id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): move test context into assertion messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for w3c fallback precedence

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tighten audit chaos cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(test): drop stray blank lines

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): ignore body-less model-info probes in langfuse precedence upstreams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep root trace/session fields on equal ids and usable caller values

The _caller_trace_field gating added for caller-metadata precedence was
presence-aware, which regressed three pre-call behaviors versus main:

- Equal ids in W3C headers and body metadata suppressed the
  traceparent/baggage branch entirely, leaving the root
  litellm_trace_id/litellm_session_id unset so router fallbacks minted a
  fresh uuid4 per request. The header branch now also fires when the
  caller value equals the header-derived id, so equal ids stamp the root
  fields exactly like main.
- A truthy but non-string metadata value (e.g. session_id 4815162342)
  counted as caller-supplied and suppressed the baggage fallback, so
  downstream str-only consumers (code interpreter sandbox reuse) got a
  new sandbox per turn. _caller_trace_field now counts only non-empty
  string values, and the generate policy falls through to generation
  when the caller value is not usable.
- On litellm_metadata routes a usable caller session id suppressed the
  missing_session_id policies while never landing on the root field, so
  the root session stayed unset. The policy now promotes the caller's
  usable session id to litellm_session_id instead of leaving it stranded.

* test(integration): non-string caller ids fall back to W3C like empty ids

The audit cell compared a headers leg against a headerless leg for
equality, which only held while a truthy non-string id suppressed the
W3C fallback on both legs. With unusable values ignored again, the
headers leg resolves to the header ids while the headerless leg cannot,
so assert the headers leg's concrete outcome instead, mirroring the
empty/null cell.

* style(proxy): trim the session promotion comment to the non-obvious why

---------

Co-authored-by: Filipe Andujar <filipeandujar@gmail.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 19:24:50 -07:00
moe-berri
dc2a14a656
feat(lens): simplify deployment and first trace setup (#45230)
* feat(lens): simplify deployment and first trace setup

* test(lens): keep setup fixtures within lint budgets

* fix(lens): preserve setup state and harden bundled storage startup

* fix(lens): handle setup recovery and Helm endpoint boundaries
2026-10-07 19:22:56 -07:00
Mateo Wang
ed9c02d2d9
fix(ollama): one stream id, 400 on non-text tool content, separated tool results, no empty message item (#45177)
* fix(ollama): one stream id, 400 on non-text tool content, separated tool results, no empty message item

The ollama/ text stream minted a fresh chatcmpl- id per chunk; every chunk now carries the one id of its response.
A tool message with int content raised a TypeError inside ollama_pt and answered 500; it now answers 400 naming the
message index and the type it carried. Merged user, tool, and function turns and the text parts of one message were
joined with no separator; they are joined with a newline. A streamed tool call with no text rebuilds to content null
instead of an empty string, and the Responses bridge no longer opens a message item on a leading empty delta, so a
/v1/responses stream on a tool call completes with the function_call alone, as the non-stream call does

* fix(ollama): defer the BadRequestError annotation so litellm imports on Python 3.10 to 3.13

* fix(ollama): read the user run lazily, type the prompt helpers, and 400 on a non-string text part

* fix(ollama): 400 on a text or image_url part that carries no value

* fix(responses): stream bridged output items at contiguous indexes and reject non-object ollama content parts

The chat-completions bridge reserved output_index 0 for the message item, so a
tool-only stream announced its function_call at index 1 with no item at 0 and the
OpenAI SDK's responses.stream() accumulator raised IndexError. Items now take the
next index in emission order, reasoning included, and response.completed lists
the output in that streamed order.

ollama_pt answers 400 for a content part that is not an object and for an
image_url object without a url string, instead of a 500 or a silent drop.

* fix(responses): list only announced items in the bridged completed snapshot

The Responses bridge drops a message item from response.completed when the
stream never announced one and its text is empty, so a tool-only stream and a
reasoning-only stream end with the items the client saw. The chunk builder's
rebuilt content goes back to what main does; aligning it with the non-stream
shape is its own change for every provider's consumers

* style(responses): sort the itertools import with the other stdlib imports

* test(integration): add ollama prompt-tools and responses bridge audit cells
2026-10-07 19:08:26 -07:00
Mateo Wang
58258409c9
refactor(proxy): expose public names for private proxy helpers (#45170)
* refactor(proxy): expose public names for private proxy helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep original class names behind public aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve internal callback filtering

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep _PROXY_ class names for managed files hooks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep old private names in package exports

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve recursive auth helper name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): align MCP limiter tests with server enforcement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): match main's MCP limiter tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep old private names bound in importing modules

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve compatibility imports through strict lint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): use exact pyright suppression in password helper test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): add reasons to compatibility import noqa comments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): update IN-list baseline for renamed helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 18:59:27 -07:00
ishaan-berri
b970e412d9
perf(lens): faster trace opens and list pages at scale (#45228)
* perf(lens): read trace spans in larger batches

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* perf(lens): evaluate the trace list page once per identity lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): queue trace reads instead of rejecting past 8 in flight

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 01:44:19 +00:00
ishaan-berri
ef694bc584
fix(ui): address usage redesign review findings (#45229)
* fix(ui): drill into an agent on the tags it was counted from, not every version

* test(ui): assert agent drill-in tags do not double count

* fix(ui): key chart series and totals so model names and repeated dates cannot collide

* test(ui): cover series key collisions and cross-year totals

* fix(ui): give stacked chart series display labels and a date tick formatter

* fix(ui): pass series labels and ISO dates to the top models chart

* fix(ui): color leaderboard rows by series label

* fix(ui): count agent coverage across all agents before the top ten

* fix(ui): key the flat cost series separately from model names

* fix(ui): show top agents only on the unfiltered admin global view

* fix(ui): skip the tag summary fetch when top agents is hidden

* fix(ui): build segmented controls on radio group for arrow key support

* test(ui): cover segmented control keyboard navigation

* chore(ui): drop descriptive styling comments

* chore(ui): drop descriptive styling comment from team user spend card

* chore(ui): drop descriptive styling comments from top model view

* chore(ui): drop descriptive palette comment

* chore(ui): drop descriptive styling comments from top key view

* chore(ui): drop descriptive comment from activity metrics

* chore(ui): drop descriptive layout comments from user agent activity
2026-10-08 01:43:15 +00:00
berriai-litellm-provider-info-sync[bot]
dcf92b44ba
fix(bedrock_mantle): lower claude sonnet 5.5 cache read price to match bedrock runtime (#45223)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 18:35:29 -07:00
Mateo Wang
9d3bf29d6d
fix(responses): keep cache breakpoints on blocks the bridge stringifies (#40032)
* fix(responses): keep cache breakpoints on requests that reach the Responses API

The chat-completions bridge rebuilt every content block without the
prompt_cache_breakpoint marker the cache-control hook had just placed on it, so
the request went out with prompt_cache_options set to explicit mode and nothing
actually marked. Explicit mode caches only what is marked, so those deployments
lost the implicit caching they were getting before the injection point was added

The native surface had a second miss. The bridging check looked the provider
config up with the still-prefixed model while the config keys on the bare one,
so every bedrock_mantle model read as having no native Responses support, and
role-targeted injection points were deferred to a chat-completions pass that
never runs for them

* fix(responses): keep prompt cache breakpoint on stringified content blocks

* test: cover the responses/ routing prefix in the native cache point case

* fix(responses): keep only the fallback-branch marker carry, main already routes the hook

#42281 landed the text, image_url and file branch carry and the implicit default
on main, and the Responses routing lookup this branch changed has no observable
effect there, so the PR shrinks to the unknown-block fallback branch and its
regression test

* fix(responses): drop a malformed prompt_cache_breakpoint under drop_params in the chat-to-Responses bridge

* fix(responses): accept the 30m prompt_cache_breakpoint ttl under drop_params

OpenAI's Responses API takes a ttl of 30m on an explicit breakpoint marker, so the drop_params validation keeps it instead of stripping it as an unknown key. Trims the bridge test docstrings to the dated vendor citation.

* test(integration): audit cells for prompt_cache_breakpoint through the chat-to-Responses bridge

Adds the /audit cells for the bridge's prompt_cache_breakpoint carry: valid and malformed markers on every
carrying block kind and on tool and assistant content, the drop_params on and off contract, the hook-injected
system marker, the Anthropic SDK path through the Responses adapter, upstream errors, idempotent spend rows,
and the chaos cells (mixed burst, upstream outage, slow streams, worker SIGKILL and a proxy restart mid burst),
all against a scripted Responses endpoint with the request it received read back by response id

* test(integration): give every scripted responses reply its own id

The prompt-cache-breakpoint audit's scripted upstream minted the response id
from the request's marker, so three identical marked requests shared one id.
The spend-log writer skips rows whose request_id already landed, and the
idempotent-logging cell saw one row for three calls on every leg. The upstream
now mints a unique id per reply like a real provider, and the cells match a
response to its request by the marker inside that id.
2026-10-07 18:28:39 -07:00
berriai-litellm-provider-info-sync[bot]
c52082e10c
fix(bedrock): lower claude sonnet 5.5 cache read price to the aws offer index (#45216)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 18:23:28 -07:00
ishaan-berri
5410752f43
feat(ui): redesign usage page with stacked model and agent charts (#45221)
* feat(ui): add shared StackedUsageChart from the model leaderboard chart

* test(ui): cover StackedUsageChart palette and tooltip rows

* feat(ui): export StackedUsageChart from shared charts

* refactor(ui): render model leaderboard chart with StackedUsageChart

* feat(ui): add usage overview series, leaderboard and totals helpers

* feat(ui): add Panel, Stat, Segmented and Sparkline usage primitives

* feat(ui): map model names to provider logos

* feat(ui): add model leaderboard rows with provider logo tiles

* feat(ui): redesign the usage overview tab

* feat(ui): group user-agent tags into named agents with logos

* test(ui): cover agent grouping against prod-shaped user-agent tags

* feat(ui): add top agents chart and table to usage overview

* feat(ui): add useTagSummary hook for tag spend and tokens

* feat(ui): wire redesigned overview, top agents and provider spend into usage

* test(ui): update usage page tests for the redesigned overview

* feat(ui): compact usage view selector with a dropdown that fits its content

* test(ui): update usage view selector tests

* feat(ui): show provider spend as a share bar over the provider table

* test(ui): assert provider share segments instead of the donut

* feat(ui): restyle team, org, tag and customer usage to match the overview

* test(ui): update entity usage tests for the restyled views

* style(ui): restyle team user spend card

* style(ui): restyle top model view

* test(ui): update top model view tests

* style(ui): restyle endpoint activity tab

* style(ui): label endpoint bars Successful and Failed

* test(ui): update endpoint bar chart tests

* style(ui): restyle endpoint line chart

* test(ui): update endpoint line chart tests

* style(ui): quiet endpoint table header

* style(ui): align usage AI chat panel typography

* style(ui): restyle key activity tab with the overview stat strip

* style(ui): replace top key toggle buttons with segmented controls

* test(ui): update top key view tests

* style(ui): restyle model and MCP activity tabs

* test(ui): update activity metrics tests

* feat(ui): restyle user agent activity and accept an initial agent filter

* test(ui): update user agent activity tests

* style(ui): outline export button in usage headers
2026-10-07 18:18:42 -07:00
mrinal-berri
0e076008e1
fix(bedrock): map Scheduled batch jobs to in_progress on retrieve (#45157)
* fix(bedrock): map scheduled batch job status to in_progress

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): cover batch retrieve status lifecycle through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): drop legacy covers marker and type job payload as Mapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 18:08:51 -07:00
Mateo Wang
c4b48f6b3f
fix(anthropic): fail closed when a token-file or inline-token federation credential is missing a rule or organization id (#45178)
* fix(anthropic): fail closed when a token-file or inline-token federation credential is missing a rule or organization id

* refactor(anthropic): flatten the federation-request helper

* test(anthropic): annotate the new WIF regression tests

* test(anthropic): audit cells for legacy federation references missing an id

Integration cells drive the seven clients, files, skills, health, Test Connect and a mixed burst against deployments whose token file or inline token reference lacks a rule id or an organization id, asserting the named 401 and no outbound call, with controls for the static key, the environment ids, the blank reference, the unauthenticated call and the complete shapes. The unit test covers the environment ids completing a token file param

* test(anthropic): strip host federation ids from the audit proxies
2026-10-07 18:08:19 -07:00
berriai-litellm-provider-info-sync[bot]
0d8d128a15
fix(vertex-ai): halve claude-sonnet-5-5 cache read price (#45217)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 18:05:58 -07:00
devin-ai-integration[bot]
9e871f2736
test: remove dead imports and helpers left behind by legacy test deletion (#45208)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 18:04:17 -07:00
D41910
87e961fad0
fix(anthropic): let /v1/messages mid-stream failures reach the proxy failure boundary (#44800)
* fix(anthropic): let /v1/messages mid-stream failures reach the proxy failure boundary

AnthropicStreamWrapper.async_anthropic_sse_wrapper swallowed every upstream
exception and emitted its own Anthropic error frame, so a provider drop on a
streamed /v1/messages call never reached the proxy's streaming boundary — no
failure spend row, no failure callbacks.

Re-raise when the stream is proxy-managed (detached failure hook armed) so
async_streaming_data_generator runs post_call_failure_hook once and serializes
the error frame; when consumed standalone (SDK litellm.messages path), run the
logging object's async failure handler and keep the client-facing error frame.

Fixes #44742

* test(anthropic): satisfy PT012 single-statement rule in mid-stream error test

* fix(anthropic): dispatch both sync and async failure callbacks

* fix(anthropic): dispatch both sync and async failure callbacks

* test(anthropic): move the regression cases into the mapped test file

* refactor(anthropic): read the public detached-failure hook and drop the explanatory comments

* test(anthropic): drive the mid-stream failure tests through a real logging object

* fix(anthropic): re-raise the provider error from the chat wrapper envelope so the router keeps its pre-content retry

* test(integration): cover mid-stream failure bookkeeping of bridged /v1/messages streams

Nine integration cells for the chat bridge: raw httpx and the Anthropic SDK (sync and async) streams that fail after the first content delta now land exactly one failure SpendLogs row carrying the provider's error class, the non-streaming 500, the response-cache twin and a client that leaves mid-stream keep their bookkeeping, a 24-stream burst lands every call id once, and the standalone SDK stream reports the provider error to failure callbacks once.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 17:42:35 -07:00
devin-ai-integration[bot]
79f62db620
feat(decisions): add the OpenAI Decisions spec types and the System One translation (#45129)
* feat(decisions): add the OpenAI Decisions spec types and the System One translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): share one DecisionsModel config and require model and usage on responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): rename the shared pydantic parent to DecisionsObjectBase

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): match the SDK on strict choice values and drop the invented min_length bounds

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): compose DecisionsRequest from model and body and read the translator top down

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): use the plural Decisions prefix only for the request and response envelopes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): import assert_never from typing_extensions for Python 3.10

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:42:28 -07:00