Commit graph

53768 commits

Author SHA1 Message Date
devin-ai-integration[bot]
abd3af422d
fix(caching): never store or serve a chat completion with no choices (#44709)
* fix(caching): never store or serve a chat completion with no choices

A provider response with empty choices was written to the response cache and served on every identical request until the TTL ended, with no provider call in between. The cache now skips storing such a response and treats an already stored one as a miss, so the next request goes back to the provider and its answer replaces the entry.

* fix(caching): skip responses with no output on the Responses API and Anthropic Messages too

* fix(caching): skip streams and stored entries that carry no output

A chat or text completion stream whose chunks carried no choice is closed
by the stream wrapper with one empty choice of its own, so the assembled
response passed the choices check and was cached. The assembled stream is
now judged on its content: a stream with no text, tool call, or other
output in any choice is never stored, on the async and sync writers alike.
The Responses API stream writer and the Anthropic Messages stream writer
apply the same no-output check before storing.

A stored entry with no output read through the worker memory tier is now
evicted from that tier on the miss, so the next read reaches Redis where
the refill lands; the text completion and messages writers only write to
Redis, and the memory copy otherwise kept missing until its own TTL.

* test(integration): response cache cells for answers without output

Deterministic cells for the response cache on every unified endpoint,
streamed and not, through the OpenAI and Anthropic SDKs and raw httpx,
plus the sync SDK paths, stale entries, malformed answers, per-request
TTLs, cache delete, and chaos (Redis stopped or paused mid burst, a
worker killed, in-memory cache mode). The scripted upstream counts only
POSTs as deployment calls, since the proxy's boot-time GET /v1/models
discovery of a config deployment is not one.

* test(caching): pin the stored entry timestamp in the worker-copy test

* test(integration): drop the restating comments from the chaos cells

* test(integration): close the breaker on the first call after the Redis restart

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 20:07:30 +00:00
ishaan-berri
5ed7ec8511
fix(lens): price agent traces by joining gen_ai.response.id to spend logs (#44738)
* fix(lens): price agent traces by joining gen_ai.response.id to spend logs

Trace spend now joins each model call to spend_logs on one key: the
span's response id (gen_ai.response.id or the id the normalizers read
from OpenInference/LangChain output) against spend_logs.response_id or
the upstream id embedded in a managed resp_ id. The litellm.call_id and
traceparent transport join paths and the per-row ownership gate are
removed; the spend SQL still restricts rows to what the reader can see.

A run with some unpriced calls now reports the sum of its priced calls
plus priced_calls, instead of an unknown total.

* test(lens): cover response id spend join and partial trace totals

* chore(lens): regenerate trace types for priced_calls

* feat(lens): show partial run cost as a lower bound with priced call count

* fix(lens): treat litellm.call_id as the same assigned call id for spend joins

The id LiteLLM assigned to a call is either the response id it returned
(gen_ai.response.id -> spend_logs.response_id) or its gateway call id
(litellm.call_id -> spend_logs.litellm_call_id). Both are exact ids the
gateway mints and logs, so the join stays one rule. Transport span
matching and the ownership gate stay removed.

* test(lens): cover litellm.call_id spend joins and restore captured totals

* feat(lens): link each priced model call to its spend log

Spans gain spend_log_request_id, the spend_logs.request_id the call was
priced from, and spend_match, which says whether a model call matched or
why not (no assigned id on the span, no spend log with that id, or an
ambiguous match). A model call span is priced from the same ids as the
run total, so its cost and the total agree.

* test(lens): cover spend log links on model call spans

* chore(lens): regenerate trace types for spend log links

* feat(lens): open the matched spend log from an LLM step

An LLM step's header now shows a Spend log chip with the matched
request id and cost; clicking it opens the request log drawer over the
run, fetched by the exact spend_logs.request_id instead of the span's
own response id. Unpriced steps say why (no assigned id on the span, or
no spend log with it). Tree rows show each model call's cost, and a
partial run cost shows its priced call count inline.

* test(lens): cover the spend log link and unmatched cost reasons

* feat(lens): show the spend log link as a bordered LiteLLM Spend Log button

* feat(lens): add a back link from the spend log drawer to the agent trace

* feat(lens): label the spend log back link Back to Lens trace with the Lens icon

* fix(lens): ignore assigned ids that name no spend log when pricing a call

An id that names no row no longer vetoes the call, so a span carrying
both a response id and a call id still prices from a spend row logged
before litellm_call_id existed. An id naming two or more rows makes the
call ambiguous, and the match reason comes from the same per-id result,
so a single matched row with no cost is reported as matched.

* fix(lens): price a trace only from spend logs in its own team

A reader with several teams could see the same assigned id in another
team's spend log; only rows from the trace's team now price it. The
user and key ownership gate stays removed.

* perf(lens): resolve each model call's spend once per trace

Model call matches are computed once when the trace is resolved and
looked up by span index, instead of scanning the model call list for
every span and walking the graph again for spans, agents and the run
total.

* fix(lens): hide a step's Cost fact only when its spend log link shows the cost

* chore(lens): drop narrative doc comments from the spend join

* fix(lens): price a model call only when its ids agree on one spend log per span

* fix(lens): keep pricing spend logs written before litellm_call_id by their request id

* fix(lens): price every attempt a model call's ids name when they agree
2026-10-06 19:55:26 +00:00
devin-ai-integration[bot]
282d733fb6
feat(auth): deny search tools by default when search_tool_deny_by_default is set (#44490)
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 12:50:26 -07:00
devin-ai-integration[bot]
abc543e701
feat(proxy): add opt-in vector_store_deny_by_default for least-privilege vector store access (#44244)
* feat(proxy): add opt-in vector_store_deny_by_default for standalone virtual keys

Adds general_settings.vector_store_deny_by_default (typed bool, default false). When enabled, a virtual key
with no team must list the requested vector store in its object_permission.vector_stores; no permission
record, null, or an empty list is denied with key_vector_store_access_denied. Omitted or false keeps the
existing behavior, including nonempty allowlist enforcement. The master key is unchanged in both modes.
Team keys and keyless callers are deferred to later increments of LIT-6035

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): require key and team vector store grants for team keys under vector_store_deny_by_default

With the flag enabled, a virtual key on a team needs both its own grant and its team's grant for every requested vector store. A missing permission record, an empty list or an unresolved team grants nothing. Dashboard session keys and the master key keep their existing behavior, and flag-off behavior is unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): require user or team vector store grants for keyless requests under vector_store_deny_by_default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover vector_store_deny_by_default through a real proxy with key, team and JWT identities

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for vector_store_deny_by_default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): cover path and file_search vector store ids under vector_store_deny_by_default

Strict mode now reads vector_store_ids and tools[].vector_store_ids from the request body without needing a vector store registry, so /v1/vector_stores/{id}/search and Responses file_search are checked. User grants load through the object permission cache, and the proxy admin user rebuild keeps object_permission_id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): broadcast user entitlement cache eviction to every worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(auth): reuse VectorStoreRegistry id extraction for vector_store_deny_by_default

Strict mode now uses get_vector_store_ids_to_run on an empty registry when none is loaded, instead of a parallel set of request-shape helpers, and vector_store_access_check documents the policy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): add strict vector store audit cells for routes, SDKs, workers and concurrency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): consolidate vector_store_deny_by_default coverage to core cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reject invalid vector_store_deny_by_default at config load and return 400 for malformed vector_store_ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): read vector store key and team grants through the object permission cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(auth): return typed permission rows in request flow vector store tests

Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(auth): validate vector store ids and tools as immutable sequences

Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 12:50:25 -07:00
devin-ai-integration[bot]
26a9b02f7b
feat(otel): let team and key Arize callbacks choose the OTLP transport (#44492)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 14:46:50 -05:00
devin-ai-integration[bot]
b6d587d23e
refactor(rust_bridge): remove rule-gated native secret-manager selection (#44906)
* refactor(rust_bridge): remove rule-gated native secret-manager selection

Mirror the cache treatment: SecretManagerRule/SecretManagerContext and the resolve_native_* binding plumbing are gone. The Rust bridge now selects the native backend from the explicitly configured client (capture_secret_manager / _SecretManagerRuntime.from_client), and get_secret_from_manager is the plain Python handler path.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): bump sharp to 0.35.5 for GHSA-wq5f-xc86-pv6w

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 19:14:55 +00:00
moe-berri
fb554d34f3
fix(lens): paginate visible conversation entries (#44878)
* fix(lens): paginate visible conversation entries

* fix(lens): stop pagination at exhausted subagents

* refactor(lens): derive pending branches without mutation

* fix(lens): bound pending ancestry work for deep traces
2026-10-06 12:10:19 -07:00
moe-berri
5147aefa8b
fix(lens): reuse trace reviews and consolidate findings across runs (#44778)
* fix(lens): reuse trace reviews and consolidate findings across runs

* chore: sync schema.prisma copies from root

* fix(lens): preserve partial reviews and bound budget admission

* fix(lens): resolve CI regressions and clarify reused runs

* test(lens): exercise review reuse in worker container smoke

* fix(lens): show actual reviews when a run finishes

* fix(lens): distinguish reuse plans and stop blocked scans

* fix(lens): retain partial findings when model requests stop

* docs(lens): clarify budget edits during active runs

* fix(lens): publish findings only after reconciliation completes

* fix(lens): guide insufficient-budget runs to budget settings

* fix(lens): renew budget holds and bound admission waits

* fix(lens): preserve provider errors during budget cleanup

* fix(ci): update PgBouncer and sharp for current builds

* fix(lens): serialize settlement and fence review checkpoints

* fix(lens): defer generated Prisma client type import

* test(lens): verify checkpoint and progress rollback together

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-10-06 12:09:56 -07:00
berriai-litellm-provider-info-sync[bot]
fa3b0a6d95
feat(vertex-ai): add gemini-nano-banana-2.1 and Nano Banana 2 priority pricing (#44892)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 11:52:07 -07:00
devin-ai-integration[bot]
de984d6ebc
fix(model_prices): correct supported_endpoints on gemini and vertex image rows (#44891)
* fix(model_prices): correct supported_endpoints on gemini and vertex image rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): drop images/generations from gemini/nano-banana-pro-preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 11:26:27 -07:00
devin-ai-integration[bot]
5cdebded90
fix(ci): skip the generated dashboard bundle in the master key guard (#44890)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 18:15:09 +00:00
berriai-litellm-provider-info-sync[bot]
d9f8dbe23a
chore(prices): add gemini/gemini-nano-banana-2.1 from the Gemini pricing page (#44868)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 18:06:35 +00:00
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
berriai-litellm-provider-info-sync[bot]
ab61410a39
chore(pricing): add azure model-router and whisper rows (#44872)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 10:54:50 -07:00
devin-ai-integration[bot]
e2c55daaaf
fix(azure): add model router flat fee to azure provider cost tracking (#44876)
* fix(azure): add model router flat fee to azure provider cost tracking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): price the azure router fee from one canonical entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): reuse the azure_ai model router fee path for azure

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure): drop model router fee unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:47:26 -07:00
devin-ai-integration[bot]
096b20b6db
chore(model_prices): remove malformed, duplicate and decommissioned palm cost map entries (#44880)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 17:33:54 +00:00
devin-ai-integration[bot]
090a4c3f24
fix(ui): read MCP submission rules from bare-array /config/list response (#44648)
* fix(ui): read MCP submission rules from bare-array /config/list response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover submission rules contract and dashboard preload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): disable MCP submission rules editor until rules load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format MCP submission integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): restore prior MCP submission rules after the rules spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): move MCP submission rules setup and cleanup into fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): use Promise.withResolvers in MCP rules loading test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): clear lint warnings in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): drop redundant JSX comments in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:09:46 -07:00
devin-ai-integration[bot]
878ba39e7c
refactor(rust): extract inference-testing crate (#44873)
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.

Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 16:33:06 +00:00
devin-ai-integration[bot]
eb385cca3e
fix(ui): restore key activity search to the top of the tab and add model activity search (#44521)
* fix(ui): restore key activity search and add model activity search

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): drop rebuilt dashboard bundle from the usage search change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format usage search tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): only show model no-match when the range has models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:27:22 -07:00
berriai-litellm-provider-info-sync[bot]
bd23e6fc3d
fix(cost-map): update together_ai Kimi-K3 and Qwen3.8-Flash prices to published rates (#44864)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 09:03:40 -07:00
devin-ai-integration[bot]
4f27e9c687
fix(responses): stop SDK retries nesting under router retries, and unbreak CircleCI integration tests (#44791)
* test(integration): budget s3 dedupe retries on the deployment so the seeded router num_retries cannot zero them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): prevent nested router retries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fix provider retry and pytest collection setup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep sdk retries on the sync router responses path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): count upstream calls instead of doubling the retry helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 15:57:33 +00:00
Yamac Eren Ay
b0fe75a917
fix(sap): correctly handle cache_control (#39122)
* fix(sap): support cache_control, drop custom content validation, add unit tests for message models

Co-authored-by: Marcel Wienöbst <marcel.wienoebst@sap.com>

* fix: changes after review

* fix tests accordingly

* tests docstrings removed

* fix: unit tests moved

* rm unnecessary import

* fix linting issue

---------

Co-authored-by: Marcel Wienöbst <marcel.wienoebst@sap.com>
2026-10-06 08:54:06 -07:00
Mateo Wang
d4a791d650
refactor(types): replace Any with proven types in 157 files (#44798)
* refactor(types): replace Any with proven types in 299 files

Clears 871 basedpyright Any errors (reportAny 6,703 to 6,240, reportExplicitAny 1,651 to 1,243) without adding a cast, an ignore or a suppression, and without touching any budget file

Most edits are annotation-only: a parameter, return or local goes from Any to object, Mapping[str, object] or the concrete type the value always held. Fourteen files validate untyped JSON once where it enters, through a module-level pydantic TypeAdapter or model_validate, and then use real types

No HTTP status, error type or response shape changes. Mistral speech and fal.ai Bria image generation now report a pydantic ValidationError instead of an AttributeError when the provider answers 2xx with a body that is not a JSON object

* refactor(types): make the config locals fix effective and trim no-op edits

The Mapping[str, object] annotation on `locals().copy()` removed no error,
because the checker narrows the variable back to the dict[str, Any] the call
returns. 21 provider config constructors now build the same copy with
dict(locals()), which the checker infers as object values under that
annotation, so each file loses one reportAny.

The same annotation is reverted in 31 other config files where it stayed a
no-op, together with the tests that were added only to cover those lines, and
the one MCP server manager line that no CI coverage shard executes is
reverted too. The pull request drops from 345 to 297 changed files.

* refactor(types): accept only int in the proxy state setter

get_proxy_state_variable is annotated to return int, but
set_proxy_state_variable still took Any, so the checker could not hold
callers to the type the getter promises. The setter now takes int, which is
what its only caller already passes.

* refactor(types): index the proxy state key so the getter returns int

* refactor(types): keep public annotations and provider error text unchanged

Restore every public return, public method parameter, public attribute and exported
alias to its annotation on main so code that type-checks against the package keeps
type-checking, and take the Mistral speech and fal.ai Bria changes back out so no
provider error message differs from main

* refactor(types): leave the Vertex RAG chunking read as it is on main

Take the chunking format validation back out of the Vertex RAG ingestion path. It needs the vertexai SDK, a storage bucket and a RAG corpus to execute, so nothing here could run it end to end, and it cleared only two errors

* test(integration): pin the validated provider boundaries on a live proxy

* test(integration): give the held burst a client that outlasts the gate

The fault cell holds a burst at the upstream for up to 60 seconds while it kills a worker, but sent the burst through the shared 15 second client, so a slow box could time the survivors out before the gate opened. The burst now goes through its own client whose timeout is twice the gate, and the gate length is one named constant.

* refactor(types): keep the license reply handling and experimental MCP signatures as they were

The license check validated the whole reply as a mapping, which changed the error text logged for a reply that is not an object. It now validates only the verify value, so every reply is handled and logged exactly as before while the value is still typed.

Three files under the experimental MCP server changed annotations on public functions and methods (three returns and three parameters). They go back to their previous content so no public signature in the diff is narrowed.

* test(integration): answer the proxy's model-list call in the OpenAI stand-ins

Every 300 seconds each proxy worker asks an OpenAI deployment for GET /v1/models. Four new cells own an OpenAI stand-in that accepted only the call under test, so a refresh landing inside a cell failed it. The stand-ins now answer that call through the suite's own helper and the cells count only the provider calls they drive.
2026-10-06 15:31:54 +00:00
devin-ai-integration[bot]
126e79c967
fix(model_prices): gemini deep research input limits, vertex flash retirement dates, azure data zone gpt-6.1-sol pricing (#44633) 2026-10-06 07:40:25 -07:00
devin-ai-integration[bot]
46d2c2a6ea
fix(otel): honor per-team Arize sampling rates in OTel v2 fan-out (#44595)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:11 -05:00
devin-ai-integration[bot]
44d5dacbaa
fix(otel): gate the Arize OTel v2 exporter on operator credentials (#44596)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:11 -05:00
devin-ai-integration[bot]
80f18e1326
fix(otel): name an Arize project on every OTel v2 Arize export (#44605)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:10 -05:00
devin-ai-integration[bot]
de74f81c69
chore(rust): prune inference deps and rewrite layering docs (#44836)
* chore(rust): prune unused inference crate dependencies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): rewrite inference layering docs for the split format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:27 -07:00
devin-ai-integration[bot]
ae35c9d775
refactor(rust): extract inference-ocr crate (#44832)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
6a55e0a8aa
refactor(rust): extract inference-chat crate (#44827)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
d7f5b40ab4
refactor(rust): extract inference-messages crate (#44818)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
785cb12f37
refactor(rust): extract inference-responses crate (#44811)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:25 -07:00
devin-ai-integration[bot]
63babf23e6
refactor(rust): extract inference-transcription crate (#44809)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
1a9b533e71
refactor(rust): rename litellm-core to litellm-inference (#44802)
* refactor(rust): rename litellm-core to litellm-inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): expose inference base API for format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move shared inference test helpers behind test-support

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
10444df3a0
test: deflake the JEV classifier select and two router tests that inherited leaked state (#44840)
* test(ui): wait for the classifier model popup before picking its option

Base UI exposes its select option asynchronously, and the helper waits for the option and its positioner to become clickable before selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: isolate two router tests from work leaked by earlier tests

Collect earlier tests' garbage before warning capture so their unawaited coroutines cannot be attributed to the target's warning assertion

Filter success-event callbacks by the request's litellm_call_id so queued logging work cannot replace the current request's captured messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:27:49 +00:00
devin-ai-integration[bot]
d4619c499a
refactor(types): replace Any with proven types in 6 files (#44491)
* refactor(types): replace Any with proven types in 14 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base parsing in sso userinfo, copilot auth and hf config lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base parsing at unproven provider seams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base delete in jwt orphan cleanup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:53:43 +00:00
berriai-litellm-provider-info-sync[bot]
18c3118eb6
fix(azure): add gpt-6-sol priority processing prices (#44829)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 01:49:46 -07:00
devin-ai-integration[bot]
5455152913
refactor: remove fresh tech debt from the 2026-10-05 window (#44821)
* refactor: remove fresh tech debt from the 2026-10-05 window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(lens): drop the review models left unused by the dead review helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 08:33:39 +00:00
devin-ai-integration[bot]
c13746c951
fix(router): price a model group from the deployments that serve it (#44732)
* fix(router): price a model group from the deployments that serve it

A model group's info read composed the group's own model_group_alias
entry into its deployments, a hop the router never takes: an alias is
resolved exactly once at request time, so a group reached as an alias
target is served by its own deployments. For the chain X -> T -> U the
price read for X included U's deployments too, and the free-model
budget waiver refused a free request to X on an over-budget key, while
GET /model_group/info reported U's providers and price for X.

The group info read now prices a group from the deployments routing
serves it with: the ones named after it, the routing group of that
name, or the wildcard route matching it when neither exists. The
budget waiver, GET /model_group/info, the rate limiters, and the
response headers all read the same set as routing. get_model_list
keeps its behavior for every other caller.

* test(router): give the paid fixtures explicit per-token prices

* test(router): call the routed-group read by name so the router coverage gate sees it

* test(integration): audit the alias chain budget waiver on every route, shape, and outage

Thirty-four cells under the management group prove an over-budget key is served through an alias chain entry at the price of the deployment that serves it, on chat, responses, and messages, sync and streamed, through the OpenAI and Anthropic SDKs and raw httpx on both replicas, with the chain middle, the reverse chain, a cost-map priced middle, a ghost middle, wildcard and routing-group targets, malformed and hostile inputs, a cached reply, a repointed alias, a provider failure, and two chaos bursts (a killed worker, a scripted outage)

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 05:31:09 +00:00
devin-ai-integration[bot]
c2c0bb583e
perf(types): defer pydantic schema builds via shared LiteLLMBaseModel (#44720)
* perf(types): defer pydantic schema builds via shared LiteLLMBaseModel

Add LiteLLMBaseModel with defer_build driven by DEFER_PYDANTIC_BUILD (default true) and move litellm and enterprise pydantic models onto it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): build deferred models created by a parent validator; keep lens worker models litellm-free

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(types): link pydantic issue on deferred-build rebuild hook

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:10:08 -07:00
devin-ai-integration[bot]
d7f69d6ba1
refactor(logging): load enterprise alerting loggers lazily so import litellm skips proxy types (#44762)
* refactor(logging): load enterprise alerting loggers lazily so import litellm skips proxy types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover lazy alerting dispatch and docs lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve litellm.proxy._types lazily on attribute access

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(imports): assert proxy types guard imports the checkout under test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:06:53 -07:00
devin-ai-integration[bot]
0b633aa9c8
refactor(types): import proxy-only types under TYPE_CHECKING in SDK modules (#44740)
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:52:59 -07:00
devin-ai-integration[bot]
b72c737fc8
refactor(types): move SpanAttributes, SpecialHeaders and AllowedModelRegion out of proxy._types (#44717)
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:40:13 -07:00
devin-ai-integration[bot]
c3e0156979
fix(ui): polish Lens runs loading, reload, and time range menu (#44789)
* fix(ui): polish Lens runs loading, reload, and time range menu

Port the dashboard-only parts of a0a275e486, 1325389624, 5e7b0afd5c, 03bec959bf, 93e6ccff8d, 3d35d9f920, dc1a2b5b60, 1d0397f1c0, 773bfef066, f3387221ea and 3d049cd4a8 from litellm_lens_server_search onto main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep the Lens timeline on the shown runs' window during a reload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop a redundant fixture comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:27:00 -07:00
devin-ai-integration[bot]
d6c4e918d8
feat(ui): lead the Lens investigation detail with a run report (#44786)
* feat(ui): lead the Lens investigation detail with a run report

Port of the dashboard changes from a43f111abe and 29a1f3175d onto main. The detail view now opens with a run report for the selected run: status, a headline, progress or the failure, then cost, duration, coverage and issues, plus a collapsed activity log. An ordered situation table picks the report's one next action (Run now, Stop run, Retry, Raise budget, Connect worker, Review issues, Monitor this), and Run now moves into the investigation actions menu

Main's live review stays as is below the report, the queue reason still shows under queued progress, and partial results keep their own warning state with the run details folded away

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Stop run on older Lens runs while another run is active

Also drop doc comments that restate the code

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:24:28 -07:00
nate-berri
b8239a9873
fix(vertex_ai): add regional endpoint uplift to gemini-3.1-flash-image (#44673)
* fix(vertex_ai): add regional endpoint uplift to gemini-3.1-flash-image

Google prices Gemini 3.1 Flash Image at 1.1x on non-global endpoints for
input, text output and image output, but the cost map row had no
regional_endpoint_uplift_multiplier, so regional calls billed at the
global rate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(integration): cover regional uplift spend for gemini-3.1-flash-image

* test(integration): read the proxy salt from the environment in the uplift spend cells

---------

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 21:00:22 -07:00
devin-ai-integration[bot]
5f9eff55f8
refactor(cache): remove dead Cache._native_cache runtime path (#44785)
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:58:19 -07:00
devin-ai-integration[bot]
a983b2a5e7
fix(responses): honor caller stream flag when provider forces SSE (internal copy of #34095) (#41235)
* fix(responses): honor caller stream flag when provider forces SSE

The Responses handlers decided whether to hand back a streaming iterator
from the provider payload's `stream` field, which chatgpt sets
unconditionally because the Codex backend only serves SSE. A caller that
sent `stream: false` therefore received a raw SSE stream on /v1/responses,
and the chat-completions bridge failed with "Unknown items in responses
API response: []" once its recovery path lost the raw SSE it reads from

Transport streaming still follows the provider payload; only the caller's
own `stream` value now decides the response shape. When the provider
forces SSE for a non-streaming caller the body is read and aggregated
through the existing path

* test(chatgpt): inject the authenticator into the responses config so handler tests never log in

* fix(responses): treat an extra_body stream flag as the caller's own and drop a redundant comment

* test(integration): audit the chatgpt caller stream flag across responses, chat and messages

---------

Co-authored-by: SeongWoon Cho <coffee@soylatte.kr>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:30:47 +00:00
berriai-litellm-provider-info-sync[bot]
e366e72502
feat(bedrock): add glm 5.3 cross-region rows and nova 2.5 sonic (#44710)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 20:25:24 -07:00
devin-ai-integration[bot]
ab3a59fe81
fix(chatgpt,github_copilot): refuse device-code login inside an event loop or worker thread (#39585)
* fix(chatgpt,github_copilot): refuse device-code login when an event loop is running

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(github_copilot): drop stray whitespace change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt): bound token refresh timeout and drop placeholder assignment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt,github_copilot): keep the token file path out of the event-loop 401 message

* fix(auth): refuse device-code login from worker threads too

/v1/messages runs its handler in an executor thread, where the
running-loop check never fires, so a chatgpt or github_copilot model
still started the interactive device-code login there and the request
hung for up to 15 minutes. The guard now also requires the main thread,
so the login only runs where a human can actually answer it.

* test(chatgpt): keep authenticator tests out of the real token directory

* fix(chatgpt): keep the 5 second connect timeout and the operator's request_timeout on the token refresh call

* test(integration): cover the device-code login guard on the proxy and the SDK

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:20:07 +00:00