* fix(lens): use async-timeout on python 3.10 for budget reservation timeouts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(deps): keep uv.lock diff to the async-timeout entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(lens): cover real request deadline expiry in reserved_budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): finish Python 3.10 timeout coverage and dependency checks
* test(lens): control event-loop time for deadline regressions
* test(tracing): include priced call count in trace fixture
* fix(lens): limit timeout compatibility changes to PR scope
---------
Co-authored-by: Moe Khalil <moe@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): never store or serve a chat completion with no choices
A provider response with empty choices was written to the response cache and served on every identical request until the TTL ended, with no provider call in between. The cache now skips storing such a response and treats an already stored one as a miss, so the next request goes back to the provider and its answer replaces the entry.
* fix(caching): skip responses with no output on the Responses API and Anthropic Messages too
* fix(caching): skip streams and stored entries that carry no output
A chat or text completion stream whose chunks carried no choice is closed
by the stream wrapper with one empty choice of its own, so the assembled
response passed the choices check and was cached. The assembled stream is
now judged on its content: a stream with no text, tool call, or other
output in any choice is never stored, on the async and sync writers alike.
The Responses API stream writer and the Anthropic Messages stream writer
apply the same no-output check before storing.
A stored entry with no output read through the worker memory tier is now
evicted from that tier on the miss, so the next read reaches Redis where
the refill lands; the text completion and messages writers only write to
Redis, and the memory copy otherwise kept missing until its own TTL.
* test(integration): response cache cells for answers without output
Deterministic cells for the response cache on every unified endpoint,
streamed and not, through the OpenAI and Anthropic SDKs and raw httpx,
plus the sync SDK paths, stale entries, malformed answers, per-request
TTLs, cache delete, and chaos (Redis stopped or paused mid burst, a
worker killed, in-memory cache mode). The scripted upstream counts only
POSTs as deployment calls, since the proxy's boot-time GET /v1/models
discovery of a config deployment is not one.
* test(caching): pin the stored entry timestamp in the worker-copy test
* test(integration): drop the restating comments from the chaos cells
* test(integration): close the breaker on the first call after the Redis restart
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(lens): price agent traces by joining gen_ai.response.id to spend logs
Trace spend now joins each model call to spend_logs on one key: the
span's response id (gen_ai.response.id or the id the normalizers read
from OpenInference/LangChain output) against spend_logs.response_id or
the upstream id embedded in a managed resp_ id. The litellm.call_id and
traceparent transport join paths and the per-row ownership gate are
removed; the spend SQL still restricts rows to what the reader can see.
A run with some unpriced calls now reports the sum of its priced calls
plus priced_calls, instead of an unknown total.
* test(lens): cover response id spend join and partial trace totals
* chore(lens): regenerate trace types for priced_calls
* feat(lens): show partial run cost as a lower bound with priced call count
* fix(lens): treat litellm.call_id as the same assigned call id for spend joins
The id LiteLLM assigned to a call is either the response id it returned
(gen_ai.response.id -> spend_logs.response_id) or its gateway call id
(litellm.call_id -> spend_logs.litellm_call_id). Both are exact ids the
gateway mints and logs, so the join stays one rule. Transport span
matching and the ownership gate stay removed.
* test(lens): cover litellm.call_id spend joins and restore captured totals
* feat(lens): link each priced model call to its spend log
Spans gain spend_log_request_id, the spend_logs.request_id the call was
priced from, and spend_match, which says whether a model call matched or
why not (no assigned id on the span, no spend log with that id, or an
ambiguous match). A model call span is priced from the same ids as the
run total, so its cost and the total agree.
* test(lens): cover spend log links on model call spans
* chore(lens): regenerate trace types for spend log links
* feat(lens): open the matched spend log from an LLM step
An LLM step's header now shows a Spend log chip with the matched
request id and cost; clicking it opens the request log drawer over the
run, fetched by the exact spend_logs.request_id instead of the span's
own response id. Unpriced steps say why (no assigned id on the span, or
no spend log with it). Tree rows show each model call's cost, and a
partial run cost shows its priced call count inline.
* test(lens): cover the spend log link and unmatched cost reasons
* feat(lens): show the spend log link as a bordered LiteLLM Spend Log button
* feat(lens): add a back link from the spend log drawer to the agent trace
* feat(lens): label the spend log back link Back to Lens trace with the Lens icon
* fix(lens): ignore assigned ids that name no spend log when pricing a call
An id that names no row no longer vetoes the call, so a span carrying
both a response id and a call id still prices from a spend row logged
before litellm_call_id existed. An id naming two or more rows makes the
call ambiguous, and the match reason comes from the same per-id result,
so a single matched row with no cost is reported as matched.
* fix(lens): price a trace only from spend logs in its own team
A reader with several teams could see the same assigned id in another
team's spend log; only rows from the trace's team now price it. The
user and key ownership gate stays removed.
* perf(lens): resolve each model call's spend once per trace
Model call matches are computed once when the trace is resolved and
looked up by span index, instead of scanning the model call list for
every span and walking the graph again for spans, agents and the run
total.
* fix(lens): hide a step's Cost fact only when its spend log link shows the cost
* chore(lens): drop narrative doc comments from the spend join
* fix(lens): price a model call only when its ids agree on one spend log per span
* fix(lens): keep pricing spend logs written before litellm_call_id by their request id
* fix(lens): price every attempt a model call's ids name when they agree
* feat(proxy): add opt-in vector_store_deny_by_default for standalone virtual keys
Adds general_settings.vector_store_deny_by_default (typed bool, default false). When enabled, a virtual key
with no team must list the requested vector store in its object_permission.vector_stores; no permission
record, null, or an empty list is denied with key_vector_store_access_denied. Omitted or false keeps the
existing behavior, including nonempty allowlist enforcement. The master key is unchanged in both modes.
Team keys and keyless callers are deferred to later increments of LIT-6035
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): require key and team vector store grants for team keys under vector_store_deny_by_default
With the flag enabled, a virtual key on a team needs both its own grant and its team's grant for every requested vector store. A missing permission record, an empty list or an unresolved team grants nothing. Dashboard session keys and the master key keep their existing behavior, and flag-off behavior is unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): require user or team vector store grants for keyless requests under vector_store_deny_by_default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover vector_store_deny_by_default through a real proxy with key, team and JWT identities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for vector_store_deny_by_default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): cover path and file_search vector store ids under vector_store_deny_by_default
Strict mode now reads vector_store_ids and tools[].vector_store_ids from the request body without needing a vector store registry, so /v1/vector_stores/{id}/search and Responses file_search are checked. User grants load through the object permission cache, and the proxy admin user rebuild keeps object_permission_id
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): broadcast user entitlement cache eviction to every worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(auth): reuse VectorStoreRegistry id extraction for vector_store_deny_by_default
Strict mode now uses get_vector_store_ids_to_run on an empty registry when none is loaded, instead of a parallel set of request-shape helpers, and vector_store_access_check documents the policy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): add strict vector store audit cells for routes, SDKs, workers and concurrency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): consolidate vector_store_deny_by_default coverage to core cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): reject invalid vector_store_deny_by_default at config load and return 400 for malformed vector_store_ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): read vector store key and team grants through the object permission cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(auth): return typed permission rows in request flow vector store tests
Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(auth): validate vector store ids and tools as immutable sequences
Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust_bridge): remove rule-gated native secret-manager selection
Mirror the cache treatment: SecretManagerRule/SecretManagerContext and the resolve_native_* binding plumbing are gone. The Rust bridge now selects the native backend from the explicitly configured client (capture_secret_manager / _SecretManagerRuntime.from_client), and get_secret_from_manager is the plain Python handler path.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(deps): bump sharp to 0.35.5 for GHSA-wq5f-xc86-pv6w
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): correct supported_endpoints on gemini and vertex image rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): drop images/generations from gemini/nano-banana-pro-preview
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure): add model router flat fee to azure provider cost tracking
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure): price the azure router fee from one canonical entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure): reuse the azure_ai model router fee path for azure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure): drop model router fee unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.
Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): restore key activity search and add model activity search
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): drop rebuilt dashboard bundle from the usage search change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format usage search tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): only show model no-match when the range has models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): budget s3 dedupe retries on the deployment so the seeded router num_retries cannot zero them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): prevent nested router retries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): fix provider retry and pytest collection setup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep sdk retries on the sync router responses path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(utils): count upstream calls instead of doubling the retry helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 299 files
Clears 871 basedpyright Any errors (reportAny 6,703 to 6,240, reportExplicitAny 1,651 to 1,243) without adding a cast, an ignore or a suppression, and without touching any budget file
Most edits are annotation-only: a parameter, return or local goes from Any to object, Mapping[str, object] or the concrete type the value always held. Fourteen files validate untyped JSON once where it enters, through a module-level pydantic TypeAdapter or model_validate, and then use real types
No HTTP status, error type or response shape changes. Mistral speech and fal.ai Bria image generation now report a pydantic ValidationError instead of an AttributeError when the provider answers 2xx with a body that is not a JSON object
* refactor(types): make the config locals fix effective and trim no-op edits
The Mapping[str, object] annotation on `locals().copy()` removed no error,
because the checker narrows the variable back to the dict[str, Any] the call
returns. 21 provider config constructors now build the same copy with
dict(locals()), which the checker infers as object values under that
annotation, so each file loses one reportAny.
The same annotation is reverted in 31 other config files where it stayed a
no-op, together with the tests that were added only to cover those lines, and
the one MCP server manager line that no CI coverage shard executes is
reverted too. The pull request drops from 345 to 297 changed files.
* refactor(types): accept only int in the proxy state setter
get_proxy_state_variable is annotated to return int, but
set_proxy_state_variable still took Any, so the checker could not hold
callers to the type the getter promises. The setter now takes int, which is
what its only caller already passes.
* refactor(types): index the proxy state key so the getter returns int
* refactor(types): keep public annotations and provider error text unchanged
Restore every public return, public method parameter, public attribute and exported
alias to its annotation on main so code that type-checks against the package keeps
type-checking, and take the Mistral speech and fal.ai Bria changes back out so no
provider error message differs from main
* refactor(types): leave the Vertex RAG chunking read as it is on main
Take the chunking format validation back out of the Vertex RAG ingestion path. It needs the vertexai SDK, a storage bucket and a RAG corpus to execute, so nothing here could run it end to end, and it cleared only two errors
* test(integration): pin the validated provider boundaries on a live proxy
* test(integration): give the held burst a client that outlasts the gate
The fault cell holds a burst at the upstream for up to 60 seconds while it kills a worker, but sent the burst through the shared 15 second client, so a slow box could time the survivors out before the gate opened. The burst now goes through its own client whose timeout is twice the gate, and the gate length is one named constant.
* refactor(types): keep the license reply handling and experimental MCP signatures as they were
The license check validated the whole reply as a mapping, which changed the error text logged for a reply that is not an object. It now validates only the verify value, so every reply is handled and logged exactly as before while the value is still typed.
Three files under the experimental MCP server changed annotations on public functions and methods (three returns and three parameters). They go back to their previous content so no public signature in the diff is narrowed.
* test(integration): answer the proxy's model-list call in the OpenAI stand-ins
Every 300 seconds each proxy worker asks an OpenAI deployment for GET /v1/models. Four new cells own an OpenAI stand-in that accepted only the call under test, so a refresh landing inside a cell failed it. The stand-ins now answer that call through the suite's own helper and the cells count only the provider calls they drive.
* chore(rust): prune unused inference crate dependencies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): rewrite inference layering docs for the split format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): rename litellm-core to litellm-inference
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): expose inference base API for format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move shared inference test helpers behind test-support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): wait for the classifier model popup before picking its option
Base UI exposes its select option asynchronously, and the helper waits for the option and its positioner to become clickable before selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: isolate two router tests from work leaked by earlier tests
Collect earlier tests' garbage before warning capture so their unawaited coroutines cannot be attributed to the target's warning assertion
Filter success-event callbacks by the request's litellm_call_id so queued logging work cannot replace the current request's captured messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 14 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base parsing in sso userinfo, copilot auth and hf config lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base parsing at unproven provider seams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base delete in jwt orphan cleanup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor: remove fresh tech debt from the 2026-10-05 window
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(lens): drop the review models left unused by the dead review helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): price a model group from the deployments that serve it
A model group's info read composed the group's own model_group_alias
entry into its deployments, a hop the router never takes: an alias is
resolved exactly once at request time, so a group reached as an alias
target is served by its own deployments. For the chain X -> T -> U the
price read for X included U's deployments too, and the free-model
budget waiver refused a free request to X on an over-budget key, while
GET /model_group/info reported U's providers and price for X.
The group info read now prices a group from the deployments routing
serves it with: the ones named after it, the routing group of that
name, or the wildcard route matching it when neither exists. The
budget waiver, GET /model_group/info, the rate limiters, and the
response headers all read the same set as routing. get_model_list
keeps its behavior for every other caller.
* test(router): give the paid fixtures explicit per-token prices
* test(router): call the routed-group read by name so the router coverage gate sees it
* test(integration): audit the alias chain budget waiver on every route, shape, and outage
Thirty-four cells under the management group prove an over-budget key is served through an alias chain entry at the price of the deployment that serves it, on chat, responses, and messages, sync and streamed, through the OpenAI and Anthropic SDKs and raw httpx on both replicas, with the chain middle, the reverse chain, a cost-map priced middle, a ghost middle, wildcard and routing-group targets, malformed and hostile inputs, a cached reply, a repointed alias, a provider failure, and two chaos bursts (a killed worker, a scripted outage)
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* perf(types): defer pydantic schema builds via shared LiteLLMBaseModel
Add LiteLLMBaseModel with defer_build driven by DEFER_PYDANTIC_BUILD (default true) and move litellm and enterprise pydantic models onto it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(types): build deferred models created by a parent validator; keep lens worker models litellm-free
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(types): link pydantic issue on deferred-build rebuild hook
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): polish Lens runs loading, reload, and time range menu
Port the dashboard-only parts of a0a275e486, 1325389624, 5e7b0afd5c, 03bec959bf, 93e6ccff8d, 3d35d9f920, dc1a2b5b60, 1d0397f1c0, 773bfef066, f3387221ea and 3d049cd4a8 from litellm_lens_server_search onto main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep the Lens timeline on the shown runs' window during a reload
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): drop a redundant fixture comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>