Commit graph

53756 commits

Author SHA1 Message Date
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
berriai-litellm-provider-info-sync[bot]
ab61410a39
chore(pricing): add azure model-router and whisper rows (#44872)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 10:54:50 -07:00
devin-ai-integration[bot]
e2c55daaaf
fix(azure): add model router flat fee to azure provider cost tracking (#44876)
* fix(azure): add model router flat fee to azure provider cost tracking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): price the azure router fee from one canonical entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): reuse the azure_ai model router fee path for azure

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure): drop model router fee unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:47:26 -07:00
devin-ai-integration[bot]
096b20b6db
chore(model_prices): remove malformed, duplicate and decommissioned palm cost map entries (#44880)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 17:33:54 +00:00
devin-ai-integration[bot]
090a4c3f24
fix(ui): read MCP submission rules from bare-array /config/list response (#44648)
* fix(ui): read MCP submission rules from bare-array /config/list response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover submission rules contract and dashboard preload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): disable MCP submission rules editor until rules load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format MCP submission integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): restore prior MCP submission rules after the rules spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): move MCP submission rules setup and cleanup into fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): use Promise.withResolvers in MCP rules loading test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): clear lint warnings in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): drop redundant JSX comments in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:09:46 -07:00
devin-ai-integration[bot]
878ba39e7c
refactor(rust): extract inference-testing crate (#44873)
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.

Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 16:33:06 +00:00
devin-ai-integration[bot]
eb385cca3e
fix(ui): restore key activity search to the top of the tab and add model activity search (#44521)
* fix(ui): restore key activity search and add model activity search

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): drop rebuilt dashboard bundle from the usage search change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format usage search tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): only show model no-match when the range has models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:27:22 -07:00
berriai-litellm-provider-info-sync[bot]
bd23e6fc3d
fix(cost-map): update together_ai Kimi-K3 and Qwen3.8-Flash prices to published rates (#44864)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 09:03:40 -07:00
devin-ai-integration[bot]
4f27e9c687
fix(responses): stop SDK retries nesting under router retries, and unbreak CircleCI integration tests (#44791)
* test(integration): budget s3 dedupe retries on the deployment so the seeded router num_retries cannot zero them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): prevent nested router retries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fix provider retry and pytest collection setup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep sdk retries on the sync router responses path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): count upstream calls instead of doubling the retry helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 15:57:33 +00:00
Yamac Eren Ay
b0fe75a917
fix(sap): correctly handle cache_control (#39122)
* fix(sap): support cache_control, drop custom content validation, add unit tests for message models

Co-authored-by: Marcel Wienöbst <marcel.wienoebst@sap.com>

* fix: changes after review

* fix tests accordingly

* tests docstrings removed

* fix: unit tests moved

* rm unnecessary import

* fix linting issue

---------

Co-authored-by: Marcel Wienöbst <marcel.wienoebst@sap.com>
2026-10-06 08:54:06 -07:00
Mateo Wang
d4a791d650
refactor(types): replace Any with proven types in 157 files (#44798)
* refactor(types): replace Any with proven types in 299 files

Clears 871 basedpyright Any errors (reportAny 6,703 to 6,240, reportExplicitAny 1,651 to 1,243) without adding a cast, an ignore or a suppression, and without touching any budget file

Most edits are annotation-only: a parameter, return or local goes from Any to object, Mapping[str, object] or the concrete type the value always held. Fourteen files validate untyped JSON once where it enters, through a module-level pydantic TypeAdapter or model_validate, and then use real types

No HTTP status, error type or response shape changes. Mistral speech and fal.ai Bria image generation now report a pydantic ValidationError instead of an AttributeError when the provider answers 2xx with a body that is not a JSON object

* refactor(types): make the config locals fix effective and trim no-op edits

The Mapping[str, object] annotation on `locals().copy()` removed no error,
because the checker narrows the variable back to the dict[str, Any] the call
returns. 21 provider config constructors now build the same copy with
dict(locals()), which the checker infers as object values under that
annotation, so each file loses one reportAny.

The same annotation is reverted in 31 other config files where it stayed a
no-op, together with the tests that were added only to cover those lines, and
the one MCP server manager line that no CI coverage shard executes is
reverted too. The pull request drops from 345 to 297 changed files.

* refactor(types): accept only int in the proxy state setter

get_proxy_state_variable is annotated to return int, but
set_proxy_state_variable still took Any, so the checker could not hold
callers to the type the getter promises. The setter now takes int, which is
what its only caller already passes.

* refactor(types): index the proxy state key so the getter returns int

* refactor(types): keep public annotations and provider error text unchanged

Restore every public return, public method parameter, public attribute and exported
alias to its annotation on main so code that type-checks against the package keeps
type-checking, and take the Mistral speech and fal.ai Bria changes back out so no
provider error message differs from main

* refactor(types): leave the Vertex RAG chunking read as it is on main

Take the chunking format validation back out of the Vertex RAG ingestion path. It needs the vertexai SDK, a storage bucket and a RAG corpus to execute, so nothing here could run it end to end, and it cleared only two errors

* test(integration): pin the validated provider boundaries on a live proxy

* test(integration): give the held burst a client that outlasts the gate

The fault cell holds a burst at the upstream for up to 60 seconds while it kills a worker, but sent the burst through the shared 15 second client, so a slow box could time the survivors out before the gate opened. The burst now goes through its own client whose timeout is twice the gate, and the gate length is one named constant.

* refactor(types): keep the license reply handling and experimental MCP signatures as they were

The license check validated the whole reply as a mapping, which changed the error text logged for a reply that is not an object. It now validates only the verify value, so every reply is handled and logged exactly as before while the value is still typed.

Three files under the experimental MCP server changed annotations on public functions and methods (three returns and three parameters). They go back to their previous content so no public signature in the diff is narrowed.

* test(integration): answer the proxy's model-list call in the OpenAI stand-ins

Every 300 seconds each proxy worker asks an OpenAI deployment for GET /v1/models. Four new cells own an OpenAI stand-in that accepted only the call under test, so a refresh landing inside a cell failed it. The stand-ins now answer that call through the suite's own helper and the cells count only the provider calls they drive.
2026-10-06 15:31:54 +00:00
devin-ai-integration[bot]
126e79c967
fix(model_prices): gemini deep research input limits, vertex flash retirement dates, azure data zone gpt-6.1-sol pricing (#44633) 2026-10-06 07:40:25 -07:00
devin-ai-integration[bot]
46d2c2a6ea
fix(otel): honor per-team Arize sampling rates in OTel v2 fan-out (#44595)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:11 -05:00
devin-ai-integration[bot]
44d5dacbaa
fix(otel): gate the Arize OTel v2 exporter on operator credentials (#44596)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:11 -05:00
devin-ai-integration[bot]
80f18e1326
fix(otel): name an Arize project on every OTel v2 Arize export (#44605)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:23:10 -05:00
devin-ai-integration[bot]
de74f81c69
chore(rust): prune inference deps and rewrite layering docs (#44836)
* chore(rust): prune unused inference crate dependencies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): rewrite inference layering docs for the split format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:27 -07:00
devin-ai-integration[bot]
ae35c9d775
refactor(rust): extract inference-ocr crate (#44832)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
6a55e0a8aa
refactor(rust): extract inference-chat crate (#44827)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
d7f5b40ab4
refactor(rust): extract inference-messages crate (#44818)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
785cb12f37
refactor(rust): extract inference-responses crate (#44811)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:25 -07:00
devin-ai-integration[bot]
63babf23e6
refactor(rust): extract inference-transcription crate (#44809)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
1a9b533e71
refactor(rust): rename litellm-core to litellm-inference (#44802)
* refactor(rust): rename litellm-core to litellm-inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): expose inference base API for format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move shared inference test helpers behind test-support

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
10444df3a0
test: deflake the JEV classifier select and two router tests that inherited leaked state (#44840)
* test(ui): wait for the classifier model popup before picking its option

Base UI exposes its select option asynchronously, and the helper waits for the option and its positioner to become clickable before selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: isolate two router tests from work leaked by earlier tests

Collect earlier tests' garbage before warning capture so their unawaited coroutines cannot be attributed to the target's warning assertion

Filter success-event callbacks by the request's litellm_call_id so queued logging work cannot replace the current request's captured messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:27:49 +00:00
devin-ai-integration[bot]
d4619c499a
refactor(types): replace Any with proven types in 6 files (#44491)
* refactor(types): replace Any with proven types in 14 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base parsing in sso userinfo, copilot auth and hf config lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base parsing at unproven provider seams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep base delete in jwt orphan cleanup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:53:43 +00:00
berriai-litellm-provider-info-sync[bot]
18c3118eb6
fix(azure): add gpt-6-sol priority processing prices (#44829)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 01:49:46 -07:00
devin-ai-integration[bot]
5455152913
refactor: remove fresh tech debt from the 2026-10-05 window (#44821)
* refactor: remove fresh tech debt from the 2026-10-05 window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(lens): drop the review models left unused by the dead review helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 08:33:39 +00:00
devin-ai-integration[bot]
c13746c951
fix(router): price a model group from the deployments that serve it (#44732)
* fix(router): price a model group from the deployments that serve it

A model group's info read composed the group's own model_group_alias
entry into its deployments, a hop the router never takes: an alias is
resolved exactly once at request time, so a group reached as an alias
target is served by its own deployments. For the chain X -> T -> U the
price read for X included U's deployments too, and the free-model
budget waiver refused a free request to X on an over-budget key, while
GET /model_group/info reported U's providers and price for X.

The group info read now prices a group from the deployments routing
serves it with: the ones named after it, the routing group of that
name, or the wildcard route matching it when neither exists. The
budget waiver, GET /model_group/info, the rate limiters, and the
response headers all read the same set as routing. get_model_list
keeps its behavior for every other caller.

* test(router): give the paid fixtures explicit per-token prices

* test(router): call the routed-group read by name so the router coverage gate sees it

* test(integration): audit the alias chain budget waiver on every route, shape, and outage

Thirty-four cells under the management group prove an over-budget key is served through an alias chain entry at the price of the deployment that serves it, on chat, responses, and messages, sync and streamed, through the OpenAI and Anthropic SDKs and raw httpx on both replicas, with the chain middle, the reverse chain, a cost-map priced middle, a ghost middle, wildcard and routing-group targets, malformed and hostile inputs, a cached reply, a repointed alias, a provider failure, and two chaos bursts (a killed worker, a scripted outage)

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 05:31:09 +00:00
devin-ai-integration[bot]
c2c0bb583e
perf(types): defer pydantic schema builds via shared LiteLLMBaseModel (#44720)
* perf(types): defer pydantic schema builds via shared LiteLLMBaseModel

Add LiteLLMBaseModel with defer_build driven by DEFER_PYDANTIC_BUILD (default true) and move litellm and enterprise pydantic models onto it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): build deferred models created by a parent validator; keep lens worker models litellm-free

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(types): link pydantic issue on deferred-build rebuild hook

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:10:08 -07:00
devin-ai-integration[bot]
d7f69d6ba1
refactor(logging): load enterprise alerting loggers lazily so import litellm skips proxy types (#44762)
* refactor(logging): load enterprise alerting loggers lazily so import litellm skips proxy types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover lazy alerting dispatch and docs lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve litellm.proxy._types lazily on attribute access

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(imports): assert proxy types guard imports the checkout under test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:06:53 -07:00
devin-ai-integration[bot]
0b633aa9c8
refactor(types): import proxy-only types under TYPE_CHECKING in SDK modules (#44740)
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:52:59 -07:00
devin-ai-integration[bot]
b72c737fc8
refactor(types): move SpanAttributes, SpecialHeaders and AllowedModelRegion out of proxy._types (#44717)
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:40:13 -07:00
devin-ai-integration[bot]
c3e0156979
fix(ui): polish Lens runs loading, reload, and time range menu (#44789)
* fix(ui): polish Lens runs loading, reload, and time range menu

Port the dashboard-only parts of a0a275e486, 1325389624, 5e7b0afd5c, 03bec959bf, 93e6ccff8d, 3d35d9f920, dc1a2b5b60, 1d0397f1c0, 773bfef066, f3387221ea and 3d049cd4a8 from litellm_lens_server_search onto main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep the Lens timeline on the shown runs' window during a reload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop a redundant fixture comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:27:00 -07:00
devin-ai-integration[bot]
d6c4e918d8
feat(ui): lead the Lens investigation detail with a run report (#44786)
* feat(ui): lead the Lens investigation detail with a run report

Port of the dashboard changes from a43f111abe and 29a1f3175d onto main. The detail view now opens with a run report for the selected run: status, a headline, progress or the failure, then cost, duration, coverage and issues, plus a collapsed activity log. An ordered situation table picks the report's one next action (Run now, Stop run, Retry, Raise budget, Connect worker, Review issues, Monitor this), and Run now moves into the investigation actions menu

Main's live review stays as is below the report, the queue reason still shows under queued progress, and partial results keep their own warning state with the run details folded away

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Stop run on older Lens runs while another run is active

Also drop doc comments that restate the code

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:24:28 -07:00
nate-berri
b8239a9873
fix(vertex_ai): add regional endpoint uplift to gemini-3.1-flash-image (#44673)
* fix(vertex_ai): add regional endpoint uplift to gemini-3.1-flash-image

Google prices Gemini 3.1 Flash Image at 1.1x on non-global endpoints for
input, text output and image output, but the cost map row had no
regional_endpoint_uplift_multiplier, so regional calls billed at the
global rate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(integration): cover regional uplift spend for gemini-3.1-flash-image

* test(integration): read the proxy salt from the environment in the uplift spend cells

---------

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 21:00:22 -07:00
devin-ai-integration[bot]
5f9eff55f8
refactor(cache): remove dead Cache._native_cache runtime path (#44785)
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:58:19 -07:00
devin-ai-integration[bot]
a983b2a5e7
fix(responses): honor caller stream flag when provider forces SSE (internal copy of #34095) (#41235)
* fix(responses): honor caller stream flag when provider forces SSE

The Responses handlers decided whether to hand back a streaming iterator
from the provider payload's `stream` field, which chatgpt sets
unconditionally because the Codex backend only serves SSE. A caller that
sent `stream: false` therefore received a raw SSE stream on /v1/responses,
and the chat-completions bridge failed with "Unknown items in responses
API response: []" once its recovery path lost the raw SSE it reads from

Transport streaming still follows the provider payload; only the caller's
own `stream` value now decides the response shape. When the provider
forces SSE for a non-streaming caller the body is read and aggregated
through the existing path

* test(chatgpt): inject the authenticator into the responses config so handler tests never log in

* fix(responses): treat an extra_body stream flag as the caller's own and drop a redundant comment

* test(integration): audit the chatgpt caller stream flag across responses, chat and messages

---------

Co-authored-by: SeongWoon Cho <coffee@soylatte.kr>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:30:47 +00:00
berriai-litellm-provider-info-sync[bot]
e366e72502
feat(bedrock): add glm 5.3 cross-region rows and nova 2.5 sonic (#44710)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 20:25:24 -07:00
devin-ai-integration[bot]
ab3a59fe81
fix(chatgpt,github_copilot): refuse device-code login inside an event loop or worker thread (#39585)
* fix(chatgpt,github_copilot): refuse device-code login when an event loop is running

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(github_copilot): drop stray whitespace change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt): bound token refresh timeout and drop placeholder assignment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt,github_copilot): keep the token file path out of the event-loop 401 message

* fix(auth): refuse device-code login from worker threads too

/v1/messages runs its handler in an executor thread, where the
running-loop check never fires, so a chatgpt or github_copilot model
still started the interactive device-code login there and the request
hung for up to 15 minutes. The guard now also requires the main thread,
so the login only runs where a human can actually answer it.

* test(chatgpt): keep authenticator tests out of the real token directory

* fix(chatgpt): keep the 5 second connect timeout and the operator's request_timeout on the token refresh call

* test(integration): cover the device-code login guard on the proxy and the SDK

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:20:07 +00:00
devin-ai-integration[bot]
64cddd6e13
test(integration): basic translation cases for the bedrock_mantle route (#44750)
* test(integration): bedrock_mantle-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use assert_translation and drop the LIT-9196 skips in the bedrock_mantle basic cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 03:17:34 +00:00
devin-ai-integration[bot]
a45c6b8f65
test(integration): keep scripted upstream connections open past the proxy's keepalive (#44695)
* test(integration): re-pin claude-sonnet-5 stream recount tokens and outlast proxy keepalive in scripted upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): derive scripted upstream keepalive from AIOHTTP_KEEPALIVE_TIMEOUT

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): hardcode scripted upstream keepalive to avoid importing litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 03:13:06 +00:00
ishaan-berri
2fd5725c04
feat(lens): add copy link button to trace header (#44749)
* feat(lens): add traceShareUrl helper for shareable trace links

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add copy link button to trace header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens): cover copy link on trace header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 03:07:56 +00:00
devin-ai-integration[bot]
402fa63366
feat(proxy): cap batch file records, daily batch uploads, and per-file downloads (#43632)
* feat(proxy): cap batch file records, daily batch uploads, and per-file downloads

Adds three opt-in limits for batch jobs, each settable in general_settings as a
per-key default and overridable in key or team metadata by a proxy admin:

max_batch_file_records rejects a purpose=batch upload with more request lines
than allowed with a 413 before it reaches the provider.

max_batch_file_uploads_per_day counts accepted batch uploads per key and per
team in a UTC day and returns 429 with Retry-After once the count is used.

max_file_downloads_per_minute counts GET /v1/files/{id}/content per key and per
team for each file in a one-minute window and returns 429 with Retry-After.

The existing admin-only guard for batch_enqueued_token_limit now covers all four
metadata keys.

* fix(proxy): keep file usage counters in their own store and gate batch limits on user creation

* fix(proxy): count keyless JWT callers per user, require positive file caps, and take the upload slot after request validation

* refactor(proxy): end file usage cap describers with an explicit return after the match

* fix(proxy): keep file usage counters when more than 200 are live without Redis

The file usage counter store used a default in-memory cache, which holds 200
entries and evicts the one that expires soonest. Without Redis, a caller got a
fresh per-file download allowance after touching about 200 other file ids in
the same minute, and a key got a fresh daily upload allowance once about 200
other keys had uploaded that day. The store now tracks up to 20,000 live
counters per worker, the same bound the login throttle uses

* test(proxy): move the file usage cap tests into the directory the proxy shard runs

Main's shard coverage check found tests/unit/proxy/openai_files_endpoints
claimed by no shard, so its tests would not run in CI. The file moves next to
the other files endpoint tests in tests/unit/proxy/openai_files_endpoint, which
the proxy-endpoints shard already runs

* fix(proxy): declare the file usage counters as rate limit calls

Main's redis producer gate requires every module that writes a shared cache to name its key family, and the file usage counters wrote theirs without one.

* test(files): audit batch file usage caps across processes, Redis outages, and config reloads

* test(files): guard the chaos cells against minute boundaries and open Redis breakers

Two chaos cells each failed once in the audit run. The restart check ran
three sequential downloads with no guard against straddling a UTC minute,
and the exact-cap probe after a Redis outage ran while both workers' Redis
circuit breakers were still open (60 s default recovery), so it counted in
per-process memory and the two workers split the cap

Every burst now carries a window guard, the chaos fixture lowers the
breaker recovery to 2 s, and the post-outage check drives a fresh key to its
cap through a one-worker sibling proxy and then expects the two-worker
candidate to refuse the whole burst, which only the shared Redis count can
produce, polled until the breakers close

* test(files): release the held uploads when the killed-worker cell fails early

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:06:36 +00:00
devin-ai-integration[bot]
12dcce8db2
refactor(tests): extract Rust cache test split (#44780)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:03:00 -07:00
nate-berri
922ea71fe5
fix(vertex_ai): apply regional endpoint uplift on image generation cost path (#44679)
* fix(vertex_ai): apply regional endpoint uplift on image generation cost path

completion_cost had vertex_location but never passed it to the image
generation cost router, so a regional_endpoint_uplift_multiplier on a
Vertex image row would be ignored. No image row carries the multiplier
yet, so nothing is misbilled today. Pass the location through to the
Vertex image calculator for both the token-based price and the
per-image fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(vertex_ai): cover regional image cost through the proxy logging path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(images): hand image_edit's vertex_location to its cost resolver

* test(integration): audit Vertex image regional uplift billing

* test(integration): require every spend row after a proxy restart

* test(integration): reject extra spend rows after a proxy restart

---------

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 20:02:22 -07:00
devin-ai-integration[bot]
c1d639afff
test(integration): basic translation cases for the vertex_ai route (#44751)
* test(integration): vertex_ai-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use assert_translation in vertex_ai-route basic translation cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:00:19 -07:00
devin-ai-integration[bot]
4a963a6c0b
fix(proxy): judge the free-model budget waiver by the group an alias routes to (#44638)
Some checks failed
Unit Tests / Build the Rust bridge (push) Waiting to run
Unit Tests / caching-local (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Blocked by required conditions
Unit Tests / enterprise-managed-files (push) Blocked by required conditions
Unit Tests / enterprise-package (push) Blocked by required conditions
Unit Tests / enterprise-routing (push) Blocked by required conditions
Unit Tests / integrations (push) Blocked by required conditions
Unit Tests / OpenAI and Meta Providers (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Blocked by required conditions
Unit Tests / Vertex AI (push) Blocked by required conditions
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Blocked by required conditions
Unit Tests / proxy-extras (push) Blocked by required conditions
Unit Tests / proxy-feature-endpoints (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Blocked by required conditions
Unit Tests / proxy-infra-root (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
VS Code Extension / vscode-extension (push) Has been cancelled
* fix(proxy): judge the free-model budget waiver by the group an alias routes to

A hidden model_group_alias can reuse the name of a real model_name. The
router serves that name from the alias target, but two of the three checks
behind the free-model budget waiver read the deployments of both the alias
target and the real model that shares the name.

So an over-budget key was served through an alias whose name belongs to a
model with an explicit $0 price when the target is $0 only by the cost map,
which the same key is refused on by its own name. The mirror case refused a
free target because the shadowed name belongs to a PTU-priced deployment.

Resolve the alias once and have the explicit-cost and PTU checks read the
routed group, the same group the price check already reads.

* test(proxy): cover the plain alias form of a shadowed free model name

The same wrong verdict exists for a plain string alias, so the unpriced-target case now runs for both alias shapes.

* fix(proxy): judge an alias chain's budget waiver by the deployments it is served from

The explicit-price and PTU checks read the alias target through
Router.get_model_list(), which follows a second alias hop when the target is
itself an alias key. The router never takes that hop, so an alias chain was
judged by a deployment the request never reaches. The checks now take the
deployments named after the routed group, or the wildcard deployment serving
it when none carries its name.

* fix(proxy): refuse the budget waiver when an alias chain is served by a priced wildcard route

* test(integration): cover the shadowing alias budget gate end to end

Adds the audit cells for a hidden alias whose name shadows an explicitly
free group: streamed SDK refusals on chat, responses and messages,
embeddings, the free wildcard and mixed-group paths, per-model budgets,
JWT and custom auth callers, cache hits, alias removal under traffic,
provider failure and fallback, a concurrent outage burst, and a worker
kill on an owned two-worker proxy

* test(integration): cite the cost-map rows the shadowing alias cells rely on

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:44:35 +00:00
devin-ai-integration[bot]
ff5084687c
fix(ci): namespace claude session ids in tracing seeds and allowlist /v1/logs on backend (#44761)
* fix(ci): namespace claude session ids in tracing seeds and allowlist /v1/logs on backend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allowlist /v1/logs on the gateway alongside /v1/traces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:39:12 -07:00
devin-ai-integration[bot]
bc2df7cdcd
fix(gemini): stop replaying thinking block signatures to Gemini (#44661)
* fix(gemini): stop replaying thinking block signatures to Gemini

A thinking block's signature has no provenance, and LiteLLM never fills it
from a Gemini response (Google signs text and functionCall parts, which ride
provider_specific_fields and the tool call id), so a Claude signature replayed
through a mixed model group reached Gemini as a thoughtSignature and Google
answered 400 Invalid thought signature on every later Gemini-served turn. The
same replay also sent the thinking text a second time as a plain text part.
The thinking text now goes out once, as the thought part built from
reasoning_content, and no part is built from thinking_blocks

* test(gemini): type the parts helper and split its comprehension

* test(gemini): cover thinking signature replay on the integration rig

Two integration files from the audit of the foreign thought signature fix: 62 wire cells asserting the model turn Google receives on chat, messages and responses across gemini and vertex_ai, streaming and not, SDK and httpx clients, the sad shapes of thinking_blocks, context caching through cachedContents, and 3 chaos cells (a concurrent burst across endpoints, upstream stream drops, a worker SIGKILL mid burst)

* test(gemini): read the integration salt from the environment

The wire test decrypted Responses ids with a literal salt; tests/integration/_support/process.py boots the proxy with LITELLM_SALT_KEY when it is set, so the test now reads the same variable with the same default

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:31:53 +00:00
devin-ai-integration[bot]
7cdb3d3760
test(integration): basic translation cases for the bedrock_invoke route (#44752)
* test(integration): bedrock_invoke-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): call assert_translation in the bedrock_invoke basic cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:31:46 -07:00
devin-ai-integration[bot]
2156d6c3c7
test(e2e): assert the sibling-replica cooldown through the router (#44706)
* test(e2e): assert the sibling-replica cooldown through the router

* test(e2e): warm the cooldown reads concurrently so every pod's read lands just before the trip

* test(e2e): send the trip right behind the warm so every pod's cooldown read is pinned to it

* test(e2e): warm every pod with a canned-answer group and trip only after every warm call answered

* test(e2e): trim the sibling cell's module docstring to what the design needs

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:12:57 +00:00