Commit graph

53110 commits

Author SHA1 Message Date
Devin AI
1a63fbfb3f style(mcp): format override entries comprehension
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 06:25:26 +00:00
Devin AI
5e2cfb2333 refactor(mcp): drop redundant wildcard check in convert_row
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 06:20:38 +00:00
Devin AI
27207bae53 fix(mcp): keep wildcard tool grants as legacy entries during conversion
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 06:19:06 +00:00
Devin AI
6d42ebd991 fix(ui): drop unused mcpAllowedToolsFor import after merge
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 04:45:04 +00:00
Devin AI
d816dc8fbc Merge branch 'main' into litellm_mcp_continuous_tool_defaults
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 04:41:54 +00:00
devin-ai-integration[bot]
c976c16a82
feat(mcp): allow ["*"] wildcard in mcp_tool_permissions to grant all current and future tools (#43108)
* feat(mcp): allow ["*"] wildcard in mcp_tool_permissions to grant all current and future tools

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): run prettier on MCPToolPermissions files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep ["*"] wildcard through toolset union and move constant to litellm.constants

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(mcp): format user_api_key_auth_mcp with ruff format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): restore wildcard ceiling and deny-all regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): treat an empty team tool list as deny-all regardless of key grants

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(mcp): keep legacy [] merge semantics, the truthiness check predates this PR

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): drop banner comment that repeats the wildcard test docstring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: joshua <joshua@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 21:34:57 -07:00
berriai-litellm-provider-info-sync[bot]
efbb3ac87e
chore(cost-map): add together-ai deprecation dates for gpt-oss-20b and gemma-4-31B-it (#43127)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-24 21:25:23 -07:00
Devin AI
1f48e23e31 fix(mcp): apply convention ceiling to echoed caller team and read overrides before field delete
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 04:22:07 +00:00
devin-ai-integration[bot]
1f8997398e
refactor(rust): extract the host coroutine into its own crate (#43129)
Move the generic async coroutine out of the host crate into litellm-coroutine,
with its requirements documented in the crate's AGENTS.md. RouteMachine becomes
CallMachine on top of it, host protocol ops move to host/src/protocol.rs, and
the messages, OCR, host-python, python-bridge and legacy callbacks crates adopt
the new types. Adds error definition rules to litellm-rust/AGENTS.md.

Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-25 04:20:06 +00:00
devin-ai-integration[bot]
0d47347ad7
fix(cost): apply a deployment's pricing override to realtime sessions (#43114)
Some checks failed
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* fix(cost): apply a deployment's pricing override to realtime sessions

Pass the resolved custom pricing model into the realtime and transcription cost paths so model_info rates and base_model on a realtime deployment are honoured instead of the model the session reported. Adds an integration test that bills a realtime turn at the deployment's configured rates

Carries the fix from #36958

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): honour audio-only and base_model realtime pricing overrides

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): keep flat per-unit prices from claiming the deployment pricing key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): keep base_model out of realtime transcription rate overrides

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): type the realtime pricing test parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): try a realtime deployment's base_model ahead of the session model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 21:03:59 -07:00
Devin AI
b7a82bb7b9 test(ui): hoist wireBody expectation under the inline-object budget
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:57:27 +00:00
Devin AI
6494f8afad test(ui): update payload and listMCPTools assertions for tool overrides
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:50:59 +00:00
yuneng-jiang
3fa02ef9fc
bump: litellm-enterprise 0.1.70 -> 0.1.71, litellm-proxy-extras 0.4.101 -> 0.4.102 (#43120) 2026-09-24 20:28:13 -07:00
Devin AI
4dd482e923 chore(mcp): regenerate openapi artifacts and sync payload expectations
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:27:51 +00:00
devin-ai-integration[bot]
f61b3c3f38
refactor(types): declare litellm-owned kwargs as typed objects and derive the lists from their fields (#42843)
* refactor(types): declare litellm-owned params in one registry

* refactor(types): re-export registry constants without redundant aliases

* refactor(types): satisfy type-discipline rules in registry projections and tests

* refactor(types): classify every registry entry and check groups against typed config models

* test(types): pin load-bearing names and exact projections in registry tests

* style(types): keep agentic projection comment within ruff format

* refactor(types): declare litellm-owned params as typed objects and derive the lists from their fields

* refactor(types): fields of the typed objects become the registry; tests use a hand-written inventory

* refactor(types): split traversal into wire_names and owned_wire_names, move rust to kwarg artifacts

rust is a module-level switch (litellm.rust) that nothing reads from a call's kwargs, so it
joins self, use_client and model_config as a registered artifact instead of a DispatchOptions
field. The field constants now import from litellm.types.litellm_params directly instead of
through a re-export in litellm.types.utils. metadata and litellm_metadata are MutableMapping
because their readers mutate them in place, and client accepts raw httpx clients

* refactor(types): own max_agentic_loops as an option and walk only nested leaves

Move max_agentic_loops from AgenticLoopState to a new AgenticLoopOptions leaf under
LiteLLMOptions, since the interception handlers read it as a deployment ceiling rather
than stamping it. Drop the owned_wire_names fallback that treated an unresolved annotation
as a direct field, which under postponed annotations silently shrank the registry. Re-export
TRUSTED_CALLBACK_VARS_FIELD and ADDRESSED_RESPONSE_ID_FIELD from types.utils so that import
path keeps working. Tests use hand-written inventories for the callback and pricing names

* refactor(types): move data_residency to call state and drop aliased re-exports

data_residency is stamped by get_litellm_params and responses.main during the
call, so it lives on CallState, not CostOptions. mock_response also accepts a
float sequence, which main.py reads for mock embeddings. The types/utils.py
re-exports become one plain import with an exact F401 suppression instead of
two X as X aliases that pushed PLC0414 over its strict-gate ceiling. Redundant
leaf docstrings and the structural artifact test are gone; the re-exported
FIELD constants are checked by identity instead

* refactor(types): project owned kwarg names once and keep pass-through extraction in request order

* refactor(types): type caching_groups from its cache reader and hoist the pass-through ownership set

caching_groups is a sequence of flat model-group sequences, which is what
Cache._get_caching_group iterates. A regression test drives the public
cache key path so two groups in one caching group share a key and a third
does not. The pass-through endpoint builds its frozenset of owned names
once at import instead of per request, reads the two metadata carriers
from the extracted mapping instead of popping them, and its extraction
mappings are read-only. Concatenation tests assert the whole derived list
and tuple, docstrings drop reader claims that nothing in the module backs

* refactor(types): read owned names live in pass-through and pin tests to literal inventories

The pass-through endpoint checks body keys against the public all_litellm_params
list at request time again, as the base does, instead of a frozenset taken at
import, so a name registered after import is still extracted. A test drives
that path with a name added after import, and another sends both metadata
carriers interleaved with provider keys and asserts the whole merged result.

retry_policy accepts the mapping form its router reader builds a RetryPolicy
from. The pricing inventory in the typed tests is a literal tuple checked
against the model's fields, the agentic compatibility test asserts type, length
and set instead of declaration order, and the typed-model overlap tests assert
the exact intersection.

* refactor(types): move model_alias_map to CallState and read the owned registry in registry order in pass-through

* refactor(types): drop restating docstrings, keep FIELD importers on types.utils, pin pass-through registry order

* fix(types): satisfy strict lint for public FIELD re-exports

* fix(types): restore clean parameter re-exports

* fix(tests): compare pass-through extraction order to registry body keys

* refactor(types): type owned request parameter leaves

* refactor(types): share routing strategy literal and tighten leaf tests

* fix(proxy): drop client-supplied proxy-stamped names from pass-through litellm_params

* refactor(proxy): name pass-through litellm key split for what it holds

* refactor(types): drop TODO markers on the kept readerless fields

* fix(types): keep deployment tag_regex and max_file_size_mb out of provider requests

* fix(types): include every routing strategy the router accepts

---------

Co-authored-by: shrey kharbanda <shreshth@berri.ai>
2026-09-24 20:18:41 -07:00
devin-ai-integration[bot]
f3cf1cdfef
fix(router): honor disable_fallbacks on mid-stream fallback (#43111)
* fix(router): honor disable_fallbacks on mid-stream fallback

The mid-stream fallback hop on chat, Responses, and Messages hardcoded
disable_fallbacks=False, so a request or key that opted out of fallbacks
still got a fallback deployment's answer when the primary's stream died
before its first chunk. Each hop now reads the opt-out from the request
the way the pre-stream path does, and the sync chat stream re-raises the
primary's error instead of re-entering the fallback chain

* test(router): prove disable_fallbacks reaches the mid-stream hop through the public entrypoints

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 20:06:53 -07:00
devin-ai-integration[bot]
d0d3b6a67e
fix(vertex_ai): stop advertising OpenAI platform-only params on Gemma and Llama routes (#43079)
* fix(vertex_ai): stop advertising OpenAI platform-only params on Gemma and Llama routes

The Anthropic /v1/messages bridge derives prompt_cache_key from Claude Code's
session id whenever the provider config advertises it, and every Vertex
OpenAI-compatible route (gemma/, openai/<endpoint>, meta/) inherited the full
OpenAI list, so the Model Garden vLLM container rejected each turn with a
pydantic extra_forbidden 400. Vertex's Llama and Gemma configs now filter one
shared list of platform-only params (prompt_cache_key, prompt_cache_retention,
safety_identifier, service_tier, store, web_search_options, modalities,
prediction, audio, max_retries) out of their supported params, so the bridge
no longer derives the key and drop_params drops an explicit one.

* fix(vertex_ai): scope the platform-param filter to self-deployed Model Garden endpoints

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 20:04:33 -07:00
Devin AI
e7456e8a0f fix(mcp): limit bare delete-verb precedence to leading or conjoined verbs
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:03:56 +00:00
Devin AI
cb1ac02814 fix(ui): outer deny precedence and prune overrides for deselected servers
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:01:35 +00:00
Devin AI
097308420e fix(mcp): overrides do not grant servers, bare delete verbs outrank read, scoped v0 conversion
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 03:01:24 +00:00
berriai-litellm-provider-info-sync[bot]
993a5b9d97
chore(cost-map): add azure retirement dates for command-r-plus and gpt-4 (#43117)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-24 19:52:16 -07:00
Devin AI
878c646e99 merge: resolve main into litellm_mcp_continuous_tool_defaults
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 02:51:29 +00:00
devin-ai-integration[bot]
1d51a8dfc3
fix(proxy): authorize key model aliases the same way as team aliases (#43049) 2026-09-24 19:51:02 -07:00
berriai-litellm-provider-info-sync[bot]
c7f15709f9
chore(cost-map): add computer-use-preview deprecation date from the openai deprecations page (#43116)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-24 19:47:08 -07:00
Devin AI
0a69a19155 fix(ui): load the full MCP tool catalog in the permissions editor
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 02:42:42 +00:00
Devin AI
94362273fe fix(mcp): wrap json CAS filters in prisma Json sentinel
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 02:26:07 +00:00
Devin AI
989ac9263d fix(mcp): serialize backfill CAS filters as plain dicts
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 02:21:00 +00:00
devin-ai-integration[bot]
cbba14682a
fix(fal_ai): price nano-banana-2 and nano-banana-pro image generations by resolution (#43101)
* fix(fal_ai): price nano-banana-2 and nano-banana-pro image generations by resolution

* fix(fal_ai): bill passthrough submits per requested image and register the resolution price keys

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 19:17:46 -07:00
devin-ai-integration[bot]
3d5660f32e
chore(cost-map): add azure retirement dates from the retired Foundry models page (#43104)
* chore(cost-map): add azure retirement dates from the retired Foundry models page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): correct azure/ada retirement date to text-embedding-ada-002 schedule

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 19:16:45 -07:00
Devin AI
5b1076d156 refactor(mcp): adopt immutable collection types for permission defaults
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 02:16:43 +00:00
devin-ai-integration[bot]
5a75f09d6d
fix(vertex_ai): surface the Gemma container's own error inside a 200 :predict response (#43075)
* fix(vertex_ai): surface the Gemma container's own error inside a 200 :predict response

* refactor(vertex_ai): move the gemma container error parser next to its adapter

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 19:00:14 -07:00
devin-ai-integration[bot]
5c0de806fb
feat(proxy): let team admins update member key budgets when enabled (#42555)
* feat(proxy): let team admins update member key budgets when enabled

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): drop casts flagged by LIT006 in member key budgets change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): send budget-only key updates when a team admin edits a member key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): block spend echo in team admin member key updates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): only send dirty budget fields in team admin member key updates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): replace class method monkeypatch with module symbol patch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 01:54:52 +00:00
devin-ai-integration[bot]
811b42ba8d
chore(cost-map): add gemini tts batch output prices from the Gemini API pricing page (#43103)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:54:15 -07:00
devin-ai-integration[bot]
eff6fc1824
chore(cost-map): add openai deprecation dates from the deprecations page (#43102)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:49:42 -07:00
devin-ai-integration[bot]
043331f95a
feat(lint): cap comprehensions at one for and one if clause (LIT014) (#42650)
* feat(lint): LIT013 caps comprehensions at one for and one if clause

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(lint): rewrap the type discipline gate rule list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): honor comprehension-ok on any line a comprehension spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): scope comprehension-ok to the innermost comprehension spanning it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): break equal-span suppression ties toward the inner comprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): let single-line and only violating comprehensions own comprehension-ok markers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(lint): type tmp_path in LIT014 tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:45:24 -07:00
devin-ai-integration[bot]
1a4a9c5ab3
fix(vertex_ai): return chunk content, extractive text, and structData from search_api vector store hits (#43100)
* fix(vertex_ai): return chunk content, extractive text, and structData from search_api vector store hits

* fix(vertex_ai): report a chunk hit's relevanceScore as the search result score

* test(vertex_ai): type the search response helper and parametrized case

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:41:32 -07:00
devin-ai-integration[bot]
57bead9842
test(e2e): pin end-user and tag attribution from Codex-style headers on /v1/responses (#43093)
* test(e2e): pin end-user and tag attribution from Codex-style headers on /v1/responses

Codex CLI has no body field for the end user, so its config.toml
http_headers attach x-litellm-customer-id or x-litellm-end-user-id plus
x-litellm-tags to every /v1/responses call. The proxy already honors
those headers on the Responses route, but nothing in the e2e stack
pinned it. The new case sends that exact wire shape with each standard
customer header and fails unless the spend row carries the end user,
the tags, the aresponses call type, and a nonzero cost.

* test(e2e): assert the customer's /customer/info total matches the header-attributed Responses row

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:33:49 -07:00
Mateo Wang
25fb7810c2
fix(spend): attribute CLI session spend to the per-user cli-session alias instead of the hashed session token (#40541)
* fix(spend): attribute CLI session spend to the per-user cli-session alias instead of the hashed session token

A CLI session token is a fresh random secret on every login, so since v1.99 each
login's spend rows carried a different sha256 hash as api_key and the usage APIs
could resolve neither key_alias nor user_email for them. Spend rows and logging
callbacks now attribute a session request to its stable alias,
cli-session-<user_id>, and the usage endpoints derive that alias and owner from
the key itself instead of scanning for a matching digest

* fix(spend): resolve the CLI session team from the user's first team in usage metadata

A cli-session key carries no team of its own in the DB, so the usage
breakdown showed team_id None for it and the export grouped it as
Unassigned. The login attaches the user's first team to the session, so
the recovery mirrors that rule for cli-session keys only.

* fix(spend): claim the session team only for a single-team user

The CLI login attaches a team on its own only when the user has exactly
one; a user in several teams picks one per login, so usage metadata for
the alias would otherwise name a team the login may not have used.

* test(pass_through): mark the mocked auth object as a plain key

The logged key follows the alias only for a session token; a bare
MagicMock reads as one, so the test names the field it relies on.

* fix(spend): attribute CLI session pass-through, queue, and managed batch spend to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): only treat the exact cli-session-<created_by> value as a batch key alias

A managed object row written by an older build can still carry the raw per-login
session token, which shares the cli-session- prefix. Matching on the prefix alone
would have surfaced that token as a trusted alias and persisted it verbatim in the
batch cost spend log, so the alias check now requires the exact per-user value and
every other prefixed value keeps going through redaction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): log proxy executed batch rows under the cli-session alias instead of the session token

_row_metadata set user_api_key from the raw bearer token while user_api_key_hash carried the alias, so the spend log redaction rejected the alias as untrusted and hashed the random session token instead

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): attribute semantic search embedding spend to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): scope /key/spend/report for a CLI session to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): use the cli-session alias for websearch spend, prometheus failure labels and the parallel limiter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(spend): drop explanatory docstrings on get_logged_api_key and attach_user_details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): only recover cli-session usage keys whose suffix is a known user

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:21:47 -07:00
devin-ai-integration[bot]
e2302be068
refactor(ocr): remove the Python OCR execution path and require the Rust route (#43081)
* refactor(ocr): remove the Python OCR execution path and require the Rust route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fmt

* refactor(ocr): tidy the native OCR passthrough binding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ocr): ruff format the azure passthrough transformation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ocr): resolve passthrough OCR costing in one Rust call

Replace passthrough_url/passthrough_transform with passthrough_response,
which matches the relayed endpoint against each Azure config's path
segments instead of building a fake request to call get_complete_url.
The binding drops the unused headers, status and api_base arguments.

Catch the ValueError/RuntimeError the binding raises so a relayed body
that is not OCR-shaped falls back to the passthrough object instead of
failing logging, and cover the relay against the real binding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ocr): drop the unused LlmProviders import from health check helpers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: drop the ocr_testing job now that tests/ocr_tests is gone

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ocr): restore the live OCR matrix and the ocr_testing job

The public litellm.ocr / aocr / Router interface is unchanged by the Rust
migration, so the live provider matrix still applies. Drops the stale VCR skip
list for the deleted test_rust_bridge.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ocr): import Final in the health check helper tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 18:18:50 -07:00
devin-ai-integration[bot]
f1ef7fc0c2
feat(rust_bridge): read secrets through Python from Rust routes and declare Rust-only routes with NO_PYTHON (#43057)
* done

* fix(rust_bridge): run Python secret reads under the caller's contextvars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust_bridge): run every blocking Python call under the caller's contextvars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:18:50 -07:00
Devin AI
dda8295af7 refactor(mcp): trim permission comments and tighten tool override wiring
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 01:14:20 +00:00
devin-ai-integration[bot]
7bdccd7371
feat(proxy): enforce tpm_limit and rpm_limit set on tag objects (#41807)
* feat(proxy): enforce tpm_limit and rpm_limit set on tag objects

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): keep tag rate limit helpers within type discipline budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop descriptive docstrings from tag rate limit helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover tag object rpm and tpm limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): type the fake tag batch helper parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): name the over-limit tag in 429 errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover tag rpm limit shared across teams, orgs and users

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): format the tag descriptor match in the v3 limiter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): drop Final annotation inside loop for pyright

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-24 18:11:57 -07:00
devin-ai-integration[bot]
67d7ac58cd
feat(ui): make the audit log detail drawer wider and resizable (#42808)
* feat(ui): make the audit log drawer wider and resizable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): scope drawer width classes to the sheet side variant

The base SheetContent variant data-[side=right]:sm:max-w-sm beats a plain
sm:max-w-none: the compound data+sm variant sorts later in the Tailwind v4
output and twMerge does not treat them as conflicting, so the sheet stays
capped at max-w-sm. That is also why the old w-[60%] sm:max-w-none on main
rendered at 384px. Expressing every width class under the same
data-[side=right] variant chain lets twMerge dedupe and makes CSS order
deterministic. Also removes the drag listeners on unmount mid-drag.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keyboard resizing and re-grab guard for the resizable drawer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): query the sheet by dialog role instead of document.querySelector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep drawer resize controls pinned and honor the 720px floor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): announce the rendered drawer width when the 720px floor overrides the stored percent

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 21:10:50 -04:00
devin-ai-integration[bot]
a76f23ac4f
fix(bedrock): map Anthropic batch row params the way real time does (#43087)
* fix(bedrock): map Anthropic batch row params the way real time does

* fix(bedrock): let a batch row's allowed_openai_params reach the mapper

* test(bedrock): assert the batch thinking value matches the real-time mapping

* fix(bedrock): keep json_mode out of Anthropic batch rows and pin route-prefixed deployments

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:09:23 -07:00
devin-ai-integration[bot]
1987133b4e
fix(router): match provider-prefixed fallback keys for bare model groups served by wildcard deployments (#43062)
* fix(router): match provider-prefixed fallback keys for bare model groups

* fix(router): infer the fallback key's provider the way routing does for bare model groups

A bare model group served by a wildcard deployment (claude-sonnet-4-6 routed to anthropic/*) now finds a fallback keyed <provider>/<group>. The provider is inferred through one shared helper, inferred_provider, which the pattern router already used inline, so the fallback lookup and routing agree on the prefix. The lookup only infers a provider when some fallback key ends in /<group>, so alias-style groups never hit the resolver

* fix(router): resolve context window and content policy fallback keys through the shared lookup

---------

Co-authored-by: Jason Dougherty <jasondoc3@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:08:56 -07:00
Devin AI
2bb1c4bf2d feat(ui): write mcp tool permissions as overrides over convention defaults
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 01:05:56 +00:00
devin-ai-integration[bot]
37f5267991
feat(cost-map): add fireworks deepseek-v4p1-flash US-only rows (#43097)
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 17:59:45 -07:00
devin-ai-integration[bot]
0d115ad8b5
chore(cost-map): sync gemini priority, flex and video token prices from the Gemini API pricing page (#43091)
* chore(cost-map): add gemini priority and flex prices to nano-banana-pro-preview and video token price to 3.1 flash live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(cost-map): add video token price to bare gemini-3.1-flash-live-preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 17:41:02 -07:00
Devin AI
766438bcc1 feat(mcp): backfill v0 tool permission rows and close toolset widening
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-25 00:39:40 +00:00
devin-ai-integration[bot]
e450f2d54c
fix(router): serve Responses turns from a sibling when the encrypted content origin has no boundary peer (#43015)
* test(integration): reproduce encrypted_content_affinity 503 when origin has no boundary peer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep encrypted content affinity turn one out of the response cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): degrade encrypted_content_affinity when the origin has no encryption-boundary peer

A routed-group candidate that is currently unavailable and shares its
(api_base, api_key) with no healthy deployment used to raise a proxy-level
503/429 from _unavailable_origin_error, even though healthy siblings in the
same model group could still serve the turn. Strip the encrypted reasoning
and dispatch to the healthy pool instead, matching the existing cross-group
behavior, and log a warning naming the origin model_id and routed group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): bound the degraded-affinity log marker and assert the strip on the wire

Address review findings: restore num_retries alongside
optional_pre_call_checks in the integration test teardown, record scenario
request bodies on the scripted upstream so the tests can assert no encrypted
reasoning reaches the sibling, and truncate the client-supplied model_id in
the degraded-dispatch warning

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): only record JSON bodies on the scripted upstream

Multipart uploads to scripted POST routes have no JSON body, so gate the
observation recording on the request content-type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): make encrypted_content_affinity runtime-toggleable so /config/update can turn it off

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 17:31:43 -07:00