Commit graph

221 commits

Author SHA1 Message Date
devin-ai-integration[bot]
54ae4c5bbf
fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (#43082)
Some checks are pending
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rerun integrations shard after unrelated gitlab prompt manager timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit azure_storage client reuse against a local Data Lake sink

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): cover the exact TTL expiry boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): wait for a rejected write before flipping the sink back

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): add azure_storage log delivery cells behind an opt-in lane

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): restart the proxy mid burst and bound the loss to the unflushed queue

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): drop the redundant stop after the owned proxy exits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): read azure_storage objects at the auth-mode-dependent layout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): install the datalake sdk in the e2e lint environment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): drop the opt-in real Azure e2e cells and their e2e-dev dependency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-30 18:50:15 -07:00
devin-ai-integration[bot]
2c3866ebb4
fix(azure_storage): name Data Lake objects without base64 padding or slashes (#43914)
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 18:17:13 -07:00
devin-ai-integration[bot]
ed4caebb65
fix(anthropic): forward the dangerous-tool-use beta to Azure AI Foundry (#43934)
Map dangerous-tool-use-2026-09-03 for azure_ai in the beta header config so the
Claude Code auto mode beta reaches Foundry instead of being stripped, matching
the anthropic, bedrock, and vertex_ai entries

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-30 16:40:11 -07:00
devin-ai-integration[bot]
c42d06fb80
fix(router): bill service tiers at catalog rates for custom-priced deployments (#43890)
* fix(router): inherit catalog service-tier rates for custom-priced deployments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): apply tier-suffixed long-context rates when only tier thresholds are set

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): cover canonical cost-map backend model resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): use descriptive names for service-tier pricing fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:06:37 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
ryan-crabbe-berri
632b69b5c8
refactor(proxy): answer every team access check with TeamAccess.allows (#43364)
* refactor(proxy): route every team-admin decision through auth/team_access.py

Move the six team-admin helpers out of common_utils, team_endpoints and
key_management_endpoints into litellm/proxy/auth/team_access.py under public
names, and point every management route and helper at them. The key routes
keep checking team admin before org admin, so a team admin whose user row is
gone still passes as before. Status codes and bodies are unchanged, which the
223-case team-admin matrix confirms at the merge base and at the tip

common_utils keeps `_is_user_team_admin` as an alias because the published
litellm-enterprise 0.1.71 wheel still imports it from there

* refactor(proxy): answer every team access check with TeamAccess.allows

Replace the six helpers in auth/team_access.py with one resolver in
litellm/proxy/management/teams/access.py. Each route passes the roles it
accepts (TEAM_OR_ORG_ADMIN or TEAM_ADMIN_ONLY), and /team/update and
/team/info rank roles through strongest_role so org admin still outranks
team admin there

The org lookup moves behind an OrgRoles protocol, implemented by
PrismaOrgRoles in management/users/service.py, and get_team_access in
management/teams/dependencies.py is the only place that reads proxy_server
globals. _check_key_admin_access keeps its name and body from main

Routes that checked org admin first now read the roster first, so a team
admin whose org lookup errors now passes on /team/delete, /team/block,
/team/unblock, member reset_spend and reset_budget, and the team callback
routes. No allowed caller is denied
2026-09-30 15:27:33 -07:00
devin-ai-integration[bot]
b41715b0c2
fix(transcription): honor base_url alias for Groq Whisper and report it as the api base (#43917)
* fix(transcription): honor base_url alias for Groq Whisper and report it as the api base

transcription() and speech() only accepted api_base, so a deployment configured with base_url leaked the alias into the provider params, which Groq rejected as an unknown param, and the request never reached the internal gateway. get_api_base() now reads the same alias so response headers and logs show the configured endpoint instead of the provider default

Resolves LIT-9071

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(speech): keep base_url after existing audio params, route Vertex speech to it, skip empty alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 15:03:23 -07:00
devin-ai-integration[bot]
e662772ad1
fix(proxy): register a UI-configured arize callback next to otel under OTel v2 (#43906)
Callbacks saved from the Admin UI reach litellm through _add_custom_logger_callback_to_specific_event, which skipped the freshly built logger whenever a callback of the same exact class was already registered. With LITELLM_OTEL_V2 enabled every OTel preset (otel, arize, ...) is an OpenTelemetryV2, so a UI-added arize was treated as a duplicate of the yaml otel callback and never attached, and no trace ever reached Arize

The exists check now compares the exact class and the logger's callback_name, so presets that share a class register side by side while a true re-registration of the same preset is still skipped

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:58:40 -07:00
devin-ai-integration[bot]
fc8f3a26bb
fix(proxy): relay Azure passthrough body model groups through the router (#43896)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:57:31 -07:00
devin-ai-integration[bot]
cbe69723b1
test(s3_v2): pin async 5xx retry through the production AsyncHTTPHandler (#43080)
* test(s3_v2): pin async 5xx retry through the production AsyncHTTPHandler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Update tests/unit/integrations/test_s3_v2.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: yucheng-berri <yucheng@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-30 14:30:27 -07:00
yuneng-jiang
e61733b170
test(e2e): align completion, SAIL, and spend-log fixtures with supported contracts (#43902) 2026-09-30 14:21:22 -07:00
devin-ai-integration[bot]
6b9766fa0c
feat(proxy): add native ROI calculator for gateway spend vs merged PRs (#43669)
* feat(proxy): add native ROI calculator for gateway spend vs merged PRs

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): serialize ROI Prisma inputs with builtin containers

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* style(proxy): format ROI calculator backend files

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix: parse fenced ROI estimates and retain completed reports

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi-calculator): correct estimator and dashboard behavior

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* chore(ui): drop next dev generated AGENTS.md block

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): chunk ROI spend user lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): show reused ROI estimates after sync

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(security): address ROI CodeQL alerts

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(proxy): make ROI calculator unit tests discoverable

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(ci): run ROI calculator tests in proxy infra shard

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi): page repository search and recover polling errors

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(roi): bring scheduled analysis and guided setup into the gateway

* fix(roi): recover interrupted syncs and resolve review findings

* fix(roi): preserve cached estimates across report scope changes

* fix(roi): normalize scheduler timestamps to UTC

* fix(roi): fence cancelled syncs and read reports from writer

* fix(roi): preserve reports during metadata outages

* refactor(roi): isolate outage validation and verify uncached retry

* fix(roi): make scheduled job registration repeatable

* style(roi): format scheduler import

* fix(roi): continue syncing accessible repositories

* fix(roi): preserve reports and identity during upstream outages

* fix(roi): persist refreshed identities for reused estimates

* perf(roi): skip writes for unchanged cached identities

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-09-30 14:16:41 -07:00
devin-ai-integration[bot]
f285229b51
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): evict email-only member caches and reset team members metric on delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new delete-team literals within the LIT002 ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve deleted-team member ids before the locked delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup

`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.

* fix(repositories): slice find_by_emails into bounded IN statements

The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.

* fix(proxy): delete a team once when /team/delete repeats its id

The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.

* test(integration): audit cells for /team/delete on large, legacy and concurrent teams

Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.

Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.

Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.

* test(integration): pin each chaos outage to a live /team/delete

The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 13:49:18 -07:00
devin-ai-integration[bot]
657bb777fa
fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599)
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages

vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(hosted_vllm): forward replayed reasoning_content only when it is a string

* test(integration): cover hosted_vllm reasoning_content replay across endpoints

* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-30 13:12:28 -07:00
devin-ai-integration[bot]
50f5cc9bbb
feat(otel v2): excluded_services opt-out for datastore spans on tenant destinations (#43278)
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep upstream support unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): excluded_services resolves from the otel callback config only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): name the otel callback logger so excluded_services owner lookup matches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert no aux datastore traces reach the tenant sink

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read bogus-start proxy log from the results dir

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read only this invocation's bogus-start proxy log

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert operator kept db spans over the whole recorded window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): split operator db-span asserts by trace scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): tolerate a bogus exclusion env when callback config wins

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep bogus exclusion env fatal when a preset parses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): hoist the preset check out of the callback loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): use a rule-scoped pyright suppression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): log and drop unknown excluded_services instead of failing boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the operator spend-writer span before checking the tenant for postgres

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): leave callback init and boot untouched when excluded_services is unset

Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): log and ignore malformed excluded_services instead of failing startup

Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): cover failed upstream calls in the excluded_services matrix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
2026-09-30 12:20:27 -07:00
devin-ai-integration[bot]
264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
yucheng-berri
e5c74cb2a6
fix(cost_calculator): stop copying optional_params into response hidden params (#43637)
* fix(cost_calculator): stop copying optional_params into response hidden params

* test(cost_calculator): assert the stored spend-log request and logging payload carry no forwarded credentials
2026-09-30 10:17:21 -07:00
devin-ai-integration[bot]
2ed9761921
fix(router): strip encrypted reasoning the pinned deployment cannot decrypt (#43781)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 09:56:27 -07:00
devin-ai-integration[bot]
9dda4d895f
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764)
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-09-30 07:32:16 -07:00
devin-ai-integration[bot]
b370996b9d
test(router): settle the shared logging worker before recording shadow callbacks (#43847)
* test(router): settle the shared logging worker before recording shadow callbacks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): always stop the shared logging worker after settling it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 04:44:43 -07:00
shrey-berri
04fa760bf2
fix(bedrock): add beta header for output config in message (#43778) 2026-09-30 00:50:09 -07:00
devin-ai-integration[bot]
314ff111e5
fix(router): carry per-request routing reads on context variables instead of public method kwargs (#43814)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 07:34:06 +00:00
devin-ai-integration[bot]
d79600987e
perf(router): honour the cooldown read interval in the routing prefetch (#43815)
Resolves LIT-9043

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 00:29:20 -07:00
shrey-berri
6cf51383bf
fix(params): filter internal traceback flag from provider requests (#43783) 2026-09-30 00:02:41 -07:00
devin-ai-integration[bot]
13d004fc5a
perf(proxy): refresh auth management objects through the request Redis pipeline (#43776)
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.

Resolves LIT-9012

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: yassin <yassin@berri.ai>
2026-09-29 17:56:31 -07:00
devin-ai-integration[bot]
9525452d37
perf(proxy): one post-call Redis pipeline per backend for spend, rate-limit, routing and response-cache writes (#43779)
Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter
increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot
release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch
is sent once the success and failure callbacks have run, or on a deadline, and pending batches are
drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request
stay direct; numeric string TTLs keep the direct-path coercion.

Resolves LIT-8883

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: yassin <yassin@berri.ai>
2026-09-29 17:26:05 -07:00
devin-ai-integration[bot]
ffb15f946f
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads (#43407)
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.

The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)

Resolves LIT-8882

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 16:42:13 -07:00
tin-berri
92c0d6f5c8
fix(router): bind Claude Code background sessions to their auto-router (#43767)
Claude Code background sessions (claude --bg) stamp x-app: cli-bg on every
request, including main-loop turns. The session router binding only
accepted x-app: cli, so a background session never bound and its
subagents' concrete-model calls bypassed the router.

The binding write already requires the requested model to resolve to a
pre-routing strategy, so background side calls naming plain models still
never bind. Since f6eff1bde0 removed the clear path, the x-app check
guarded nothing else.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 15:17:14 -07:00
devin-ai-integration[bot]
24a7e89738
perf(responses): run aresponses through the async wrapper so the cache is read once (#43769)
* perf(responses): run aresponses through the async wrapper so the cache is read once

aresponses() sets kwargs["aresponses"] = True and runs the decorated sync
responses() on an executor, but _is_async_request() did not recognise that
flag, so the sync wrapper did a second cache lookup on the executor thread
with a differently ordered cache-key input. Every /v1/responses request paid
two cache GETs against two different keys. Recognising aresponses in
_is_async_request() leaves the async wrapper as the only cache reader and
writer for the async path, one GET per request, same key on read and write

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): let a responses cache entry cover aresponses so responses-only configs keep caching

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:09:22 -07:00
devin-ai-integration[bot]
2d034bb35b
perf(proxy): hold one spend counter batch across admission and across post-call accounting (#43369)
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.

Resolves LIT-8881

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:05:33 -07:00
devin-ai-integration[bot]
d2a574b791
perf(router): fetch cooldown state and usage counters in one Redis round trip (#43320)
* perf(router): fetch cooldown state and usage counters in one Redis round trip

The cooldown filter (CooldownCache) and usage-based-routing-v2 selection
(LowestTPMLoggingHandler_v2) each issued their own MGET on every request
because they live in different objects. RoutingReadBatch fetches both key
sets through DualCache.async_batch_get_cache_shared while the healthy
deployments are resolved and hands the usage slice to the strategy, so
selection does not read again. Each cache keeps its own memory tier,
throttling, reservation rollback and circuit-breaker handling, and the
strategy falls back to its own read when the prefetch does not cover its
keys. simple-shuffle keeps reading only cooldowns.

aresponses no longer issues a second, blocking response-cache read from
the worker thread that runs the sync wrapper.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep per-cache tier failures inside the shared batch read

Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared
so a failing tier degrades that cache's read to None the way async_batch_get_cache does,
instead of escaping into routing. Drop the aresponses sync-cache guard: for native
Responses models the worker-thread read is the one whose key matches the write, so
skipping it broke cached /v1/responses replays.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(router): rename usage key builder so the async cache-call check reads it as a key helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): narrow daily-report cache values before numeric comparison

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): type the shared batch-read helpers and merge Redis results without mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: fix import sort in test_dual_cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): flatten shared batch read keys without a stacked comprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 13:52:20 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
devin-ai-integration[bot]
1bfa3d4fa6
fix(model-prices): align Azure, Bedrock, Copilot, Gemini, Groq, OpenAI and OpenRouter entries with official docs (#43598)
* fix(model-prices): correct azure/eu/gpt-6-astra to Data Zone rates

Co-authored-by: rain <1504569896@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): align groq, gemini and openai entries with official docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): roll in verified Vertex, Gemini, OpenRouter and Azure AI registry fixes

Absorbs the fields from #43609, #43666, #43671 and #43644 that match the provider's own docs or price API today, and adds a cost test for the azure/eu/gpt-6-astra Data Zone tiers

Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): add Copilot, Bedrock Kimi K3, Gemini Robotics and OpenRouter values from official sources

Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: rain <1504569896@qq.com>
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
2026-09-29 12:45:22 -07:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
yucheng-berri
5a5e563938
test(proxy): classify every credential-bearing param for the canary suite (#43298)
* test(proxy): classify every credential-bearing param for the canary suite

* test(proxy): classify gcs_path_service_account as secret, run registry in auth-checks shard, check slot ids at import

* test(proxy): use one generic slot id for callback and request-body credential params

* test(proxy): name a canary slot only for params an integration test plants

* test(proxy): classify the SigNoz callback params

* test(proxy): move the slot sync note into the module docstring

* test(proxy): model unplanted credential params as their own classification

* test(security): classify request-body api_key as unplanted until D1 exists; check registry slots against the harness

* test(security): classify request-body api_key under slot D1

* test(security): classify Langfuse and Datadog callback secrets under slots C1 and C3
2026-09-29 18:11:12 +00:00
Itai Modiano
0c553f0398
feat(guardrails): send a configured gateway_name from noma_v2 to Noma (#43678)
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma

The noma_v2 guardrail accepts a gateway_name param, falling back to the
NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty
it goes out as a top-level gateway_name field on /litellm/guardrail. The
param works for both guardrail: noma_v2 and guardrail: noma with use_v2,
and it is appended after the existing constructor params so positional
callers keep their meaning

* chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name

The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so
the lazy snapshot and the generated dashboard types need regenerating

* Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-29 10:20:40 -07:00
devin-ai-integration[bot]
d46304900f
fix(router): stream /v1/messages lifecycle frames live when no fallback can take over (#43600)
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over

The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.

Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): mirror every dispatcher fallback path in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): skip already-tried order levels in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(router): credit the #39566 branch this fix supersedes

Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
2026-09-29 09:22:05 -07:00
fedaeho
85dc7cb62e
fix(proxy): resolve model_group_alias in the zero-cost budget predicate (#43512)
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.

The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.

`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.

Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.

Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:46:33 -07:00
Chase
0fb93ed9de
fix(vertex_ai): forward the per-turn-control beta for per-message output_config (#43558)
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.

Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
2026-09-28 21:35:08 -07:00
Ankit Jha
60fca8298e
fix(otel): send cache and reasoning tokens in langfuse usage_details (#43553)
* fix(otel): send cache and reasoning tokens in langfuse usage_details

The OTel V2 Langfuse mapper only sent input, output and total, so cache reads, cache writes and reasoning tokens never reached Langfuse. Emit them as input_cached_tokens, input_cache_creation and output_reasoning_tokens, and send input/output net of those buckets so Langfuse does not price the same tokens twice.

Fixes #43542

* fix(otel): drop redundant comments from the usage_details change
2026-09-28 21:30:56 -07:00
hsm207
319b08b4b1
fix(google_genai): preserve proxy_server_request in completion adapter (#43536) 2026-09-28 21:29:55 -07:00
devin-ai-integration[bot]
39d14bd855
feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers (#43641)
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): drive the router request test through an httpx MockTransport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): assert tool definitions reach the router upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:08:16 -07:00
devin-ai-integration[bot]
b49661064f
feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5 (#43647)
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 01:44:13 +00:00
devin-ai-integration[bot]
ce25856424
feat(mcp): scan and pin upstream tool descriptions (#43283)
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(mcp): scan and pin upstream tool descriptions

Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.

* chore: sync schema.prisma copies from root

* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes

* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending

The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.

* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper

* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift

* refactor(mcp): trim the tool catalog guard docstrings to one line

* test(mcp): cover guarded discovery boundaries and response definitions

* fix(mcp): bound discovery guardrail concurrency per catalog

* fix(mcp): scan tool catalogs in bounded parallel batches

* fix(mcp): hide pinned catalogs from restricted management views

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-28 18:38:49 -07:00
devin-ai-integration[bot]
5e38a08741
feat(cache): select Rust caching through explicit cache objects (#43601)
* refactor(cache): organize v2 cache as a package

* docs: clarify experimental v2 guidance

* fix(cache): verify cache-hit accounting and preserve logging metadata

* refactor(cache): separate execution facts from host accounting

* refactor(rust): build messages routes with named dependencies

* wip

* fix(cache): preserve facade policy and preflight fallback

* refactor(cache): defer shared Python logging changes

* test(gateway-inference): allow dead code in shared test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cache): key prepared requests and honor facade controls

* feat(cache): use Python caches from Rust Messages inference

* refactor(cache): separate native and Python cache adapters

* refactor(cache): enforce shared composition and adapter boundaries

* fix(cache): let Python key delegated Rust Messages entries

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 00:01:44 +00:00
devin-ai-integration[bot]
5df502b360
feat(providers): add Prism provider (internal copy of #40914) (#41961)
* feat(providers): add Prism provider

* fix(providers): complete Prism registration

* feat(providers): expose Prism responses and messages

* feat(providers): add DeepSeek V4.1 Flash to Prism

* test(providers): exercise Prism endpoint requests

* fix(providers): align Prism pricing and limits with the live catalog

deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models

* test(prism): assert cost-map invariants instead of pinning catalog facts

* test(prism): derive the asserted model list from the cost map instead of pinning it

* test(prism): capture requests through respx instead of appending to a list and swapping the client transport

---------

Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
2026-09-28 16:07:09 -07:00
devin-ai-integration[bot]
e4190d86a6
refactor(rust): centralize host execution and compose callbacks (#43515)
* refactor(rust): extract litellm-host-native as the shared Rust host driver

Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): interrupt the machine when the in-process stream consumer fails

Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): separate the machine contract from coroutine execution

* auth update

* refactor(rust): use standard flow control for host requests

* style(rust): keep host driver imports formatted

* chores

* mostly relocation

* refactor(rust): separate interceptors from queued observers

* refactor(rust): centralize legacy callback mappings and lifecycle

* docs: define Python host boundaries and migration plan

* refactor: enforce Python host and bridge boundaries

* refactor(rust): separate operations from callback composition

* refactor(rust): compose SDK policy through call hooks

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:20:27 +00:00
yuneng-jiang
37be82e45e
test(unit): stop test modules from putting their own directory on sys.path (#43421)
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
2026-09-28 12:20:20 -07:00
devin-ai-integration[bot]
2e5034016f
fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys (#43587)
* feat(anthropic): add Claude Sonnet 5.5

Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys

Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:12:31 +00:00
devin-ai-integration[bot]
c60c714278
test(rust): enforce shared upstream error contract in wheel checks (#43520)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 08:15:41 -07:00