Commit graph

53321 commits

Author SHA1 Message Date
devin-ai-integration[bot]
d79600987e
perf(router): honour the cooldown read interval in the routing prefetch (#43815)
Resolves LIT-9043

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 00:29:20 -07:00
shrey-berri
6cf51383bf
fix(params): filter internal traceback flag from provider requests (#43783) 2026-09-30 00:02:41 -07:00
yuneng-jiang
d02ff435bf
test(bedrock): accept regional aliases that inherit Converse routing (#43785)
* test(bedrock): accept regional aliases that inherit Converse routing

* test(bedrock): cover regional alias metadata independently of catalog
2026-09-29 23:30:06 -07:00
yuneng-jiang
ba6d6d1a95
test(ci): repair MCP Responses and budget fixtures (#43788)
* test(ci): repair MCP Responses and budget fixtures

* test(auth): verify delegated budget changes persist
2026-09-29 23:29:52 -07:00
devin-ai-integration[bot]
b71f02dbcf
fix(ui): keep MCP permissions visible after key, team and MCP server saves (#43810)
* fix(ui): keep MCP permissions visible after key, team and MCP server saves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type the object_permission include as a prisma TypedDict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): do not block key save confirmation on cache refetch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 22:49:20 -07:00
yuneng-jiang
7d9cc28dce
test(ci): refresh retired OpenAI tool-call models (#43676) 2026-09-29 21:53:59 -07:00
berriai-litellm-provider-info-sync[bot]
cd0ac30881
fix(cost-map): add deprecation_date to two together_ai nvidia rows (#43809)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 21:24:04 -07:00
devin-ai-integration[bot]
61a73c59b0
fix(proxy): look up hashed key names with two spend log rows per key (#43656)
* fix(proxy): look up hashed key names with two spend log rows per key

The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.

* fix(proxy): cap each spend log name probe at 100 rows per key

* fix(proxy): bound the newest-row probe at where the oldest probe stopped

The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.

* test(integration): add spend log alias probe cells for the daily activity routes

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 20:21:55 -07:00
devin-ai-integration[bot]
8afabe81f1
fix(ui): surface x-litellm-call-id in Logs search, table and drawer (#42436)
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate api types for spend logs search description

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop redundant comments from the call id logs helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(e2e): format logs call id helper and spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer

The Request ID cell shows only request_id again. When the row's litellm_call_id
differs, the cell tooltip lists it as x-litellm-call-id with its own copy button,
and the drawer header labels the second line x-litellm-call-id: instead of the
call id caption. Stacking two ids in every row made the column noisy for the
common case where the viewer only needs the row they searched for.

* test(e2e): cover the Request ID tooltip hover and copy path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: poll the clipboard after the tooltip copy and drop a jsdom aside

The e2e read navigator.clipboard right after the click, so a slow async write
could fail the check even though copy works. The unit test's fireEvent choice
(jsdom has no layout, so a real pointer move off the trigger closes the tooltip
before the click lands) is documented here instead of inline.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 02:16:56 +00:00
yuneng-jiang
82eb7405f5
chore(deps): bump pyjwt, moment and brace-expansion to clear osv-scan (#43792)
pyjwt 2.13.0 -> 2.14.0 (uv.lock only, pyproject floor unchanged), moment
2.30.1 -> 2.31.0 and the brace-expansion override 5.0.9 -> 5.0.12 in the
dashboard. oauthlib's only fixed release (4.0.0, 2026-09-28) is still inside
the 3-day uv exclude-newer cooldown, so its two findings are ignored until
2026-10-02
2026-09-29 19:11:39 -07:00
yuneng-jiang
d098b02ed9
fix(auth): give UI/CLI session tokens their own AES-GCM context and header-safe shape (#43790)
Some checks are pending
Unit Tests / misc (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context

UI and CLI session tokens are now always encrypted with AES-256-GCM and a
fixed session associated-data value, and the session-token check only accepts
AES-GCM values carrying that same value. Stored secrets keep their current
encryption and decrypt unchanged, so nothing needs migrating.

encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20
cannot bind associated data, so an AAD-bound value is always written as
AES-256-GCM, and an AAD-bound decrypt refuses the legacy format.

Session tokens issued before the upgrade stop validating, so UI and CLI users
sign in once more after upgrading.

* test(e2e): cover real SSO login through the dashboard and the lite CLI

Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts
against a live Keycloak stack. The dashboard spec checks that the SSO
session authorizes the Virtual Keys and Models data requests. The CLI
spec runs a real lite login in an isolated HOME with the keyring
disabled, then lists models and sends one chat completion with the
stored session. The main Playwright config now ignores oidc/.

* fix(auth): encode UI/CLI session tokens as unpadded base64url

Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them

Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token

* fix(auth): prefix UI/CLI session tokens with litellm_login_

A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs

The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-29 18:42:24 -07:00
yuneng-jiang
46f7775157
bump: litellm-enterprise 0.1.71 -> 0.1.72, litellm-proxy-extras 0.4.102 -> 0.4.103, litellm 1.104.0 -> 1.105.0 (#43789) 2026-09-30 01:35:34 +00:00
joshua-berri
6684256136
feat(agents): add identity storage and validation contracts (#43720)
* feat(agents): identity storage and contracts

* fix(agents): cache positive identity lookups with fresh policy checks

* test(agents): include identity attribution in spend fixture

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-29 17:59:21 -07:00
devin-ai-integration[bot]
13d004fc5a
perf(proxy): refresh auth management objects through the request Redis pipeline (#43776)
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.

Resolves LIT-9012

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: yassin <yassin@berri.ai>
2026-09-29 17:56:31 -07:00
berriai-litellm-provider-info-sync[bot]
cae179e655
fix(bedrock): set gpt-6.1-sol max output tokens to 131072 (#43782)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 17:44:48 -07:00
devin-ai-integration[bot]
9525452d37
perf(proxy): one post-call Redis pipeline per backend for spend, rate-limit, routing and response-cache writes (#43779)
Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter
increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot
release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch
is sent once the success and failure callbacks have run, or on a deadline, and pending batches are
drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request
stay direct; numeric string TTLs keep the direct-path coercion.

Resolves LIT-8883

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: yassin <yassin@berri.ai>
2026-09-29 17:26:05 -07:00
devin-ai-integration[bot]
c129ea4fc9
fix(mcp): scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved (#43608)
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates

Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys

An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.

The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.

Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep OAuth metadata generations only while a fetch is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 17:25:39 -07:00
devin-ai-integration[bot]
ffb15f946f
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads (#43407)
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.

The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)

Resolves LIT-8882

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 16:42:13 -07:00
berriai-litellm-provider-info-sync[bot]
e7460f1cff
chore(cost-map): take azure context limits from models-sold-directly (#43759)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 16:07:08 -07:00
devin-ai-integration[bot]
f5a1c9f1f1
fix(proxy): recover session key owners from daily spend for usage attribution (#43642)
* fix(proxy): recover daily spend key owners

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): simplify daily spend owner recovery

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): format daily activity metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover recovered owner metadata merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): bound the daily spend owner lookup with the statement timeout

* test(integration): audit the daily activity key owner fallback on every usage route

Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart

The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 15:36:09 -07:00
berriai-litellm-provider-info-sync[bot]
56a63b4b29
feat(bedrock): add openai.gpt-6.1-sol us geo cris and Mantle rows (#43763)
Co-authored-by: kerry-berri <kerry@berri.ai>
2026-09-29 22:29:45 +00:00
tin-berri
92c0d6f5c8
fix(router): bind Claude Code background sessions to their auto-router (#43767)
Claude Code background sessions (claude --bg) stamp x-app: cli-bg on every
request, including main-loop turns. The session router binding only
accepted x-app: cli, so a background session never bound and its
subagents' concrete-model calls bypassed the router.

The binding write already requires the requested model to resolve to a
pre-routing strategy, so background side calls naming plain models still
never bind. Since f6eff1bde0 removed the clear path, the x-app check
guarded nothing else.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 15:17:14 -07:00
devin-ai-integration[bot]
24a7e89738
perf(responses): run aresponses through the async wrapper so the cache is read once (#43769)
* perf(responses): run aresponses through the async wrapper so the cache is read once

aresponses() sets kwargs["aresponses"] = True and runs the decorated sync
responses() on an executor, but _is_async_request() did not recognise that
flag, so the sync wrapper did a second cache lookup on the executor thread
with a differently ordered cache-key input. Every /v1/responses request paid
two cache GETs against two different keys. Recognising aresponses in
_is_async_request() leaves the async wrapper as the only cache reader and
writer for the async path, one GET per request, same key on read and write

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): let a responses cache entry cover aresponses so responses-only configs keep caching

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:09:22 -07:00
devin-ai-integration[bot]
2d034bb35b
perf(proxy): hold one spend counter batch across admission and across post-call accounting (#43369)
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.

Resolves LIT-8881

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:05:33 -07:00
devin-ai-integration[bot]
d2a574b791
perf(router): fetch cooldown state and usage counters in one Redis round trip (#43320)
* perf(router): fetch cooldown state and usage counters in one Redis round trip

The cooldown filter (CooldownCache) and usage-based-routing-v2 selection
(LowestTPMLoggingHandler_v2) each issued their own MGET on every request
because they live in different objects. RoutingReadBatch fetches both key
sets through DualCache.async_batch_get_cache_shared while the healthy
deployments are resolved and hands the usage slice to the strategy, so
selection does not read again. Each cache keeps its own memory tier,
throttling, reservation rollback and circuit-breaker handling, and the
strategy falls back to its own read when the prefetch does not cover its
keys. simple-shuffle keeps reading only cooldowns.

aresponses no longer issues a second, blocking response-cache read from
the worker thread that runs the sync wrapper.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep per-cache tier failures inside the shared batch read

Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared
so a failing tier degrades that cache's read to None the way async_batch_get_cache does,
instead of escaping into routing. Drop the aresponses sync-cache guard: for native
Responses models the worker-thread read is the one whose key matches the write, so
skipping it broke cached /v1/responses replays.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(router): rename usage key builder so the async cache-call check reads it as a key helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): narrow daily-report cache values before numeric comparison

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): type the shared batch-read helpers and merge Redis results without mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: fix import sort in test_dual_cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): flatten shared batch read keys without a stacked comprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 13:52:20 -07:00
berriai-litellm-provider-info-sync[bot]
27c110cb71
feat(bedrock): add openai gpt-6.1-sol global and base rows (#43758)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 13:00:01 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
devin-ai-integration[bot]
273489824a
refactor(rust): orchestrate Messages route execution (#43719)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 19:54:12 +00:00
devin-ai-integration[bot]
fb74957ddd
fix(guardrails): enable explicit PANW MCP output scanning (#43109)
* fix(guardrails): declare post_mcp_call for PANW Prisma AIRS

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): exercise post_mcp_call_hook dispatch in PANW tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-catalog): add fal_ai resolution-tiered image cost fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep post_mcp_call opt-in for PANW Prisma AIRS

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-catalog): add fal_ai resolution-tiered image cost keys to cost map schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: joshua <joshua@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-29 12:50:08 -07:00
devin-ai-integration[bot]
1bfa3d4fa6
fix(model-prices): align Azure, Bedrock, Copilot, Gemini, Groq, OpenAI and OpenRouter entries with official docs (#43598)
* fix(model-prices): correct azure/eu/gpt-6-astra to Data Zone rates

Co-authored-by: rain <1504569896@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): align groq, gemini and openai entries with official docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): roll in verified Vertex, Gemini, OpenRouter and Azure AI registry fixes

Absorbs the fields from #43609, #43666, #43671 and #43644 that match the provider's own docs or price API today, and adds a cost test for the azure/eu/gpt-6-astra Data Zone tiers

Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): add Copilot, Bedrock Kimi K3, Gemini Robotics and OpenRouter values from official sources

Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: rain <1504569896@qq.com>
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
2026-09-29 12:45:22 -07:00
tin-berri
3b2a447fae
fix(autorouter): compare historical and new savings consistently (#43348)
* fix(autorouter): compare historical and new savings consistently

* fix(autorouter): reject comparisons if request counts changed

* fix(router): restore eligible LLM and classification breakdown

* fix(router): avoid ambiguous baseline labels for partial comparisons
2026-09-29 12:40:19 -07:00
berriai-litellm-provider-info-sync[bot]
f4a7c04d99
chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page (#43745)
* chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page

Price-Sync: litellm-providers

* feat(cost): support openai ultrafast tier fields in the model catalog

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry <kerry@berri.ai>
2026-09-29 19:05:57 +00:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
berriai-litellm-provider-info-sync[bot]
0fe4028cd9
fix(cost-map): lower fireworks up-to-4b size tier to the pricing page price (#43740)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 11:14:14 -07:00
yucheng-berri
5a5e563938
test(proxy): classify every credential-bearing param for the canary suite (#43298)
* test(proxy): classify every credential-bearing param for the canary suite

* test(proxy): classify gcs_path_service_account as secret, run registry in auth-checks shard, check slot ids at import

* test(proxy): use one generic slot id for callback and request-body credential params

* test(proxy): name a canary slot only for params an integration test plants

* test(proxy): classify the SigNoz callback params

* test(proxy): move the slot sync note into the module docstring

* test(proxy): model unplanted credential params as their own classification

* test(security): classify request-body api_key as unplanted until D1 exists; check registry slots against the harness

* test(security): classify request-body api_key under slot D1

* test(security): classify Langfuse and Datadog callback secrets under slots C1 and C3
2026-09-29 18:11:12 +00:00
devin-ai-integration[bot]
bda2763f2c
chore(cost-map): add azure and openrouter gpt-6.1-sol rows (#43744)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 18:06:14 +00:00
yucheng-berri
336c7c0849
test(integration): sweep proxy logs, metrics, a Datadog intake and the Logs drawer for credential canaries (#43306)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): sweep proxy logs, metrics, a gzip Datadog intake and the Logs drawer for credential canaries

* test(e2e): treat an unset prompt-storage setting as unset and restore it

* test(integration): name the Datadog sink slot G1d

* test(e2e): search the Logs page for base64 forms of the deployment key

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): pass the resolved deployment id to the Datadog route sweep

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-29 17:57:28 +00:00
yucheng-berri
cede93e826
test(integration): request-path credential canary slots D1-D4 (#43307)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): request-path credential canary slots D1-D4

* test(integration): read the Logs drawer and spend-log filter for failed request rows

* test(integration): check the marker in each failed row's spend-log filter; run header slots on chat-family routes

* test(integration): run D2 on embeddings again; only the client-header slot runs on chat-family routes

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): pass the slot deployment's model_info id to the route sweep

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-29 10:52:13 -07:00
yucheng-berri
b3dcf8208d
test(integration): callback credential canary slots C1-C3 and D5 (#43630)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown

* test(integration): callback credential canary slots C1-C3 and D5

Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter.

* test(integration): give the callback sink waits a wider bound

* test(integration): sweep provider requests for callback credentials
2026-09-29 10:49:30 -07:00
berriai-litellm-provider-info-sync[bot]
a3552c451b
chore(cost-map): add openai gpt-6.1-sol from the pricing page (#43738)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 10:34:04 -07:00
Itai Modiano
0c553f0398
feat(guardrails): send a configured gateway_name from noma_v2 to Noma (#43678)
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma

The noma_v2 guardrail accepts a gateway_name param, falling back to the
NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty
it goes out as a top-level gateway_name field on /litellm/guardrail. The
param works for both guardrail: noma_v2 and guardrail: noma with use_v2,
and it is appended after the existing constructor params so positional
callers keep their meaning

* chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name

The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so
the lazy snapshot and the generated dashboard types need regenerating

* Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-29 10:20:40 -07:00
berriai-litellm-provider-info-sync[bot]
d2cbc94fc6
feat(cost-map): add baseten DeepSeek-V4.1-Flash-Fast (#43735)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 10:09:11 -07:00
devin-ai-integration[bot]
d46304900f
fix(router): stream /v1/messages lifecycle frames live when no fallback can take over (#43600)
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over

The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.

Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): mirror every dispatcher fallback path in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): skip already-tried order levels in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(router): credit the #39566 branch this fix supersedes

Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
2026-09-29 09:22:05 -07:00
devin-ai-integration[bot]
66db132627
refactor(rust): add shared llms wire type derives (#43730) 2026-09-29 09:19:33 -07:00
devin-ai-integration[bot]
684a1edd44
docs(security): point readers to the security announcements mailing list signup (#43713)
* docs(security): point readers to the security announcements mailing list signup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(security): formalize the security announcements wording

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(security): tighten the best-effort sentence

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 13:15:37 +00:00
devin-ai-integration[bot]
9dfa42dcde
refactor(types): replace Any with proven types in 7 files (#43704)
* refactor(types): replace Any with proven types in 11 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert Any changes that broke existing callers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): drop prompt factory helper wrappers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover typing sweep surfaces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tighten sweep audit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 06:12:58 -07:00
devin-ai-integration[bot]
7f95b5f361
refactor: clean up fresh tech debt from 2026-09-28 (#43674)
* refactor: clean up fresh tech debt from 2026-09-28

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: group leaderboard rows in one pass and wrap docstring at 120

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 02:15:12 -07:00
fedaeho
85dc7cb62e
fix(proxy): resolve model_group_alias in the zero-cost budget predicate (#43512)
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.

The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.

`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.

Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.

Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:46:33 -07:00
Flexomatic81
7b2cbf6e7f
fix(cost-map): add tool calling and reasoning flags, correct max output for nebius DeepSeek-V4.1-Flash (#43588) 2026-09-28 21:40:01 -07:00
Chase
0fb93ed9de
fix(vertex_ai): forward the per-turn-control beta for per-message output_config (#43558)
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.

Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
2026-09-28 21:35:08 -07:00