Commit graph

53843 commits

Author SHA1 Message Date
devin-ai-integration[bot]
740d0435a8
test(integration): scope the team-scoped models upstream check to its own model (#45106)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 11:09:56 -07:00
tin-berri
0338498067
fix(router): preserve native baseline identity and accounting (#44960)
* fix(router): preserve native baseline identity and accounting

* fix(router): preserve injected system caches in native baselines

* fix(router): prepare native baselines through shared request owners

* fix(router): capture native baseline fields from the provider schema

* refactor(router): reuse native provider parameter discovery

* fix(router): keep long native baselines and abstain after compaction

Message history no longer counts against the settings snapshot budget, so long
and non-ASCII native sessions keep modeled baselines. Selected-tier compaction
now abstains because the baseline would otherwise inherit the compacted history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(router): judge implicit caching against the selected request

The implicit-cache guard compared the selected response's cache usage with the
projected baseline's breakpoints, so selected-tier cache markers made a usable
unmarked baseline plan look like unexplained caching. The guard now checks the
selected wire request. Also removes a stamp-reuse branch that could never run
because routing clears the stamp first; every pass already captures caller
settings from fresh kwargs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 11:00:31 -07:00
ryan-crabbe-berri
3538e87e45
test(e2e): tag the remaining quota_management tests and record budget and spend client steps (#44966)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag quota_management tests with Subject metadata and record budget client steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause
2026-10-07 10:33:12 -07:00
ryan-crabbe-berri
c943d650f4
test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps (#44964)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps

* test(e2e): leave the batches cleanup harness unit tests untagged

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): name every driven model on the vllm batch, prompt caching and complexity router subjects
2026-10-07 10:32:58 -07:00
devin-ai-integration[bot]
77fc3315e5
fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts (#43878)
* fix(caching): keep tool calls and tool results in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): keep semantic tool prompt helpers within lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep structured function_call_output text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split Responses text-field collection to stay within complexity budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): tag each tool result with the position of the call it answers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): encode tool result position and output together so tool text cannot forge result tags

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): expect encoded tool result record in qdrant semantic prompt parity case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(caching): embed every semantic cache prompt field except volatile ones

Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"

This reverts commit 86c82b949b.

* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"

This reverts commit 39efb5d9da.

* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"

This reverts commit 5aed3ab3de.

* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split semantic cache prompt extraction by API format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): pick the Responses text field without a Final inside a loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): write the semantic cache prompt builders as plain loops

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): drop formatting-only churn from the redis semantic cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(integration): drop the caching group wiring that main already carries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): recurse into tool_result content in the semantic cache prompt helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): check the max_messages cap on the shared exact-cache proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): read list-form function_call_output text in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"

This reverts commit 0665296bf1.

* style(caching): suppress C901 on the Responses input walker instead of splitting it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:32:39 +00:00
ryan-crabbe-berri
79209b92a1
test(e2e): tag guardrails and logging tests with Subject metadata and record client steps (#44963)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps

* test(e2e): leave the guardrails and logging harness unit tests untagged

* test(e2e): let the inner create_model step name the guardrail backend deployment

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): declare the default guardrail backend model on the tests that drive it
2026-10-07 10:32:25 -07:00
ryan-crabbe-berri
d7cdc88c66
test(e2e): tag management tests with Subject metadata and record management client steps (#44962)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag management tests with Subject metadata and record management client steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): keep the prompt out of the chat_status step so polled retries collapse
2026-10-07 10:32:09 -07:00
ryan-crabbe-berri
0ae55bdf7a
test(e2e): tag claude_code cells with Subject metadata and record CLI driver steps (#44961)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag claude_code tests with Subject metadata and record CLI driver steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): decorate run_claude directly so the label gate discovers its step
2026-10-07 10:31:54 -07:00
ryan-crabbe-berri
ade17902a0
test(e2e): tag llm_translation tests with Subject metadata and record harness steps (#44950)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag llm_translation tests with Subject metadata and record harness steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): declare the realtime param tuples Final
2026-10-07 10:30:37 -07:00
devin-ai-integration[bot]
736da28ac1
fix(tests): match the lowercased bind error in the owned-proxy port-race retry (#45097)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 10:30:28 -07:00
berriai-litellm-provider-info-sync[bot]
086bcd2a47
fix(bedrock): take context and output limits from the Bedrock model cards (#45091)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 10:03:04 -07:00
devin-ai-integration[bot]
1e88d9955d
fix(proxy): opt-in Redis hash-tag grouping for v3 rate limiter (#45085)
Add litellm_settings.force_redis_hash_tag_grouping so Redis endpoints that enforce cluster slot rules behind a standalone protocol (Redis Enterprise clustering policy) group multi-key scripts by slot like RedisClusterCache, and return grouped batch values in the caller's key order.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 11:37:43 -05:00
berriai-litellm-provider-info-sync[bot]
aeec8703a8
fix(bedrock): add 2027-03-30 deprecation date to DeepSeek R1 and Qwen3 Coder rows (#45089)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 09:24:21 -07:00
devin-ai-integration[bot]
e2971e0af4
refactor(llms): expose public names for private provider helpers (#45037)
* refactor(litellm): migrate private usage in llms

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve Bedrock batch signature marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): migrate llms private usage symbols

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve UUID Watsonx project IDs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): retarget llms mocks to public names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): limit llms changes to renames and forwarders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): keep GCS mock client patching private Vertex auth methods

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): assert forwarder arguments and type forwarder signatures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 08:49:25 -07:00
berriai-litellm-provider-info-sync[bot]
da1053f068
fix(cost-map): add us data residency multiplier to anthropic claude-opus-5-5 (#45078)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 08:41:03 -07:00
berriai-litellm-provider-info-sync[bot]
e1bdd42726
fix(bedrock): update GovCloud OpenAI prices, add GPT-6 Astra ultrafast tier and Titan Image v2 EOL (#45077)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 08:32:23 -07:00
devin-ai-integration[bot]
2b8e53b5ca
fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls (#45053)
* fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ollama): separate replayed tool calls from text and tighten parser typing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 11:05:50 -04:00
devin-ai-integration[bot]
736ff14f11
fix(health): attribute background health check results to their own deployment (#44982)
* fix(health): attribute background health check results to their own deployment

Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(health): annotate locals with Final and split nested comprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 10:00:50 -04:00
devin-ai-integration[bot]
91e9b1f06b
fix(auth): clear the recent-miss user memo when /user/new creates the user (#45020)
* fix(auth): clear the recent-miss user memo when /user/new creates the user

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(auth): pin the miss memo window and drop the class patch in the new_user regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): pin a first-time SSO user's first message and second sign-in on one worker

* test(integration): cover /user/new clearing the user-miss memo on the creating worker

The fake IdP now signs in the subject a login_hint names, so cells on the shared one-worker proxy can each use a fresh user. New cells: a plain-key miss followed by /user/new is budgeted at once (same on both legs, the auth prefetch loads the row) and the admin-created user's first SSO sign-in inside the window completes (500 at the callback before the fix). The first-sign-in cell moved onto the shared one-worker proxy fixture

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 06:45:41 -07:00
devin-ai-integration[bot]
cb138ba92f
refactor(types): replace Any with proven types in 16 files (#45029)
* refactor(types): replace Any with proven types in 16 files

* fix(types): import TypedDict from typing_extensions for pydantic on 3.10

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 04:31:51 -07:00
devin-ai-integration[bot]
498a3e1a67
test(proxy): make the stagger-offset and linear-dedup guards independent of runner identity and load (#45031)
* test(proxy): inject the cleanup job stagger offset so the runtime sync test no longer depends on the runner's host and pid

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(policy_engine): count resolver calls instead of timing them so the linear dedup guard is deterministic under CI load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(policy_engine): annotate the line-count guard's locals as Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 04:07:36 -07:00
devin-ai-integration[bot]
17a982d6df
ci: give the database-backed CircleCI jobs their own Postgres (#45007)
* ci: give installing_litellm jobs their own Postgres

* ci: give the entrypoint jobs their own Postgres sidecar

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 02:43:54 -07:00
devin-ai-integration[bot]
7b439d9138
refactor: replace fresh getattr string access and a restating comment (#45017)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 02:08:22 -07:00
berriai-litellm-provider-info-sync[bot]
2c67ae90bd
fix(azure): date claude-opus-5-5 and sonnet-5-5 and fill azure/eu/gpt-6-astra limits (#45018)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 01:19:38 -07:00
devin-ai-integration[bot]
d364d5e7ca
refactor: expose public hidden_params accessors (#44668)
* refactor: expose public hidden_params accessors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: keep hidden params writes on duck-typed responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: move hidden params helpers to core utils

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve hidden params in duck-typed cost responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover hidden params accessors across response types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve hidden params for dynamic responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: expose logs through backend component

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "fix: expose logs through backend component"

This reverts commit 621e5ff25e.

* test: cover dict hidden params helper path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve hidden params storage on item writes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover hidden params accessors for OpenAI types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover hidden params helper and response paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: complete hidden params accessor migration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): tighten hidden parameter validation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): read HiddenParams model storage through get_hidden_params

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): transfer hidden params storage instead of the mapping view

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): expose only set HiddenParams keys through the mapping view

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(litellm): read HiddenParams view keys without dumping the model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): assert HiddenParams extras through model_extra

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): copy streaming hidden params into a plain dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 00:01:33 -07:00
devin-ai-integration[bot]
bce020f8d8
ci: run integration-mcp on an xlarge machine (#44608)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 06:34:49 +00:00
devin-ai-integration[bot]
5555e7f593
test(mcp): set server_id on openapi local fake tools (#44998)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-06 23:28:58 -07:00
devin-ai-integration[bot]
5ffad371b0
revert(docker): unpin openssl-3.6-dev in the pgbouncer-builder stage (#44997)
This reverts commit b91b2fc138 (#44991).

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-06 23:28:27 -07:00
berriai-litellm-provider-info-sync[bot]
1cdc0ecbf3
fix(vertex-ai): correct input_cost_per_second on gemini-3.5 transcribe-live and live-translate (#45002)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 23:22:31 -07:00
devin-ai-integration[bot]
f8ddb20965
fix(ui): render team_metadata_schema keys as fixed labels (#41482)
* fix(ui): render team_metadata_schema keys as fixed labels

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): update team metadata schema tests for fixed labels

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): derive team metadata schema labels from live key values

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 23:21:09 -07:00
devin-ai-integration[bot]
43cf7e7493
fix(ci): stop deferred pydantic builds leaking caller locals and add the missing Lens FK migration (#44981)
* fix(ci): stop deferred pydantic builds leaking caller locals and add missing Lens FK migration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(types): drop narrating comment from caller-locals regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): guard the Lens review FK migration with DO blocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(types): run the caller-locals regression in-process

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): scope the Lens review FK guards to their table

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 23:20:54 -07:00
devin-ai-integration[bot]
68c0a972a7
fix(realtime): skip guardrail VAD session.update injection for transcription sessions (#44843)
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(realtime): flag transcription sessions from the route intent and backend events only

A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail

* fix(realtime): flag transcription sessions from provider-transformed session events

* test(integration): cover transcription sessions skipping the VAD auto-response injection

Adds the realtime transcript guardrail audit cells: transcription sessions on all three
realtime routes, the OpenAI SDK, the beta protocol, the whisper default deployment, the Azure
GA path over a TLS scripted upstream, and Meta Muse push-to-talk sessions keep the client's
session.update verbatim and get their transcript, while voice sessions keep the injected
create_response gate and a client-declared transcription type no longer bypasses it. Sad,
edge, and chaos cells cover malformed session fields, duplicate and older backend session
events, unauthenticated upgrades, repeated sessions, an upstream outage under open sessions,
and a worker kill with a proxy restart.

The scripted upstream now answers session.update the way the vendor does (session.updated,
or the missing turn_detection.type and session-type errors), records every websocket frame,
serves the Azure and Muse realtime paths, and can run over TLS from an owned upstream.

---------

Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 22:42:00 -07:00
devin-ai-integration[bot]
308e575287
test(mcp): run the static root path issuer discovery test in-process (#44988)
* test(mcp): run the static root path issuer discovery test in-process

The test spawned a fresh interpreter with a 60 s deadline to import litellm
and the MCP discovery router cold, so on a loaded box it died with
subprocess.TimeoutExpired before any assertion ran. The discovery routes
bake SERVER_ROOT_PATH into their paths when the module executes, so the
test now reloads that one module under the gateway env, restores its
namespace afterwards, and asserts on the same four discovery documents.
PROXY_BASE_URL now names an origin distinct from the test client's, so the
assertions fail when it stops being honored.

* test(mcp): type the gateway discovery fixture and its test

* test(mcp): mark the registry fill the fixture hands the test

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 22:19:06 -07:00
devin-ai-integration[bot]
bd70e6ddb6
fix(bedrock): forward each Nova Sonic assistant sentence once over the realtime API (#44987)
* test(bedrock): live e2e asserting nova sonic realtime delivers each assistant sentence once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): drop ticket reference from nova sonic e2e docstring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): forward each Nova Sonic assistant sentence once over the realtime API

Nova 2 Sonic sends every assistant text block twice, a SPECULATIVE preview
next to the audio and a FINAL transcript once the audio turn has ended. The
realtime bridge forwarded both, so voice clients rendered each sentence twice
and the FINAL copies opened extra responses after response.done, the last of
which never closed. FINAL assistant text blocks are now dropped whole, so a
turn carries each sentence once inside the one response with its audio

* test(bedrock): tag the live Nova Sonic test and type its helpers

* test(bedrock): pin that the Nova Sonic barge-in marker is dropped with its FINAL block

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 22:02:31 -07:00
devin-ai-integration[bot]
dfe4df8c8a
test: move whole-unit legacy test files into tests/unit and delete dead skips (#44807)
* test: delete unconditionally skipped legacy tests

* test: move whole-unit legacy test files into tests/unit

* test: keep moved legacy tests free of import-time global state

* test: keep the module-level invocation scan pointed at tests/local_testing

* ci: drop the agent_testing CircleCI job emptied by the move

* test: fix moved-test isolation and router coverage

* test: add Tinyfish search package marker

* test: isolate moved tests from logger state leaks

* test: isolate Helicone logging fixture state

* test: isolate Vertex pass-through credentials between moved tests

* test: cancel S3 periodic flush tasks started by moved tests

* ci: restore CircleCI assistant test selection after move

* test: prevent Datadog datetime import shadowing

* ci: drop the litellm_assistants_api_testing CircleCI job emptied by the move

* test: deduplicate imports in rebased unit tests

* test: remove duplicate passthrough router patch import

* test: remove moved legacy source files after rebase

* test: align moved tests with rebased main

* test: carry main's legacy-file edits into moved destinations

* test: make the moved cost map fallback tests assert the fetch and the backup

The four fallback cases only checked the result was non-empty, so they still
passed with integrity validation disabled. They now inject a mock client, assert
one fetch happened, and assert the result is exactly the local backup with the
fallback reason recorded.

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 05:02:28 +00:00
berriai-litellm-provider-info-sync[bot]
15eb898c2d
fix(azure): sync azure openai audio, realtime and image prices with azure pricing page (#44992)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 22:02:00 -07:00
berriai-litellm-provider-info-sync[bot]
22c4b8948b
feat(bedrock): add kimi k3 india cross-region row (#44990)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 21:54:14 -07:00
devin-ai-integration[bot]
b91b2fc138
fix(docker): pin openssl-3.6-dev in the pgbouncer-builder stage (#44991)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-06 21:50:04 -07:00
berriai-litellm-provider-info-sync[bot]
8adb8f3e69
fix(azure): set kimi-k2.7-code deprecation date from the models list API (#44986)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 20:59:44 -07:00
devin-ai-integration[bot]
62dee3d730
feat(claude_code_gateway): issue rotating refresh tokens and a revocation endpoint (#44635)
* feat(claude_code_gateway): issue rotating refresh tokens and a revocation endpoint

The Claude Code gateway's device-code grant now returns a refresh token
alongside the session JWT, so a Claude Code session renews itself before
the JWT expires instead of forcing the user back through the browser sign-in.
grant_type=refresh_token re-mints the JWT from the live user row, rotates the
refresh token, and refuses a replayed, foreign, or identity-only token with
invalid_grant. The discovery document now advertises an RFC 7009 revocation
endpoint, which Claude Code's /logout calls with both tokens, so sign-out
burns the refresh token. The refresh token is the MCP gateway's sealed
session refresh token bound to the fixed client id claude_code, so both
front doors share one single-use record.

* fix(claude_code_gateway): mint the whole credential before claiming the device code

A refresh token that failed to mint answered 500 after the device code was
already claimed and the login deleted, so the client could not redeem the
completed sign-in again. The response is now built first, and the code is
claimed only when it is a 200.

* fix(claude-code-gateway): keep device sign-in working when session signing is unusable

* feat(claude-code-gateway): end the whole refresh chain on a replay or a revocation

* fix(mcp-gateway): fail the single-use peek closed on a Redis fault

* fix(mcp-gateway): read the single-use marker under the cache namespace

* fix(mcp-gateway): refuse a replayed refresh token without ending its chain

A refresh token presented a second time is refused as already used and
nothing else happens to the chain it was rotated from. Claude Code renews
from its in-memory copy of the credential, so a second terminal on the
same machine presents the token the first terminal already rotated and
then recovers from the shared credential file; ending the chain there
would sign both terminals out at every expiry. Only a revocation ends a
chain, so the 503-on-unrecorded-chain-ending path and its tests go away.

* fix(mcp-gateway): end the refresh chain before burning the revoked token

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 20:57:50 -07:00
devin-ai-integration[bot]
e98bbc2f8e
fix(exceptions): map upstream 402 to PaymentRequiredError and cool down 402 deployments (#44879)
* fix(exceptions): keep upstream 402 status and cool down 402 deployments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(exceptions): map 402 to PaymentRequiredError subclass of BadRequestError

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(exceptions): annotate PaymentRequiredError methods

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): skip 402 cooldown on single-deployment model groups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(exceptions): single prefix and 402 fallback response for PaymentRequiredError

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic): map billing_error to PaymentRequiredError regardless of status

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): honor explicit allowed-fails policy for single-deployment 402s

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(anthropic): drop billing_error body mapping to PaymentRequiredError

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(exceptions): type the PaymentRequiredError constructor parameters

* test(integration): cover 402 PaymentRequiredError mapping and cooldown

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 20:35:13 -07:00
joshua-berri
d1cfe17518
feat(mcp): add portable catalog pagination (#44446)
Some checks failed
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / Vertex AI (push) Blocked by required conditions
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / Build the Rust bridge (push) Waiting to run
Unit Tests / caching-local (push) Blocked by required conditions
Unit Tests / core-utils (push) Blocked by required conditions
Unit Tests / enterprise-managed-files (push) Blocked by required conditions
Unit Tests / enterprise-package (push) Blocked by required conditions
Unit Tests / enterprise-routing (push) Blocked by required conditions
Unit Tests / integrations (push) Blocked by required conditions
Unit Tests / OpenAI and Meta Providers (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / proxy-extras (push) Blocked by required conditions
Unit Tests / proxy-feature-endpoints (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Blocked by required conditions
Unit Tests / proxy-infra-root (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(mcp): add portable authorized catalog pagination

* fix(mcp): preserve routing and isolation across catalog pages

* fix(mcp): preserve bare calls after complete catalog pages

* test(mcp): verify bare calls with cached database catalog revisions

* test(mcp): cover refreshed authority and failed catalog continuations

* refactor(mcp): narrow pagination interfaces and preserve listed metadata

* fix(mcp): publish listed metadata only after successful pagination

* fix(mcp): defer bare routes until aggregate listing succeeds

* fix(mcp): refresh key access groups on catalog pages

* fix(mcp): accept read-only optional catalog headers

* fix(mcp): preserve consolidated listing metadata after restack

* fix(mcp): propagate unexpected catalog continuation failures

* fix(mcp): preserve key grants when a user record is absent

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-06 18:54:50 -07:00
devin-ai-integration[bot]
4909bd9e8c
fix(ci): generate a master key for the migration startup jobs (#44978)
The migration_startup_tests job runs pytest over tests/e2e/migrations with
PYTHONPATH=tests/e2e, so it loads tests/e2e/conftest.py, which has required
LITELLM_MASTER_KEY at import since #44718. That PR gave every other e2e job a
"Generate LiteLLM master key" step but not this one, so all five scheduled
migration jobs died before collection with KeyError: 'LITELLM_MASTER_KEY'.
Give the job the same step the other jobs have.

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 18:45:17 -07:00
devin-ai-integration[bot]
51c49038ef
fix(ui): give Lens traces a flush toolbar layout (#44972)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 18:38:42 -07:00
joshua-berri
950da7c2ec
fix(mcp): keep worker MCP configurations consistent via a catalog revision (#42568)
* fix(mcp): refresh server catalog for each gateway operation

* fix(mcp): reject configuration changes during scoped dispatch

* fix(mcp): refresh shared catalog state for each operation

* fix(mcp): refresh catalog before native alias routing

* docs(mcp): clarify native route resolution order

* test(mcp): align streaming fixture and generated API documentation

* fix(mcp): preserve discovery published during catalog refresh

* fix(mcp): coalesce queued catalog reads without a stale window

* fix(mcp): reconcile concurrent route changes when publishing catalog

* test(mcp): provide catalog scope in post-call hook fixtures

* fix(mcp): retain valid live routes and handlers during refresh

* fix(mcp): keep refreshed OpenAPI operation membership authoritative

* test(mcp): preserve logging fixtures after catalog integration

* fix(mcp): isolate transport test admission state and exhaust OAuth outcomes

* refactor(mcp): narrow catalog consistency change to ticket scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(mcp): gate catalog refresh on a database revision marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): refresh waiters that observed a newer catalog revision

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): awaitable catalog revision doubles in prisma mocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): satisfy catalog refresh lint and type gates

* test(mcp): provide catalog scope in toolset fixtures

* test(mcp): exercise temporary OAuth through catalog operations

* fix(mcp): share active catalog snapshots across discovery tasks

* fix(mcp): retain discovered routes across cached catalog operations

* fix(mcp): preserve local tool ownership across discovery

* refactor(mcp): satisfy tightened immutability lint ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): authorize local handlers by registered server ownership

* fix(mcp): check registered handler freshness and isolate fixtures

* fix(mcp): read catalog revisions and snapshots from the writer

* fix(mcp): resolve access groups from the operation catalog

* fix(mcp): preserve empty access group restrictions

* fix(mcp): retain rediscovered routes across concurrent updates

* fix(mcp): enforce registered ownership for local tool dispatch

* fix(mcp): preserve issuer discovery across catalog snapshots

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 18:21:29 -07:00
Simon
48c88bae04
fix(responses): preserve prompt cache reuse in chat bridge (#42281)
* fix(responses): preserve prompt cache breakpoints in chat bridge

* fix(responses): preserve multimodal cache breakpoints

* fix(cache): retain implicit lookup with injected breakpoints

* test(cache): expect implicit responses lookup

* test(cache): expect implicit chat lookup
2026-10-06 17:50:20 -07:00
moyai-devin-berriai[bot]
5dd3ff1714
feat(mistral): add Mistral Large 4 (Le Chonk) support (#44870)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
2026-10-07 00:35:19 +00:00
berriai-litellm-provider-info-sync[bot]
a5a86cdda7
chore(cost-map): sync openrouter prices from the models API (#44975)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 17:20:09 -07:00
devin-ai-integration[bot]
47afd1ba0b
test: fix shared-provider discovery, Codex catalog size and generated master key mismatches in CircleCI suites (#44905)
* test(integration): isolate Codex catalog and provider discovery fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep model discovery constant import-safe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: align test keys and CI env with generated master keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 00:11:25 +00:00
devin-ai-integration[bot]
c3a23fe499
refactor: expose core private helpers under public names (#44871)
* refactor: expose core private symbols with compatibility aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* types: narrow core migration diagnostics

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve runtime behavior in core symbol migration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve core private usage migration behavior

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve private value rebinding compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore optional imports and cover public helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: exempt router property from call coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: align recursive detector ignore names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 00:05:27 +00:00