Commit graph

53869 commits

Author SHA1 Message Date
devin-ai-integration[bot]
fa2c8984ba
test: move the unit half of 126 mixed legacy files into tests/unit (#45090)
* test: move the unit half of 126 mixed legacy files into tests/unit

* test: restore litellm globals that moved tests set

* test: finalize migration test cleanup

* test: restore original bodies of moved legacy tests

The move into tests/unit had rewritten 612 test bodies, and some of the rewrites dropped assertions. Each moved test now carries its original body from the legacy file, with only the imports, helpers, fake provider credentials and monkeypatched env it needs to run under tests/unit

test_timeout_streaming goes back to tests/local_testing because it needs the fake OpenAI endpoint server. The image payload fixture moves with its only user, and two tests that leaked global state (a registered model cost entry and queued logging tasks) are now isolated

* test: drop module imports shadowed by restored local imports

* test: assert on LiteLLM output in no-assertion moved tests and isolate leaks

Twenty no-assertion candidates get one assertion on the value LiteLLM returns, with the original lines unchanged. Four tests go back to their legacy files because they only check types or imports, write into the working directory, or cannot assert without a body change

Two moved tests leaked globals into later tests in the same worker, so monkeypatch fixtures now restore the retry-after header parser and the end user cost tracking flags

* test: drain queued logging tasks before the Phoenix span test

The moved Phoenix test counted spans from logging tasks that earlier tests had queued, so the drain fixture moves to tests/unit/conftest.py and both it and the Datadog batch test use it. test_factory_function goes back to its legacy file because its returned wrapper calls the real Assistants API and cannot be asserted on without a body change

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 14:07:43 -07:00
devin-ai-integration[bot]
127278f951
feat(guardrails): add logging_only_scope to observe input, output, or both (#43695)
* feat(guardrails): add logging_only_scope to observe one direction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(guardrails): remove unrelated test churn

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): allow native lifecycle logging-only scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate api types for logging_only_scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): remove callbacks when scope validation fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep output-only scans when request copy fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(ui): configure logging_only_scope on guardrails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep guardrails enforcing when logging_only_scope is invalid at load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): avoid inline guardrail test fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): support directional scope in PATCH and provider UI

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): format guardrail files for frontend lint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): reset unsupported directional scope selections

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): normalize logging-only scope and sanitize warning logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): normalize scope in shared guardrail field

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(api): sync guardrail schema artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): tolerate invalid stored logging-only scopes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add logging_only_scope integration audit cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): fix scope test lint issues

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): extend K4 chaos test timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): restore stored guardrail row verbatim on rejected patch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): satisfy collection lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): validate masked params through TypeAdapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): tolerate invalid stored params on reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep constructor coercion in tolerant params parser

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(ui): edit logging_only_scope in the custom code guardrail modal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): extract custom code logging-only scope control

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep no-op PUT and logging_only scan-failure semantics stable

Three regressions from the logging_only_scope feature, fixed while keeping
the input/output/both selection working:

1. PUT with byte-identical litellm_params no longer forces a teardown +
   re-init of the live callback. reject_invalid_logging_only_scope now
   only validates; re-init still happens exactly when params/name change.
   Previously a description-only PUT re-appended the callback at the END
   of litellm.callbacks, reordering guardrails: with a BLOCK guardrail
   created before a MASK one, the mask started winning and blocked
   requests started succeeding (400 -> PUT 200 -> 200). Invalid unchanged
   scopes are still rejected with 422 without touching the live instance.

2. CustomGuardrail._scan_logged_call no longer swallows input-scan
   exceptions per branch: a raising or BLOCKED input scan aborts the
   logging_only hook again (one policy call, one verdict) for guardrails
   that never selected a logging_only_scope. Explicit input/output/both
   selection keeps selecting which scans run.

3. A failed PATCH now rolls the DB row back to exactly the stored raw
   litellm_params (a legacy 4-key row stays 4 keys) instead of expanding
   it to a full LitellmParams dump; pinned with a test. This matches the
   PUT rollback shape and is the faithful rollback.

* refactor(guardrails): type directional-scope provider list as a tuple

* fix(lint): stay within the basedpyright budget

- drop a Final annotation assigned inside the validation loop
  (reportGeneralTypeIssues over ceiling by one)
- rename the PR-introduced _configured_event_hooks to public
  configured_event_hooks; its cross-module import added the two
  reportPrivateUsage errors that pushed the rule over its ceiling

* fix(guardrails): keep the response verdict for explicit logging_only_scope=both on input-scan failure

Greptile P1: an explicitly configured both-direction observer asked for a
verdict on each direction, so a failed request scan must not silently drop
the response verdict. The implicit default (logging_only_scope None) keeps
the abort semantics of a logging_only hook whose scan raised, which is the
base behavior the earlier fix restored.

Also drops two test comments that restated their assertions (Greptile P2).

* test(guardrails): align logging_only_scope integration rows with abort-on-input-failure and encrypted params

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep the response scan when the request copy fails for explicit both scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 14:03:33 -07:00
devin-ai-integration[bot]
eb9355d11a
feat(ui): configure Anthropic workload identity federation from the dashboard (#44889)
* feat(ui): configure Anthropic workload identity federation from the dashboard

Add Credential and Edit Credential offer workload identity federation for Anthropic, Add Model creates a federated credential and attaches it, and the credentials table marks federated credentials. Editing a credential now sends only the values the admin changed, and switching provider no longer leaves the previous provider's default base URL on screen

* fix(ui): keep a credential edit to what the admin set in the federation form

* fix(ui): lock the provider in the federation dialog opened from Add Model

* fix(ui): require one federation id when the identity source is the proxy environment

* test(credentials): cover the Anthropic federation dashboard and credential routes

Integration cells for the credential routes every dashboard shape writes (round trips, PATCH set and delete, malformed bodies, non-admin refusals, the token-file allowlist and exchange-host checks, every identity source through chat and messages against a scripted exchange, concurrent writes across two workers and a worker kill mid burst), plus Playwright specs for the Add Credential, Edit Credential and Add Model federation flows and the team-admin view. The owned proxies boot with a 2 s config reload so both workers serve a stored credential inside the fixture budget.

* test(e2e): type the federation spec's captured bodies and clean up the Add Model deployment by its created id

captureRequestBody and postAsMaster take a type parameter instead of returning Record<string, any>, the spec names the credential and model write shapes it captures, and the Add Model cell reads the deployment id from the /model/new response right after the click so a later failing check no longer leaves the deployment behind.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 13:38:58 -07:00
devin-ai-integration[bot]
8caf271e6c
feat(lens): flag traces with global System 1 signals (#45094)
* feat(lens): add global trace signals

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(lens): show global trace signals in Settings and Traces

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): resolve lens signal router at call time

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(lens): signal library, all signal pills and an always-on Signals column

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): tidy signal setup and stop trace ids crowding signal pills

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): address review on signal scans, claims and settings drafts

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): resume signal scan at a partially consumed page

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): keep custom signals separate from library signals

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): give toggled library signals their own row keys

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): restore signal toggle-off checks

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): map signal spans and remove recursive helpers

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 13:27:22 -07:00
devin-ai-integration[bot]
d474b433cf
ci(codeql): run the default suite on full scans and security-extended on PRs (#45149)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 13:24:28 -07:00
ryan-crabbe-berri
a7cf9c6c86
fix(openai-compat): send provider attribution headers on the default SDK path (+ Perplexity) (#44291)
* fix(openai-compat): send provider attribution headers on the default SDK path

Provider-specific headers set in validate_environment are never sent for
OpenAI-compatible providers on the default OpenAI SDK path, which doesn't
call it. Novita's X-Novita-Source has been silently missing as a result.

Add BaseConfig.get_attribution_headers() and merge it into the outbound
headers in _complete_custom_openai, which feeds both the SDK and the
experimental http-handler paths. Caller headers win, case-insensitively.

* feat(perplexity): send X-Pplx-Integration attribution header

Ports the change from #38565 onto the attribution-header hook so it is
sent on the default SDK path too.

Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>

* test: capture attribution headers in-process instead of over a socket

Address review: unit tests now drive litellm.completion into an httpx
transport rather than a local HTTP server; type the header helper as
dict[str, str]; bind the merged headers to a Final instead of
rebinding headers in _complete_custom_openai.

* style: drop explanatory comments flagged by review

---------

Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
2026-10-07 13:21:27 -07:00
devin-ai-integration[bot]
dbdf555da2
test(e2e): cover ollama and ollama_chat on chat completions, responses and messages (#45079)
* test(e2e): cover ollama and ollama_chat on chat completions, responses and messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): declare Subject metadata and check streamed tool call ids in the ollama suite

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 20:17:35 +00:00
ishaan-berri
7921716f39
feat(lens-ui): show findings ranked by priority with frequency and highlighted evidence (#45143)
* feat(lens-ui): compute how often a finding hits sampled traces per day

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): copy a finding for an agent as markdown

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add affected, unaffected and quote highlight color tokens

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add a frequency card with stacked affected traces per day

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(lens-ui): keep the issue brief title out of the page heading outline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): lay out a finding as summary, fix, frequency and highlighted examples

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show findings as a dated list with percent affected beside the open finding

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover frequency and highlighted quotes on a finding

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): follow findings into the split list and example cards

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* style(lens-ui): take finding chart and quote colors from the dashboard theme

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add a shared priority dot and pill for findings

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* style(lens-ui): soften the frequency card and show its date range

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): rank findings under high, medium and low priority headings

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* style(lens-ui): show finding priority, label quotes by content and collapse extra examples

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): prove findings are grouped and ordered by priority

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover finding priority, quote labels and example collapsing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 12:57:04 -07:00
devin-ai-integration[bot]
a9b9700790
perf(lens): bound single trace reads by the sampled start time (#45088)
* perf(lens): bound single trace reads by the sampled start time

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): allow unused query fixture field in load tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): pass start_time in every lens content and evidence test

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 12:50:35 -07:00
devin-ai-integration[bot]
c877e0f055
fix(logging): bound data URI regex so base64 truncation stays linear (#45132)
* fix(logging): bound data URI regex so base64 truncation stays linear

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover whitespace-free data: prefixes in data URI regex regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 12:32:04 -07:00
devin-ai-integration[bot]
8f85de740f
perf(lens): prune ClickHouse partitions when sampling and sample in one pass (#45087)
* perf(lens): prune ClickHouse partitions when sampling and sample in one pass

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): allow unused query fixture field in load tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): shrink the sample page when a response exceeds the read limit

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): qualify request sample window columns

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 12:24:25 -07:00
devin-ai-integration[bot]
65dcf43257
perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens (#45095)
* perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): apply the due_at index concurrently on its own and default legacy rows to due

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): pass repository to claim lifecycle tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): page past unsupported due lenses and declare the full due_at index

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): claim due lenses in a loop instead of recursion

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 12:24:01 -07:00
berriai-litellm-provider-info-sync[bot]
d174c43518
fix(azure): use integer prompt_cache_min_tokens for azure_ai/claude-haiku-5-5 (#45114)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 19:13:32 +00:00
ryan-crabbe-berri
40a9b959a6
test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata (#44965)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause
2026-10-07 12:11:55 -07:00
ANKITA SAHNI
aa3cd70c18
feat(helm): allow custom labels, annotations, command and args on migrationJob (#42242)
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-10-07 19:10:12 +00:00
berriai-litellm-provider-info-sync[bot]
f0031e9a7e
fix(vertex-ai): correct claude-haiku-5-5 thinking and forced tool flags (#45116)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 19:05:10 +00:00
devin-ai-integration[bot]
810a3106e1
test(integration): hold the wire barrier until the test releases it or the wire closes (#45124)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 12:01:12 -07:00
devin-ai-integration[bot]
1c6f714187
fix(realtime): run transcript guardrails on raw-path transcription sessions with a transcription-safe block (#44844)
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(realtime): run transcript guardrails on raw-path transcription sessions with a transcription-safe block

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(realtime): assert transcription session keeps transcribing after a guardrail block

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(realtime): only expect a follow-up transcript when the block keeps the session open

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(realtime): type the transcription block regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(realtime): flag transcription sessions from the route intent and backend events only

A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail

* fix(realtime): flag transcription sessions from provider-transformed session events

* test(realtime): cover transcript guardrail blocks on transcription sessions

---------

Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 11:59:34 -07:00
Krrish Dholakia
5cd1126870
feat(model_prices): add claude-haiku-5-5 model pricing (#45108)
* feat(model_prices): add claude-haiku-5-5 model pricing and capabilities

Source: https://docs.anthropic.com/en/docs/about-claude/models

* fix(model_prices): allow above_100k_tokens tier keys in the cost map schemas

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 18:52:35 +00:00
ryan-crabbe-berri
6352613d1b
test(e2e): let the realtime send step accept input audio buffer frames (#45127) 2026-10-07 18:51:20 +00:00
devin-ai-integration[bot]
8283033c0e
test(integration): turn the model-info refresh off in every integration proxy config (#45104)
The shared integration proxy configs left the five-minute /v1/models refresh
on, so every proxy booted from them sent a GET /v1/models to each
openai-compatible deployment at boot and every 300 s after, including the
per-test deployments other cells register against their own scripted wires.
A scripted provider that asserts on the exact requests it receives then saw a
GET it never scripted, mid-test.

Set disable_model_info_refresh in proxy_config.yaml,
coordination_redis_proxy_config.yaml, and oci_proxy_test_config.yaml, and
drop the three per-test copies of that setting from the observability cells
that each loaded the stock config and set it themselves.

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 11:46:34 -07:00
ryan-crabbe-berri
44bc1656e7
fix(prompt-caching): default injected /v1/messages breakpoints to implicit lookup (#45121) 2026-10-07 18:42:33 +00:00
berriai-litellm-provider-info-sync[bot]
2121322983
fix(cost-map): lower anthropic claude-sonnet-5-5 cache read price (#45113)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 11:39:28 -07:00
devin-ai-integration[bot]
31b90ebb47
test(observability): wait for warm-up spend rows in capped and count only the post-wipe half in X4 (#45006)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 11:29:05 -07:00
devin-ai-integration[bot]
0586289817
test(mcp): fix stale bridge-hook, applied-guardrails and pagination-revoke integration tests (#45008)
* test(mcp): fix stale bridge-hook, applied-guardrails and pagination-revoke integration tests

* test(mcp): clear the spare direct grant when restoring the access-group policy

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 11:28:14 -07:00
ahamedshaik16
2c1847f8a2
fix(prometheus): add model_group label to end-to-end latency metrics (#44860)
litellm_llm_api_latency_metric, litellm_llm_api_time_to_first_token_metric,
litellm_request_total_latency_metric, and litellm_deployment_latency_per_output_token
previously carried requested_model/litellm_model_name/model_id but not model_group,
so pooled-deployment latency couldn't be grouped by model pool on dashboards --
only the proxy-overhead-only metrics (litellm_overhead_latency_metric and friends)
had model_group. All four metrics read enum_values.model_group through the existing
prometheus_label_factory plumbing, so no new parameter threading was needed, just
the label-list addition.

Co-authored-by: ahamedshaik16 <24526479+ahamedshaik16@users.noreply.github.com>
2026-10-07 11:18:45 -07:00
devin-ai-integration[bot]
740d0435a8
test(integration): scope the team-scoped models upstream check to its own model (#45106)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 11:09:56 -07:00
tin-berri
0338498067
fix(router): preserve native baseline identity and accounting (#44960)
* fix(router): preserve native baseline identity and accounting

* fix(router): preserve injected system caches in native baselines

* fix(router): prepare native baselines through shared request owners

* fix(router): capture native baseline fields from the provider schema

* refactor(router): reuse native provider parameter discovery

* fix(router): keep long native baselines and abstain after compaction

Message history no longer counts against the settings snapshot budget, so long
and non-ASCII native sessions keep modeled baselines. Selected-tier compaction
now abstains because the baseline would otherwise inherit the compacted history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(router): judge implicit caching against the selected request

The implicit-cache guard compared the selected response's cache usage with the
projected baseline's breakpoints, so selected-tier cache markers made a usable
unmarked baseline plan look like unexplained caching. The guard now checks the
selected wire request. Also removes a stamp-reuse branch that could never run
because routing clears the stamp first; every pass already captures caller
settings from fresh kwargs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 11:00:31 -07:00
ryan-crabbe-berri
3538e87e45
test(e2e): tag the remaining quota_management tests and record budget and spend client steps (#44966)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag quota_management tests with Subject metadata and record budget client steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause
2026-10-07 10:33:12 -07:00
ryan-crabbe-berri
c943d650f4
test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps (#44964)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps

* test(e2e): leave the batches cleanup harness unit tests untagged

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): name every driven model on the vllm batch, prompt caching and complexity router subjects
2026-10-07 10:32:58 -07:00
devin-ai-integration[bot]
77fc3315e5
fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts (#43878)
* fix(caching): keep tool calls and tool results in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): keep semantic tool prompt helpers within lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep structured function_call_output text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split Responses text-field collection to stay within complexity budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): tag each tool result with the position of the call it answers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): encode tool result position and output together so tool text cannot forge result tags

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): expect encoded tool result record in qdrant semantic prompt parity case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(caching): embed every semantic cache prompt field except volatile ones

Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"

This reverts commit 86c82b949b.

* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"

This reverts commit 39efb5d9da.

* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"

This reverts commit 5aed3ab3de.

* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split semantic cache prompt extraction by API format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): pick the Responses text field without a Final inside a loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): write the semantic cache prompt builders as plain loops

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): drop formatting-only churn from the redis semantic cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(integration): drop the caching group wiring that main already carries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): recurse into tool_result content in the semantic cache prompt helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): check the max_messages cap on the shared exact-cache proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): read list-form function_call_output text in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"

This reverts commit 0665296bf1.

* style(caching): suppress C901 on the Responses input walker instead of splitting it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:32:39 +00:00
ryan-crabbe-berri
79209b92a1
test(e2e): tag guardrails and logging tests with Subject metadata and record client steps (#44963)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps

* test(e2e): leave the guardrails and logging harness unit tests untagged

* test(e2e): let the inner create_model step name the guardrail backend deployment

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): declare the default guardrail backend model on the tests that drive it
2026-10-07 10:32:25 -07:00
ryan-crabbe-berri
d7cdc88c66
test(e2e): tag management tests with Subject metadata and record management client steps (#44962)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag management tests with Subject metadata and record management client steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): keep the prompt out of the chat_status step so polled retries collapse
2026-10-07 10:32:09 -07:00
ryan-crabbe-berri
0ae55bdf7a
test(e2e): tag claude_code cells with Subject metadata and record CLI driver steps (#44961)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag claude_code tests with Subject metadata and record CLI driver steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): decorate run_claude directly so the label gate discovers its step
2026-10-07 10:31:54 -07:00
ryan-crabbe-berri
ade17902a0
test(e2e): tag llm_translation tests with Subject metadata and record harness steps (#44950)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* test(e2e): tag llm_translation tests with Subject metadata and record harness steps

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause

* test(e2e): declare the realtime param tuples Final
2026-10-07 10:30:37 -07:00
devin-ai-integration[bot]
736da28ac1
fix(tests): match the lowercased bind error in the owned-proxy port-race retry (#45097)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 10:30:28 -07:00
berriai-litellm-provider-info-sync[bot]
086bcd2a47
fix(bedrock): take context and output limits from the Bedrock model cards (#45091)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 10:03:04 -07:00
devin-ai-integration[bot]
1e88d9955d
fix(proxy): opt-in Redis hash-tag grouping for v3 rate limiter (#45085)
Add litellm_settings.force_redis_hash_tag_grouping so Redis endpoints that enforce cluster slot rules behind a standalone protocol (Redis Enterprise clustering policy) group multi-key scripts by slot like RedisClusterCache, and return grouped batch values in the caller's key order.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 11:37:43 -05:00
berriai-litellm-provider-info-sync[bot]
aeec8703a8
fix(bedrock): add 2027-03-30 deprecation date to DeepSeek R1 and Qwen3 Coder rows (#45089)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 09:24:21 -07:00
devin-ai-integration[bot]
e2971e0af4
refactor(llms): expose public names for private provider helpers (#45037)
* refactor(litellm): migrate private usage in llms

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve Bedrock batch signature marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): migrate llms private usage symbols

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve UUID Watsonx project IDs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): retarget llms mocks to public names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): limit llms changes to renames and forwarders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): keep GCS mock client patching private Vertex auth methods

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): assert forwarder arguments and type forwarder signatures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 08:49:25 -07:00
berriai-litellm-provider-info-sync[bot]
da1053f068
fix(cost-map): add us data residency multiplier to anthropic claude-opus-5-5 (#45078)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 08:41:03 -07:00
berriai-litellm-provider-info-sync[bot]
e1bdd42726
fix(bedrock): update GovCloud OpenAI prices, add GPT-6 Astra ultrafast tier and Titan Image v2 EOL (#45077)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 08:32:23 -07:00
devin-ai-integration[bot]
2b8e53b5ca
fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls (#45053)
* fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ollama): separate replayed tool calls from text and tighten parser typing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 11:05:50 -04:00
devin-ai-integration[bot]
736ff14f11
fix(health): attribute background health check results to their own deployment (#44982)
* fix(health): attribute background health check results to their own deployment

Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(health): annotate locals with Final and split nested comprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 10:00:50 -04:00
devin-ai-integration[bot]
91e9b1f06b
fix(auth): clear the recent-miss user memo when /user/new creates the user (#45020)
* fix(auth): clear the recent-miss user memo when /user/new creates the user

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(auth): pin the miss memo window and drop the class patch in the new_user regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): pin a first-time SSO user's first message and second sign-in on one worker

* test(integration): cover /user/new clearing the user-miss memo on the creating worker

The fake IdP now signs in the subject a login_hint names, so cells on the shared one-worker proxy can each use a fresh user. New cells: a plain-key miss followed by /user/new is budgeted at once (same on both legs, the auth prefetch loads the row) and the admin-created user's first SSO sign-in inside the window completes (500 at the callback before the fix). The first-sign-in cell moved onto the shared one-worker proxy fixture

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 06:45:41 -07:00
devin-ai-integration[bot]
cb138ba92f
refactor(types): replace Any with proven types in 16 files (#45029)
* refactor(types): replace Any with proven types in 16 files

* fix(types): import TypedDict from typing_extensions for pydantic on 3.10

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 04:31:51 -07:00
devin-ai-integration[bot]
498a3e1a67
test(proxy): make the stagger-offset and linear-dedup guards independent of runner identity and load (#45031)
* test(proxy): inject the cleanup job stagger offset so the runtime sync test no longer depends on the runner's host and pid

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(policy_engine): count resolver calls instead of timing them so the linear dedup guard is deterministic under CI load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(policy_engine): annotate the line-count guard's locals as Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 04:07:36 -07:00
devin-ai-integration[bot]
17a982d6df
ci: give the database-backed CircleCI jobs their own Postgres (#45007)
* ci: give installing_litellm jobs their own Postgres

* ci: give the entrypoint jobs their own Postgres sidecar

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-07 02:43:54 -07:00
devin-ai-integration[bot]
7b439d9138
refactor: replace fresh getattr string access and a restating comment (#45017)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 02:08:22 -07:00
berriai-litellm-provider-info-sync[bot]
2c67ae90bd
fix(azure): date claude-opus-5-5 and sonnet-5-5 and fill azure/eu/gpt-6-astra limits (#45018)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 01:19:38 -07:00