* fix(openai-compat): send provider attribution headers on the default SDK path
Provider-specific headers set in validate_environment are never sent for
OpenAI-compatible providers on the default OpenAI SDK path, which doesn't
call it. Novita's X-Novita-Source has been silently missing as a result.
Add BaseConfig.get_attribution_headers() and merge it into the outbound
headers in _complete_custom_openai, which feeds both the SDK and the
experimental http-handler paths. Caller headers win, case-insensitively.
* feat(perplexity): send X-Pplx-Integration attribution header
Ports the change from #38565 onto the attribution-header hook so it is
sent on the default SDK path too.
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test: capture attribution headers in-process instead of over a socket
Address review: unit tests now drive litellm.completion into an httpx
transport rather than a local HTTP server; type the header helper as
dict[str, str]; bind the merged headers to a Final instead of
rebinding headers in _complete_custom_openai.
* style: drop explanatory comments flagged by review
---------
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test(e2e): cover ollama and ollama_chat on chat completions, responses and messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): declare Subject metadata and check streamed tool call ids in the ollama suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens-ui): compute how often a finding hits sampled traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): copy a finding for an agent as markdown
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add affected, unaffected and quote highlight color tokens
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a frequency card with stacked affected traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens-ui): keep the issue brief title out of the page heading outline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): lay out a finding as summary, fix, frequency and highlighted examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): show findings as a dated list with percent affected beside the open finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover frequency and highlighted quotes on a finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): follow findings into the split list and example cards
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): take finding chart and quote colors from the dashboard theme
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a shared priority dot and pill for findings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): soften the frequency card and show its date range
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): rank findings under high, medium and low priority headings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): show finding priority, label quotes by content and collapse extra examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): prove findings are grouped and ordered by priority
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover finding priority, quote labels and example collapsing
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): bound single trace reads by the sampled start time
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): allow unused query fixture field in load tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass start_time in every lens content and evidence test
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(logging): bound data URI regex so base64 truncation stays linear
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): cover whitespace-free data: prefixes in data URI regex regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): apply the due_at index concurrently on its own and default legacy rows to due
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass repository to claim lifecycle tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): page past unsupported due lenses and declare the full due_at index
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): claim due lenses in a loop instead of recursion
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): run transcript guardrails on raw-path transcription sessions with a transcription-safe block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): assert transcription session keeps transcribing after a guardrail block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): only expect a follow-up transcript when the block keeps the session open
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): type the transcription block regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): flag transcription sessions from the route intent and backend events only
A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail
* fix(realtime): flag transcription sessions from provider-transformed session events
* test(realtime): cover transcript guardrail blocks on transcription sessions
---------
Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The shared integration proxy configs left the five-minute /v1/models refresh
on, so every proxy booted from them sent a GET /v1/models to each
openai-compatible deployment at boot and every 300 s after, including the
per-test deployments other cells register against their own scripted wires.
A scripted provider that asserts on the exact requests it receives then saw a
GET it never scripted, mid-test.
Set disable_model_info_refresh in proxy_config.yaml,
coordination_redis_proxy_config.yaml, and oci_proxy_test_config.yaml, and
drop the three per-test copies of that setting from the observability cells
that each loaded the stock config and set it themselves.
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(mcp): fix stale bridge-hook, applied-guardrails and pagination-revoke integration tests
* test(mcp): clear the spare direct grant when restoring the access-group policy
---------
Co-authored-by: yuneng <yuneng@berri.ai>
litellm_llm_api_latency_metric, litellm_llm_api_time_to_first_token_metric,
litellm_request_total_latency_metric, and litellm_deployment_latency_per_output_token
previously carried requested_model/litellm_model_name/model_id but not model_group,
so pooled-deployment latency couldn't be grouped by model pool on dashboards --
only the proxy-overhead-only metrics (litellm_overhead_latency_metric and friends)
had model_group. All four metrics read enum_values.model_group through the existing
prometheus_label_factory plumbing, so no new parameter threading was needed, just
the label-list addition.
Co-authored-by: ahamedshaik16 <24526479+ahamedshaik16@users.noreply.github.com>
* fix(router): preserve native baseline identity and accounting
* fix(router): preserve injected system caches in native baselines
* fix(router): prepare native baselines through shared request owners
* fix(router): capture native baseline fields from the provider schema
* refactor(router): reuse native provider parameter discovery
* fix(router): keep long native baselines and abstain after compaction
Message history no longer counts against the settings snapshot budget, so long
and non-ASCII native sessions keep modeled baselines. Selected-tier compaction
now abstains because the baseline would otherwise inherit the compacted history.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(router): judge implicit caching against the selected request
The implicit-cache guard compared the selected response's cache usage with the
projected baseline's breakpoints, so selected-tier cache markers made a usable
unmarked baseline plan look like unexplained caching. The guard now checks the
selected wire request. Also removes a stamp-reuse branch that could never run
because routing clears the stamp first; every pass already captures caller
settings from fresh kwargs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag quota_management tests with Subject metadata and record budget client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps
* test(e2e): leave the batches cleanup harness unit tests untagged
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): name every driven model on the vllm batch, prompt caching and complexity router subjects
* fix(caching): keep tool calls and tool results in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): keep semantic tool prompt helpers within lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep structured function_call_output text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split Responses text-field collection to stay within complexity budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): tag each tool result with the position of the call it answers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): encode tool result position and output together so tool text cannot forge result tags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): expect encoded tool result record in qdrant semantic prompt parity case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(caching): embed every semantic cache prompt field except volatile ones
Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"
This reverts commit 86c82b949b.
* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"
This reverts commit 39efb5d9da.
* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"
This reverts commit 5aed3ab3de.
* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split semantic cache prompt extraction by API format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): pick the Responses text field without a Final inside a loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): write the semantic cache prompt builders as plain loops
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): drop formatting-only churn from the redis semantic cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(integration): drop the caching group wiring that main already carries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): recurse into tool_result content in the semantic cache prompt helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): check the max_messages cap on the shared exact-cache proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): read list-form function_call_output text in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"
This reverts commit 0665296bf1.
* style(caching): suppress C901 on the Responses input walker instead of splitting it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps
* test(e2e): leave the guardrails and logging harness unit tests untagged
* test(e2e): let the inner create_model step name the guardrail backend deployment
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the default guardrail backend model on the tests that drive it
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag management tests with Subject metadata and record management client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): keep the prompt out of the chat_status step so polled retries collapse
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag claude_code tests with Subject metadata and record CLI driver steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): decorate run_claude directly so the label gate discovers its step
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag llm_translation tests with Subject metadata and record harness steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the realtime param tuples Final
Add litellm_settings.force_redis_hash_tag_grouping so Redis endpoints that enforce cluster slot rules behind a standalone protocol (Redis Enterprise clustering policy) group multi-key scripts by slot like RedisClusterCache, and return grouped batch values in the caller's key order.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ollama): separate replayed tool calls from text and tighten parser typing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(health): attribute background health check results to their own deployment
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(health): annotate locals with Final and split nested comprehension
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): clear the recent-miss user memo when /user/new creates the user
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(auth): pin the miss memo window and drop the class patch in the new_user regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): pin a first-time SSO user's first message and second sign-in on one worker
* test(integration): cover /user/new clearing the user-miss memo on the creating worker
The fake IdP now signs in the subject a login_hint names, so cells on the shared one-worker proxy can each use a fresh user. New cells: a plain-key miss followed by /user/new is budgeted at once (same on both legs, the auth prefetch loads the row) and the admin-created user's first SSO sign-in inside the window completes (500 at the callback before the fix). The first-sign-in cell moved onto the shared one-worker proxy fixture
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(types): replace Any with proven types in 16 files
* fix(types): import TypedDict from typing_extensions for pydantic on 3.10
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject the cleanup job stagger offset so the runtime sync test no longer depends on the runner's host and pid
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(policy_engine): count resolver calls instead of timing them so the linear dedup guard is deterministic under CI load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(policy_engine): annotate the line-count guard's locals as Final
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: give installing_litellm jobs their own Postgres
* ci: give the entrypoint jobs their own Postgres sidecar
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>