A deferred LiteLLM model whose first use happens on two threads at once could lose its freshly built validator to the second thread's rebuild, so GenericLiteLLMParams.model_validate handed back a CredentialLiteLLMParams and the request failed 400 on use_litellm_proxy. LiteLLMBaseModel.model_rebuild now runs under one process-wide re-entrant lock, so a thread arriving mid-build waits for the finished validator instead of rebuilding over it
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ci): pass NativeCall to transcription bridge fakes in rust bridge tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): assert the broadened e2e harness and basedpyright diff gates from #45172
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust-bridge): pin every NativeCall field in transcription bridge fakes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep prompt_cache_breakpoint markers in the chat to responses bridge
Preserve cache-breakpoint markers through the bridge for supported models and drop them for models without breakpoint support
Co-authored-by: Simon Sorg <simonsorg13@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor base_model when gating bridge cache breakpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): avoid recursive cache-breakpoint stripping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): justify bridge stripping casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep cache breakpoints out of non-bridge converter callers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): read prompt_cache_breakpoint without validating content blocks
The marker read ran every chat content block through
TypeAdapter(dict[str, object]).validate_python, which rejects dict blocks
with non-string keys that chat completion callers passing Python dicts
could previously send; the request then failed with a pydantic
ValidationError on the bridge keep path, the strip path, and the
image/file conversions alike. item is already isinstance-narrowed to a
dict at every read site, so read the marker with dict.get directly and
drop the adapter.
Adds a regression test covering text/image_url/file blocks carrying
non-string keys on both the keep (gpt-5.6) and strip (gpt-4o) paths.
* fix(responses): cast content block before reading prompt_cache_breakpoint
basedpyright flags the raw dict.get read as reportUnknownArgumentType
(+2 against the error budget); cast the isinstance-narrowed block to
dict[str, object] first, matching the strip helpers' cast-ok idiom.
* style(responses): ruff-format the marker-read cast
* fix(responses): hoist one cast-ok content block read for the type gates
A Final assignment inside the conversion loop trips
reportGeneralTypeIssues, and per-site casts trip the LIT006 budget;
read the marker through one cast-narrowed local instead.
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Simon Sorg <simonsorg13@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): add denied_passthrough_routes deny list for custom pass-through endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): harden denied_passthrough_routes against non-admin clears and dot-segment paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for typed denied_passthrough_routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): match trailing-slash deny entries, block null metadata from dropping denies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): close bulk update, encoded ?/# and ordering gaps in passthrough deny list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): check denied pass-through routes against the path the forwarder sends upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type matching_denied_passthrough_route metadata mappings
* fix(proxy): treat a `/` deny entry as denying every pass-through route
Also types the new deny-list tests and drops their get_server_root_path mock in favour of unsetting SERVER_ROOT_PATH.
* test(proxy): type the deny-list tests and drop unrelated test reformatting
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng He <yucheng@berri.ai>
* perf(lens): add a 2 second live signal sweep over recently finished traces
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): run the live and backlog signal sweeps side by side
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): cover the live sweep window and backlog draining
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens): keep known signal results on screen and poll every 2s while runs wait
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): cover signal polling speed and results surviving list changes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens): show Checking instead of Queued and leave clean runs blank
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): expect Checking for runs waiting on signals
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): cover cross-worker cache eviction for user tpm/rpm limit updates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): broadcast user cache eviction when tpm_limit or rpm_limit changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): return tpm_limit and rpm_limit from /v2/user/info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): edit user tpm and rpm limits from the user edit form
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): clear user tpm_limit and rpm_limit when sent as null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover user tpm/rpm limit updates across proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover user rate-limit routes and bulk updates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): reject unsafe rate limit integers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): cover user rate limit seed and saved state
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): type user endpoint test doubles
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert: drop rate limit upper bound
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): validate user rate limits with zod and clear user edit lint warnings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): rename user edit schema and name its input and output types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop redundant event loop yields in rate limit eviction test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stop re-wrapping a bridged stream's MidStreamFallbackError
A chat completion served over the Responses API already gets its mid-stream
error wrapped by the Responses iterator. The chat stream wrapper wrapped it a
second time, and the router unwraps one layer, so the client saw the inner
sentinel (message prefixed litellm.MidStreamFallbackError, type null) instead
of the provider's RateLimitError. The chat wrapper now re-raises an already
wrapped error untouched.
* test(streaming): type the bridged stream regression test's locals
* fix(streaming): rebuild a bridged mid-stream error with the outer wrapper's bookkeeping
* test(integration): cover the bridged stream error typing on every surface
Checked-in audit cells for the Responses bridge: chat completions through the OpenAI SDK and httpx, /v1/messages through httpx and the Anthropic SDK, native /v1/responses, the litellm and Router SDK stream paths, and a chaos file with a mixed burst, a worker SIGKILL and a proxy SIGTERM against an owned two-worker proxy. Every cell scripts the provider through a wire server and asserts the caller's body, the upstream's received requests and the spend row by id.
* test(integration): read bridged chat content through one helper
* test(integration): build the bridged fallback config without mutating the loaded yaml
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(guardrails): extend Akto guardrail to responses, MCP tools, attachments and masking
- post_call now waits for Akto and blocks or masks the reply instead of only logging it
- pre_mcp_call and post_mcp_call check MCP tool arguments and tool results
- attached images, audio and files are sent to Akto's file check
- streamed replies are checked every streaming_sampling_rate chunks
- context_source routes traffic to Akto's endpoint or agentic policies
- tags carry user email, team alias and key alias for attribution
* fix(guardrails): harden Akto MCP detection, keep AGENTIC default, block unmappable masking
- MCP handling trusts the logger's call type, so request body keys can't skip the prompt check
- context_source defaults to AGENTIC, as before
- masking that also hits text we can't write back now blocks
- move tests to tests/unit and rename attachments.py to akto_attachments.py
- regenerate the OpenAPI snapshot and dashboard types
* fix(guardrails): ignore a client-sent response in Akto output checks
- a "response" field sent in the request body is no longer scanned in place of the model's reply
- MCP tool calls in the reply are still checked in that case
- split match or-patterns so CodeQL can follow the bound names
- cover a masked payload that is not JSON and drop unused test imports
* refactor(guardrails): read Akto attachment fields directly instead of pattern captures
* fix(guardrails): check Akto attachments in both messages and input
* fix(guardrails): check every Akto attachment source and keep more prompt text in scope
- check all of a file or image block's sources (file_data, file_url, file_id), since providers pick different ones
- send Anthropic search_result blocks to the file check as text
- keep document title and context, and legacy functions, in the checked request
- take the client IP from the proxy's requester_ip_address before client forwarding headers
- read litellm_params identity only from server-side call details
* fix(guardrails): never drop an Akto attachment the text check removed
- optional metadata (filename, title, format, media type) that isn't a string is ignored instead of failing the block
- an attachment block that still can't be read blocks the request
- search_result text is checked once, as a file, instead of also in the text check
* fix(guardrails): strip only what the Akto file check sends from the text check
- the text check keeps every attachment field except the ones the file check sends
- document title/context and search_result source/title go to the file check as text, since the /v1/messages text check drops them
- accept every image shape LiteLLM forwards (image_url or url, string or object) and check each source
- ignore blocks whose type is not a string instead of failing the file check
* fix(guardrails): keep model-visible text in the Akto text check
- document title/context, text documents and search_result stay in the text check, so no Akto backend skips them
- on /v1/messages the text check reads the messages Anthropic receives, with the guardrail's skip/scan scoping applied
- a client-sent "response" key can only add reply checks, never skip recording or MCP tool-call checks
- a "messages" key on the Responses API can't replace its input in the text check
- the recorded IP comes only from the proxy's requester_ip_address
* fix(guardrails): keep AktoGuardrail positional args backward compatible
* fix(guardrails): keep zero Akto timeouts working and record every stream check
A guardrail_timeout of 0 used to fall back to the default; the new ge=1 made the config invalid, so the proxy dropped the guardrail. Zero settings now fall back to the defaults again.
The end-of-stream check can be skipped when the last sampled check covered the reply, so mid-stream checks now record, like base.
* fix(guardrails): use defaults for non-positive Akto timeouts and sampling rate
A zero or negative guardrail_timeout, file_guardrail_timeout or streaming_sampling_rate used to reach the HTTP call or the stream cadence. They now fall back to the defaults, like unset values.
* feat(lens): carry lens.source attributes on trace span rows
* feat(lens): read lens.source attributes in trace_spans query
* feat(lens): read lens.source attributes in trace_page_spans query
* feat(lens): read lens.source attributes in trace_span_batch query
* feat(lens): read lens.source attributes in trace_list_span_batch query
* feat(lens): decode lens.source columns from clickhouse span rows
* feat(lens): add run source to the trace summary contract
* feat(lens): export RunSource from litellm-traces
* feat(lens): resolve the https run source from the root span
* test(lens): cover run source resolution and https guard
* test(lens): add source fields to capture fixture rows
* test(lens): round-trip source fields on span row contract
* test(lens): add source fields to trace cache test rows
* test(lens): add source fields to trace cache read fixtures
* test(lens): add source fields to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with run source
* chore(lens): regenerate trace json schema with run source
* chore(lens): regenerate trace page json schema with run source
* chore(ui): regenerate api types with trace run source
* feat(lens): add run source link with hover card
* test(lens): cover run source app detection and url guard
* feat(lens): show the run source next to the trace name
* feat(lens): label the run source link, e.g. Slack thread
* test(lens): assert run source link labels
* feat(lens): move the run source link into the trace stats row
* feat(lens): add a typed run source, e.g. slack or teams
* feat(lens): export RunSourceType
* feat(lens): carry the source type on trace span rows
* feat(lens): resolve the run source type, defaulting to custom
* feat(lens): read agent.source attributes in trace_spans query
* feat(lens): read agent.source attributes in trace_page_spans query
* feat(lens): read agent.source attributes in trace_span_batch query
* feat(lens): read agent.source attributes in trace_list_span_batch query
* feat(lens): decode the source type from clickhouse span rows
* test(lens): cover run source type parsing
* test(lens): add source type to capture fixture rows
* test(lens): round-trip source type on span row contract
* test(lens): add source type to trace cache test rows
* test(lens): add source type to trace cache read fixtures
* test(lens): add source type to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with source type
* chore(lens): regenerate trace json schema with source type
* chore(lens): regenerate trace page json schema with source type
* chore(ui): regenerate api types with run source type
* feat(lens): pick the run source logo and label from its type
* test(lens): cover run source labels by type and slack logo
* feat(lens): show the run source as a Source stat with the app name
* test(lens): assert run source app names
* feat(lens): place the Source stat before Duration
Add @step labels to the HttpTransport methods, the poll and wait helpers and the boot helpers that did real IO without recording a step, so a test that reaches the proxy through them no longer reports an empty or gappy step timeline in the JUnit report.
* test(e2e): move the harness self-tests out of tests/e2e
The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead
* fix(ci): point the providers replay controls at tests/e2e_harness
The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension
* style(tests): mark the raw-HTTP check's bindings Final
* refactor(rust): separate Messages settings from capability inputs
* refactor(rust-bridge): unify Messages OCR and Responses call inputs
* refactor(rust-bridge): share NativeCall across inference entrypoints
* fix(rust-bridge): preserve Responses URL aliases and public test inputs
* feat(types): declare above_100k_tokens price fields on ModelInfoBase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 186f6c81a0)
* fix(router): mirror above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 75c9c3ff1e)
* chore(ui): regenerate schema.d.ts for above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 53163d2184)
* refactor(types): mark above_100k ModelInfoBase fields ReadOnly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit febed3199a)
* fix(model_prices): bill claude-haiku-5-5 long prompts on every provider and in batch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 218c00cf5a)
* fix(cost): pass *_above_Nk_tokens_batches rates through get_model_info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 37b6f2fba0)
* fix(model_prices): allow disabling thinking and forced tool use on claude-haiku-5-5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 6e91e0ed36)
* test(model_prices): cite the vendor source for claude-haiku-5-5 capability flags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 1f97bb27b4)
* test(model_prices): type and tidy the claude-haiku-5-5 config tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 3d2526035f)
* feat(bedrock): add claude haiku 5.5 over 100k token tier
Price-Sync: litellm-providers
(cherry picked from commit ee3822c29c)
* feat(model_prices): add openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): dedupe above_100k pricing keys from text merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): bill vertex claude-haiku-5-5 prompts over 100k tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): add adaptive thinking and cache minimum to openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
* test: move the unit half of 126 mixed legacy files into tests/unit
* test: restore litellm globals that moved tests set
* test: finalize migration test cleanup
* test: restore original bodies of moved legacy tests
The move into tests/unit had rewritten 612 test bodies, and some of the rewrites dropped assertions. Each moved test now carries its original body from the legacy file, with only the imports, helpers, fake provider credentials and monkeypatched env it needs to run under tests/unit
test_timeout_streaming goes back to tests/local_testing because it needs the fake OpenAI endpoint server. The image payload fixture moves with its only user, and two tests that leaked global state (a registered model cost entry and queued logging tasks) are now isolated
* test: drop module imports shadowed by restored local imports
* test: assert on LiteLLM output in no-assertion moved tests and isolate leaks
Twenty no-assertion candidates get one assertion on the value LiteLLM returns, with the original lines unchanged. Four tests go back to their legacy files because they only check types or imports, write into the working directory, or cannot assert without a body change
Two moved tests leaked globals into later tests in the same worker, so monkeypatch fixtures now restore the retry-after header parser and the end user cost tracking flags
* test: drain queued logging tasks before the Phoenix span test
The moved Phoenix test counted spans from logging tasks that earlier tests had queued, so the drain fixture moves to tests/unit/conftest.py and both it and the Datadog batch test use it. test_factory_function goes back to its legacy file because its returned wrapper calls the real Assistants API and cannot be asserted on without a body change
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(guardrails): add logging_only_scope to observe one direction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(guardrails): remove unrelated test churn
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): allow native lifecycle logging-only scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate api types for logging_only_scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): remove callbacks when scope validation fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep output-only scans when request copy fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): configure logging_only_scope on guardrails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep guardrails enforcing when logging_only_scope is invalid at load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): avoid inline guardrail test fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): support directional scope in PATCH and provider UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): format guardrail files for frontend lint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): reset unsupported directional scope selections
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): normalize logging-only scope and sanitize warning logs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): normalize scope in shared guardrail field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(api): sync guardrail schema artifacts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate invalid stored logging-only scopes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): add logging_only_scope integration audit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): fix scope test lint issues
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): extend K4 chaos test timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): restore stored guardrail row verbatim on rejected patch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): satisfy collection lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): validate masked params through TypeAdapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate invalid stored params on reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep constructor coercion in tolerant params parser
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): edit logging_only_scope in the custom code guardrail modal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): extract custom code logging-only scope control
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep no-op PUT and logging_only scan-failure semantics stable
Three regressions from the logging_only_scope feature, fixed while keeping
the input/output/both selection working:
1. PUT with byte-identical litellm_params no longer forces a teardown +
re-init of the live callback. reject_invalid_logging_only_scope now
only validates; re-init still happens exactly when params/name change.
Previously a description-only PUT re-appended the callback at the END
of litellm.callbacks, reordering guardrails: with a BLOCK guardrail
created before a MASK one, the mask started winning and blocked
requests started succeeding (400 -> PUT 200 -> 200). Invalid unchanged
scopes are still rejected with 422 without touching the live instance.
2. CustomGuardrail._scan_logged_call no longer swallows input-scan
exceptions per branch: a raising or BLOCKED input scan aborts the
logging_only hook again (one policy call, one verdict) for guardrails
that never selected a logging_only_scope. Explicit input/output/both
selection keeps selecting which scans run.
3. A failed PATCH now rolls the DB row back to exactly the stored raw
litellm_params (a legacy 4-key row stays 4 keys) instead of expanding
it to a full LitellmParams dump; pinned with a test. This matches the
PUT rollback shape and is the faithful rollback.
* refactor(guardrails): type directional-scope provider list as a tuple
* fix(lint): stay within the basedpyright budget
- drop a Final annotation assigned inside the validation loop
(reportGeneralTypeIssues over ceiling by one)
- rename the PR-introduced _configured_event_hooks to public
configured_event_hooks; its cross-module import added the two
reportPrivateUsage errors that pushed the rule over its ceiling
* fix(guardrails): keep the response verdict for explicit logging_only_scope=both on input-scan failure
Greptile P1: an explicitly configured both-direction observer asked for a
verdict on each direction, so a failed request scan must not silently drop
the response verdict. The implicit default (logging_only_scope None) keeps
the abort semantics of a logging_only hook whose scan raised, which is the
base behavior the earlier fix restored.
Also drops two test comments that restated their assertions (Greptile P2).
* test(guardrails): align logging_only_scope integration rows with abort-on-input-failure and encrypted params
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep the response scan when the request copy fails for explicit both scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): configure Anthropic workload identity federation from the dashboard
Add Credential and Edit Credential offer workload identity federation for Anthropic, Add Model creates a federated credential and attaches it, and the credentials table marks federated credentials. Editing a credential now sends only the values the admin changed, and switching provider no longer leaves the previous provider's default base URL on screen
* fix(ui): keep a credential edit to what the admin set in the federation form
* fix(ui): lock the provider in the federation dialog opened from Add Model
* fix(ui): require one federation id when the identity source is the proxy environment
* test(credentials): cover the Anthropic federation dashboard and credential routes
Integration cells for the credential routes every dashboard shape writes (round trips, PATCH set and delete, malformed bodies, non-admin refusals, the token-file allowlist and exchange-host checks, every identity source through chat and messages against a scripted exchange, concurrent writes across two workers and a worker kill mid burst), plus Playwright specs for the Add Credential, Edit Credential and Add Model federation flows and the team-admin view. The owned proxies boot with a 2 s config reload so both workers serve a stored credential inside the fixture budget.
* test(e2e): type the federation spec's captured bodies and clean up the Add Model deployment by its created id
captureRequestBody and postAsMaster take a type parameter instead of returning Record<string, any>, the spec names the credential and model write shapes it captures, and the Add Model cell reads the deployment id from the /model/new response right after the click so a later failing check no longer leaves the deployment behind.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(openai-compat): send provider attribution headers on the default SDK path
Provider-specific headers set in validate_environment are never sent for
OpenAI-compatible providers on the default OpenAI SDK path, which doesn't
call it. Novita's X-Novita-Source has been silently missing as a result.
Add BaseConfig.get_attribution_headers() and merge it into the outbound
headers in _complete_custom_openai, which feeds both the SDK and the
experimental http-handler paths. Caller headers win, case-insensitively.
* feat(perplexity): send X-Pplx-Integration attribution header
Ports the change from #38565 onto the attribution-header hook so it is
sent on the default SDK path too.
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test: capture attribution headers in-process instead of over a socket
Address review: unit tests now drive litellm.completion into an httpx
transport rather than a local HTTP server; type the header helper as
dict[str, str]; bind the merged headers to a Final instead of
rebinding headers in _complete_custom_openai.
* style: drop explanatory comments flagged by review
---------
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test(e2e): cover ollama and ollama_chat on chat completions, responses and messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): declare Subject metadata and check streamed tool call ids in the ollama suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens-ui): compute how often a finding hits sampled traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): copy a finding for an agent as markdown
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add affected, unaffected and quote highlight color tokens
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a frequency card with stacked affected traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens-ui): keep the issue brief title out of the page heading outline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): lay out a finding as summary, fix, frequency and highlighted examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): show findings as a dated list with percent affected beside the open finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover frequency and highlighted quotes on a finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): follow findings into the split list and example cards
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): take finding chart and quote colors from the dashboard theme
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a shared priority dot and pill for findings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): soften the frequency card and show its date range
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): rank findings under high, medium and low priority headings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): show finding priority, label quotes by content and collapse extra examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): prove findings are grouped and ordered by priority
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover finding priority, quote labels and example collapsing
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): bound single trace reads by the sampled start time
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): allow unused query fixture field in load tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass start_time in every lens content and evidence test
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(logging): bound data URI regex so base64 truncation stays linear
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): cover whitespace-free data: prefixes in data URI regex regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): apply the due_at index concurrently on its own and default legacy rows to due
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass repository to claim lifecycle tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): page past unsupported due lenses and declare the full due_at index
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): claim due lenses in a loop instead of recursion
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): run transcript guardrails on raw-path transcription sessions with a transcription-safe block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): assert transcription session keeps transcribing after a guardrail block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): only expect a follow-up transcript when the block keeps the session open
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): type the transcription block regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): flag transcription sessions from the route intent and backend events only
A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail
* fix(realtime): flag transcription sessions from provider-transformed session events
* test(realtime): cover transcript guardrail blocks on transcription sessions
---------
Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The shared integration proxy configs left the five-minute /v1/models refresh
on, so every proxy booted from them sent a GET /v1/models to each
openai-compatible deployment at boot and every 300 s after, including the
per-test deployments other cells register against their own scripted wires.
A scripted provider that asserts on the exact requests it receives then saw a
GET it never scripted, mid-test.
Set disable_model_info_refresh in proxy_config.yaml,
coordination_redis_proxy_config.yaml, and oci_proxy_test_config.yaml, and
drop the three per-test copies of that setting from the observability cells
that each loaded the stock config and set it themselves.
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(mcp): fix stale bridge-hook, applied-guardrails and pagination-revoke integration tests
* test(mcp): clear the spare direct grant when restoring the access-group policy
---------
Co-authored-by: yuneng <yuneng@berri.ai>
litellm_llm_api_latency_metric, litellm_llm_api_time_to_first_token_metric,
litellm_request_total_latency_metric, and litellm_deployment_latency_per_output_token
previously carried requested_model/litellm_model_name/model_id but not model_group,
so pooled-deployment latency couldn't be grouped by model pool on dashboards --
only the proxy-overhead-only metrics (litellm_overhead_latency_metric and friends)
had model_group. All four metrics read enum_values.model_group through the existing
prometheus_label_factory plumbing, so no new parameter threading was needed, just
the label-list addition.
Co-authored-by: ahamedshaik16 <24526479+ahamedshaik16@users.noreply.github.com>
* fix(router): preserve native baseline identity and accounting
* fix(router): preserve injected system caches in native baselines
* fix(router): prepare native baselines through shared request owners
* fix(router): capture native baseline fields from the provider schema
* refactor(router): reuse native provider parameter discovery
* fix(router): keep long native baselines and abstain after compaction
Message history no longer counts against the settings snapshot budget, so long
and non-ASCII native sessions keep modeled baselines. Selected-tier compaction
now abstains because the baseline would otherwise inherit the compacted history.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(router): judge implicit caching against the selected request
The implicit-cache guard compared the selected response's cache usage with the
projected baseline's breakpoints, so selected-tier cache markers made a usable
unmarked baseline plan look like unexplained caching. The guard now checks the
selected wire request. Also removes a stamp-reuse branch that could never run
because routing clears the stamp first; every pass already captures caller
settings from fresh kwargs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag quota_management tests with Subject metadata and record budget client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag router, batches and mcp tests with Subject metadata and record client steps
* test(e2e): leave the batches cleanup harness unit tests untagged
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): name every driven model on the vllm batch, prompt caching and complexity router subjects
* fix(caching): keep tool calls and tool results in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): keep semantic tool prompt helpers within lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep structured function_call_output text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split Responses text-field collection to stay within complexity budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): tag each tool result with the position of the call it answers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): encode tool result position and output together so tool text cannot forge result tags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): expect encoded tool result record in qdrant semantic prompt parity case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(caching): embed every semantic cache prompt field except volatile ones
Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"
This reverts commit 86c82b949b.
* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"
This reverts commit 39efb5d9da.
* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"
This reverts commit 5aed3ab3de.
* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split semantic cache prompt extraction by API format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): pick the Responses text field without a Final inside a loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): write the semantic cache prompt builders as plain loops
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): drop formatting-only churn from the redis semantic cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(integration): drop the caching group wiring that main already carries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): recurse into tool_result content in the semantic cache prompt helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): check the max_messages cap on the shared exact-cache proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): read list-form function_call_output text in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"
This reverts commit 0665296bf1.
* style(caching): suppress C901 on the Responses input walker instead of splitting it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps
* test(e2e): leave the guardrails and logging harness unit tests untagged
* test(e2e): let the inner create_model step name the guardrail backend deployment
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the default guardrail backend model on the tests that drive it
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag management tests with Subject metadata and record management client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): keep the prompt out of the chat_status step so polled retries collapse