* feat(decisions): add the OpenAI Decisions spec types and the System One translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): share one DecisionsModel config and require model and usage on responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): rename the shared pydantic parent to DecisionsObjectBase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): match the SDK on strict choice values and drop the invented min_length bounds
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): compose DecisionsRequest from model and body and read the translator top down
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): use the plural Decisions prefix only for the request and response envelopes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): import assert_never from typing_extensions for Python 3.10
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vertex_ai): dial the multi-region Live API host for realtime sessions
A realtime deployment with vertex_location us or eu dialed
{location}-aiplatform.googleapis.com, which Vertex does not serve for a
multi-region, so the WebSocket handshake came back 404. The URL builder now
resolves its host through the shared Vertex host resolver, which already
maps us and eu to aiplatform.{geo}.rep.googleapis.com and leaves regions
and global unchanged.
* test(vertex_ai): annotate the realtime multi-region URL tests
* fix(vertex_ai): skip the TLS argument when the realtime api_base is plain ws://
The Vertex realtime session and health-check connects passed the shared
TLS context to every websocket dial, so an api_base the transformation
already maps from http:// to ws:// failed with 'ssl argument is
incompatible with a ws:// URI' before any frame was sent. The OpenAI
realtime handler already skips the argument for ws://; the shared helper
now does the same for the Vertex path
* test(integration): cover the Vertex realtime multi-region host through the proxy
Fourteen cells on the scripted Gemini Live upstream: api_base overrides for
us, eu, us-central1 and global complete a turn; malformed locations are
refused before any dial; a missing location defaults to us-central1; the
realtime health check handshakes the override host and reports a malformed
location without dialing; a refused handshake reaches the client and the
next session connects; twelve concurrent sessions across locations each
reach the upstream once; an upstream outage closes every open session and
the proxy recovers
* test(realtime): type the health-check helpers and flatten the burst order
* test(http_handler): type the realtime TLS test signatures
* fix(scheduler): remove a request's queue entry once it stops waiting
Requests admitted while a healthy deployment existed never left the priority
queue, and neither did requests cancelled while waiting. The stale entries
blocked later requests during cooldown and, with Redis, made add_request
raise a TypeError on queues read back as JSON lists
Both scheduling paths now share one polling helper that removes the entry in
a finally block, whether the request was admitted, timed out or cancelled
Related to #43059
* test(router): allowlist _wait_for_scheduler_turn in the router coverage check
The coverage script only counts direct calls in test files. The helper is
exercised through prioritized acompletion and atext_completion in
test_router.py, like the other allowlisted entries
* fix(scheduler): admit healthy requests before reading the queue, and clean up cancelled enqueues
poll() raised on an empty queue before checking for healthy deployments. With
the cleanup now rewriting the queue after every admission, a concurrent write
from another replica can erase a waiting request's entry, and that request
then failed while a deployment was healthy. poll() now admits as soon as a
deployment is healthy and only reads the queue during cooldown
add_request also moved inside the try block, so a request cancelled while its
queue write is in flight still has its entry removed
* refactor(scheduler): move the wait loop into Scheduler.wait_for_turn
The router passes a healthy-deployments callable into the scheduler, so tests
inject a Scheduler directly instead of replacing the router's scheduler
attribute. The cancelled-mid-enqueue test moves to test_scheduler.py with the
other scheduler tests, and poll() takes the deployments as a Sequence since it
only checks whether any are healthy
* test(scheduler): move scheduler tests into tests/unit
* fix(scheduler): finish the queue removal when cancellation is delivered again during cleanup
A second cancel, or an anyio cancel scope that re-cancels on every await,
interrupted remove_request mid-write and left the entry in Redis.
* fix(scheduler): read queue entries back from redis as tuples
* fix(scheduler): admit a request whose queue entry vanished during cooldown
* fix(scheduler): re-enqueue a request whose queue entry vanished during cooldown
* test(integration): cover priority scheduler queue cleanup across instances
Adds tests/integration/routing/test_priority_scheduler_queue_cleanup.py: 27 cells
against the real proxy (two workers, Postgres, Redis, a scripted upstream) for every
prioritized surface (/v1/chat/completions, /v1/completions, /queue/chat/completions,
streaming and not, OpenAI SDK sync and async, raw httpx), the non-integer priority
pass-through, the in-memory queue on one proxy, two proxies sharing a Redis queue
(served requests leave no entry, a dead replica's entry is skipped or expires, a
waiter that times out or disconnects removes only itself), a Redis outage mid burst
and a SIGKILLed worker. Each cell asserts the caller's response, the upstream's
requests by marker and the Redis queue contents. On the merge base the cross-instance
cells fail with the list-of-lists TypeError and the in-memory cell with 408s behind
the leaked entry; on this branch every cell passes
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(rust-bridge): share field and response marshaling
* test(rust-bridge): isolate response factory test modules per case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust-bridge): document shared marshaling helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* test: delete 77 legacy tests covered by e2e, unable to fail, or dead in CI
* test: keep router helper tests and coverage ignore list unchanged
* test: keep cohere error handling tests
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(lens): add trace_agents rollup query
* feat(lens): register the trace_agents read query
* feat(lens): add trace_agents params and row types
* feat(lens): dispatch the trace_agents query
* feat(lens): export trace_agents wire schemas
* test(lens): pin the trace_agents query name
* test(lens): cover trace_agents scope, window and failure counts
* feat(lens): cap the agent list size
* feat(lens): declare the trace_agents bridge query
* feat(lens): read trace agents from clickhouse storage
* feat(lens): add trace agent response models
* feat(lens): list agents within the reader's trace scope
* feat(lens): add GET /v1/traces/agents
* chore(lens): regenerate trace models with trace_agents
* chore(lens): regenerate trace types with trace_agents
* chore(lens): regenerate read query name schema
* chore(lens): add trace agent row schema
* chore(lens): add trace agents params schema
* test(lens): cover the trace agents route
* test(lens): cover agent listing scope and timestamps
* chore(ui): regenerate api types with trace agents route
* feat(lens): derive trace agent types from the schema
* feat(lens): fetch the agent list from the traces api
* feat(lens): roll up demo runs into agents
* feat(lens): serve the agent list in demo data
* feat(lens): load agents seen in the last two weeks
* feat(lens): remember the selected agent per browser
* feat(lens): add the agent picker
* feat(lens): wire agent selection into the lens header
* feat(lens): show the agent picker next to the lens title
* refactor(lens): drop the toolbar agent filter in favor of the header picker
* test(lens): remove tests for the toolbar agent filter
* test(lens): cover agent resolution and rollup
* test(lens): cover scoping, switching and remembering the agent
* test(lens): read only the create request in guided setup
* test(lens): assert the toolbar agent filter is gone
* test(lens): read stubbed requests without their abort signal
A deferred LiteLLM model whose first use happens on two threads at once could lose its freshly built validator to the second thread's rebuild, so GenericLiteLLMParams.model_validate handed back a CredentialLiteLLMParams and the request failed 400 on use_litellm_proxy. LiteLLMBaseModel.model_rebuild now runs under one process-wide re-entrant lock, so a thread arriving mid-build waits for the finished validator instead of rebuilding over it
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ci): pass NativeCall to transcription bridge fakes in rust bridge tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): assert the broadened e2e harness and basedpyright diff gates from #45172
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust-bridge): pin every NativeCall field in transcription bridge fakes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep prompt_cache_breakpoint markers in the chat to responses bridge
Preserve cache-breakpoint markers through the bridge for supported models and drop them for models without breakpoint support
Co-authored-by: Simon Sorg <simonsorg13@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor base_model when gating bridge cache breakpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): avoid recursive cache-breakpoint stripping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): justify bridge stripping casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep cache breakpoints out of non-bridge converter callers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): read prompt_cache_breakpoint without validating content blocks
The marker read ran every chat content block through
TypeAdapter(dict[str, object]).validate_python, which rejects dict blocks
with non-string keys that chat completion callers passing Python dicts
could previously send; the request then failed with a pydantic
ValidationError on the bridge keep path, the strip path, and the
image/file conversions alike. item is already isinstance-narrowed to a
dict at every read site, so read the marker with dict.get directly and
drop the adapter.
Adds a regression test covering text/image_url/file blocks carrying
non-string keys on both the keep (gpt-5.6) and strip (gpt-4o) paths.
* fix(responses): cast content block before reading prompt_cache_breakpoint
basedpyright flags the raw dict.get read as reportUnknownArgumentType
(+2 against the error budget); cast the isinstance-narrowed block to
dict[str, object] first, matching the strip helpers' cast-ok idiom.
* style(responses): ruff-format the marker-read cast
* fix(responses): hoist one cast-ok content block read for the type gates
A Final assignment inside the conversion loop trips
reportGeneralTypeIssues, and per-site casts trip the LIT006 budget;
read the marker through one cast-narrowed local instead.
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Simon Sorg <simonsorg13@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): add denied_passthrough_routes deny list for custom pass-through endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): harden denied_passthrough_routes against non-admin clears and dot-segment paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for typed denied_passthrough_routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): match trailing-slash deny entries, block null metadata from dropping denies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): close bulk update, encoded ?/# and ordering gaps in passthrough deny list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): check denied pass-through routes against the path the forwarder sends upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type matching_denied_passthrough_route metadata mappings
* fix(proxy): treat a `/` deny entry as denying every pass-through route
Also types the new deny-list tests and drops their get_server_root_path mock in favour of unsetting SERVER_ROOT_PATH.
* test(proxy): type the deny-list tests and drop unrelated test reformatting
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng He <yucheng@berri.ai>
* perf(lens): add a 2 second live signal sweep over recently finished traces
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): run the live and backlog signal sweeps side by side
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): cover the live sweep window and backlog draining
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens): keep known signal results on screen and poll every 2s while runs wait
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): cover signal polling speed and results surviving list changes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens): show Checking instead of Queued and leave clean runs blank
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): expect Checking for runs waiting on signals
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): cover cross-worker cache eviction for user tpm/rpm limit updates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): broadcast user cache eviction when tpm_limit or rpm_limit changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): return tpm_limit and rpm_limit from /v2/user/info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): edit user tpm and rpm limits from the user edit form
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): clear user tpm_limit and rpm_limit when sent as null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover user tpm/rpm limit updates across proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover user rate-limit routes and bulk updates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): reject unsafe rate limit integers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): cover user rate limit seed and saved state
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): type user endpoint test doubles
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert: drop rate limit upper bound
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): validate user rate limits with zod and clear user edit lint warnings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): rename user edit schema and name its input and output types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop redundant event loop yields in rate limit eviction test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stop re-wrapping a bridged stream's MidStreamFallbackError
A chat completion served over the Responses API already gets its mid-stream
error wrapped by the Responses iterator. The chat stream wrapper wrapped it a
second time, and the router unwraps one layer, so the client saw the inner
sentinel (message prefixed litellm.MidStreamFallbackError, type null) instead
of the provider's RateLimitError. The chat wrapper now re-raises an already
wrapped error untouched.
* test(streaming): type the bridged stream regression test's locals
* fix(streaming): rebuild a bridged mid-stream error with the outer wrapper's bookkeeping
* test(integration): cover the bridged stream error typing on every surface
Checked-in audit cells for the Responses bridge: chat completions through the OpenAI SDK and httpx, /v1/messages through httpx and the Anthropic SDK, native /v1/responses, the litellm and Router SDK stream paths, and a chaos file with a mixed burst, a worker SIGKILL and a proxy SIGTERM against an owned two-worker proxy. Every cell scripts the provider through a wire server and asserts the caller's body, the upstream's received requests and the spend row by id.
* test(integration): read bridged chat content through one helper
* test(integration): build the bridged fallback config without mutating the loaded yaml
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(guardrails): extend Akto guardrail to responses, MCP tools, attachments and masking
- post_call now waits for Akto and blocks or masks the reply instead of only logging it
- pre_mcp_call and post_mcp_call check MCP tool arguments and tool results
- attached images, audio and files are sent to Akto's file check
- streamed replies are checked every streaming_sampling_rate chunks
- context_source routes traffic to Akto's endpoint or agentic policies
- tags carry user email, team alias and key alias for attribution
* fix(guardrails): harden Akto MCP detection, keep AGENTIC default, block unmappable masking
- MCP handling trusts the logger's call type, so request body keys can't skip the prompt check
- context_source defaults to AGENTIC, as before
- masking that also hits text we can't write back now blocks
- move tests to tests/unit and rename attachments.py to akto_attachments.py
- regenerate the OpenAPI snapshot and dashboard types
* fix(guardrails): ignore a client-sent response in Akto output checks
- a "response" field sent in the request body is no longer scanned in place of the model's reply
- MCP tool calls in the reply are still checked in that case
- split match or-patterns so CodeQL can follow the bound names
- cover a masked payload that is not JSON and drop unused test imports
* refactor(guardrails): read Akto attachment fields directly instead of pattern captures
* fix(guardrails): check Akto attachments in both messages and input
* fix(guardrails): check every Akto attachment source and keep more prompt text in scope
- check all of a file or image block's sources (file_data, file_url, file_id), since providers pick different ones
- send Anthropic search_result blocks to the file check as text
- keep document title and context, and legacy functions, in the checked request
- take the client IP from the proxy's requester_ip_address before client forwarding headers
- read litellm_params identity only from server-side call details
* fix(guardrails): never drop an Akto attachment the text check removed
- optional metadata (filename, title, format, media type) that isn't a string is ignored instead of failing the block
- an attachment block that still can't be read blocks the request
- search_result text is checked once, as a file, instead of also in the text check
* fix(guardrails): strip only what the Akto file check sends from the text check
- the text check keeps every attachment field except the ones the file check sends
- document title/context and search_result source/title go to the file check as text, since the /v1/messages text check drops them
- accept every image shape LiteLLM forwards (image_url or url, string or object) and check each source
- ignore blocks whose type is not a string instead of failing the file check
* fix(guardrails): keep model-visible text in the Akto text check
- document title/context, text documents and search_result stay in the text check, so no Akto backend skips them
- on /v1/messages the text check reads the messages Anthropic receives, with the guardrail's skip/scan scoping applied
- a client-sent "response" key can only add reply checks, never skip recording or MCP tool-call checks
- a "messages" key on the Responses API can't replace its input in the text check
- the recorded IP comes only from the proxy's requester_ip_address
* fix(guardrails): keep AktoGuardrail positional args backward compatible
* fix(guardrails): keep zero Akto timeouts working and record every stream check
A guardrail_timeout of 0 used to fall back to the default; the new ge=1 made the config invalid, so the proxy dropped the guardrail. Zero settings now fall back to the defaults again.
The end-of-stream check can be skipped when the last sampled check covered the reply, so mid-stream checks now record, like base.
* fix(guardrails): use defaults for non-positive Akto timeouts and sampling rate
A zero or negative guardrail_timeout, file_guardrail_timeout or streaming_sampling_rate used to reach the HTTP call or the stream cadence. They now fall back to the defaults, like unset values.
* feat(lens): carry lens.source attributes on trace span rows
* feat(lens): read lens.source attributes in trace_spans query
* feat(lens): read lens.source attributes in trace_page_spans query
* feat(lens): read lens.source attributes in trace_span_batch query
* feat(lens): read lens.source attributes in trace_list_span_batch query
* feat(lens): decode lens.source columns from clickhouse span rows
* feat(lens): add run source to the trace summary contract
* feat(lens): export RunSource from litellm-traces
* feat(lens): resolve the https run source from the root span
* test(lens): cover run source resolution and https guard
* test(lens): add source fields to capture fixture rows
* test(lens): round-trip source fields on span row contract
* test(lens): add source fields to trace cache test rows
* test(lens): add source fields to trace cache read fixtures
* test(lens): add source fields to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with run source
* chore(lens): regenerate trace json schema with run source
* chore(lens): regenerate trace page json schema with run source
* chore(ui): regenerate api types with trace run source
* feat(lens): add run source link with hover card
* test(lens): cover run source app detection and url guard
* feat(lens): show the run source next to the trace name
* feat(lens): label the run source link, e.g. Slack thread
* test(lens): assert run source link labels
* feat(lens): move the run source link into the trace stats row
* feat(lens): add a typed run source, e.g. slack or teams
* feat(lens): export RunSourceType
* feat(lens): carry the source type on trace span rows
* feat(lens): resolve the run source type, defaulting to custom
* feat(lens): read agent.source attributes in trace_spans query
* feat(lens): read agent.source attributes in trace_page_spans query
* feat(lens): read agent.source attributes in trace_span_batch query
* feat(lens): read agent.source attributes in trace_list_span_batch query
* feat(lens): decode the source type from clickhouse span rows
* test(lens): cover run source type parsing
* test(lens): add source type to capture fixture rows
* test(lens): round-trip source type on span row contract
* test(lens): add source type to trace cache test rows
* test(lens): add source type to trace cache read fixtures
* test(lens): add source type to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with source type
* chore(lens): regenerate trace json schema with source type
* chore(lens): regenerate trace page json schema with source type
* chore(ui): regenerate api types with run source type
* feat(lens): pick the run source logo and label from its type
* test(lens): cover run source labels by type and slack logo
* feat(lens): show the run source as a Source stat with the app name
* test(lens): assert run source app names
* feat(lens): place the Source stat before Duration
Add @step labels to the HttpTransport methods, the poll and wait helpers and the boot helpers that did real IO without recording a step, so a test that reaches the proxy through them no longer reports an empty or gappy step timeline in the JUnit report.
* test(e2e): move the harness self-tests out of tests/e2e
The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead
* fix(ci): point the providers replay controls at tests/e2e_harness
The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension
* style(tests): mark the raw-HTTP check's bindings Final
* refactor(rust): separate Messages settings from capability inputs
* refactor(rust-bridge): unify Messages OCR and Responses call inputs
* refactor(rust-bridge): share NativeCall across inference entrypoints
* fix(rust-bridge): preserve Responses URL aliases and public test inputs
* feat(types): declare above_100k_tokens price fields on ModelInfoBase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 186f6c81a0)
* fix(router): mirror above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 75c9c3ff1e)
* chore(ui): regenerate schema.d.ts for above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 53163d2184)
* refactor(types): mark above_100k ModelInfoBase fields ReadOnly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit febed3199a)
* fix(model_prices): bill claude-haiku-5-5 long prompts on every provider and in batch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 218c00cf5a)
* fix(cost): pass *_above_Nk_tokens_batches rates through get_model_info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 37b6f2fba0)
* fix(model_prices): allow disabling thinking and forced tool use on claude-haiku-5-5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 6e91e0ed36)
* test(model_prices): cite the vendor source for claude-haiku-5-5 capability flags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 1f97bb27b4)
* test(model_prices): type and tidy the claude-haiku-5-5 config tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 3d2526035f)
* feat(bedrock): add claude haiku 5.5 over 100k token tier
Price-Sync: litellm-providers
(cherry picked from commit ee3822c29c)
* feat(model_prices): add openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): dedupe above_100k pricing keys from text merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): bill vertex claude-haiku-5-5 prompts over 100k tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): add adaptive thinking and cache minimum to openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
* test: move the unit half of 126 mixed legacy files into tests/unit
* test: restore litellm globals that moved tests set
* test: finalize migration test cleanup
* test: restore original bodies of moved legacy tests
The move into tests/unit had rewritten 612 test bodies, and some of the rewrites dropped assertions. Each moved test now carries its original body from the legacy file, with only the imports, helpers, fake provider credentials and monkeypatched env it needs to run under tests/unit
test_timeout_streaming goes back to tests/local_testing because it needs the fake OpenAI endpoint server. The image payload fixture moves with its only user, and two tests that leaked global state (a registered model cost entry and queued logging tasks) are now isolated
* test: drop module imports shadowed by restored local imports
* test: assert on LiteLLM output in no-assertion moved tests and isolate leaks
Twenty no-assertion candidates get one assertion on the value LiteLLM returns, with the original lines unchanged. Four tests go back to their legacy files because they only check types or imports, write into the working directory, or cannot assert without a body change
Two moved tests leaked globals into later tests in the same worker, so monkeypatch fixtures now restore the retry-after header parser and the end user cost tracking flags
* test: drain queued logging tasks before the Phoenix span test
The moved Phoenix test counted spans from logging tasks that earlier tests had queued, so the drain fixture moves to tests/unit/conftest.py and both it and the Datadog batch test use it. test_factory_function goes back to its legacy file because its returned wrapper calls the real Assistants API and cannot be asserted on without a body change
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(guardrails): add logging_only_scope to observe one direction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(guardrails): remove unrelated test churn
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): allow native lifecycle logging-only scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate api types for logging_only_scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): remove callbacks when scope validation fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep output-only scans when request copy fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): configure logging_only_scope on guardrails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep guardrails enforcing when logging_only_scope is invalid at load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): avoid inline guardrail test fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): support directional scope in PATCH and provider UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): format guardrail files for frontend lint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): reset unsupported directional scope selections
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): normalize logging-only scope and sanitize warning logs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): normalize scope in shared guardrail field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(api): sync guardrail schema artifacts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate invalid stored logging-only scopes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): add logging_only_scope integration audit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): fix scope test lint issues
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): extend K4 chaos test timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): restore stored guardrail row verbatim on rejected patch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): satisfy collection lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): validate masked params through TypeAdapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): tolerate invalid stored params on reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep constructor coercion in tolerant params parser
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): edit logging_only_scope in the custom code guardrail modal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): extract custom code logging-only scope control
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep no-op PUT and logging_only scan-failure semantics stable
Three regressions from the logging_only_scope feature, fixed while keeping
the input/output/both selection working:
1. PUT with byte-identical litellm_params no longer forces a teardown +
re-init of the live callback. reject_invalid_logging_only_scope now
only validates; re-init still happens exactly when params/name change.
Previously a description-only PUT re-appended the callback at the END
of litellm.callbacks, reordering guardrails: with a BLOCK guardrail
created before a MASK one, the mask started winning and blocked
requests started succeeding (400 -> PUT 200 -> 200). Invalid unchanged
scopes are still rejected with 422 without touching the live instance.
2. CustomGuardrail._scan_logged_call no longer swallows input-scan
exceptions per branch: a raising or BLOCKED input scan aborts the
logging_only hook again (one policy call, one verdict) for guardrails
that never selected a logging_only_scope. Explicit input/output/both
selection keeps selecting which scans run.
3. A failed PATCH now rolls the DB row back to exactly the stored raw
litellm_params (a legacy 4-key row stays 4 keys) instead of expanding
it to a full LitellmParams dump; pinned with a test. This matches the
PUT rollback shape and is the faithful rollback.
* refactor(guardrails): type directional-scope provider list as a tuple
* fix(lint): stay within the basedpyright budget
- drop a Final annotation assigned inside the validation loop
(reportGeneralTypeIssues over ceiling by one)
- rename the PR-introduced _configured_event_hooks to public
configured_event_hooks; its cross-module import added the two
reportPrivateUsage errors that pushed the rule over its ceiling
* fix(guardrails): keep the response verdict for explicit logging_only_scope=both on input-scan failure
Greptile P1: an explicitly configured both-direction observer asked for a
verdict on each direction, so a failed request scan must not silently drop
the response verdict. The implicit default (logging_only_scope None) keeps
the abort semantics of a logging_only hook whose scan raised, which is the
base behavior the earlier fix restored.
Also drops two test comments that restated their assertions (Greptile P2).
* test(guardrails): align logging_only_scope integration rows with abort-on-input-failure and encrypted params
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep the response scan when the request copy fails for explicit both scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): configure Anthropic workload identity federation from the dashboard
Add Credential and Edit Credential offer workload identity federation for Anthropic, Add Model creates a federated credential and attaches it, and the credentials table marks federated credentials. Editing a credential now sends only the values the admin changed, and switching provider no longer leaves the previous provider's default base URL on screen
* fix(ui): keep a credential edit to what the admin set in the federation form
* fix(ui): lock the provider in the federation dialog opened from Add Model
* fix(ui): require one federation id when the identity source is the proxy environment
* test(credentials): cover the Anthropic federation dashboard and credential routes
Integration cells for the credential routes every dashboard shape writes (round trips, PATCH set and delete, malformed bodies, non-admin refusals, the token-file allowlist and exchange-host checks, every identity source through chat and messages against a scripted exchange, concurrent writes across two workers and a worker kill mid burst), plus Playwright specs for the Add Credential, Edit Credential and Add Model federation flows and the team-admin view. The owned proxies boot with a 2 s config reload so both workers serve a stored credential inside the fixture budget.
* test(e2e): type the federation spec's captured bodies and clean up the Add Model deployment by its created id
captureRequestBody and postAsMaster take a type parameter instead of returning Record<string, any>, the spec names the credential and model write shapes it captures, and the Add Model cell reads the deployment id from the /model/new response right after the click so a later failing check no longer leaves the deployment behind.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(openai-compat): send provider attribution headers on the default SDK path
Provider-specific headers set in validate_environment are never sent for
OpenAI-compatible providers on the default OpenAI SDK path, which doesn't
call it. Novita's X-Novita-Source has been silently missing as a result.
Add BaseConfig.get_attribution_headers() and merge it into the outbound
headers in _complete_custom_openai, which feeds both the SDK and the
experimental http-handler paths. Caller headers win, case-insensitively.
* feat(perplexity): send X-Pplx-Integration attribution header
Ports the change from #38565 onto the attribution-header hook so it is
sent on the default SDK path too.
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test: capture attribution headers in-process instead of over a socket
Address review: unit tests now drive litellm.completion into an httpx
transport rather than a local HTTP server; type the header helper as
dict[str, str]; bind the merged headers to a Final instead of
rebinding headers in _complete_custom_openai.
* style: drop explanatory comments flagged by review
---------
Co-authored-by: Saleh Alghusson <1331721+qirh@users.noreply.github.com>
* test(e2e): cover ollama and ollama_chat on chat completions, responses and messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): declare Subject metadata and check streamed tool call ids in the ollama suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens-ui): compute how often a finding hits sampled traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): copy a finding for an agent as markdown
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add affected, unaffected and quote highlight color tokens
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a frequency card with stacked affected traces per day
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(lens-ui): keep the issue brief title out of the page heading outline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): lay out a finding as summary, fix, frequency and highlighted examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): show findings as a dated list with percent affected beside the open finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover frequency and highlighted quotes on a finding
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): follow findings into the split list and example cards
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): take finding chart and quote colors from the dashboard theme
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add a shared priority dot and pill for findings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): soften the frequency card and show its date range
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): rank findings under high, medium and low priority headings
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(lens-ui): show finding priority, label quotes by content and collapse extra examples
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): prove findings are grouped and ordered by priority
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens-ui): cover finding priority, quote labels and example collapsing
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): bound single trace reads by the sampled start time
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): allow unused query fixture field in load tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass start_time in every lens content and evidence test
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(logging): bound data URI regex so base64 truncation stays linear
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): cover whitespace-free data: prefixes in data URI regex regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): apply the due_at index concurrently on its own and default legacy rows to due
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass repository to claim lifecycle tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): page past unsupported due lenses and declare the full due_at index
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): claim due lenses in a loop instead of recursion
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag a2a, access_control, other, secret_manager and migrations tests with Subject metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): run transcript guardrails on raw-path transcription sessions with a transcription-safe block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): assert transcription session keeps transcribing after a guardrail block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): only expect a follow-up transcript when the block keeps the session open
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(realtime): type the transcription block regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): flag transcription sessions from the route intent and backend events only
A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail
* fix(realtime): flag transcription sessions from provider-transformed session events
* test(realtime): cover transcript guardrail blocks on transcription sessions
---------
Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>