* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(bedrock): apply ruff format to invoke messages stream passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(bedrock): drop drive-by reformat of existing invoke messages tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): collect streamed chunks into a tuple in passthrough regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): give the passthrough regression test a 10s first-chunk budget
* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]
* test(bedrock): take the gated byte stream's chunks as an immutable Sequence
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 975bd28549)
The cascade zeroed end-user spend with a single update_many whose where
clause enumerated every dependent user id. Prisma compiles that IN-list
into one prepared statement carrying one bind variable per customer, and
PostgreSQL caps a statement at 32,767 of them. Once a shared budget had
more dependents than that the statement could not be parsed at all, so
the atomic cascade rolled back, budget_reset_at never advanced, and the
tier stayed due on every later tick forever. Customers sitting at their
cap were blocked indefinitely with only a recurring log line to show for
it.
End users now match on budget_id like every other gated table, plus a
NULL-budget_id branch for the implicitly created rows that carry no link
and ride the default tier. The statement's bind count now tracks the
number of expiring tiers rather than the customer population, so a reset
costs the same whether a budget has ten dependents or a million.
Fixes#40564
Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
(cherry picked from commit 760043b533)
Backport of #41607 to stable/1.101.x.
Cherry-picked from deb9d8aedd (main).
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Backport of #42048 to stable/1.101.x.
Cherry-picked from 7966f50c34 (main). The safeguards backport maps the dangerous-tool-use-2026-09-03 beta for Bedrock, which Claude Opus 4.5 on Bedrock Invoke rejects as an invalid beta flag, so the all-beta-headers Bedrock cases run on Claude Fable 5.1 as they do on main.
The picked TypedDict fields use read-only Sequence[Mapping[str, object]] annotations and the picked Vertex test carries a test-quality-ok marker, so stable/1.101.x's LIT001, LIT012 and TQ008 budgets hold. Static typing only, no runtime change.
Hand-ported to stable/1.101.x from 47b2479c94 on main (fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag), the one prerequisite the #42288 handler tests need; the rest of that commit stays on main.
Backport of #42288 to stable/1.101.x.
Cherry-picked from merge commit fc82f6e8fa (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
The test-quality gate counts every patch of a litellm internal against a ceiling, and the three patches this test needs pushed it over. The callback imports increment_spend_counters, update_cache and proxy_logging_obj from proxy_server inside its own body, so there is no seam to inject fakes through; every other test in this file uses the same three patches for the same reason
Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
The proxy cost callback zeroed response_cost whenever cache_hit was true. That rule dates from Jan 2024 when it was the only place cache hits were priced. The logging layer has priced the LLM share at 0 on a cache hit since Aug 2024, and since guardrail cost joined the standard logging payload the proxy-side zeroing has thrown away a real provider charge: a pre_call guardrail runs before the cache is consulted, so a cached response still cost whatever the guardrail billed. Drop the redundant zeroing so the payload's response_cost, which is already LLM 0 + guardrail cost, reaches spend logs, daily tables and budgets untouched
Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
The Anthropic and Together AI /v1/messages streaming tests required at
least two content_block_delta events. How many deltas a reply is split into
is the provider's choice, and Haiku answers a short count in one or two, so
the assertion failed on provider variance with no change in the proxy: four
of the day's full runs on the PR e2e gate went red on it on 2026-09-05.
The harness now stamps when each SSE event reached the client
(StreamingResponse.stream_event_arrivals, index-aligned with stream_events,
with the clock injectable so the reader has a unit test). Both tests ask for
a reply long enough to take seconds to generate and require the first
content delta to land at least STREAM_MIN_LEAD_SECONDS before message_stop.
A relayed stream shows a lead of about two seconds. A proxy that buffered
the response delivers every event in one burst and fails every time, which
a whole-response buffering relay in front of a live proxy confirmed. The
event-grammar assertions are unchanged.
Replay hands the proxy its recorded chunks back to back, so timing says
nothing there. The assertion is gated on provider_paces_stream() and replay
proves the grammar only, which tests/e2e/CLAUDE.md now says.
* fix(mcp): scan and mask MCP tool call arguments in unified guardrails
A guardrail configured with mode pre_mcp_call was handed only a synthetic
tool definition (name plus an empty parameters schema), so it never saw the
argument values it was configured to inspect, and any rewrite it returned was
discarded. Detection could not fire and masking could not take effect, while
the applied-guardrails metadata still reported the guardrail as having run.
Pass every string leaf of the tool call arguments as texts, and fold the
guardrail's rewritten leaves back into modified_arguments, which is the channel
the MCP call path reads to decide what to send upstream. The leaf walk reuses
the json_string_leaves / with_json_string_leaves helpers the tool result path
already uses, so both directions share one bounded traversal.
Two guardrails running concurrently under run_in_parallel scan the same payload
snapshot, so each returns a full replacement derived from the original leaf.
Rewrites of the same leaf to different values are rejected rather than silently
losing one redaction; a leaf that already holds this guardrail's own replacement
is convergent and still masks, which is what the bundled content filter does
when it rewrites the arguments itself as well as through texts.
* fix(mcp): annotate guardrail argument rewrites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): isolate MCP guardrail callback state
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: ratchet LIT010 budget after merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): remove duplicate Bedrock hook parameter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): fail closed when guardrail rewrites cannot be mapped to MCP arguments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): patch the guardrail translation mappings cache where staging now keeps it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): serve Prometheus /metrics from a separate process via --prometheus_metrics_port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy): ruff format prometheus_metrics_server
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): fail fast when the separate metrics server cannot start and force the multiproc dir whenever it is enabled
- wait for the child's /health before starting uvicorn; raise a ClickException if it exits first (port in use)
- create PROMETHEUS_MULTIPROC_DIR whenever --prometheus_metrics_port is set, so DB-configured prometheus callbacks work
- honour lowercase prometheus_multiproc_dir; validate the port before spawning
- cover main() entry point, readiness, bind failure and wildcard-host probing in tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): pin metrics-server readiness to the child pid so another service on the port cannot pass the health check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): probe metrics-server readiness through the shared HTTPHandler instead of bare httpx.get
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): serve only /metrics on the prometheus metrics port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): validate metrics server CLI args with pydantic instead of typing.cast
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): satisfy metrics server lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Anthropic now returns the 64-token 'count to 20' reply in one to three content_block_delta events, measured directly against api.anthropic.com and through proxies at 7672399 and 49a1145 alike, so the incrementality assertion (at least two deltas) failed in litellm-e2e builds 125, 130 and the 278 rerun with no proxy change behind it. A 'count to 100' reply at max_tokens 400 arrived in five to fifty deltas across every measured run
Custom code guardrails could only allow(), block(reason) or modify(). This adds flag(reason, metadata={}) which lets the request or response through unchanged and records a guardrail_flagged entry carrying the guardrail name, configured mode, evaluated input_type (request or response), reason and structured metadata. The new status is threaded through the request-level guardrail_status aggregation, the Guardrails Monitor rollup (flagged_count), Request Logs (action=flagged, most severe phase wins when a guardrail runs pre and post call) and the Request Logs detail view in the dashboard, which now renders FLAGGED with warning styling instead of falling into FAILED.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): cover presidio post_call, tool_permission, and weave logging cells
Five registry cells in Logging & Guardrails had no covering test. Each one now
has a live scenario read back from the real destination:
- guardrail.presidio.post_call.masks: an output-scoped Presidio guardrail
anonymizes the PII the model repeats back. The prompt also asks for the
address's local part, which Presidio does not mask, so one response proves the
model saw the raw address (no pre-call masking) while the address itself comes
back as <EMAIL_ADDRESS>
- guardrail.tool_permission.pre_call.blocks / .allows: an allow-list of one tool.
A request declaring an unlisted tool is rejected 400 naming it; a request
declaring the permitted tool is served and carries
x-litellm-applied-guardrails, so the allow half cannot pass by the guardrail
never running
- logging.niche_integrations.success.logs_spend / .failure.logs_spend: a
key-scoped weave_otel callback delivers to the real Weave project, read back
through Weave's query API. Success asserts exactly one call whose
llm.response.cost equals the x-litellm-response-cost header; failure asserts
one ERROR-status call naming the provider exception and carrying no cost
Logging & Guardrails coverage goes 24/59 to 29/59. No registry rows are added.
* test(e2e): make the tool-permission allow case deterministic and scope the Weave read-back
Review follow-ups on the coverage PR.
- the allow scenario forced the outcome to depend on whether the model felt like
calling an optional tool, and checked for the tool name as a substring of the
whole body, which a prose mention would satisfy. It now sends
tool_choice="required" and asserts the parsed response carries exactly one tool
call, for the permitted tool
- the Weave read-back queried the newest 200 calls of a shared project and
filtered client-side, so busy traffic could push the target out of the window
and read as a delivery failure. The query now scopes server-side to the
litellm_request op and to calls started after the request, and pages through
the window with offset
- the reader builds its results as tuples instead of accumulating into lists
Also unblocks the lint gate: `basedpyright tests/e2e` runs only on PRs that touch
tests/e2e, and it has been failing on staging for three FakeItem arguments in
test_junit_properties.py. The stand-in now goes through one typed adapter that
says why, so the gate is green without touching junit_properties.py itself.
* test(e2e): scope the presidio post_call guardrail to email and phone
Running the suite three times in a row caught a real flake: Presidio's broader
recognizers sometimes claim the email's local part as an NRP entity, so the
answer came back as `<NRP>\n<EMAIL_ADDRESS>\n<PHONE_NUMBER>` and the assertion
that the raw local part survives failed. That token is what tells output masking
apart from input masking, so it has to survive.
The post_call guardrail now registers pii_entities_config for EMAIL_ADDRESS and
PHONE_NUMBER only, which is also the narrower thing the scenario means. Verified
against the exact marker that failed, plus two others.
* test(e2e): mark weave logging cells stage red
* test(e2e): use per-test stage red skips for the weave logging cells