Commit graph

15252 commits

Author SHA1 Message Date
devin-ai-integration[bot]
d34f5b3a4e fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (#43956)
The integration test hunks are dropped because tests/integration/observability/test_straiker_v3_platform.py is not on this line

(cherry picked from commit a3a7650569)
2026-10-01 01:25:34 +00:00
PhimmStraiker
edc85e5721
feat(guardrails): straiker guardrail speaks the v3 platform API (/api/v3/detect) (#41880)
* feat(guardrails): speak the Straiker v3 platform API (/api/v3/detect)

The Straiker guardrail posted a webhook envelope to /api/v1/detect/webhook.
The v3 platform exposes /api/v3/detect instead, and its integration keys
(sk_agt_…) are rejected by the v1 route with an empty 401, so a tenant on
the v3 platform could not run this guardrail at all. Measured on a
customer gateway on 2026-09-17 after they rotated to a v3 key.

v3 parses the gateway's own traffic server-side, the same contract as
Straiker's unified Kong plugin. So on v3 the guardrail relays: the
request phase posts the provider body LiteLLM received (Anthropic
Messages or OpenAI chat), the response phase posts
{straiker_phase, sse, model, request}, the answer beside the request it
answers, and Straiker derives prompt, answer, agent and archetype. Both
phases also carry the flat prompt / app_response pair: a gateway-mode
integration key scores only the flat pair and an api-mode key only the
relayed body, each ignoring the other, so one payload serves whichever
key the console issued and it is one turn either way (measured on tenant
123, both key modes, 2026-09-18).

- api_version: "v1" | "v3", unset follows the key prefix, so a v3 key
  needs no extra configuration. Explicit override still wins.
- The relayed body is an allowlist of provider fields. The hook sees the
  client body merged with proxy state: `deployment` carries the resolved
  provider credential and `proxy_server_request` the client's own
  Authorization header. Neither travels. Identity survives as the
  metadata subset Straiker's LiteLLM adapter reads.
- Identity never sends a proxy placeholder. `default_user_id` and the
  master-key alias were being forwarded as a user and became the
  session's identity on the platform.
- Headers: x-tool: litellm (ingress), x-straiker-phase, x-straiker-user,
  and x-claude-code-session-id forwarded when the client sent it.
- Verdict: hookSpecificOutput.permissionDecision on the gateway envelope,
  `action` on the flat one; block on block/deny, and on a non-empty
  blocked_by as a backstop. A detect-mode control reads NONE.
- An error status from Straiker is now a webhook failure. LiteLLM's HTTP
  client raises on any non-2xx and the retry loop caught only connection
  errors, so a 401 or 503 from Straiker escaped the guardrail as an
  exception and was relayed raw to the client, bypassing fail_open /
  fail_closed. Retryable statuses retry; the rest are final.
- v1 is unchanged: same envelope, same X-Straiker-Webhook-Format header.

Tests: 15 new, fixtures from the request dict a hook sees on 1.98.0 and
the verdict envelopes the v3 platform returned on 2026-09-18. Each fix
was mutation-checked (handling removed, the test fails). Live: the same
eight-case battery (chat, /v1/messages, streaming, tool call; benign,
injection, PII) passes on a gateway-mode and an api-mode key, blocks at
pre_call with the tenant's block message, and lands under the declared
agent with the end user attributed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(guardrails): name the agent per application on v3 (x-s6r-agent)

One integration key can front several applications. Straiker enumerates them
as separate agents when the turn names one, which is what the unified Kong
plugin sends as x-s6r-agent. Without it every application on a gateway
collapses onto a single agent.

- Forwards a client-supplied x-s6r-agent.
- New `agent_ref` config names one agent for a route when the client sends
  nothing. The client wins, matching Kong's precedence.
- Neither set: no header, and the platform derives the agent from the traffic.

Verified live on tenant 123 against an integration whose connector is
`gateway`: three distinct values minted three observed agents, and a turn
with no hint derived one from the traffic shape. An integration whose
connector is `custom-agent` declares its agent, so every turn attributes to
that one agent and the hint is ignored (agent_ref_source: attested).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(guardrails): which v3 shape is scored depends on the connector, not the key mode

The earlier comment said a gateway-mode key scores only the flat pair. Re-measured
on tenant 123 across all three integration types with one injection prompt:

  custom-agent connector (Add Agent)  raw body ignored   flat prompt scored
  gateway connector                   raw body scored    flat prompt scored
  api mode                            raw body scored    flat prompt ignored

Behaviour unchanged: the payload already carries both shapes, which is why it works
on every type. Comment only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(guardrails): send exactly what the unified Kong plugin sends on v3

The v3 platform parses the gateway's traffic itself and derives agent,
archetype and identity from it. The earlier commits added to the relayed
body (a flat prompt / app_response pair, source, user_name) and to the
headers (x-tool, x-straiker-phase, x-straiker-user). None of that is in
the Kong v0.12 contract, and traffic through this guardrail was not
classifying by shape the way the same traffic through Kong does. Match
Kong byte for byte and leave classification to the platform.

Request phase: the provider body, plus session_id and
original.processed.Meta.user. Response phase: {straiker_phase, sse,
model, request} plus the same two. No flat fields, no phase or user
headers, no x-tool.

Session id follows Kong's precedence: the client's x-claude-code-session-id,
then the session LiteLLM resolved, then an md5 of system prompt + first
message so a conversation that states no session still groups across its
replays.

Routing hints complete the Kong set: x-s6r-agent (client header, else
`agent_ref`), and new `client` (x-s6r-client) and `format_hint`
(x-s6r-format) config, both optional.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(guardrails): sort imports in the v3 session test

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): send a streamed Messages answer back in the Messages shape on v3

On a streamed /v1/messages call the proxy rebuilds the answer as a chat
completion before the post-call hook runs, and that is what the plugin put in
the response envelope's sse field. Straiker's coding-agent reader parses a
Messages answer, so a Claude Code turn relayed this way came back
coding_agent/claude with no session and zero events scored: the model's tool
calls were never screened on the response phase. Captured live on 2026-09-18
against tenant 123, a real Claude Code Bash tool call through the proxy.

The proxy's own Anthropic adapter turns the rebuilt answer back into a Messages
response when the call arrived on the anthropic_messages route, which is what a
transport relay forwards. Chat completions calls keep the chat completion shape
and a buffered Messages answer is relayed untouched.

The regression test's fixture is the chat completion the proxy actually built
for that captured turn. After the fix the same turn scores on the response
phase (session resolved, one event, the Bash tool_use block present).

* style(straiker): ruff format the v3 guardrail and its tests

* refactor(straiker): one attempt per call in the webhook retry loop

The HTTPStatusError branch added for v3 duplicated the non-200 branch and put
_post_webhook over the strict complexity ceiling. One attempt is now its own
method that returns the verdict or a failure marked retryable, and the loop only
decides whether to try again. Behaviour is unchanged: retryable statuses and
transport errors retry, everything else is final.

* fix(straiker): name Claude Code's client and agent on v3 so its session lands under one coding agent

Straiker types a gateway turn as a coding agent from the "You are Claude Code"
preamble, which only the main agent turns carry. Claude Code's title and
topic-detection sidecars have their own system prompts, so they resolved by
shape as autonomous, and because they share the session id with the main turns
the whole session was filed under Autonomous rather than under a coding agent.
Kong does not hit this because its plugin config names the client and agent on
every call.

The User-Agent (claude-cli/...) is on every call including the sidecars, so the
plugin now reads it and sends x-s6r-client: claude plus, when the route names no
agent, x-s6r-agent: "Claude (LiteLLM)". A client-supplied x-s6r-agent or the
agent_ref config still wins. Verified live on tenant 123: a real Claude Code
session now lands as one coding_agent labelled "Claude (LiteLLM)" with its turns
scored, where before it split across Autonomous.

Identity: the key's own user (email then id) now outranks the end user the
request named. LiteLLM resolves Claude Code's hashed metadata.user_id as the end
user when nothing better is set, so a per-user key was being shadowed by a
session token. The key is the authenticated principal, the way a Kong consumer
is, so it wins; the request end user is the fallback.

* refactor(straiker): build the v3 request, envelope and headers as frozen mappings

The v3 builders seeded dicts and grew them, which the type-discipline gate
counts as mutable accumulators. Each is now one expression over a tuple of
pairs, frozen with MappingProxyType, and the JSON encoder unwraps a frozen
mapping through a default. The session seed and the verdict parser no longer
rebind locals. The wire is unchanged: 36 live calls through the proxy on this
commit carry the same fields, shapes, headers and identities as before, with
no mappingproxy text in any body.

* fix(straiker): satisfy basedpyright on the v3 builders

The frozen-mapping refactor left a shadowed headers local, a Mapping handed to
an HTTP client that takes a dict, an unguarded optional response, a turn id
typed object, and a redundant isinstance on already-typed texts. No behaviour
change: 4 live calls (chat, Messages, Bedrock, injection) return 200 with the
expected verdicts on this commit.

* fix(straiker): type the v3 config fields at the initializer and keep the verbose log as JSON

The four v3 routing fields (api_version, agent_ref, client, format_hint)
travelled through the untyped kwargs passthrough, which basedpyright counts
against the budget. They are now validated through a small Pydantic model at
the initializer and passed by name.

The verbose log serialized the frozen payload with default=str, which printed
a Python repr instead of JSON once the builders returned MappingProxyType.
Every serializer now unwraps a frozen mapping first. A test asserts the logged
payload parses as JSON and carries the identity; mutating the log site back to
default=str fails it.

* fix(straiker): address review findings on the v3 relay

Text completions relay their prompt: `prompt`, `suffix`, `echo` and `best_of`
join the provider allowlist, so /v1/completions traffic is screened.

The route's `agent_ref` now outranks the caller's `x-s6r-agent` header. The
header is caller-supplied, and letting it beat a pinned route would let any key
file its traffic under another application's agent and controls. On a route
that names nothing the header still names the application, which is how
several applications enumerate behind one key.

Credentials inside `tools` and `mcp_servers` (an OpenAI `mcp` tool's `headers`,
Anthropic's `authorization_token`) are replaced with `[redacted]` before the
body leaves the proxy, on both phases and in the verbose log. Detection reads
tool names, descriptions and schemas, never these.

A 200 whose body is valid JSON but not an object now reports an invalid
schema and follows the failure policy instead of raising out of the hook.

Comments that restated a constant are gone. Tests cover each change and the
failure paths (unreadable error body, client exceptions, missing response,
unmodellable request, session seeds from Anthropic block shapes); every fix
fails its test when reverted.

* fix(straiker): scrub tool credentials one level deep, without recursion

* fix(straiker): scrub only the fields that carry a credential, never a schema

The credential set is now the three fields that actually hold one on a tools
or mcp_servers entry (headers, authorization, authorization_token), read one
level deep. A function tool whose parameter schema defines a token, headers or
api_key property is relayed exactly as sent; a test pins that, and fails
against the recursive version.

* test(straiker): use example.com identities; drop a comment that restated its branch

* fix(straiker): present a legacy completion as the chat exchange it is

Straiker scores chat on both phases of a gateway turn but has no reader for a
text_completion answer: the request phase of a /v1/completions call was
scored and the response phase was refused with 501, whether or not the call
named an agent. A completion is one user turn and one assistant turn, so both
phases now present that exchange: the prompt becomes the single user message
and the TextCompletionResponse becomes a chat completion. Measured through the
proxy on this commit, both phases return 200 and score, and the derived
session is shared between them.

The derived session seed accepts the tuple the conversion produces; the test
pins the session on both phases and fails against the list-only check. The
unreachable "parsed is None" branch is folded into the failure branch, and a
malformed tools value is shown to relay as sent.

* fix(straiker): screen a completions prompt as the text the model receives

LiteLLM's /v1/completions accepts a string, a list of strings, a list of
token ids or a list of token-id lists, and decodes token ids with the
text-davinci-003 tokenizer before calling the model. The relay now renders
the prompt the same way, one user message per prompt, so a pre-tokenized
prompt is screened as the text it stands for rather than as digit strings.
A prompt in a shape this cannot render (empty, mixed, or with no tokenizer
available) is relayed untouched instead of being replaced with something
else. Tests cover all four accepted shapes and six unrenderable ones.

* fix(straiker): seed the derived session on the preamble and the first user turn

An OpenAI chat body carries its system prompt as messages[0], and the derived
session seeded on the Anthropic `system` field plus messages[0] with no role
check. For that shape the seed was the system prompt twice and the first user
turn never counted, so every unnamed conversation behind one system prompt
collapsed into one Straiker session. The seed now takes the preamble from
wherever the API puts it (`system`, `instructions`, or a leading system or
developer message) and the first message with role `user`, else a Responses
`input` string, else `prompt`. Two conversations sharing a system prompt are
two sessions again; a replayed conversation stays one.

* fix(straiker): seed the derived session on every text block of the first turn

A user turn that opens with an image or a document block and carries its
text later seeded the session on an empty string, so two different
conversations under the same preamble shared one Straiker session. Read
every text block of the turn instead of only the first block. A plain
string or a single text block seeds exactly as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(straiker): cover the tokenizer fallback, a textless first turn and Responses instructions

Three branches of the v3 relay had no test: a token-id prompt relayed as
sent when the tokenizer cannot be fetched, a first user turn with no text
seeding the session on the preamble alone, and a Responses API body
seeding on its instructions and first input turn. Each test fails when
its branch is mutated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): seed the derived session on the principal as well as the conversation

Straiker skips turns it has already scored for a session. The derived
session hashed the system prompt and the first user turn alone, so two
users who opened a conversation with the same words shared one session,
and the second user's copy of an attack came back as a replay: unscored
and allowed. Measured live on 2026-09-20: the first user's SSN turn was
blocked (`social_security_number`, scored=2), the second user's identical
turn was allowed (`controls: []`, replayed=2).

The principal now joins the seed. Explicit session ids, the Claude Code
header and LiteLLM's own session are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): derive the session id with sha256 and drop comments that restated constants

The derived session now hashes the principal, and CodeQL flags MD5 over an
identity as a weak hash on sensitive data. SHA-256 truncated to the same
32 hex characters keeps the id shape. Comments that only labelled the
allowlist groups or restated a constant are removed; the two that explain
a non-obvious choice stay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): keep a blocked conversation blocked when it is replayed

Straiker de-duplicates turns it has already scored per session and
answers a replay `allow`, whatever the first verdict was. A client that
resends a blocked request, or grows the conversation past the blocked
turn, was let through: measured on 2026-09-20, `block` then `allow,
events_replayed=2` for the same session and body, and Claude Code's
automatic retry after the 400 turned a blocked poisoned-file read into
a pass.

The guardrail now remembers, per session, a fingerprint of every
conversation it blocked (a bounded, day-long in-memory cache) and blocks
a request that repeats or extends one without asking again. A different
session with the same words is a new conversation and is scored afresh.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): scope the block memory by session or principal, never by content alone

A request with no derivable session keyed the replay memory on the
conversation fingerprint alone, so one caller's block could answer
another caller's identical request. The memory is now scoped by the
session, else by the principal, and a request with neither is not
remembered at all.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): remember only a block that names a control, never one that comes from state

The replay memory kept every block, including one the platform returns
because a kill switch is engaged (`action: block` with `blocked_by: []`).
An administrator lifting the kill switch then left the conversation
refused by the remembered copy: measured on 2026-09-21, traffic stayed
blocked after `POST /inventory/agents/{id}/restore` returned `engaged:
false`.

The same words are the same attack tomorrow, so a control-named block is
still worth remembering; state is not ours to cache. The parsed verdict
now carries `blocked_by` so the two can be told apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Phimmasone Phonpaseuth <PhimmStraiker@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e0af9917a1)
2026-09-30 16:27:27 -07:00
mateo-berri
7cc6268587
fix(guardrails): stop the Javelin api_version default leaking into Azure Content Safety (#41941)
(cherry picked from commit 385932b4e3)
2026-09-30 16:27:27 -07:00
Yuneng Jiang
9161d7f9d6
refactor(auth): bind UI/CLI session tokens to their own AES-GCM context
Backport of BerriAI/litellm-private#5 (12981f93d3) onto stable/1.101.x, applied as
the PR's net diff so main-only intermediate refactors stay out.

The dashboard and lite CLI SSO specs under tests/e2e/ui/oidc are left out
because this line has no OIDC e2e harness to run them.
2026-09-29 15:17:24 -07:00
devin-ai-integration[bot]
59eedca7b1
fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker (#42607)
* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): apply ruff format to invoke messages stream passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): drop drive-by reformat of existing invoke messages tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): collect streamed chunks into a tuple in passthrough regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): give the passthrough regression test a 10s first-chunk budget

* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]

* test(bedrock): take the gated byte stream's chunks as an immutable Sequence

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 975bd28549)
2026-09-23 12:05:51 -07:00
yuneng-jiang
a03bc3e927
Merge pull request #42635 from BerriAI/litellm_backport_stable_1_101_x_bp_40639_1101
chore(release): backport #40639 to stable/1.101.x
2026-09-23 12:01:22 -07:00
Mateo Wang
432215e619
Merge pull request #42596 from BerriAI/litellm_cherrypick_1_101_x
feat(typesafe): backport #41607 to stable/1.101.x for v1.101.1
2026-09-22 20:01:33 -07:00
mateo
7e6d9e89ed fix: satisfy stable/1.101.x lint and api-sync gates for the jev backport 2026-09-23 02:21:16 +00:00
devin-ai-integration[bot]
347ec67498 feat(openrouter): price typesafe/jev-1.13 and add an openrouter decisions pass-through (#42301)
Backport of #42301 to stable/1.101.x.
Cherry-picked from 1106b16745 (main).
2026-09-23 02:21:15 +00:00
ryan-crabbe-berri
173f71d120 fix(reset_budget_job): reset end users by budget link, not by user id
The cascade zeroed end-user spend with a single update_many whose where
clause enumerated every dependent user id. Prisma compiles that IN-list
into one prepared statement carrying one bind variable per customer, and
PostgreSQL caps a statement at 32,767 of them. Once a shared budget had
more dependents than that the statement could not be parsed at all, so
the atomic cascade rolled back, budget_reset_at never advanced, and the
tier stayed due on every later tick forever. Customers sitting at their
cap were blocked indefinitely with only a recurring log line to show for
it.

End users now match on budget_id like every other gated table, plus a
NULL-budget_id branch for the implicitly created rows that carry no link
and ride the default tier. The statement's bind count now tracks the
number of expiring tiers rather than the customer population, so a reset
costs the same whether a budget has ten dependents or a million.

Fixes #40564

Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
(cherry picked from commit 760043b533)
2026-09-23 01:12:36 +00:00
moe-berri
4bfd03e3de feat(auto-router): add JEV classifier alongside LLM classifier
Backport of #41886 to stable/1.101.x.
Cherry-picked from a83773cfa5 (main).
2026-09-22 23:49:19 +00:00
moe-berri
02f386a82e fix(proxy): enforce virtual key budgets for JEV test routing
Backport of #41879 to stable/1.101.x.
Cherry-picked from 1e161f516c (main).
2026-09-22 23:15:11 +00:00
Yassin Kortam
8c30093c5a feat(guardrails): add TypeSafe Jev relevance-based compaction guardrail
Backport of #41757 to stable/1.101.x.
Cherry-picked from 2edda5aec3 (main).
2026-09-22 23:12:35 +00:00
yuneng-jiang
eb0eca5451 fix(proxy): forward every method on the typesafe pass-through route
Backport of #41723 to stable/1.101.x.
Cherry-picked from 34718f0da6 (main).
2026-09-22 23:12:01 +00:00
Mateo Wang
2f63249816 feat(router): add TypeSafe Jev as a complexity router classifier
Backport of #41615 to stable/1.101.x.
Cherry-picked from cf42b607c3 (main).
2026-09-22 23:09:35 +00:00
mateo-berri
081f65ea44 feat(typesafe): add TypeSafe Jev passthrough with logging and cost tracking
Backport of #41607 to stable/1.101.x.
Cherry-picked from deb9d8aedd (main).

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:34:24 +00:00
kerry
d4ca007b6c fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1
Backport of #42048 to stable/1.101.x.
Cherry-picked from 7966f50c34 (main). The safeguards backport maps the dangerous-tool-use-2026-09-03 beta for Bedrock, which Claude Opus 4.5 on Bedrock Invoke rejects as an invalid beta flag, so the all-beta-headers Bedrock cases run on Claude Fable 5.1 as they do on main.
2026-09-22 12:54:53 -07:00
mateo-berri
3c6a66dc48 chore(types): keep the backported safeguards annotations within the line's budgets
The picked TypedDict fields use read-only Sequence[Mapping[str, object]] annotations and the picked Vertex test carries a test-quality-ok marker, so stable/1.101.x's LIT001, LIT012 and TQ008 budgets hold. Static typing only, no runtime change.
2026-09-22 11:31:16 -07:00
mateo-berri
2414f8fcf0 test: add the local_beta_headers_config fixture the safeguards tests use
Hand-ported to stable/1.101.x from 47b2479c94 on main (fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag), the one prerequisite the #42288 handler tests need; the rest of that commit stays on main.
2026-09-22 10:38:09 -07:00
mateo-berri
85bf3c8380 fix(anthropic): forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages
Backport of #42288 to stable/1.101.x.
Cherry-picked from merge commit fc82f6e8fa (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
2026-09-22 10:16:55 -07:00
Yassin Kortam
64e011ccab fix(anthropic): forward safeguards and anthropic-beta unchanged on native /v1/messages
Backport of #42152 to stable/1.101.x.
Cherry-picked from merge commit e912ebe999 (litellm_claude_code_safeguards_passthrough).
2026-09-22 10:16:26 -07:00
Mateo Wang
aec3026860
Merge pull request #41870 from BerriAI/litellm_bedrock_openai_gpt_min_max_tokens
fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse

(cherry picked from commit a6e3a72ed8)
2026-09-21 14:18:30 -07:00
mateo-berri
c9ed764c58 fix(license): let a wildcard allowed_features license grant the auto_router feature
Backport of #41684 to stable/1.101.x.
Cherry-picked from c2fbb11dca (litellm_wildcard_license_auto_router).
2026-09-17 16:27:10 -07:00
Yuneng Jiang
abfb2cf329
test(load): backport Redis chaos qualification to rc/1.101.0 (#40482)
(cherry picked from commit 8e4f2abb40)
2026-09-12 12:28:44 -07:00
Yuneng Jiang
c1a834baaf
fix(caching): backport Redis breaker recovery guards to rc/1.101.0 (#40624)
(cherry picked from commit ff4b558243)
2026-09-12 12:28:18 -07:00
Yuneng Jiang
ca4a61304f
fix(redis): backport quiet breaker refusals to rc/1.101.0 (#40620)
(cherry picked from commit ef1a37795c)
2026-09-12 12:27:57 -07:00
Yuneng Jiang
c324a2b3c5
perf(proxy): backport pipelined spend counters to rc/1.101.0 (#40371)
(cherry picked from commit 996ee5635a)
2026-09-12 12:27:18 -07:00
Yuneng Jiang
bcddffcbdc
fix(otel): backport auth spans and callback merge to rc/1.101.0
Cherry-pick PR #40335, restoring the Datadog auth span and last-wins callback credential merging.

(cherry picked from commit 43a1b2992a)
2026-09-08 18:28:11 -07:00
yuneng-jiang
f0b6d66b84
Merge pull request #40325 from BerriAI/litellm_backport_37667_rc_1_101_0
feat(team): backport team admin callbacks to rc/1.101.0 (#37667)
2026-09-08 16:42:17 -07:00
Yuneng Jiang
b6e73cccdb
feat(team): backport team admin callbacks to rc/1.101.0
Backport #37667 without conflict resolution or implementation changes

(cherry picked from commit 90731576e3)
2026-09-08 16:35:28 -07:00
Yuneng Jiang
d8b9fcf399
chore(ui): backport dashboard dependencies to rc/1.101.0
Backport #40312 without conflict resolution or implementation changes

(cherry picked from commit 8280d7ca9d, mainline parent 1)
2026-09-08 16:35:28 -07:00
Yuneng Jiang
09976043ac
feat(otel): backport tenant trace destinations to rc/1.101.0
Backport #39654 without conflict resolution or implementation changes

(cherry picked from commit 3165951e6d)
2026-09-08 16:29:16 -07:00
yuneng-jiang
935bb5260f
feat: backport MongoDB sidecar to rc/1.101.0
(cherry picked from commit b8d573c5f9)
2026-09-08 15:52:58 -07:00
yuneng-jiang
45cf1a7ef1
Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…"
This reverts commit c091dd4608.
2026-09-05 16:07:09 -07:00
yuneng-jiang
1b25132863
Merge pull request #39953 from BerriAI/litellm_/litellm-e2e-flaky-test-2159ae
test(e2e): judge /v1/messages streaming on the clock, not on the provider's delta count
2026-09-05 16:04:45 -07:00
ryan-crabbe-berri
63156a7bd6 test(proxy): explain the proxy_server patches in the cache-hit regression test
The test-quality gate counts every patch of a litellm internal against a ceiling, and the three patches this test needs pushed it over. The callback imports increment_spend_counters, update_cache and proxy_logging_obj from proxy_server inside its own body, so there is no seam to inject fakes through; every other test in this file uses the same three patches for the same reason

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-05 15:44:37 -07:00
ryan-crabbe-berri
acddd21860 fix(proxy): keep guardrail cost in spend on cache hits
The proxy cost callback zeroed response_cost whenever cache_hit was true. That rule dates from Jan 2024 when it was the only place cache hits were priced. The logging layer has priced the LLM share at 0 on a cache hit since Aug 2024, and since guardrail cost joined the standard logging payload the proxy-side zeroing has thrown away a real provider charge: a pre_call guardrail runs before the cache is consulted, so a cached response still cost whatever the guardrail billed. Drop the redundant zeroing so the payload's response_cost, which is already LLM 0 + guardrail cost, reaches spend logs, daily tables and budgets untouched

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-05 15:32:02 -07:00
Mateo Wang
bf51dea36b
Merge pull request #39862 from BerriAI/litellm_lit_6992_cohere_parse
feat(ocr): add Cohere Parse support for cohere and azure_ai
2026-09-05 15:16:16 -07:00
yuneng-jiang
6a4fb2bbe8
Merge pull request #39938 from BerriAI/litellm_e2e_vertex_cache_first_call
test(e2e): prove Vertex context caching on the first cold call and on the spend row
2026-09-05 15:10:15 -07:00
Yuneng Jiang
cd976624d1
test(e2e): drop the explanatory sentence from the StreamingResponse docstring 2026-09-05 14:53:12 -07:00
Yuneng Jiang
b55a4317a6
test(e2e): annotate new stream-timing locals as Final and trim the docstrings 2026-09-05 14:49:56 -07:00
Yuneng Jiang
d56affa814
test(e2e): judge /v1/messages streaming on the clock, not on the provider's delta count
The Anthropic and Together AI /v1/messages streaming tests required at
least two content_block_delta events. How many deltas a reply is split into
is the provider's choice, and Haiku answers a short count in one or two, so
the assertion failed on provider variance with no change in the proxy: four
of the day's full runs on the PR e2e gate went red on it on 2026-09-05.

The harness now stamps when each SSE event reached the client
(StreamingResponse.stream_event_arrivals, index-aligned with stream_events,
with the clock injectable so the reader has a unit test). Both tests ask for
a reply long enough to take seconds to generate and require the first
content delta to land at least STREAM_MIN_LEAD_SECONDS before message_stop.
A relayed stream shows a lead of about two seconds. A proxy that buffered
the response delivers every event in one burst and fails every time, which
a whole-response buffering relay in front of a live proxy confirmed. The
event-grammar assertions are unchanged.

Replay hands the proxy its recorded chunks back to back, so timing says
nothing there. The assertion is gated on provider_paces_stream() and replay
proves the grammar only, which tests/e2e/CLAUDE.md now says.
2026-09-05 14:44:21 -07:00
ryan-crabbe-berri
1745d74293
Merge pull request #39853 from BerriAI/litellm_guardrail_usage_cost_ui
feat(ui): show guardrail usage units and cost on the Guardrails Monitor
2026-09-05 14:41:11 -07:00
yuneng-jiang
9222a4de2d
Merge pull request #39946 from BerriAI/litellm_e2e_messages_stream_delta_count
test(e2e): stream a longer /v1/messages reply so the delta-count pin has margin
2026-09-05 14:26:55 -07:00
devin-ai-integration[bot]
a46a076b2a
fix(proxy): reject ambiguous name or alias keys in mcp_tool_permissions on write (#39947)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 14:22:26 -07:00
Yassin Kortam
9832d6e4a6
fix(mcp): scan and mask MCP tool call arguments in unified guardrails (#35142)
* fix(mcp): scan and mask MCP tool call arguments in unified guardrails

A guardrail configured with mode pre_mcp_call was handed only a synthetic
tool definition (name plus an empty parameters schema), so it never saw the
argument values it was configured to inspect, and any rewrite it returned was
discarded. Detection could not fire and masking could not take effect, while
the applied-guardrails metadata still reported the guardrail as having run.

Pass every string leaf of the tool call arguments as texts, and fold the
guardrail's rewritten leaves back into modified_arguments, which is the channel
the MCP call path reads to decide what to send upstream. The leaf walk reuses
the json_string_leaves / with_json_string_leaves helpers the tool result path
already uses, so both directions share one bounded traversal.

Two guardrails running concurrently under run_in_parallel scan the same payload
snapshot, so each returns a full replacement derived from the original leaf.
Rewrites of the same leaf to different values are rejected rather than silently
losing one redaction; a leaf that already holds this guardrail's own replacement
is convergent and still masks, which is what the bundled content filter does
when it rewrites the arguments itself as well as through texts.

* fix(mcp): annotate guardrail argument rewrites

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): isolate MCP guardrail callback state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet LIT010 budget after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): remove duplicate Bedrock hook parameter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): fail closed when guardrail rewrites cannot be mapped to MCP arguments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): patch the guardrail translation mappings cache where staging now keeps it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 20:51:25 +00:00
devin-ai-integration[bot]
80839bb33c
feat(proxy): serve Prometheus /metrics from a separate process via --prometheus_metrics_port (#39889)
* feat(proxy): serve Prometheus /metrics from a separate process via --prometheus_metrics_port

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): ruff format prometheus_metrics_server

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): fail fast when the separate metrics server cannot start and force the multiproc dir whenever it is enabled

- wait for the child's /health before starting uvicorn; raise a ClickException if it exits first (port in use)
- create PROMETHEUS_MULTIPROC_DIR whenever --prometheus_metrics_port is set, so DB-configured prometheus callbacks work
- honour lowercase prometheus_multiproc_dir; validate the port before spawning
- cover main() entry point, readiness, bind failure and wildcard-host probing in tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): pin metrics-server readiness to the child pid so another service on the port cannot pass the health check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): probe metrics-server readiness through the shared HTTPHandler instead of bare httpx.get

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): serve only /metrics on the prometheus metrics port

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): validate metrics server CLI args with pydantic instead of typing.cast

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): satisfy metrics server lint gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 13:26:09 -07:00
Yuneng Jiang
d6bc8fe289
test(e2e): ask the streamed /v1/messages pin for a reply long enough to span several deltas
Anthropic now returns the 64-token 'count to 20' reply in one to three content_block_delta events, measured directly against api.anthropic.com and through proxies at 7672399 and 49a1145 alike, so the incrementality assertion (at least two deltas) failed in litellm-e2e builds 125, 130 and the 278 rerun with no proxy change behind it. A 'count to 100' reply at max_tokens 400 arrived in five to fifty deltas across every measured run
2026-09-05 13:23:38 -07:00
ryan-crabbe-berri
0be8bb98b0 Merge branch 'litellm_internal_staging' into litellm_guardrail_usage_cost_ui 2026-09-05 13:09:45 -07:00
devin-ai-integration[bot]
5df0e12e0f
feat(guardrails): add non-blocking flag() verdict to custom code guardrails (#39728)
Custom code guardrails could only allow(), block(reason) or modify(). This adds flag(reason, metadata={}) which lets the request or response through unchanged and records a guardrail_flagged entry carrying the guardrail name, configured mode, evaluated input_type (request or response), reason and structured metadata. The new status is threaded through the request-level guardrail_status aggregation, the Guardrails Monitor rollup (flagged_count), Request Logs (action=flagged, most severe phase wins when a guardrail runs pre and post call) and the Request Logs detail view in the dashboard, which now renders FLAGGED with warning styling instead of falling into FAILED.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 13:08:03 -07:00