Commit graph

49384 commits

Author SHA1 Message Date
yuneng-jiang
7809eacb8b
test(e2e/ui): cover the Logs page filter drawer (#39056)
The Logs page had coverage for opening a request and for the End User filter,
but nothing for the filters an on-call engineer actually reaches for: whose
key made the request, and which requests failed.

Each test mints its own keys and asserts on request ids it generated itself,
so a filter that quietly does nothing fails on the other key's row still being
on screen rather than passing because our own row happens to be there.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 12:34:10 -07:00
Sean Yasnogorodski
8a4ba78869
feat(guardrails): add Alice guardrail (#38898)
* feat(guardrails): add Alice by ActiveFence guardrail

Adds `guardrail: alice` — policy-based guardrails for prompts and model
responses, evaluated against ActiveFence's Alice.

What makes this different from the other providers: Alice evaluates against
policies configured per *application*, and a proxy typically fronts several of
them, so the application cannot be a static config value. It is named on the
LiteLLM virtual key instead:

    curl $PROXY/key/generate -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
      -d '{"key_alias": "payments-bot",
           "metadata": {"alice_app_id": "payments-bot"}}'

read via `CustomGuardrail._get_admin_metadata`, with `key_alias` as the
fallback. That helper is what makes it trustworthy: it reads whichever metadata
holder the proxy wrote the authenticated key's values into — which differs by
route — and the proxy strips caller-supplied `user_api_key_*` from both, so a
caller cannot point its own traffic at an application with laxer policies than
the one its key was issued for. A request whose key names no application is
refused rather than evaluated against a guess.

Implements `apply_guardrail` only, so pre_call, during_call, post_call and
streaming all come from UnifiedLLMGuardrails. Blocks with
GuardrailRaisedException; masks by substituting Alice's redacted text; a MASK
carrying no replacement blocks rather than passing the original through. A
verdict reporting `errors[]` is treated as a failure, not a pass — otherwise a
half-evaluated message would be allowed. `unreachable_fallback` (already on
LitellmParams) chooses fail-closed or fail-open on transport failure.

Config:

    guardrails:
      - guardrail_name: alice
        litellm_params:
          guardrail: alice
          mode: [pre_call, post_call]
          api_key: os.environ/ALICE_API_KEY

21 tests in tests/test_litellm/proxy/guardrails/guardrail_hooks/test_alice.py
cover registration, credential resolution, the app-id ladder including the
forged-metadata case, every verdict, and both unreachable policies.

No new LitellmParams field, so no schema.d.ts regeneration is needed.

* refactor(guardrails): post to Alice's LiteLLM endpoint and forward verbatim

Switches from `/v2/evaluate/message` — Alice's single-text endpoint — to
`/v2/evaluate/litellm`, which takes the hook's arguments as they arrive and
answers with a verdict.

That inverts where the work happens, and shrinks this plugin accordingly. It
now selects nothing and renames nothing: it posts `{input_type, inputs,
request_data}` and enforces `{verdict, categories, correlation_id, message,
replacements}`. Which parts of a conversation are worth evaluating, and how a
verdict is reached, are decided by Alice — so changing either is a change on
their side rather than a LiteLLM upgrade for every user.

The app-id resolution this plugin carried is gone with it. Alice reads the
application off the authenticated key's metadata itself, from the payload it is
handed, so the ladder here was duplicating a decision the far side already
makes. The security property is unchanged and still comes from the proxy
stripping caller-supplied `user_api_key_*` before a guardrail sees the request.

Masking is now positional — the far side chose which texts it was answering
for, so it says which by index. Only `texts` is written; a new
`structured_messages` object would make the chat translation layer skip the
`texts` write-back and silently drop the edits. A mask that lands nowhere
blocks rather than passing the original through.

`request_data` carries live Python objects (an OpenTelemetry span among them),
so `_json_safe` copies it into something serialisable by a mechanical rule
rather than a field list — a list drifts from what the far side needs, a rule
cannot. Serialising naively raises, and that error would read as "guardrail
unavailable" on every request.

26 tests, covering verbatim forwarding, each verdict, positional masking, the
`structured_messages` identity trap, both unreachable policies, and the
serialiser's handling of unserialisable values and cycles.

* fix(alice guardrail): satisfy lint and code-quality CI gates

- Bound _json_safe's recursion and register it in recursive_detector's
  ignore list (it already caps depth and dedupes cycles by id, matching
  the repo's established pattern for legitimate bounded recursion).
- Clear ruff-strict budget breaches: annotate __init__'s return type,
  raise TypeError (not ValueError) for a bad response body, type
  _json_safe's payload as object instead of Any, and file-scope-ignore
  ANN401 for **kwargs (forwarding it as object broke the call into
  CustomGuardrail.__init__, confirmed via basedpyright).
- Clear type-discipline budget breaches: suppress the construction/
  annotation checks on one-shot HTTP payloads, the module-level
  guardrail registries, and _json_safe's bounded accumulator; narrow
  AliceVerdict's list fields to tuples and _evaluate's request_data to
  Mapping[str, object] where nothing downstream mutates them.

* test(alice guardrail): assert the guardrail actually registers

The registration test called init_guardrails_v2 and asserted nothing, so it
passed whether or not the guardrail was ever registered — TQ001 in the
test-quality gate, and a fair catch: a test that cannot fail is not covering
the thing it names.

Now asserts exactly one AliceGuardrail lands in litellm.callbacks under the
configured name.

This surfaced only after the ruff-strict and type-discipline gates stopped
failing ahead of it; the lint job runs its gates in sequence, so an earlier
failure masks every later one.

* fix(alice guardrail): reach 100% patch coverage, drop the ActiveFence naming

Codecov flagged 10 uncovered lines, all of them error paths — which is where a
guardrail most needs covering, since each one decides whether traffic flows
unscreened.

Two of the ten turned out to be dead rather than untested, and are removed:

- `except GuardrailRaisedException: raise` in apply_guardrail. `_evaluate`
  raises httpx errors, Timeout and TypeError, never that — so the clause could
  never fire.
- the trailing `json.dumps` probe in `_json_safe`. Everything json.dumps
  handles natively is caught by the isinstance branches above (a dict or list
  subclass included), so anything reaching the bottom — bytes, datetime, an
  OpenTelemetry span — cannot cross the wire regardless. It now says so and
  returns None.

The rest are now tested: a timeout, 502/503/504 as unreachable, a 4xx as NOT
unreachable (a rejected credential is our misconfiguration, not an outage, and
must not fail open), a non-object response body, and a model whose model_dump
raises.

Also drops "by ActiveFence" throughout — the product is Alice — and points the
header at alice.io. `ui_friendly_name` is now "Alice", which is the key
guardrailLogoMap and the garden card look up, so all three moved together.

* fix(alice guardrail): strip caller credentials, widen unreachable detection, block partial MASK

Addresses PR review: request_data no longer forwards secret_fields.raw_headers or
the root api_key to Alice (the caller's Authorization token in the clear otherwise);
HTTP 500, malformed JSON, and a non-object body now route through the configured
unreachable_fallback instead of raising raw, so fail_open still fails open on those;
a MASK verdict with even one out-of-range replacement now blocks entirely instead of
silently letting the rest through unmasked. Also tightens request_data's type and
documents the known streaming-mask limitation on the class.

* fix(alice guardrail): strip credentials at any depth, stop filtering on texts

secret_fields/api_key/headers/provider_specific_header can appear nested
under proxy_server_request, metadata, litellm_metadata, and their
requester_metadata/body sub-paths in a real captured payload — a
top-level-only strip missed all of those. _json_safe now drops these keys
by name wherever they occur during serialization, so a new nesting path
can't reintroduce the leak.

apply_guardrail also stopped skipping the call whenever texts was empty,
even when tool_calls/images/structured_messages carried content — that
was the plugin making a selection decision Alice's design says belongs on
the far side. It now only skips when none of the selectable fields have
anything in them.

* fix(alice guardrail): route an undecodable response body through the fallback

`response.json()` raises UnicodeDecodeError when the body carries bytes that
are not valid UTF-8, and that escaped the except clause: UnicodeDecodeError is
a *sibling* of json.JSONDecodeError under ValueError, not a subclass of it, so
naming only JSONDecodeError left it uncaught. Both fallback modes surfaced a
raw decoding error instead of applying unreachable_fallback — which for a
fail_open deployment meant a hard failure where it had asked for an allow.

Named explicitly rather than widening to ValueError, so the clause still says
which three conditions it means. Tested under both policies.
2026-09-01 12:33:39 -07:00
mateo-berri
cf4738c3b7 style: format s3 vectors transformation and rag endpoints 2026-09-01 12:32:04 -07:00
mateo-berri
a0cf82b359 Merge origin/litellm_internal_staging into litellm_anthropic_stream_model_alias 2026-09-01 12:31:37 -07:00
Mateo Wang
4517e5e613
Merge pull request #39148 from BerriAI/litellm_claude_fable_5_1
feat(models): add Claude Fable 5.1 across Anthropic, Bedrock, Vertex AI, and Azure AI
2026-09-01 12:30:37 -07:00
Devin AI
24ee419c85 fix(models): registry audit 2026-09-01 for openai realtime, mistral aliases, voyage, xai, fireworks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:28:04 +00:00
mateo-berri
6012f893fa Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_s3_vectors_search 2026-09-01 12:27:29 -07:00
mateo-berri
753f3705d0 Merge litellm_internal_staging (4c3ef9ae0a) into litellm_fix_s3_vectors_search 2026-09-01 12:27:23 -07:00
mateo-berri
fa581931f7 test(llm_translation): expect unpaired server tool calls to replay as tool_use 2026-09-01 12:27:18 -07:00
Mateo Wang
695d943745
Merge pull request #39104 from BerriAI/litellm_decrease_anys_opus5_r3
refactor(types): replace Any with precise types across 73 modules
2026-09-01 12:26:52 -07:00
mateo
bba951c5eb fix(bedrock): drop client_metadata for ARNs that hide the model family
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:26:19 +00:00
mateo-berri
4875872fe5 fix(cost): treat a non-mapping off_peak_pricing value as never off-peak
A bare string or list under off_peak_pricing in YAML passed the truthy
guard and crashed _is_off_peak with AttributeError, breaking cost
calculation for that deployment. Malformed pieces of the block are
documented to not match rather than error, so guard the block itself the
same way and bill standard rates.
2026-09-01 12:25:39 -07:00
mateo
d6005a1876 merge: resolve conflict with litellm_internal_staging in anthropic transformation tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:23:47 +00:00
mateo-berri
f8298fa35c fix(bedrock): stop Converse crashing on bearer-token auth without SigV4 credentials
Since the Rust core handoff in #37241, BedrockConverseLLM.completion read
access_key, secret_key and token off the boto3 credentials before asking
the Rust gate whether it wanted the call. On a deployment that only sets
AWS_BEARER_TOKEN_BEDROCK boto3 resolves no credentials, so every Converse
call through /v1/chat/completions and /v1/responses failed with
"'NoneType' object has no attribute 'access_key'", with or without the
Rust opt-in

Bearer auth resolves no SigV4 principal at all, and both the Python and
the Rust path read the bearer token themselves, so only hand the
principal keys down when boto3 actually resolved one

get_request_headers now accepts credentials=None and raises botocore's
NoCredentialsError when neither a bearer token nor a principal exists
instead of handing SigV4Auth a None

Fixes #38579

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFDpYC45u9p8eKd4aBATnS
2026-09-01 12:22:15 -07:00
yuneng-jiang
0cf236bebb
test(e2e/ui): cover creating, testing and deleting a guardrail (#39053)
* test(e2e/ui): cover creating, testing and deleting a guardrail

The Guardrails page had no browser coverage. The RC checklist covers it by
hand against a live Presidio, which is why it has always been skipped in CI.

These drive the LiteLLM content filter instead, which runs inside the proxy,
so the whole flow is exercised without a third-party moderation service. The
create test does not stop at the table row: it sends a prompt carrying the
keyword it just banned and asserts the gateway refuses it, then sends a clean
prompt through the same guardrail and asserts it is served.

* test(e2e/ui): delete the guardrails these tests create

Review caught the fixtures being left behind. Guardrails are database rows
that show up in the table and in the playground's list, so a run that leaves
them changes what the next run sees.

Also trims the comments that restated what the helpers already say.

* test(e2e/ui): fail the run when guardrail teardown does not delete

Review caught the afterEach discarding the DELETE response, so a failed
cleanup finished quietly and left the guardrail for the next run to trip on.

* test(e2e/ui): wait for a new guardrail to reach the request path

The wizard test drove one chat completion immediately after creating the
guardrail and required a 400. A trace from the deployed stack shows the
record is stored correctly (blocked_words, action BLOCK, block_on_violation)
and the call six seconds later is still served unguarded, so the first
request can land before the proxy picks the guardrail up.

Polls the same call to the same 400 instead, which keeps the assertion and
lets the refresh land. If it never blocks, this stays red, which is what we
want it to say.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 12:21:45 -07:00
mateo-berri
8d0e7aee2f test(router): cover _inherit_builtin_base_rates_for_off_peak directly
The router_code_coverage gate only counts calls made from test files with
router in the filename, so the helper needs direct unit tests beside the
other inheritance helpers in test_router_model_cost_isolation.py: fills
missing base rates from the builtin entry, leaves explicit rates alone,
and no-ops without a block or for an unmapped backend model.
2026-09-01 12:20:09 -07:00
Mateo Wang
c50d83ece2
Merge pull request #39070 from BerriAI/litellm_bedrock_invoke_native_structured_output
fix(bedrock): forward native structured outputs on Invoke instead of silently inlining the schema
2026-09-01 12:19:15 -07:00
mateo-berri
e12e5e34f4 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_headroom_ccr_streaming_responses
# Conflicts:
#	tests/test_litellm/litellm_core_utils/test_streaming_handler.py
2026-09-01 12:19:09 -07:00
Mateo Wang
435433fa07
Merge pull request #39149 from BerriAI/litellm_qwencloud_provider_aliases
feat(dashscope): add QwenCloud and Qwen AI Platform provider aliases
2026-09-01 12:18:05 -07:00
yuneng-jiang
33004d2f0c
test(e2e/ui): cover the Budgets page create, edit and delete flows (#39052)
* test(e2e/ui): cover the Budgets page create, edit and delete flows

The Budgets page had no browser coverage at all, so an admin creating or
editing a spend cap through the UI was only exercised by hand at RC time.

Each test reads the budget back from /budget/list, a different route from
the one the table renders, so a row that only exists in the table's cache
does not pass. The edit test pins the rate limits an unrelated spend-cap
edit has no business touching.

* test(e2e/ui): trim comments that restate the test steps

Review flagged the explanatory comments as restating ordinary setup rather
than explaining anything. Keeps the two that carry the regression rationale
for an assertion and drops the rest.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 12:17:02 -07:00
mateo-berri
c21e895fe2 fix(proxy): handle CRLF and CR SSE frame terminators and flush held tail in anthropic stream restamper 2026-09-01 12:16:27 -07:00
mateo-berri
b0751169eb fix(cost): bill off-peak rates for deployments that set only off_peak_pricing
Cost lookup selects the deployment-scoped cost map entry only when custom
pricing is detected and the entry carries a base pricing field. A deployment
whose model_info held nothing but off_peak_pricing failed both conditions, so
its schedule was silently ignored and every request billed at the shared
backend rate.

use_custom_pricing_for_model now also treats deployment-scoped pricing fields
in the metadata model_info as custom pricing, and the router inherits the
backend model's built-in base token rates onto such an entry at registration,
which also lets cache pricing inheritance apply. Regression tests cover the
registration, the detection, and the costed request end to end.
2026-09-01 12:14:46 -07:00
mateo-berri
2adae6b475 fix(prometheus): pass through router-originated labels when no proxy router exists 2026-09-01 12:14:27 -07:00
mateo-berri
8ccbd82bd2 fix: authorize custom-logger spend payload against its own owner
The custom-logger detail branch reads the payload from cold storage, which is
written independently of the spend-log table and can outlive its row. The DB
owner pre-check then has nothing to verify for an id lookup that matches no row,
so a foreign tenant's stored payload could be returned. Authorize the returned
payload against the owner recorded inside it (metadata user/team id), failing
closed when none is recorded. Also fold the three identical 403 raises into one
helper.
2026-09-01 12:13:26 -07:00
Mateo Wang
f3dbd253be
Merge pull request #38106 from Timik232/bugfix/streaming-stable-response-id
fix(streaming): keep response id stable across streamed chunks
2026-09-01 12:11:32 -07:00
Devin AI
b96121efb1 refactor(anthropic): build adaptive output_config in one shot in legacy thinking helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:10:22 +00:00
mateo-berri
98ea5eaab4 fix(responses): correlate streamed tool call events on normalized item ids 2026-09-01 12:08:22 -07:00
mateo-berri
24a49726c2 Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD 2026-09-01 12:06:38 -07:00
mateo
d816b75dd4 feat(bedrock): gate forced tool_choice on supports_forced_tool_use in converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:05:03 +00:00
mateo
6513f5c539 test(utils): allow supports_forced_tool_use in model prices schema test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:05:03 +00:00
milan
ed4343a026 fix(bedrock): scope client_metadata drop to anthropic converse models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:04:32 +00:00
Devin AI
b8dd27a77f style(anthropic): keep mutable-ok annotation within ruff format width
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:00:46 +00:00
Mateo Wang
4c3ef9ae0a
Merge pull request #39023 from BerriAI/litellm_add_azure_deepseek_v4_flash_0731
feat: add Azure AI DeepSeek V4 Flash 0731 pricing
2026-09-01 11:54:25 -07:00
yuneng-jiang
75f0a22fc6
Merge pull request #39130 from BerriAI/litellm_dark_mode_skill_detail
fix(ui): render the skill detail page with theme tokens
2026-09-01 11:52:57 -07:00
mateo-berri
0042493bca Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_invoke_native_structured_output 2026-09-01 11:50:05 -07:00
mateo
a92ca6cfde chore(models): regenerate model prices schema for supports_forced_tool_use
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:48:26 +00:00
ryan-crabbe-berri
d9f7f9ea16 feat(ui): add search to Agent Hub tab and admin agents table
Ports the Model Hub search to the AI Hub Agent Hub tab and the admin
/agents toolbar as a client-side filter over agent name and description.
Extracts the hub search matching into utils/searchUtils and fixes the
public Model Hub rendering the whole catalog when a search matches
nothing (LIT-5230)
2026-09-01 11:46:58 -07:00
ryan-crabbe-berri
01de283715
Merge pull request #39142 from BerriAI/litellm_osv_browserslist
build(deps): bump browserslist to 4.28.8 to clear osv-scan
2026-09-01 11:46:26 -07:00
yuneng-jiang
2b1bd20834
Merge pull request #31125 from BerriAI/litellm_/stoic-jones-7de871
feat(proxy): default to the v2 migration resolver, keep v1 as an opt-out
2026-09-01 11:46:13 -07:00
mateo
3c9ce458fd feat(anthropic): gate forced tool_choice for Fable 5.1 behind supports_forced_tool_use
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:46:03 +00:00
mateo-berri
4ce2b4d5b4 Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD 2026-09-01 11:45:45 -07:00
Mateo Wang
caeb5d181e
Merge pull request #39074 from BerriAI/litellm_fix_websearch_tool_selection_test
test(websearch): register configured search tool in pre-request hook test
2026-09-01 11:45:29 -07:00
mateo-berri
24393be4a6 Merge branch 'litellm_internal_staging' into bugfix/streaming-stable-response-id 2026-09-01 11:45:27 -07:00
Devin AI
2063c29f5d fix(anthropic): upgrade legacy thinking to adaptive on adaptive-only models for chat and Bedrock Converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:45:26 +00:00
yuneng-jiang
82dd36c1a4
Merge pull request #39146 from BerriAI/litellm_revert_websearch_search_tool_validation
revert: restore search tool fallback when no router is configured
2026-09-01 11:38:11 -07:00
mateo-berri
62a42b4b47 refactor(dashscope): wrap long error message strings in common_utils 2026-09-01 11:34:23 -07:00
mateo-berri
61cef45d8e fix: re-verify ownership on custom-logger payload branch for id lookups 2026-09-01 11:33:20 -07:00
mateo
0610331aa1 test(reasoning-effort-grid): bump the cell count for the four new Fable 5.1 cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:28:33 +00:00
mateo-berri
c01ef712a2 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_headroom_ccr_streaming_responses
# Conflicts:
#	tests/test_litellm/responses/litellm_completion_transformation/test_streaming_iterator_transformation.py
2026-09-01 11:27:03 -07:00
Yuneng Jiang
c65dfd5d37
refactor(websearch): build the search tool lists as tuples
The restored code seeded two mutable lists, which trips LIT002 now that
the type-discipline budget has ratcheted past what they cost

Build both in one shot as tuples and widen the parameter to Sequence so
the single caller still type checks. No behavior change, both are only
ever read
2026-09-01 11:26:02 -07:00