Commit graph

47406 commits

Author SHA1 Message Date
tin-berri
d8ca43a800
feat(complexity_router): let the LLM classifier see request images (#39825)
The classifier scores extracted text, so a turn whose complexity lives in
its image is invisible to it: a screenshot of a stack trace classifies on
its caption, and an image-only turn flattens to empty text and never
reaches the classifier at all.

classifier_llm_config.vision opts in, off by default, with max_images
bounding what one turn can add. Images are still dropped when the
classifier model is declared supports_vision false. Anthropic and
Responses image parts are rewritten into chat-completions dialect before
they reach the classifier call, since /v1/messages hands the pre-routing
hook its own dialect untranslated.

The local scorer no longer short-circuits heuristic_first or hybrid on a
turn carrying forwarded images, because it reads text alone and its
confidence describes a request it has only partly seen.
2026-09-04 18:59:50 -07:00
mateo-berri
9c068117e7 chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
moe-berri
d1fd3a3457 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_shadow_eval_judge_output_cap
# Conflicts:
#	tests/test_litellm/integrations/test_shadow_eval_logger.py
2026-09-04 18:57:19 -07:00
mateo-berri
a61bead287 test(store_model_in_db): accept both 400 shapes in the unknown-model spend log test 2026-09-04 18:52:17 -07:00
mateo-berri
ee50f2bc44 test(e2e/batches): run the list assertion when a batch completes before cancel
The completed-batch early return skipped both the cancel and the list
assertion while the lifecycle's covers markers still credited both cells.
List does not depend on the batch being cancellable, so it now runs either
way; cancel on a completed batch stays a documented vacuous pass
2026-09-04 18:50:47 -07:00
Mateo Wang
2f90a264f6
Merge pull request #39827 from BerriAI/litellm_azure_gpt_6_astra
feat(cost-map): add azure/gpt-6-astra and azure/us/gpt-6-astra Foundry pricing
2026-09-04 18:48:51 -07:00
ryan-crabbe-berri
53d7b45e87
Merge pull request #38703 from BerriAI/litellm_fix_stale_team_on_user_row
fix(team_endpoints): let member_delete clear a team left on the user row
2026-09-04 18:48:07 -07:00
mateo-berri
d4fd658891 fix(datadog_llm_obs): keep guardrail_cost_by_unit on redacted spans 2026-09-04 18:47:48 -07:00
moe-berri
639b3f4f62
Merge pull request #39809 from BerriAI/litellm_stall_escalation
feat(router): auto-escalate stalled complexity-router tasks
2026-09-04 18:44:25 -07:00
moe-berri
2c3c7dd1a6
feat(shadow_eval): judge tool-call turns instead of dropping or erroring on them (#39818)
* fix(shadow_eval): tell a tool-call shadow reply apart from an empty one

Both arrive at the attempt row as the same 'shadow router returned an empty
response', because _chat_final_text returns empty for a tool-final turn by
design and for a reply that genuinely carried no text. Those are different
things: an arm that chose a tool where the real model wrote prose is a
divergence a text judge cannot score, and the sampling side already drops the
real arm's tool-final turns for exactly that reason, so the shadow side reads
as a fault where the real side reads as a filter. A job that is almost all
'empty response' gives no way to tell a tool-happy arm from a broken one.

The error now names which of the two happened, and carries the finish_reason
and the routed model so the row says what the arm was doing. Every varying
part sits behind the first semicolon: operators read these by grouping on the
error text, and interpolating the model into the leading sentence would make
each row its own group.

The outcome stays 'error'. Whether a tool-call reply should instead be its own
non-judged outcome, excluded from the loss rate the way the real arm's
tool-final turns already are, needs the four aggregation predicates that spell
judged as outcome != 'error' rewritten, and a decision on how to surface the
new bucket. That is a separate change.

* fix(shadow_eval): read the tool name of a custom tool call

A custom tool call carries its name under custom.name with no function key,
so every one of them reported as tool=unnamed.

* feat(shadow_eval): judge tool calls instead of dropping the turn

A turn where either arm called a tool was discarded before it could be
compared: the real arm's at sampling, the shadow arm's as an error row. On
agentic traffic that is most of the traffic, so a job set to sample 10% was
sampling 10% of the prose-only slice. Tool calls now serialize to text on
every surface and are judged like any other response, and the judge is told
a tool call is not a defect so it scores the choice rather than the shape.

* feat(shadow_eval): show the judge what tools were available

Both arms were offered the same tools, but the judge only ever saw the
chosen call in isolation, with no way to tell whether a better tool existed
or the arguments matched what the tool expects. Threads the request's tool
definitions (name and description only) into the judge prompt, capped and
omitted entirely on turns that offered none.

* fix(shadow_eval): read a custom tool definition's name from custom, not function

A chat-completions custom tool definition nests name and description under
custom, mirroring how a custom tool call nests them (openai.types.chat.
ChatCompletionCustomToolParam). Reading only function rendered every one as
unnamed, telling the judge nothing about what it was.
2026-09-04 18:41:47 -07:00
yucheng-berri
853fed824e
fix(bedrock): stop sending toolConfig tool definitions to guardrails on passthrough converse (#39281)
Bedrock passthrough Converse routes flattened every non-empty string under
toolConfig.tools into the guardrail INPUT texts, so tool names, tool
descriptions and JSON-schema strings (object, property names, titles, type
names, enum values) each arrived as a separate guardrail item. A request whose
only prompt was one benign user message could be blocked outright because a
denied term appeared in an app-authored tool definition.

Tool definitions are now excluded from the extracted texts, matching every
other guardrail translation handler, which carries tool definitions in the
structured tools input rather than in texts. Caller content stays scanned:
message text, toolUse.input, toolResult content and json, and
additionalModelRequestFields are unchanged.

Resolves LIT-5797
2026-09-04 18:40:35 -07:00
mateo-berri
3b9568bcdd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_7015_migration_job_node_selector 2026-09-04 18:40:25 -07:00
mateo-berri
7351911b53 fix(proxy): refuse OpenAI websocket passthrough on every enforced model allowlist and propagate the DB opt-in 2026-09-04 18:39:56 -07:00
mateo-berri
1748dd81a7 docs(e2e/batches): say the unified Bedrock lifecycle lists with plain GET /v1/batches 2026-09-04 18:39:28 -07:00
moe-berri
88ada40cda fix(type-checking): satisfy the basedpyright budget gate for auto-router compression
Two fixes for the zero-headroom basedpyright budget:

- arm_pre_call's data parameter is dict[str, object], not MutableMapping: the
  latter is itself banned by LIT001 with no benefit, and it mismatched every
  dict-typed helper (get_or_create_metadata_bucket, resolve_structured_messages,
  _get_tags_from_request_kwargs), which is what the budget was actually flagging.
- Router.async_pre_routing_hook computed pre_routing_hook_response in one shot
  instead of reassigning a Final-annotated local.

The remaining two reportArgumentType hits are pre-existing: LiteLLM_Params(**merged)
in _create_deployment_object already fails this check for all ~165 of its other
fields, since the merged dict's value type is partly untyped/float; adding two new
string fields to the model just grows that existing pile by two. Suppressed at the
one call site with a reason, since fixing the root typing is out of scope here.
2026-09-04 18:37:38 -07:00
Mateo Wang
77e27b1866
Merge pull request #39780 from BerriAI/litellm_/goofy-bohr-6cd011
fix(proxy): strip every TypedDict qualifier before numeric form-field detection
2026-09-04 18:36:54 -07:00
Mateo Wang
3d08daecfe
Merge pull request #39729 from amasen02/fix/end-user-budget-reset-cache-invalidation-39726
fix(proxy): invalidate end-user spend counter and cache on budget reset (#39726)
2026-09-04 18:36:05 -07:00
mateo-berri
f022b5eda8 test(e2e/batches): assert Bedrock batch cancel and list in the lifecycle
Bedrock batch cancel (StopModelInvocationJob) and the managed list view
both work through the proxy since LIT-4774, but the batches e2e still
gated them off and the coverage registry claimed no cell for either.
Flip can_cancel/can_list for the Bedrock provider, assert cancel the
same way the OpenAI leg does, add the two registry cells the gates
select, and update COVERAGE.md
2026-09-04 18:30:55 -07:00
mateo-berri
2674934e45 feat(helm): render nodeSelector, tolerations, and affinity on the componentized chart migrations Job 2026-09-04 18:28:53 -07:00
mateo-berri
c4759a9d98 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_techdebt_20260903 2026-09-04 18:27:43 -07:00
yucheng-berri
e2741b5643
fix(datadog_llm_obs): keep the guardrail audit record under message redaction (#39702)
* fix(datadog_llm_obs): keep the guardrail audit record under message redaction

Redaction nulled `guardrail_information` on the span whole, so an operator
running `turn_off_message_logging` (or a caller sending
`x-litellm-enable-message-redaction`) lost the record of which guardrails ran,
what they returned, and what they masked. Four of the record's fields can quote
the prompt; the rest report what the guardrail decided without reproducing it.

Replace only those four, the way
`_sanitize_guardrail_information_for_spend_logs` already does for spend logs,
and declare the field list once in `litellm/types/utils.py` so both readers
share it.

* fix(datadog_llm_obs): keep a lone guardrail record, and test through the span

Review round 1.

A guardrail that writes the metadata key itself leaves a single record where
the type says list, which Prometheus already normalizes at
`_guardrail_overhead_seconds`. Redaction dropped that shape and the latency
extraction raised on it, so the span was lost outright. Normalize once and use
it in both places.

The new tests now drive `create_llm_obs_payload` instead of reading the module's
private helpers and the record's declared field names.
2026-09-04 18:24:16 -07:00
mateo-berri
f27699a1c9 test(store_model_in_db): assert the 400 contract in the unknown-model spend log test 2026-09-04 18:23:36 -07:00
Mateo Wang
41c8c2f410
Merge pull request #39763 from BerriAI/litellm_fireworks_tool_choice_short_name
fix(fireworks_ai): resolve tool_choice and reasoning support for short model names
2026-09-04 18:22:25 -07:00
yuneng-jiang
e733ca1065
Merge pull request #39811 from BerriAI/litellm_/mongodb-vector-store-e4ff63
feat(vector_stores): add a MongoDB vector store provider for Atlas and self-managed deployments
2026-09-04 18:19:03 -07:00
mateo-berri
a2d5215a4f fix(proxy): gate the OpenAI websocket passthrough behind an explicit opt-in 2026-09-04 18:14:44 -07:00
mateo-berri
cd113c3a2e ci: allowlist the bounded _unqualified qualifier peel in the recursion detector 2026-09-04 18:13:51 -07:00
mateo-berri
270452db79 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/goofy-bohr-6cd011 2026-09-04 18:09:57 -07:00
yassin
35ae1ca151 test(cli): drop redundant comment in pi arg ordering test
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
LiteLLM Rust / release wheel (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 01:01:41 +00:00
mateo-berri
4ffd2ffb25 test(proxy): parametrize the stale end-user counter case so no Final local sits in a loop 2026-09-04 18:00:40 -07:00
mateo-berri
b4fd63f621 chore(proxy): annotate the new spend-counter test locals and correct the floor comments 2026-09-04 17:53:45 -07:00
mateo-berri
c02f198a4a fix(azure): responses none-effort temperature gate reads the azure/ cost-map entry 2026-09-04 17:45:06 -07:00
mateo
aedaf0d5a7 chore: ratchet lint budgets after merging litellm_internal_staging
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 00:45:00 +00:00
moe-berri
dc63428395 refactor(auto-router compression): satisfy the LIT001/LIT002 type-discipline gate
The gate has no headroom, so the new module had to stop introducing mutable
collections rather than spend budget on them:

- the marker lookup falls back to () and drops an `or {}` that isinstance
  already covered
- the suppression list is stored as the tuple it was built as; the read side
  in custom_guardrail accepts list or tuple, since JSON round-trips it to a list
- the snapshot holds MappingProxyType entries, so it is immutable at rest and
  _snapshot_messages can hand back the stored tuple with no defensive copy
- arm_pre_call returns None instead of echoing back the dict it mutates in place
- _suppressed_by_auto_router_compression takes a Mapping, which is all it reads

The four remaining mutable spots are external contracts, each suppressed with
the reason: the pre-routing hook protocol types messages as list[dict], the
metadata["guardrails"] key is extended by litellm_pre_call_utils via an
isinstance(..., list) check, apply_guardrail takes a dict it writes stats into,
and pydantic's model_copy takes a dict.
2026-09-04 17:42:52 -07:00
mateo-berri
89086db282 fix(proxy): floor end-user budget checks on the DB row after a reset
The reset job evicts the cached end-user object only from its own worker's
in-memory cache (plus Redis), so every other uvicorn worker and replica keeps
the pre-reset spend for up to user_api_key_cache_ttl (60s by default). Those
workers pass that stale spend as fallback_spend, and since the authoritative
floor read returned None for spend:end_user: keys, get_current_spend handed
the stale value straight back and the end user kept getting 429 after the
rollover on every worker but the one that ran the reset.

The floor read now consults LiteLLM_EndUserTable.spend for end-user counters,
the same way keys, teams, users, and orgs already read their rows. It runs only
when the shared counter sits below the cached spend (a reset or a Redis
restart) and stays behind the existing 5s in-process marker, so the normal
request path still does no DB read. Cold end-user counters keep seeding from
the cached object rather than the row, so from_db is unchanged for them.
2026-09-04 17:39:46 -07:00
yassin
7a9f466657 fix(cli): satisfy pi type discipline budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 00:37:04 +00:00
mateo
2e3c58c0fc chore: merge litellm_internal_staging into litellm_techdebt_20260903
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 00:35:26 +00:00
yucheng-berri
59d42d36e6
fix(headroom): bound the /v1/compress and /v1/retrieve calls with a timeout (#39527)
* fix(headroom): bound the /v1/compress and /v1/retrieve calls with a timeout

The headroom guardrail builds its client with get_async_httpx_client(GuardrailCallback)
and no params, and passes no timeout on either outbound call. That client's read, write
and pool legs are 600s (litellm.request_timeout when set explicitly, default 6000s), so
an unreachable or stalled compression service holds the caller's pre-call request open
for the whole window before unreachable_fallback ever runs. Because the client is shared
with every other no-params guardrail, each stalled call also pins a pooled connection for
the same window, so a saturated pool makes unrelated requests block on the pool leg.

Bound both calls at 60s by default, honoring litellm_params.timeout when set (the field
already exists and documents itself as the per-guardrail API timeout; headroom accepted
it and ignored it). The connect leg stays at the http_handler default, or the configured
budget when that is shorter, so a dead host still fails fast.

Live on a proxy against a stalled /v1/compress: 600.4s -> 60.2s before the 502, and 5.2s
with timeout: 5 configured.

* fix(headroom): reject non-finite timeouts and trim the timeout commentary

`timeout: .inf` on a Headroom guardrail reached httpx and the aiohttp transport
raised OverflowError, so every request came back as a raw 500 instead of going
through unreachable_fallback. Reject non-finite values the same way as
non-positive ones, and cut the comments and docstrings back to what the code
does not already say.
2026-09-04 17:34:25 -07:00
ryan-crabbe-berri
be76dfad9c
Merge pull request #35853 from BerriAI/litellm_retry_policy_503
fix(router): resolve retry_policy by exception hierarchy, add ServiceUnavailableErrorRetries and DefaultRetries
2026-09-04 17:34:00 -07:00
Mateo Wang
6960a42008
Merge pull request #36575 from BerriAI/litellm_mcp_stateless_follow_up_zdr
fix(responses/mcp): keep follow-up calls stateless when store=false
2026-09-04 17:33:27 -07:00
mateo-berri
51514b9123 fix(cost-map): azure/gpt-6-astra accepts reasoning_effort none on Foundry 2026-09-04 17:29:55 -07:00
yassin
b57cee65cc fix(cli): write pi models atomically and privately
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 00:22:42 +00:00
yassin
48849a6f36 Merge origin/litellm_internal_staging into litellm_lite_pi
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 00:22:11 +00:00
moe-berri
e273cf301f fix(ci): satisfy ruff format, prettier, and eslint max-lines gates
- ruff format on auto_router_compression.py (a long comprehension wrapped
  across three lines instead of one)
- prettier on buildAutoRouterCompression.ts and the two test files it touched
- ComplexityRouterConfig.tsx crossed the 800-line eslint max-lines ceiling
  once the compression accordion entry landed. Extracted TierRowSelect into
  its own file (already self-contained, used only within this file and
  PlanModeOverrideControls) and simplified CompressionControls' props to a
  single state/onChange pair instead of six individual callbacks, moving the
  per-field derivation into the component that already owns this state shape
2026-09-04 17:21:04 -07:00
moe-berri
aa3d59d086
fix(anthropic): never carry cache_control on translated thinking blocks (#39815)
* fix(anthropic): never carry cache_control on translated thinking blocks

The /v1/messages adapter built every thinking and redacted_thinking block with
cache_control=content.get("cache_control", {}), so a block the client never
marked still came out carrying an empty cache_control. anthropic_messages_pt
replays thinking blocks verbatim and first, so that value landed at content[0]
of the outbound assistant message and Anthropic rejected the request with
messages.N.content.0.thinking.cache_control: Extra inputs are not permitted.

Anthropic's schema has no cache_control on either block type, so there is
nothing to gate or translate here, only to stop copying. Every sibling block
type already routes through _add_cache_control_if_applicable; these two were
the only ones setting the key unconditionally.

This is reachable from any caller that round-trips Anthropic messages through
the OpenAI shape, which is why shadow eval saw it on a majority of sampled
Claude Code turns while the same traffic served natively was fine.

* test(anthropic): assert the outbound wire body for redacted thinking blocks
2026-09-04 17:18:59 -07:00
Mateo Wang
da09976c16
Merge pull request #39572 from BerriAI/litellm_spend_logs_keep_service_account_key_readable
fix(spend-tracking): keep internal service-account key names readable in spend logs
2026-09-04 17:17:39 -07:00
devin-ai-integration[bot]
a6b7384094
feat(cli): sync OpenCode models from /v1/models in lite opencode (#39789)
* feat(cli): sync OpenCode models from /v1/models in lite opencode

lite opencode now fetches the proxy's /v1/models with the resolved key and
hands OpenCode an OPENCODE_CONFIG_CONTENT declaring a litellm provider
(@ai-sdk/openai-compatible, proxy /v1 base URL, {env:OPENAI_API_KEY}) with one
model entry per listed chat model, so the model picker mirrors the proxy
without a hand-maintained opencode.json

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): sync OpenCode models only after the key check passes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 17:16:48 -07:00
devin-ai-integration[bot]
c373645e21
fix(proxy): recognize opencode's bare x-session-id header for session affinity (#39802)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 17:16:45 -07:00
moe-berri
d0a8006737 fix(auto-router compression): tag-scoped markers now take precedence over untagged
An untagged marker (no tags key or empty tags list) was matching every request
because requested.issuperset(frozenset()) is always true. When an alias carried
multiple markers, the loop tried tag-matched markers first, but an untagged one
could still match the tag-match query, and then the first one with a policy would
be returned. Now only markers with a non-empty tags list can match via the
tag-specific lookup; untagged markers are tried only after all tag-specific ones.

Regression test added: test_tag_scoped_marker_takes_precedence_over_untagged
fails with the old code.

Also removed unused Any import per greptile's typing note.
2026-09-04 17:03:01 -07:00
mateo-berri
3202963f25 feat(cost-map): add azure/gpt-6-astra and azure/us/gpt-6-astra Foundry pricing 2026-09-04 16:57:40 -07:00
mateo-berri
4b24e3b727 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mcp_stateless_follow_up_zdr 2026-09-04 16:56:47 -07:00