Commit graph

449 commits

Author SHA1 Message Date
mateo
4da12795fc fix: filter deployment default API key limits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 21:50:45 +00:00
mateo
3e3e4d6970 fix(anthropic): use native structured output for claude-fable-5-1 on Vertex AI and Bedrock Invoke
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:01:03 +00:00
mateo
d6005a1876 merge: resolve conflict with litellm_internal_staging in anthropic transformation tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:23:47 +00:00
mateo
3c9ce458fd feat(anthropic): gate forced tool_choice for Fable 5.1 behind supports_forced_tool_use
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:46:03 +00:00
mateo
ab549a8da3 test(fallbacks): use an unmapped fable id now that 5.1 is in the cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:25:18 +00:00
Tin Chi Lo
a27e12367e fix(bedrock): forward native structured outputs on Invoke instead of silently inlining the schema 2026-08-31 23:19:22 -07:00
Mateo Wang
99703a30f0
Merge pull request #36008 from nuernber/litellm_bedrock_messages_disconnect_billing
fix(anthropic_messages): drain upstream in a detached pump so client …
2026-08-31 16:54:41 -07:00
mateo-berri
a99f62d1bc fix(anthropic_messages): park deferred billing before the end-of-stream sentinel
At end of drain the pump enqueued the sentinel first and picked the
billing mode from client_detached afterward, so a client that consumed
the sentinel and tore the relay down before the pump resumed (possible
whenever the sentinel enqueue hit a full queue) had its fully delivered
response billed through the teardown path, skipping the proxy's
post-response hook. Bill or park before the sentinel goes out, and let
an unconsumed sentinel fall back to dispatching the parked billing.
2026-08-31 16:46:12 -07:00
mateo-berri
46e090d2f3 fix(anthropic_messages): bill partial spend when a queued pump error is never consumed
When the upstream errors while the client is still connected, the pump
forwards the exception through the relay queue so the proxy's failure
handling re-raises it. If the client disconnects before consuming that
queued exception, neither the failure hook nor billing ran and the spend
row was lost. The pump now waits for client detach and, if the exception
was never consumed, salvages partial spend like the post-disconnect
error path.

Also rewrites the bedrock disconnect logging test to the detached-pump
contract: billing fires after the upstream drain completes, not
synchronously at aclose().
2026-08-31 12:47:26 -07:00
mateo-berri
0e78c5bff7 fix(anthropic_messages): dispatch deferred spend logging when the client disconnects mid-relay
When the pump finishes draining while the client is still connected,
billing is deferred to the proxy's post-response hook, which only fires
on a normally completed response. A client disconnect before the relay
consumed the queued tail tore the generator down past that hook, so the
request logged no spend at all. The relay teardown now dispatches the
stored deferred billing whenever it never reached the end-of-stream
sentinel.

Also drops the live pass_through_tests script: that CI job runs against
a fixed config with no Bedrock model or AWS credentials, so it could
only fail there. The scenario is covered by unit tests on the
relay/pump seam.
2026-08-31 10:17:27 -07:00
nuernber
79fd2f4872 test(anthropic): cover ANTHROPIC_MESSAGES_MAX_DETACHED_STREAM_DRAINS=0 fallback to partial billing 2026-08-31 09:15:36 -07:00
nuernber
e1fece511a test(anthropic): fix PT012 lint violation in upstream-error regression test 2026-08-31 09:11:51 -07:00
nuernber
95a2586228 Merge remote-tracking branch 'origin/main' into litellm_bedrock_messages_disconnect_billing
# Conflicts:
#	basedpyright-code-budget.json
#	litellm/llms/anthropic/experimental_pass_through/messages/streaming_iterator.py
#	tests/test_litellm/llms/anthropic/experimental_pass_through/messages/test_streaming_iterator.py
#	type-discipline-budget.json
2026-08-31 08:58:37 -07:00
feng.tsai
bb51c121cf docs: reference the source union by type instead of a line number
The line number went stale when the base moved.
2026-08-31 12:21:24 +08:00
samtsai15
0e7562dbc6 test(guardrails): cover every Anthropic image source shape in the extractor's own suite
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.

Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].

The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.

Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:37:51 +08:00
Mateo Wang
1f5e76155b
Merge pull request #38836 from BerriAI/litellm_fix_messages_effort_budget_cap
fix(anthropic): cap reasoning_effort thinking budget below max_tokens on /v1/messages
2026-08-29 16:45:18 -07:00
mateo-berri
71a951691a fix(anthropic): cap reasoning_effort thinking budget below max_tokens on /v1/messages
A deployment carrying reasoning_effort in its litellm_params on the
/v1/messages passthrough mapped the effort to a legacy thinking block
whose budget_tokens was forwarded as is, so any request whose max_tokens
sat at or below that budget was rejected upstream with a 400. The mapped
budget now runs through the same cap the adaptive-to-legacy branch and
the chat path already use: it is clamped to max_tokens - 1, and dropped
with a warning when even the minimum budget cannot fit.

The cap helper becomes public since three call sites outside
AnthropicConfig use it.
2026-08-29 15:24:05 -07:00
tin-berri
36ea28b092
fix(anthropic): emit signature-only thinking blocks on the /v1/messages bridge (#38809) 2026-08-29 15:04:22 -07:00
mateo-berri
1a5a856e3e fix(guardrails): defer native /v1/messages stream logging until post_call scans finish 2026-08-28 16:05:45 -07:00
tin-berri
39e5b0c2d1
fix(anthropic): drop and self-heal empty thinking blocks on /v1/messages (#38625)
* fix(anthropic): drop and self-heal empty thinking blocks on /v1/messages

* test(anthropic): pin early-signature carry across the blank thinking chunk skip
2026-08-27 21:41:56 -07:00
tin-berri
49e6081978
fix(anthropic): resolve /v1/messages effort tiers through the capability owner (#38492)
* fix(anthropic): resolve /v1/messages effort tiers through the capability owner

The bridge normalizer read three supports_*_reasoning_effort booleans of its own, so it
answered "which levels does this deployment take" independently of the resolver behind
/model_group/info. The two disagreed: a proxy advertising kimi-k3 max forwarded high.

Degrade against resolve_supported_reasoning_efforts instead, with the chains as a declared
table. When no step of a chain is accepted, the fallback is read off that same resolved set
rather than assumed, since an entry naming its levels outright can exclude the tiers the
per-level flags treat as unconditional. none is never chosen as that fallback, being an off
switch rather than a tier, and a deployment accepting no tier at all keeps the floor every
deployment degraded to before.

* test(anthropic): pin the normalized effort at the /v1/messages request boundary

The existing coverage stopped at normalize_reasoning_effort_value, so nothing failed if the
handler dropped or overwrote the normalized tier on its way into completion_kwargs. Drive
_prepare_completion_kwargs instead and assert on the kwargs handed to acompletion, in both the
string and the dict effort shapes, including the provider-prefixed model name the handler is
actually called with.

Against the pre-fix normalizer the fallback case fails, and against the baseline before a map
entry could declare its levels 7 of the 12 fail, so the boundary is pinned rather than restated.
2026-08-27 19:13:13 -07:00
tin-berri
1ab6fd89d2
fix(anthropic): carry the effort tier only where the target declares reasoning_effort (#38592)
The /v1/messages bridge decided a Claude target could take `reasoning_effort` from
the model name, which says nothing about the params the provider in front of it
accepts. Snowflake serves Claude over the Anthropic dialect and declares `thinking`
alone, so `get_optional_params` raised `UnsupportedParamsError` before the request
reached the wire: every adaptive request carrying an effort tier turned a 200 into
a 400 for all seven of its Claude entries.

The tier is now offered only where the target declares the param, reading the same
`get_supported_openai_params` the sibling `_supports_prompt_cache_key` reads twelve
lines up. A target declaring neither carrier keeps its bare `thinking` block, which
is what this bridge sent before it carried a tier at all.

Without a resolved provider the tier stays behind rather than being offered blind.
Resolving one from the model's prefix instead would run an OAuth device flow for
github_copilot and chatgpt, blocking for minutes, and one of the two callers in that
position is a logging callback. The copilot case is pinned by a test.
2026-08-28 00:41:46 +00:00
tin-berri
44d84360fb
fix(anthropic): carry the adaptive effort tier to every bridged Claude target (#38533)
/v1/messages forwarded `thinking` verbatim for a Claude-family model and then returned,
carrying `output_config.effort` only when the model string started with a Bedrock prefix.
Every other bridged provider got a bare adaptive thinking block, so the caller's effort did
nothing: max and minimal produced byte-identical upstream bodies.

Send those targets the tier as `reasoning_effort`, which is the param they take. Bedrock keeps
taking `output_config`, since the two are not interchangeable there: an application inference
profile ARN resolves to no chat config, so `reasoning_effort` is dropped and the tier vanishes,
and a provider that rebuilds `output_config` from it overwrites a caller-set `thinking.display`
on the way. The tier stays a plain string, the summary already travelling inside the forwarded
`thinking` block. Adaptive with no tier, and budgeted thinking, both stay exactly as they were.
2026-08-27 15:48:42 -07:00
tin-berri
30ff3723b2
feat(model_prices): let a map entry declare its exact reasoning_effort levels (#38481)
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.

Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.

The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
2026-08-27 15:38:01 -07:00
Mateo Wang
4ef1c28877
Merge pull request #38431 from BerriAI/litellm_fix_messages_native_tools
fix(anthropic-adapter): pass provider-native and OpenAI-format tools through on /v1/messages
2026-08-27 15:24:07 -07:00
Mateo Wang
649dc23d6a
Merge pull request #38465 from BerriAI/litellm_lit6103_tool_reference_passthrough
fix(anthropic): carry tool_reference tool results through the guardrail translation round trip
2026-08-27 15:23:51 -07:00
Devin AI
63d7920f8b refactor: dedupe server_tool_use web search reads and type fresh test locals
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-27 08:01:52 +00:00
mateo-berri
e26ea0bc95 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit6103_tool_reference_passthrough 2026-08-26 23:07:11 -07:00
mateo-berri
99af9ad9eb fix(anthropic): carry tool_reference tool results through the guardrail translation round trip 2026-08-26 22:35:46 -07:00
mateo-berri
6d7fafa347 fix(anthropic): close hybrid tool-name allowlist gap and keep native tools through guardrails 2026-08-26 20:37:11 -07:00
mateo-berri
815fa0ff08 fix(anthropic_adapter): carry web search usage into /v1/messages cost breakdown
For non-Anthropic models served over /v1/messages, the outer wrapper recomputes
cost over the adapter-translated Anthropic response dict. That dict dropped every
web search usage signal, so the recompute overwrote the correct cost breakdown
with a token-only one: x-litellm-response-cost-tool-usage read 0.0 and
x-litellm-response-cost-original excluded the search cost, while the total kept it.

The adapter now maps web search request counts (from Usage.server_tool_use or
Gemini's prompt_tokens_details) into usage.server_tool_use.web_search_requests,
matching the Anthropic API shape, and the Gemini web search cost calculator falls
back to server_tool_use when prompt_tokens_details carries no count. The shared
get_web_search_requests helper is now public since five modules consume it.

Resolves LIT-6288
2026-08-26 18:14:10 -07:00
mateo-berri
f7af44a505 fix(anthropic-adapter): pass provider-native and OpenAI-format tools through on /v1/messages 2026-08-26 17:55:15 -07:00
mateo-berri
ece187ea24 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_anthropic_fast_mode_cache_and_response_speed
# Conflicts:
#	tests/test_litellm/test_cost_calculator.py
2026-08-26 12:10:33 -07:00
Mateo Wang
8dc17e808a
Merge pull request #38240 from BerriAI/devin_ai_anthropic_messages_missing_key
fix(anthropic): raise missing-credential error on /v1/messages passthrough
2026-08-26 12:06:37 -07:00
mateo-berri
c330d466f3 fix(anthropic): normalize oauth authorization header to one canonical casing 2026-08-26 11:45:51 -07:00
mateo-berri
a0d1fef89d fix(anthropic): scale cache costs by fast mode and trust served speed 2026-08-26 11:24:12 -07:00
mateo-berri
2068066d69 fix(anthropic): detect oauth Authorization header case-insensitively 2026-08-26 11:16:17 -07:00
Mateo Wang
d46bcde2b5
Merge pull request #37922 from BerriAI/litellm_techdebt_20260822
chore(typing): roll up the daily tech debt cleanups from Aug 20 to Aug 26
2026-08-26 10:54:50 -07:00
Mateo Wang
95285c3433
Merge pull request #38211 from eugene-yao-zocdoc/litellm_anthropic_responses_strictness_pr
fix(anthropic-responses): preserve structured output strictness
2026-08-26 10:39:45 -07:00
mateo-berri
c449f11451 fix(anthropic): detect client credential headers case-insensitively on /v1/messages 2026-08-26 10:35:39 -07:00
Devin AI
2eedcb62ce Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_techdebt_20260822 2026-08-26 07:54:51 +00:00
Devin AI
055b6f6f69 chore: merge litellm_internal_staging into rolling techdebt branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-26 07:44:58 +00:00
mateo-berri
5cfc1608f9 fix(anthropic): map only provider failures on the /v1/messages boundary 2026-08-26 00:31:18 -07:00
mateo-berri
3fe65029bf fix(exceptions): map anthropic 403 to PermissionDeniedError
Now that /v1/messages routes provider failures through exception_type, an
Anthropic permission_error fell through the anthropic branch to the generic
APIConnectionError and reached the client as a 500 where the raw exception
used to answer 403. Map 403 to PermissionDeniedError so the status survives
on every route.
2026-08-26 00:09:10 -07:00
mateo-berri
666648d58c fix(otel): map /v1/messages provider errors before failure logging 2026-08-25 23:31:05 -07:00
Mateo Wang
d0c527f3bb
Merge pull request #36245 from BerriAI/litellm_fix_headroom_stream_leak
fix(anthropic): buffer streamed responses carrying server-fulfilled tools so retrieval tool calls never reach the client
2026-08-25 17:48:31 -07:00
mateo-berri
8d877980e0 fix(caching): carry has_buffered_provider_output through the anthropic stream cache writer 2026-08-25 17:31:46 -07:00
mateo-berri
1b37397633 fix(router): forward the hold-back keepalive ping live and carry the withheld-output flag through the stream wrappers
The router's pre-content ping filter dropped AgenticAnthropicStreamingIterator's
hold-back keepalive, so a held-back turn sent the client nothing until the buffer
settled. A ping that no lifecycle frame precedes is now forwarded live, since a
fallback's message_start can still follow it without overlapping lifecycles

The proxy's cancel-refund guard checked isinstance against the iterator, but the
proxy only ever sees it behind FallbackAwareAnthropicMessagesStream and
AnthropicMessagesStreamingResponse, so a disconnect during hold-back refunded the
budget reservation anyway. Both wrappers now forward a duck-typed
has_buffered_provider_output flag, and the router wrapper follows a fallback
source so the flag tracks the stream actually being consumed
2026-08-25 16:39:14 -07:00
Deepanshu Lulla
6684bb343b
perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic (#36610)
* perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic

Vertex's handle_accumulated_json_chunk and Anthropic's
_handle_accumulated_json_chunk each independently accumulated SSE fragments
into a JSON envelope with self.accumulated_json += fragment. Because the
attribute holds a live reference, CPython copies the whole prior buffer on
every fragment, making buffer assembly O(n^2) in total payload size.
Anthropic additionally had no completeness heuristic at all and called
json.loads on the whole buffer after every fragment, and could wedge forever
on two concatenated envelopes.

Add JSONFragmentAccumulator in litellm_core_utils/: fragments append to a
list in O(1), a could_close_json heuristic lets callers skip the join+parse
entirely until a value could plausibly be complete, and pop_next_value peels
one JSON value off the front of the buffer at a time using
json.JSONDecoder().raw_decode, keeping any unconsumed remainder instead of
failing on concatenated values. Migrate both providers onto it; Anthropic's
__next__/__anext__ end-of-stream handlers now delegate to
_handle_accumulated_json_chunk(is_final=True) instead of duplicating the
parse-and-reset logic inline.

Fixes #31861.

* test(streaming): close diff-coverage gaps in JSONFragmentAccumulator migration

Codecov flagged 11 uncovered lines in the migration: the accumulated_json
setter, __next__/__anext__'s end-of-stream drain branches in both providers,
and the pop_next_value "not found" path when a buffer's newest fragment ends
in "}" but is genuinely incomplete (an inner object closed, the outer one
didn't). Add targeted tests for each.

* fix(streaming): make JSONFragmentAccumulator's completeness heuristic O(1)

could_close_json rescanned every preceding blank fragment on each call, so a
hostile upstream that sends malformed JSON (never closing) followed by many
blank keepalive fragments could drive that scan, and the join+parse it
gates, to O(n^2) total. Track the last non-blank fragment's trailing byte
incrementally in append/pop_next_value/set instead of rescanning the buffer.

Reported by automated review on PR #36610.

* fix(streaming): make JSONFragmentAccumulator.pop_next_value O(1) per call

pop_next_value previously rebuilt the full remaining string and sliced a
new remainder on every call, so draining N concatenated JSON values
already sitting in one buffer cost O(n^2) total. Replace the rebuild-and-
slice with a materialize-once cursor: pending fragments are joined into
the buffer only when new ones have arrived since the last pop, and
consumed values are dropped by advancing an offset instead of copying
the remaining string.

* test(streaming): make JSONFragmentAccumulator drain regression test CI-stable

The 80k-value drain test used an absolute ms budget that flaked on a
busier CI runner (233.5ms vs a 150ms budget calibrated on a quiet
machine). Replace it with a doubling-ratio check: draining twice as many
concatenated values should take roughly 2x as long for O(n), not
~4x for O(n^2), and that ratio holds regardless of machine speed.

* test(streaming): suppress TQ002 on the append-laziness spy test

A test-quality gate (TQ002: don't assert only that a mock was called)
landed upstream since this branch's last rebase and now flags
test_append_never_calls_raw_decode. The test verifies append() defers
all decoding to pop_next_value, which has no caller-observable proxy
other than spying on the stdlib call it must avoid making.

* test(vertex): spy on raw_decode instead of json.loads in accumulator regression tests

Post-migration to the shared JSONFragmentAccumulator, Vertex's decode
path goes through json.JSONDecoder.raw_decode, not json.loads. The two
O(n^2)/partial-fragment regression tests still patched json.loads, which
that path never calls, so both passed unconditionally regardless of
whether the underlying implementation regressed. Verified by simulating
an eager-reparse regression: both tests now fail against it and pass
against the correct implementation.

* fix(streaming): widen JSONFragmentAccumulator's whitespace skip to match str.strip()

pop_next_value's whitespace skip only matched json.decoder.WHITESPACE's
ASCII set, narrower than str.strip() (Unicode-aware) which the O(1)-cursor
rewrite replaced. A non-ASCII separator like U+00A0 between two
concatenated JSON values on one SSE line made raw_decode fail on it, and
the buffer never advanced past that byte again, permanently stranding
everything after it for the rest of the stream. Use str.isspace() to
match str.strip()'s tolerance.

---------

Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-25 16:15:47 -07:00
mateo-berri
4ab654cbc5 merge: litellm_internal_staging into litellm_fix_headroom_stream_leak 2026-08-25 16:04:33 -07:00