The source drops the rejected field in place, so a payload shared across
tests could in principle be consumed by whichever case ran first. It does
not happen today, because the request is copied before the transform runs,
and the cases pass in reverse and async-first order alike. Building the
payload per call costs nothing and keeps that true if the copy ever goes.
Azure AI is the only provider that retries a 422 inside the translation
layer: when the endpoint rejects a field, litellm drops that field and sends
the request again, up to twice. That is the difference between a customer's
tool call working and coming back as a hard 400, and none of it was covered.
The retry loop in llm_http_handler.py is 13,419 lines of source against a
0.20 test-to-source ratio, and nothing exercised this path at all.
Drives real litellm.completion and litellm.acompletion calls against a
recorded Azure AI endpoint, so the assertions read the bytes that actually
went over the wire rather than a mock's call list. Nothing internal is
patched: respx fakes the HTTP boundary and the provider config, retry loop
and serialization are all the real ones.
Pins:
- a tool field the endpoint rejects is dropped and the call retried, and the
caller gets a normal completion
- the retry changes only the field the provider named
- a provider that keeps rejecting stops after exactly two attempts
- a rejection the provider cannot fix is not retried at all
- an extra input outside a tool is retried only when drop_params was asked for
Mutating the source confirms these bite: raising the retry cap from 2 to 3,
and making the tool-level field check always return False, each turn the
suite red.
The async cases pin the transport to httpx, because the aiohttp default
carries its own transport that an httpx-level fake cannot intercept. Without
that the two async tests reached the real Azure endpoint and failed on a 401.
Adds streaming, async, /v1/responses, and /v1/messages coverage for the
Together AI overhaul (#38233, #38248, #38230, #38265, #38275), plus the
legacy api.together.xyz host and TOGETHER_AI_API_BASE through
litellm.completion. Each new test fails under a one-line mutation of the
merged code.
Type the GPT-5.x reasoning payload with a ReadOnly TypedDict so the dict
literal satisfies the type-discipline budget, and drop the now-redundant
thinking pop (the thinking mapping is already skipped for these models).
Update the cross-region capability test to expect reasoning_effort offered
and thinking/output_config withheld for GPT-5.x on Converse.
The router's pre-content ping filter dropped AgenticAnthropicStreamingIterator's
hold-back keepalive, so a held-back turn sent the client nothing until the buffer
settled. A ping that no lifecycle frame precedes is now forwarded live, since a
fallback's message_start can still follow it without overlapping lifecycles
The proxy's cancel-refund guard checked isinstance against the iterator, but the
proxy only ever sees it behind FallbackAwareAnthropicMessagesStream and
AnthropicMessagesStreamingResponse, so a disconnect during hold-back refunded the
budget reservation anyway. Both wrappers now forward a duck-typed
has_buffered_provider_output flag, and the router wrapper follows a fallback
source so the flag tracks the stream actually being consumed
Stop advertising thinking/output_config as supported for OpenAI GPT-5.x and
skip the thinking mapping for these models, so a request combining thinking
with reasoning_effort can no longer leak a thinking block into
additionalModelRequestFields regardless of parameter order, which Bedrock
rejects with unknown_parameter.
OpenAI GPT-5.x models on Bedrock Converse expect reasoning effort under
additionalModelRequestFields as {"reasoning": {"effort": ...}}. They were
falling into the Anthropic branch and emitting a `thinking` block, which
Converse rejects with unknown_parameter.
The bedrock_converse gpt-5.6 entries were also missing supports_reasoning,
so reasoning_effort was dropped before mapping. Setting the flag lets the
existing config-driven supported-params path accept it, rather than adding
another model-name branch.
* perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic
Vertex's handle_accumulated_json_chunk and Anthropic's
_handle_accumulated_json_chunk each independently accumulated SSE fragments
into a JSON envelope with self.accumulated_json += fragment. Because the
attribute holds a live reference, CPython copies the whole prior buffer on
every fragment, making buffer assembly O(n^2) in total payload size.
Anthropic additionally had no completeness heuristic at all and called
json.loads on the whole buffer after every fragment, and could wedge forever
on two concatenated envelopes.
Add JSONFragmentAccumulator in litellm_core_utils/: fragments append to a
list in O(1), a could_close_json heuristic lets callers skip the join+parse
entirely until a value could plausibly be complete, and pop_next_value peels
one JSON value off the front of the buffer at a time using
json.JSONDecoder().raw_decode, keeping any unconsumed remainder instead of
failing on concatenated values. Migrate both providers onto it; Anthropic's
__next__/__anext__ end-of-stream handlers now delegate to
_handle_accumulated_json_chunk(is_final=True) instead of duplicating the
parse-and-reset logic inline.
Fixes#31861.
* test(streaming): close diff-coverage gaps in JSONFragmentAccumulator migration
Codecov flagged 11 uncovered lines in the migration: the accumulated_json
setter, __next__/__anext__'s end-of-stream drain branches in both providers,
and the pop_next_value "not found" path when a buffer's newest fragment ends
in "}" but is genuinely incomplete (an inner object closed, the outer one
didn't). Add targeted tests for each.
* fix(streaming): make JSONFragmentAccumulator's completeness heuristic O(1)
could_close_json rescanned every preceding blank fragment on each call, so a
hostile upstream that sends malformed JSON (never closing) followed by many
blank keepalive fragments could drive that scan, and the join+parse it
gates, to O(n^2) total. Track the last non-blank fragment's trailing byte
incrementally in append/pop_next_value/set instead of rescanning the buffer.
Reported by automated review on PR #36610.
* fix(streaming): make JSONFragmentAccumulator.pop_next_value O(1) per call
pop_next_value previously rebuilt the full remaining string and sliced a
new remainder on every call, so draining N concatenated JSON values
already sitting in one buffer cost O(n^2) total. Replace the rebuild-and-
slice with a materialize-once cursor: pending fragments are joined into
the buffer only when new ones have arrived since the last pop, and
consumed values are dropped by advancing an offset instead of copying
the remaining string.
* test(streaming): make JSONFragmentAccumulator drain regression test CI-stable
The 80k-value drain test used an absolute ms budget that flaked on a
busier CI runner (233.5ms vs a 150ms budget calibrated on a quiet
machine). Replace it with a doubling-ratio check: draining twice as many
concatenated values should take roughly 2x as long for O(n), not
~4x for O(n^2), and that ratio holds regardless of machine speed.
* test(streaming): suppress TQ002 on the append-laziness spy test
A test-quality gate (TQ002: don't assert only that a mock was called)
landed upstream since this branch's last rebase and now flags
test_append_never_calls_raw_decode. The test verifies append() defers
all decoding to pop_next_value, which has no caller-observable proxy
other than spying on the stdlib call it must avoid making.
* test(vertex): spy on raw_decode instead of json.loads in accumulator regression tests
Post-migration to the shared JSONFragmentAccumulator, Vertex's decode
path goes through json.JSONDecoder.raw_decode, not json.loads. The two
O(n^2)/partial-fragment regression tests still patched json.loads, which
that path never calls, so both passed unconditionally regardless of
whether the underlying implementation regressed. Verified by simulating
an eager-reparse regression: both tests now fail against it and pass
against the correct implementation.
* fix(streaming): widen JSONFragmentAccumulator's whitespace skip to match str.strip()
pop_next_value's whitespace skip only matched json.decoder.WHITESPACE's
ASCII set, narrower than str.strip() (Unicode-aware) which the O(1)-cursor
rewrite replaced. A non-ASCII separator like U+00A0 between two
concatenated JSON values on one SSE line made raw_decode fail on it, and
the buffer never advanced past that byte again, permanently stranding
everything after it for the rest of the stream. Use str.isspace() to
match str.strip()'s tolerance.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(rerank): emit latency and cost headers on /rerank
Thread the logging object into the rerank httpx calls and pass hidden_params through to get_custom_headers, so x-litellm-overhead-duration-ms, x-litellm-response-duration-ms, x-litellm-response-cost, x-litellm-call-id and the LITELLM_DETAILED_TIMING x-litellm-timing-* headers show up on rerank like they do on chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rerank): keep zero response cost in the /rerank cost header
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: assign the new rerank endpoint tests to the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: suppress TQ008 on the rerank header tests with reasons
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Tools passed for Together models with no supports_function_calling entry in
model_prices_and_context_window.json were rejected with UnsupportedParamsError
by default and silently dropped under drop_params, which made models emit tool
calls as plain text. Fail open instead: pass tool params through with a warning
and let Together validate. Models the registry explicitly marks as not
supporting function calling keep the loud contract: raise by default, drop with
a warning under drop_params.
anthropic_messages goes through _ageneric_api_call_with_fallbacks rather
than _acompletion, so its returned streaming iterator was never wrapped
by the chat-completions fallback handler. A retriable SSE event: error
frame (overloaded_error, internal_server_error) from a native
Anthropic/Bedrock passthrough passed through to the client unchanged,
and a MidStreamFallbackError raised by the completion-bridge path's
CustomStreamWrapper propagated unhandled.
Add _aanthropic_messages_streaming_iterator, mirroring
_acompletion_streaming_iterator: it detects a retriable SSE error event
via the new parse_anthropic_error_event helper, raises
MidStreamFallbackError once real generated content (a content_block_delta
frame) has not yet reached the caller, and re-enters the Router's
fallback chain. A MidStreamFallbackError raised directly by the source
iterator (the completion-bridge path) is gated the same way via its own
is_pre_first_chunk flag. The raised MidStreamFallbackError carries a
status-coded original_exception built from the parsed error type, so
status_code/cooldown logic sees the real 429/500/503/etc. instead of a
hardcoded 503.
Lifecycle/bookkeeping frames (message_start, content_block_start, ping,
...) never disqualify a fallback attempt by themselves, since Anthropic
routinely sends message_start before an overload error - but they are
buffered rather than forwarded immediately, since forwarding one and
then appending a fallback attempt's own message_start would produce two
overlapping message lifecycles on one SSE stream. Buffered frames flush,
in order, once real content arrives or the stream ends without error.
Once real content has streamed, or the error is a non-retriable 4xx, the
chunk (or exception) is forwarded as-is rather than starting a second
lifecycle. Content and error coalesced into a single physical read are
handled the same way: once the client has genuinely received the content
(bundled in that same forwarded chunk), no fallback is attempted. A
`ping` keepalive is dropped outright before any real content arrives
(it recurs indefinitely on a slow-starting connection and carries
nothing worth buffering), and the pre-content lifecycle buffer is capped
at MAX_BUFFERED_PRE_CONTENT_ANTHROPIC_CHUNKS, forcing an early commit to
the primary stream so a hostile or pathological upstream can't grow it
without bound. is_anthropic_ping_chunk only matches a chunk whose every
event: line is event: ping, so a ping coalesced with real content or a
retriable error into one physical transport chunk is never dropped.
The fallback request kwargs also deep-copy nested litellm_metadata/metadata
(matching the Responses API path) so the primary attempt's
deployment-specific fields never leak into the fallback request, and the
fallback deployment's own provider headers are merged onto the wrapper's
_hidden_params so they still reach the client/logging pipeline. A
fallback that resolves to a non-streaming response (e.g. an agentic
tool-use interception loop) is synthesized into a real Anthropic SSE
event sequence via the new anthropic_messages_response_as_sse_events
helper, instead of yielding a raw dict into the byte stream - including
a trailing signature_delta for a thinking block, and a message_start
whose stop_reason/stop_sequence/output_tokens stay null/zero the way a
real stream's does instead of leaking the completed response's final
state.
Resolves#24004