Commit graph

2292 commits

Author SHA1 Message Date
mateo-berri
910c41f154 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_branchless_provider_status_mapping
# Conflicts:
#	tests/test_litellm/litellm_core_utils/test_exception_mapping_utils.py
2026-08-26 00:55:42 -07:00
mateo-berri
53037c34ed fix(exceptions): map upstream status codes for providers with no exception_type branch 2026-08-26 00:47:28 -07:00
mateo-berri
5cfc1608f9 fix(anthropic): map only provider failures on the /v1/messages boundary 2026-08-26 00:31:18 -07:00
mateo-berri
3fe65029bf fix(exceptions): map anthropic 403 to PermissionDeniedError
Now that /v1/messages routes provider failures through exception_type, an
Anthropic permission_error fell through the anthropic branch to the generic
APIConnectionError and reached the client as a 500 where the raw exception
used to answer 403. Map 403 to PermissionDeniedError so the status survives
on every route.
2026-08-26 00:09:10 -07:00
mateo-berri
666648d58c fix(otel): map /v1/messages provider errors before failure logging 2026-08-25 23:31:05 -07:00
Mateo Wang
ea6ac3dd99
Merge pull request #38283 from BerriAI/litellm_together_regression_tests
test(together_ai): regression suite across chat, responses, and messages surfaces
2026-08-25 17:58:14 -07:00
Mateo Wang
d0c527f3bb
Merge pull request #36245 from BerriAI/litellm_fix_headroom_stream_leak
fix(anthropic): buffer streamed responses carrying server-fulfilled tools so retrieval tool calls never reach the client
2026-08-25 17:48:31 -07:00
mateo-berri
602bf1c0ac test(together_ai): close the injected async client after the streaming test 2026-08-25 17:45:54 -07:00
mateo-berri
7c3da429c7 test(together_ai): regression suite across chat, responses, and messages surfaces
Adds streaming, async, /v1/responses, and /v1/messages coverage for the
Together AI overhaul (#38233, #38248, #38230, #38265, #38275), plus the
legacy api.together.xyz host and TOGETHER_AI_API_BASE through
litellm.completion. Each new test fails under a one-line mutation of the
merged code.
2026-08-25 17:40:41 -07:00
mateo-berri
8d877980e0 fix(caching): carry has_buffered_provider_output through the anthropic stream cache writer 2026-08-25 17:31:46 -07:00
mateo-berri
1b37397633 fix(router): forward the hold-back keepalive ping live and carry the withheld-output flag through the stream wrappers
The router's pre-content ping filter dropped AgenticAnthropicStreamingIterator's
hold-back keepalive, so a held-back turn sent the client nothing until the buffer
settled. A ping that no lifecycle frame precedes is now forwarded live, since a
fallback's message_start can still follow it without overlapping lifecycles

The proxy's cancel-refund guard checked isinstance against the iterator, but the
proxy only ever sees it behind FallbackAwareAnthropicMessagesStream and
AnthropicMessagesStreamingResponse, so a disconnect during hold-back refunded the
budget reservation anyway. Both wrappers now forward a duck-typed
has_buffered_provider_output flag, and the router wrapper follows a fallback
source so the flag tracks the stream actually being consumed
2026-08-25 16:39:14 -07:00
Mateo Wang
1ff615c335
Merge pull request #38275 from BerriAI/litellm_together_chat_template_kwargs
fix(together_ai): strip internal thinking fields from outbound messages, keep reasoning_content
2026-08-25 16:15:50 -07:00
Deepanshu Lulla
6684bb343b
perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic (#36610)
* perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic

Vertex's handle_accumulated_json_chunk and Anthropic's
_handle_accumulated_json_chunk each independently accumulated SSE fragments
into a JSON envelope with self.accumulated_json += fragment. Because the
attribute holds a live reference, CPython copies the whole prior buffer on
every fragment, making buffer assembly O(n^2) in total payload size.
Anthropic additionally had no completeness heuristic at all and called
json.loads on the whole buffer after every fragment, and could wedge forever
on two concatenated envelopes.

Add JSONFragmentAccumulator in litellm_core_utils/: fragments append to a
list in O(1), a could_close_json heuristic lets callers skip the join+parse
entirely until a value could plausibly be complete, and pop_next_value peels
one JSON value off the front of the buffer at a time using
json.JSONDecoder().raw_decode, keeping any unconsumed remainder instead of
failing on concatenated values. Migrate both providers onto it; Anthropic's
__next__/__anext__ end-of-stream handlers now delegate to
_handle_accumulated_json_chunk(is_final=True) instead of duplicating the
parse-and-reset logic inline.

Fixes #31861.

* test(streaming): close diff-coverage gaps in JSONFragmentAccumulator migration

Codecov flagged 11 uncovered lines in the migration: the accumulated_json
setter, __next__/__anext__'s end-of-stream drain branches in both providers,
and the pop_next_value "not found" path when a buffer's newest fragment ends
in "}" but is genuinely incomplete (an inner object closed, the outer one
didn't). Add targeted tests for each.

* fix(streaming): make JSONFragmentAccumulator's completeness heuristic O(1)

could_close_json rescanned every preceding blank fragment on each call, so a
hostile upstream that sends malformed JSON (never closing) followed by many
blank keepalive fragments could drive that scan, and the join+parse it
gates, to O(n^2) total. Track the last non-blank fragment's trailing byte
incrementally in append/pop_next_value/set instead of rescanning the buffer.

Reported by automated review on PR #36610.

* fix(streaming): make JSONFragmentAccumulator.pop_next_value O(1) per call

pop_next_value previously rebuilt the full remaining string and sliced a
new remainder on every call, so draining N concatenated JSON values
already sitting in one buffer cost O(n^2) total. Replace the rebuild-and-
slice with a materialize-once cursor: pending fragments are joined into
the buffer only when new ones have arrived since the last pop, and
consumed values are dropped by advancing an offset instead of copying
the remaining string.

* test(streaming): make JSONFragmentAccumulator drain regression test CI-stable

The 80k-value drain test used an absolute ms budget that flaked on a
busier CI runner (233.5ms vs a 150ms budget calibrated on a quiet
machine). Replace it with a doubling-ratio check: draining twice as many
concatenated values should take roughly 2x as long for O(n), not
~4x for O(n^2), and that ratio holds regardless of machine speed.

* test(streaming): suppress TQ002 on the append-laziness spy test

A test-quality gate (TQ002: don't assert only that a mock was called)
landed upstream since this branch's last rebase and now flags
test_append_never_calls_raw_decode. The test verifies append() defers
all decoding to pop_next_value, which has no caller-observable proxy
other than spying on the stdlib call it must avoid making.

* test(vertex): spy on raw_decode instead of json.loads in accumulator regression tests

Post-migration to the shared JSONFragmentAccumulator, Vertex's decode
path goes through json.JSONDecoder.raw_decode, not json.loads. The two
O(n^2)/partial-fragment regression tests still patched json.loads, which
that path never calls, so both passed unconditionally regardless of
whether the underlying implementation regressed. Verified by simulating
an eager-reparse regression: both tests now fail against it and pass
against the correct implementation.

* fix(streaming): widen JSONFragmentAccumulator's whitespace skip to match str.strip()

pop_next_value's whitespace skip only matched json.decoder.WHITESPACE's
ASCII set, narrower than str.strip() (Unicode-aware) which the O(1)-cursor
rewrite replaced. A non-ASCII separator like U+00A0 between two
concatenated JSON values on one SSE line made raw_decode fail on it, and
the buffer never advanced past that byte again, permanently stranding
everything after it for the rest of the stream. Use str.isspace() to
match str.strip()'s tolerance.

---------

Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-25 16:15:47 -07:00
mateo-berri
4ab654cbc5 merge: litellm_internal_staging into litellm_fix_headroom_stream_leak 2026-08-25 16:04:33 -07:00
mateo-berri
8a61b286d1 test(together_ai): annotate preserved-thinking test helper and capture list 2026-08-25 16:03:43 -07:00
devin-ai-integration[bot]
bb22742025
fix(rerank): emit latency and cost headers on /rerank (#35419)
* fix(rerank): emit latency and cost headers on /rerank

Thread the logging object into the rerank httpx calls and pass hidden_params through to get_custom_headers, so x-litellm-overhead-duration-ms, x-litellm-response-duration-ms, x-litellm-response-cost, x-litellm-call-id and the LITELLM_DETAILED_TIMING x-litellm-timing-* headers show up on rerank like they do on chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rerank): keep zero response cost in the /rerank cost header

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: assign the new rerank endpoint tests to the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: suppress TQ008 on the rerank header tests with reasons

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-08-25 15:54:25 -07:00
mateo-berri
1ba66b1555 fix(together_ai): strip internal thinking fields from outbound messages, pin chat_template_kwargs passthrough 2026-08-25 15:48:58 -07:00
Mateo Wang
04818a3554
Merge pull request #38267 from BerriAI/litellm_fix_responses_user_document_drop
fix(anthropic): carry user-content document blocks through the /v1/messages responses bridge
2026-08-25 15:20:30 -07:00
Mateo Wang
e4ff44f623
Merge pull request #38265 from BerriAI/litellm_together_tools_passthrough
fix(together_ai): pass tools through for models missing from the registry
2026-08-25 15:12:08 -07:00
Mateo Wang
dc40377959
Merge pull request #38261 from BerriAI/litellm_fix_responses_tool_result_document_drop
fix(anthropic): carry tool_result document blocks through the /v1/messages responses bridge
2026-08-25 15:09:23 -07:00
mateo-berri
61ee2f5ecb fix(anthropic): carry user-content document blocks through the /v1/messages responses bridge 2026-08-25 15:00:57 -07:00
Mateo Wang
9962f4fba8
Merge pull request #38251 from BerriAI/litellm_fix_tool_result_document_drop
fix(anthropic): translate tool_result document blocks in the /v1/messages bridge
2026-08-25 14:54:01 -07:00
mateo-berri
ec365a3c12 fix(together_ai): pass tools through for models missing from the registry
Tools passed for Together models with no supports_function_calling entry in
model_prices_and_context_window.json were rejected with UnsupportedParamsError
by default and silently dropped under drop_params, which made models emit tool
calls as plain text. Fail open instead: pass tool params through with a warning
and let Together validate. Models the registry explicitly marks as not
supporting function calling keep the loud contract: raise by default, drop with
a warning under drop_params.
2026-08-25 14:38:34 -07:00
Deepanshu Lulla
75c4565dde
fix(cerebras): add max_retries and extra_headers to get_supported_openai_params (#36601)
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-25 14:10:56 -07:00
mateo-berri
c3c9e3a919 test(anthropic): pin dropped file-id and empty-url document sources to string output 2026-08-25 13:58:37 -07:00
Yassin Kortam
c70b911122
fix(router): support mid-stream fallback for anthropic_messages route type (#38153)
anthropic_messages goes through _ageneric_api_call_with_fallbacks rather
than _acompletion, so its returned streaming iterator was never wrapped
by the chat-completions fallback handler. A retriable SSE event: error
frame (overloaded_error, internal_server_error) from a native
Anthropic/Bedrock passthrough passed through to the client unchanged,
and a MidStreamFallbackError raised by the completion-bridge path's
CustomStreamWrapper propagated unhandled.

Add _aanthropic_messages_streaming_iterator, mirroring
_acompletion_streaming_iterator: it detects a retriable SSE error event
via the new parse_anthropic_error_event helper, raises
MidStreamFallbackError once real generated content (a content_block_delta
frame) has not yet reached the caller, and re-enters the Router's
fallback chain. A MidStreamFallbackError raised directly by the source
iterator (the completion-bridge path) is gated the same way via its own
is_pre_first_chunk flag. The raised MidStreamFallbackError carries a
status-coded original_exception built from the parsed error type, so
status_code/cooldown logic sees the real 429/500/503/etc. instead of a
hardcoded 503.

Lifecycle/bookkeeping frames (message_start, content_block_start, ping,
...) never disqualify a fallback attempt by themselves, since Anthropic
routinely sends message_start before an overload error - but they are
buffered rather than forwarded immediately, since forwarding one and
then appending a fallback attempt's own message_start would produce two
overlapping message lifecycles on one SSE stream. Buffered frames flush,
in order, once real content arrives or the stream ends without error.
Once real content has streamed, or the error is a non-retriable 4xx, the
chunk (or exception) is forwarded as-is rather than starting a second
lifecycle. Content and error coalesced into a single physical read are
handled the same way: once the client has genuinely received the content
(bundled in that same forwarded chunk), no fallback is attempted. A
`ping` keepalive is dropped outright before any real content arrives
(it recurs indefinitely on a slow-starting connection and carries
nothing worth buffering), and the pre-content lifecycle buffer is capped
at MAX_BUFFERED_PRE_CONTENT_ANTHROPIC_CHUNKS, forcing an early commit to
the primary stream so a hostile or pathological upstream can't grow it
without bound. is_anthropic_ping_chunk only matches a chunk whose every
event: line is event: ping, so a ping coalesced with real content or a
retriable error into one physical transport chunk is never dropped.

The fallback request kwargs also deep-copy nested litellm_metadata/metadata
(matching the Responses API path) so the primary attempt's
deployment-specific fields never leak into the fallback request, and the
fallback deployment's own provider headers are merged onto the wrapper's
_hidden_params so they still reach the client/logging pipeline. A
fallback that resolves to a non-streaming response (e.g. an agentic
tool-use interception loop) is synthesized into a real Anthropic SSE
event sequence via the new anthropic_messages_response_as_sse_events
helper, instead of yielding a raw dict into the byte stream - including
a trailing signature_delta for a thinking block, and a message_start
whose stop_reason/stop_sequence/output_tokens stay null/zero the way a
real stream's does instead of leaking the completed response's final
state.

Resolves #24004
2026-08-25 13:40:46 -07:00
mateo-berri
5998c0d1d2 fix(anthropic): carry tool_result document blocks through the /v1/messages responses bridge 2026-08-25 13:32:26 -07:00
Mateo Wang
ece03ceafe
Merge pull request #38248 from BerriAI/litellm_together_chat_config
fix(together_ai): route chat completions through a dedicated TogetherAIChatConfig
2026-08-25 12:55:31 -07:00
Mateo Wang
c6082c5255
Merge pull request #37897 from BerriAI/litellm_reasoning_effort_capability_v2
feat(router): per-group supported reasoning efforts with the max level
2026-08-25 12:41:21 -07:00
ryan-crabbe-berri
3db6c5ab18
Merge pull request #37770 from BerriAI/litellm_fix_model_router_spend_log_model
fix(proxy): store the actual selected model in spend logs for Azure Model Router
2026-08-25 12:35:35 -07:00
mateo-berri
17845b4fb0 fix(anthropic): translate tool_result document blocks in the /v1/messages bridge 2026-08-25 12:20:51 -07:00
mateo-berri
5b2d1874b0 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_reasoning_effort_capability_v2
# Conflicts:
#	tests/test_litellm/test_router.py
2026-08-25 12:07:05 -07:00
mateo-berri
636ae5d3c7 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mantle_codex_input_items 2026-08-25 11:56:13 -07:00
Mateo Wang
c543461297
Merge pull request #38225 from BerriAI/litellm_fix_bedrock_mantle_gpt5_context_window
fix(model_prices): raise bedrock_mantle gpt-5.6 max_input_tokens to Mantle's enforced 1050000
2026-08-25 11:51:19 -07:00
Mateo Wang
18108ecc24
Merge pull request #38231 from BerriAI/litellm_fix_bedrock_mantle_passthrough_invoke
fix(bedrock_mantle): register a Bedrock runtime passthrough config so /bedrock/model/<deployment>/invoke works
2026-08-25 11:51:03 -07:00
mateo-berri
62ec3b6116 fix(together_ai): route chat completions through a dedicated TogetherAIChatConfig 2026-08-25 11:44:06 -07:00
mateo-berri
e8bdbcd1cf fix(bedrock_mantle): parse converse passthrough bodies with the converse shape config for logging 2026-08-25 10:56:55 -07:00
ryan-crabbe-berri
c04b5dba32 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_model_router_spend_log_model
# Conflicts:
#	tests/test_litellm/litellm_core_utils/test_litellm_logging.py
#	tests/test_litellm/llms/azure_ai/chat/test_azure_ai_transformation.py
2026-08-25 10:55:53 -07:00
ryan-crabbe-berri
0ea6f5e159 fix(azure_ai): stamp the model router's selected model instead of matching on the model name
The model Azure Model Router served was recovered by checking whether the text
"model_router" or "model-router" appeared in a model string. Spend logs applied that
check to the litellm model path, where the route prefix guarantees a match, but the
proxy applied it to the client's model group alias, which carries no prefix. A model
group named anything else therefore lost the selected model in both the response and
the spend row.

AzureModelRouterConfig now stamps the served model onto _hidden_params, and the spend
log payload and the proxy's response restamping read that stamp. The name heuristic
survives as a fallback for callers with no response in hand, routed through
get_azure_ai_route so it lives in one place.
2026-08-25 10:50:11 -07:00
Mateo Wang
751976db8b
Merge pull request #38229 from BerriAI/litellm_vertex_ai_interactions
feat(vertex_ai): add native Vertex AI Interactions API support
2026-08-25 10:50:02 -07:00
mateo-berri
6be000f1f3 test(bedrock_mantle): type _repo_cost_map return instead of bare dict 2026-08-25 10:23:12 -07:00
mateo-berri
68ad575fc2 fix(bedrock_mantle): register a Bedrock runtime passthrough config so /bedrock/model/<deployment>/invoke works 2026-08-25 10:10:38 -07:00
Mateo Wang
9dff9cdd9a
Merge pull request #37979 from BerriAI/litellm_lit5714_adaptive_thinking_display
fix(anthropic/bedrock): request summarized adaptive thinking for reasoning_effort and use provider thinking token counts
2026-08-25 09:58:03 -07:00
mateo-berri
ed28581d79 fix(bedrock_mantle): normalize Codex input item types Mantle rejects
Mantle 400s ("Invalid 'input': value did not match any expected variant")
on the Codex history item types agent_message, context_compaction, and
local_shell_call, killing every Codex multi-agent session on the first
sub-agent turn. Rewrite agent_message into an assistant output_text message
(preserving encrypted_content slot payloads, which carry the plaintext task
through Mantle), context_compaction into Mantle's supported compaction
spelling, and local_shell_call into the function_call its recorded
function_call_output already pairs with.
2026-08-25 09:57:25 -07:00
mateo-berri
530dab32b9 feat(vertex_ai): add native Vertex AI Interactions API support 2026-08-25 09:55:13 -07:00
mateo-berri
f583151a5b fix(model_prices): raise bedrock_mantle gpt-5.6 max_input_tokens to Mantle's enforced 1050000 2026-08-25 09:53:19 -07:00
mateo-berri
7dc5a1682d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_reasoning_effort_capability_v2 2026-08-25 09:28:55 -07:00
Anmol Jaiswal
bb27bfd9a7
fix(http_handler): dispose aiohttp session when AsyncHTTPHandler is finalized without a running loop (#36670)
* fix(http_handler): dispose aiohttp session when finalized without a running loop

AsyncHTTPHandler.__del__ can only schedule an async close when a running
event loop exists at finalization time; in any other context (worker
threads whose loop has closed, sync contexts, interpreter shutdown) the
RuntimeError from get_running_loop() is swallowed and the underlying
aiohttp ClientSession is abandoned to GC, emitting 'Unclosed client
session' / 'Unclosed connector' warnings.

This is the disposal gap left after the recycle-time fix: clients created
for short-lived event loops (the loop-id-keyed LLM client cache mints one
handler per loop) are never recycled - they live and die with their loop,
and their finalization is precisely the loop-less case.

Fix:
- no running loop: fall back to the connector's synchronous teardown via
  LiteLLMAiohttpTransport._mark_connector_closed - the same finalizer-safe
  path used for dead-loop recycles - honoring _owns_session so a shared
  session is never closed.
- running loop: keep the async close, but hold a strong reference to the
  scheduled task until it completes (a bare create_task() result may be
  collected before running), mirroring _background_close_tasks.

Tests: loop-less finalization closes a dead-loop session; running-loop
finalization registers and drains the close task; the sync fallback
respects session ownership. All three fail without the fix.

* lint: conform new finalizer code to the type-discipline budget

Final on the five never-rebound locals (LIT010); the class-level task
registry keeps its mutable set with the sanctioned mutable-ok reason,
mirroring the aiohttp transport's registry (LIT001).

* lint: reasoned pyright ignore on the cross-class teardown call

The handler deliberately reuses the transport's finalizer-safe connector
teardown; no public seam exists and an async close can never run at
loop-less finalization. Clears the net-new reportPrivateUsage the
basedpyright budget gate flagged once the LIT stage passed.

* fix(http_handler): retrieve exceptions from finalizer close tasks

A bare discard done-callback dropped the task without consuming its
exception, so a failing aclose() emitted "Task exception was never
retrieved" at GC, the same noise class this path exists to remove.
Mirror the transport's _on_close_task_done: discard, early-return on
cancellation, retrieve and debug-log the exception.

* fix(http_handler): dispose foreign-loop sessions instead of scheduling aclose on the live loop

GC on a live loop (e.g. the app's) of a handler whose session belongs to
another, possibly dead, loop scheduled aclose() on the current loop, the
cross-loop path the transport refuses. Route both that case and the
loop-less case through the transport's lifecycle-aware
_close_recycled_session, which picks async close on the session's own
loop, threadsafe handoff, or the synchronous connector teardown.

Regression test: a dead-loop session collected while another loop runs
is disposed without scheduling anything on that loop.

* chore: retrigger CI (test_mcp_logging payload-order flake, also failed on litellm_spendlogs_fallback_metadata minutes earlier)

* test(mcp): select the MCP tool-call payload instead of the last-delivered one

TestMCPLogger kept a single last-writer slot; an async success event from
another call (a mocked acompletion whose log task lands late) races the
MCP event for it, so the cost assertions intermittently read the wrong
payload. This PR's finalizer change shifts task interleaving on the loop
and tips that latent race over (also seen on an unrelated PR minutes
earlier). Collect call_type=call_mcp_tool payloads in their own list and
assert on those.

* test(mcp): MCPLoggerHook inherits the order-independent payload capture

It duplicated TestMCPLogger's init and success handler verbatim; the
hook test reads the same MCP payload selection, so subclass instead.
2026-08-25 08:12:10 -07:00
Mateo Wang
da91d4b6c9
Merge pull request #38119 from BerriAI/litellm_bing_grounding_search_provider
feat(search): add Grounding with Bing Search (bing_grounding) as a search provider
2026-08-24 19:33:23 -07:00
mateo-berri
7d0df4a062 fix(videos): forward uploaded source file on /v1/videos/edits to the provider
The video edit endpoint parsed the multipart body but dropped the uploaded
source video, only normalizing it to an id. When a raw file is uploaded it now
flows through videos.main -> the http handler -> the provider transform, which
emits multipart/form-data with the source video as a file part, matching the
official OpenAI SDK's videos.edit wire format. Edit-by-id still egresses JSON.
2026-08-24 15:34:21 -07:00