Gate web search like OpenAI (output/annotations/web_search_requests).
xAI uses server_side_tool_usage_details only for per-call cost math, with
web_search_requests mirrored in llms/xai for existing gate compatibility.
Treat positive web_search_calls as a web-search signal in built-in tool
cost gating, and mirror counts onto prompt_tokens_details.web_search_requests
when attaching xAI tool usage details so charges are not skipped.
Use usage.server_side_tool_usage_details.web_search_calls at $5/1k calls
instead of legacy num_sources_used/web_search_requests. Preserve tool usage
details through Responses usage transform for accurate response cost.
HTTPHandler.post and AsyncHTTPHandler.post call raise_for_status before returning, so the status_code != 200 branches after the create and cancel POSTs could never run. Non-2xx already surfaces as httpx.HTTPStatusError from inside the client. The checks after GETs stay: the get helpers return without raising. Tests that faked a non-raising POST response are replaced by HTTPStatusError propagation coverage.
* fix(websearch): restore snippet text in native web_search_tool_result blocks (LIT-5315)
The build_web_search_tool_result_block method copied url/title/page_age but
hardcoded encrypted_content to empty string, never reading SearchResult.snippet.
This left every native block content-free, forcing clients to web_fetch each
result to recover evidence—the reported symptom.
The Anthropic spec carries page text only in encrypted_content (an opaque
server-issued blob we cannot mint), so snippet is emitted as an additive key
alongside the spec fields. encrypted_content stays empty rather than holding
plaintext, which would assert encryption semantics that don't hold.
The anthropic SDK's BaseModel sets extra='allow', so the additive snippet key
survives SDK parsing. litellm has no typed model for web_search_result at all,
so nothing drops it internally. Turn-2 replay behavior is unaffected: the
empty encrypted_content already exists today.
Tests:
- Updated test_shape_with_results to assert snippet present
- Added test_snippet_carried_for_every_result to cover multi-result ordering
- Added test_missing_snippet_degrades_to_empty_string for edge case
- Mutation check: reverting source-only yields 3 test failures, restored to 117 passed
Fixes: LIT-5315
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(websearch): make synthesized web_search blocks replayable by native clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(websearch): flatten a resultless replayed search block so Bedrock accepts the next turn
The flatten added for LIT-5315 bails when the replayed web_search_tool_result
carries an empty content list, but that is exactly what the interceptor emits
when a search legitimately returns nothing and when a search raises. The block
survived into the outbound body, Bedrock rejected the tag, and the conversation
died on the following turn just as it did before the flatten existed.
An empty content list has no encrypted_content to respect and no evidence to
preserve, so it flattens safely, and its paired server_tool_use goes with it.
The rendered text now says so explicitly rather than emitting a bare header.
Adds the multi-turn replay coverage that existed nowhere: the outbound Bedrock
invoke body is asserted free of both block types, parametrized over the
results-present and resultless cases, and built from the interceptor's own
builder so the fixture cannot drift from what it emits.
Resolves LIT-5320
* test(websearch): pin flatten idempotency for the agentic-loop re-entry
The agentic loop re-enters the same /v1/messages entry point for its follow-up
call and hands it the original client history, so the flatten runs again over
already-flattened messages once per iteration. Bedrock always takes that path,
since its config reports web search as natively handled and the short-circuit
is skipped.
A pass that appended the rendered text instead of replacing the block would
duplicate the evidence on every iteration and re-ship the unsupported tag, and
no existing single-pass test sees it. Mutation checked: keeping the original
block alongside the rendered text fails this test on its own.
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
#35978 stopped the pooled A2A client replaying one upstream's Set-Cookie to
another by installing a blocking policy on that client's httpx cookie jar. That
covers only one of the two jars on the request path. AiohttpTransport is the
default transport unless it is explicitly disabled, and the aiohttp ClientSession
behind it keeps its own cookie jar which no httpx-level assertion can observe, so
the leak is still live on the default path: a live proxy on that commit still
delivers agent-alpha's session cookie to agent-beta's card fetch and JSON-RPC
call.
The reason it looked fixed is that aiohttp's default CookieJar is built with
unsafe=False and refuses to store cookies for IP hosts, so a proof addressed to
127.0.0.1 comes back clean whether or not that jar is blocked.
Cookie persistence is now blocked where the clients are built rather than at one
call site: blocked_cookie_jar() gives every httpx client, async and sync, a jar
whose DefaultCookiePolicy(allowed_domains=()) rejects every domain in both
directions, and both ClientSession constructions litellm owns, the transport's
session factory and the proxy's shared startup session, get a DummyCookieJar.
LiteLLM reads a response cookie nowhere, and an explicitly supplied Cookie header
still goes out, so passthrough forwarding and an agent's extra_headers are
unaffected. The A2A-scoped policy #35978 added is removed, since it is now dead.
The two suites that drive the aiohttp session factory synchronously mock
ClientSession because a real one needs a running event loop; DummyCookieJar has
the same requirement, so they mock it for the same reason.
The end-to-end extras test now injects a MagicMock(spec=HTTPHandler) via
the client parameter instead of patching post on a real handler, and the
docstrings on the new regression tests are removed, addressing the
remaining Greptile review feedback.
Mirror the chat extras translation for /v1/completions, adapted to the
typed OpenAI SDK: anything completions.create() rejects (reasoning_effort,
response_format, fireworks-native extras) rides inside extra_body, which
the SDK merges server-side. Top-level reasoning_effort and response_format
are moved into extra_body (they raised TypeError before), truncate
aliases, chat_template_kwargs effort keys, and guided_* resolve into
extra_body fields, and the strip set removes the rest. Verified live:
/v1/completions rejects prompt_truncate_len, so both truncate names are
stripped on this path rather than renamed.
The http handler merges extra_body after transform_request, so a
response_format nested in an explicit extra_body would silently clobber
the explicit top-level response_format. Drop the nested copy with a
debug log so the top-level value wins, closing the precedence hole in
the guided-param native-wins path.
The constraint was 1.0, so anything above that was silently clamped down.
SCX accepts [0.0, 2.0), verified live against both GLM-5.2 and Qwen3.8
Max: 1.5, 1.99 and 1.999 all return 200, while 2.0 returns 400 with
"Temperature should be in [0.0, 2.0)"
Since the clamp is an inclusive min(), 2.0 cannot be the ceiling or it
would pass through a value the endpoint rejects. 1.99 is the practical
maximum
The clamp test now pins both ends: 2.5 comes back as 1.99, and 1.7 rides
through untouched where it used to be flattened to 1.0
Replaces the five launch models with the two that SCX.ai now leads on.
Both are live on api.scx.ai and both were verified against it for tool
calling, json_object and json_schema output, reasoning, prompt caching,
and, for Qwen3.8 Max, image input
Pricing follows SCX's published USD rates. GLM-5.2 lands at $0.55/M input
and $1.9255/M output, tracking the recent GLM-5.2 market repricing;
Qwen3.8 Max at $1.815/M and $5.4461/M sits under the only other seller of
that model, and is the first Qwen3.8 Max entry in the catalog
Also corrects a metadata bug the removed entries carried: they set
max_tokens equal to max_input_tokens, conflating the context window with
the output cap. Both new entries declare a max_output_tokens of 131072,
which is what the endpoint's own validator enforces
The Add Model placeholder moves to scx-ai/GLM-5.2 now that MiniMax-M2.7
is no longer in the catalog
Bare fireworks_ai/<slug> only resolved to accounts/fireworks/models/<slug>,
so Fireworks routers (served at accounts/fireworks/routers/<id>, e.g.
glm-latest and firerouter) could not be reached without passing the full
resource id. Add a shared resolve_fireworks_resource_name helper that maps an
explicit routers/<id> or models/<id> segment to the right resource path, keeps
the existing -fast router heuristic, and defaults bare slugs to models/ for
backward compatibility. Wire it into both the chat and text-completion
transforms, which had drifted (completion lacked router handling entirely)
Add test_abort_upstream_logs_warning_when_aclose_raises: verifies that _abort_upstream swallows and logs any exception raised by the upstream's aclose() method instead of propagating it.
Add test_enqueue_for_client_returns_false_when_already_detached: verifies that _enqueue_for_client returns False immediately without touching the queue when client_detached is already set before the call.
Add test_enqueue_for_
Strip net-new inline # blocks from streaming_iterator.py, the unit test file,
and the live-proxy regression test to comply with the no-new-comments rule.
Add test_async_sse_wrapper_aborts_upstream_when_detached_drain_cap_reached:
verifies that when the detached-drain cap is already full, the pump calls
aclose() on the upstream so the provider stops generating and billing
instead of continuing to stream while we record only the partial prefix.
Also fixes LIT001 (bare dict in AsyncIterator union) by replacing dict
with Mapping[str, object] across all three stream-type annotations, and
adds the required LIT003 reason strings to the three noqa: BLE001 directives.
The relay queue was unbounded, so a client reading a long stream more slowly
than Bedrock produced it let the pump accumulate every pending SSE chunk in
memory, and detached post-disconnect drains had no concurrency bound, so an
authenticated client could open many large streams and read slowly to pin
unbounded worker state.
Bound the relay queue and make the pump apply backpressure while the client is
connected (it blocks on a full queue, racing the disconnect signal), so a slow
reader throttles the upstream read exactly as the old direct yield did. Cap how
many detached drains run at once; over the cap a disconnected pump bills what it
collected instead of draining further. Detached-drain lifetime is otherwise
bounded by the upstream stream/read timeout. Both limits are tunable via env.
The detached pump previously caught every upstream exception (Bedrock read,
decode, provider-response, or chunk-conversion error) and terminated the
client stream normally, masking the original provider exception and its
status so downstream failure handling never ran.
Now, when the upstream fails while the client is still connected, forward the
original exception through the queue so the client-facing generator re-raises
it and the proxy's failure handling (status code, post_call_failure_hook)
runs unchanged. Only when the client has already disconnected, where there is
no one to propagate to and no failure hook will fire, fall back to salvaging
partial spend from the collected chunks.
On the /v1/messages -> bedrock/ invoke streaming path a client disconnect
raises CancelledError inside the httpx socket read, which unwinds the whole
upstream generator chain before any finally can drain it. Bedrock keeps
generating and billing the full response, so spend tracking logged only the
truncated partial the client drained (output tokens ~1-15 vs the real count)
and undercounted against AWS invocation logs.
Move the upstream read into a detached background task that fully drains the
provider stream to its terminal message_delta/message_stop and bills there.
The client-facing generator only relays chunks off a queue, so a disconnect
tears down the relay but not the pump. A client_detached event stops
enqueueing after disconnect so the queue can't grow unbounded.
Map the remaining gateway-documented effort keys: thinking as an alias
for enable_thinking (enable_thinking wins when both are present),
reasoning_budget to an integer reasoning_effort (skipped when thinking
is explicitly off), and low_effort=true to reasoning_effort=low (budget
wins when both are set). guided_json and guided_choice response_format
wrappers now include the name field (response and choice) to match the
gateway wire shape.
Under scan_only_tool_results, legacy OpenAI function-role messages now count as tool results, and duplicate names among guardrail-returned tools keep only the first occurrence. CustomGuardrail.structured_messages_cover_full_request lets CrowdStrike AIDR declare that its writeback already rebuilds the whole conversation, so handlers install it as-is instead of merging it into the full message list a second time and duplicating out-of-scope rows. Lint budget ceilings ratchet down to match the tree
OpenAI emits a reasoning output item on every reasoning turn, but only emits
reasoning_summary_text deltas when a summary was requested and actually
produced. The Anthropic /v1/messages Responses stream adapter opened the
thinking content block eagerly on response.output_item.added, so a summary-less
reasoning item surfaced as {"type": "thinking", "thinking": ""}. Clients persist
that in their session transcript and replay it on the next turn; an Anthropic
model then rejects the request with "each thinking block must contain thinking",
which is what users hit when a resumed session falls back to the default
Anthropic model.
Open the thinking block on the first non-empty summary delta instead, and only
emit content_block_stop for items that actually have an open block.
Resolves conflicts from the LIT010/LIT011 Final-enforcement lint pass
landing on litellm_internal_staging after this branch diverged. Keeps
this PR's behavior changes (float support in _UrlEncodableParams,
in-place results truncation + header stashing in
transform_search_response) and adopts the upstream Final annotations
and updated _TINYFISH_RESULT_CAP comment.
Gate the OpenAI handler's tools forwarding behind scan_only_tool_results,
matching the Anthropic handler, so a tool-results-only scan can no longer
evaluate or rewrite trusted function definitions.
When a guardrail returns a replacement structured_messages list, substitute
the returned messages back into the positions their scoped originals came
from instead of installing the scoped list as the whole conversation, so
out-of-scope messages (system prompt, prior turns) survive redaction on
both the OpenAI and Anthropic paths.
A cached HTTPHandler hands its raw httpx.Client out to consumers that keep it
for the process lifetime. When the shared client cache expires the entry on its
TTL or evicts it under the 200-entry cap, nothing references the handler, so it
is collected and its finalizer closed the client those consumers still hold.
Langfuse ingestion then failed silently on the SDK's background flush thread
until the process restarted.
A finalizer running proves only that nothing references the handler; it proves
nothing about the client. Both handlers now close the client during finalization
only when they built it and are still its sole referrer, so an unshared client is
still released promptly and a handed-out one is left alone. That keeps the
pooled-socket reclamation the finalizer was providing, which measures identical
to base over 2000 handler create-and-drop cycles.
Explicit close() stays, now gated on _owns_client so the wrapper never closes a
caller-injected client, and __aexit__ routes through it.
LangFuseLogger also keeps a reference to the handler whose client it hands the
SDK. Previously that handler was a local that went out of scope immediately,
leaving the client reachable only from the SDK. It still shares the cached
client, so no extra clients are created per logger.