Commit graph

1959 commits

Author SHA1 Message Date
mateo-berri
280844c2f5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_xai_web_search_cost_billing 2026-08-11 17:23:00 -07:00
mateo-berri
52f9b4a6e1 fix(xai): keep Responses API usage schema while billing web search
Drop the transform overrides that swapped response.usage to the chat
shape, which broke the /v1/responses client contract. Provider extras
like server_side_tool_usage_details already survive validation via
ResponseAPIUsage extra fields, so the shared usage bridge now carries
them onto the bridged chat Usage generically. The web_search_call
output gate also reads dict output items, since items that fail SDK
validation stay plain dicts, and the chat path gains billing tests.
2026-08-11 17:16:43 -07:00
Mateo Wang
7b34a6c632
Merge pull request #36507 from BerriAI/litellm_bedrock_converse_thinking_stream_fix
fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge
2026-08-11 12:23:29 -07:00
Mateo Wang
5c704055c6
Merge pull request #36502 from BerriAI/litellm_bedrock_invoke_tool_search_fix
fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages
2026-08-11 12:21:46 -07:00
mateo-berri
7d4488d2e8 refactor(bedrock): read tool search support from the model map
Record supports_tool_search on the Bedrock Claude entries in both cost
map files and have _supports_tool_search_on_bedrock read it first via
the provider-resolved capability lookup, keeping the name patterns as a
fallback for ARNs and ids the map cannot resolve. Threads the flag
through ModelInfoBase and drops a dated remark from the pattern list
2026-08-11 10:06:04 +00:00
mateo-berri
0f41365c34 fix(bedrock): forward output_config effort for application inference profile ARNs 2026-08-11 09:11:22 +00:00
Mateo Wang
1a8cd8a078
Merge pull request #34290 from eugene-yao-zocdoc/litellm_fix_anthropic_midturn_interjections
fix(anthropic): preserve midturn system corrections
2026-08-10 22:55:09 -07:00
mateo-berri
b9b200b348 fix(anthropic): keep tool exchanges intact around midturn system write-back 2026-08-10 21:43:42 -07:00
mateo-berri
929ee52b87 fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge
Claude Code drives Opus 4.7 with thinking {"type": "adaptive"} plus
output_config {"effort": "max"}. The anthropic-to-openai adapter
forwarded thinking verbatim for Claude models but dropped output_config,
and Bedrock Converse streams zero reasoningContent blocks for adaptive
thinking without an effort tier. Forward the effort subset of
output_config for Bedrock targets, accept it in the converse supported
params, and map it with the model's effort ceiling applied. Re-enable
the skipped e2e compat cell that catches this
2026-08-11 03:23:16 +00:00
mateo-berri
b5eed5e526 test(bedrock): cover the aws request param merge guard in the unit suite 2026-08-10 20:08:20 -07:00
mateo-berri
bc98c67028 fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages
Bedrock InvokeModel rejects tool_search_tool_* tool types unless the
request body carries the tool-search-tool-2025-10-19 beta. The model
allowlist gating that beta omitted Haiku 4.5 (and Opus 4.7, supported
since launch per live verification), so every tool-search request on
those models got a Bedrock 400. Add both to the allowlist and re-enable
the e2e compat cell that caught it.
2026-08-11 02:46:27 +00:00
mateo-berri
78dd81c014 Merge origin/litellm_internal_staging into litellm_fix_anthropic_midturn_interjections 2026-08-10 19:36:35 -07:00
ryan-crabbe-berri
20354bfcdc
fix(bedrock): reject Anthropic server-side web_search tool with actionable error (#36473)
* fix(bedrock): reject Anthropic server-side web_search tool with actionable error

Bedrock's Anthropic Messages endpoints cannot execute Anthropic's
server-side web_search tool, so forwarding it returns an opaque
"The provided request is not valid" 400 from Bedrock. Fail fast in the
invoke transform with an error that names the unsupported tool, the
model, and links the web search interception docs as the fix.

* refactor(bedrock): address review nits on web_search guard typing
2026-08-10 15:59:18 -07:00
Mateo Wang
9de3315dad
Merge pull request #36403 from BerriAI/litellm_model_registry_deprecation_audit
fix(model_prices): refresh deprecation dates, correct xAI pricing and add missing provider models
2026-08-10 11:26:41 -07:00
Alex Shtof
280c95ccb0
fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 (#35669)
* fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2

* ci: empty commit

---------

Co-authored-by: Alexander Shtoff <alexander.shtoff@tii.ae>
2026-08-10 10:18:13 -07:00
mateo
b5b663fd8b fix(xai): bill the above-200k tier at exactly 200k prompt tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-10 14:13:05 +00:00
Devin AI
9456564b3c fix(model_prices): refresh deprecation dates and xAI pricing from provider docs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-10 13:11:55 +00:00
Yang Yang
749a8b0701 fix(xai): keep chat Usage through Responses completions bridge
xAI already converts Responses usage to chat Usage so web_search_calls survive
cost tracking. The chat completions bridge then re-ran the Responses usage
transform and crashed on missing input_tokens. Pass through already-chat Usage
and chat-shaped dumps instead
2026-08-08 16:49:09 -07:00
Yang Yang
a9277b4b6e style: black format xAI cost calculator tests 2026-08-08 16:22:06 -07:00
Yang Yang
74100989a2 revert: remove xAI-specific web search gate from shared cost tracking
Gate web search like OpenAI (output/annotations/web_search_requests).
xAI uses server_side_tool_usage_details only for per-call cost math, with
web_search_requests mirrored in llms/xai for existing gate compatibility.
2026-08-08 16:22:06 -07:00
Yang Yang
8ad0a57387 style: drop unused pytest import in xAI responses tests 2026-08-08 16:21:49 -07:00
Yang Yang
478118ac36 test(xai): expand cost_calculator coverage for web search helpers
Add unit tests for apply_server_side_tool_usage_details_to_usage edge
cases and model_info-driven web_search per-call pricing fallbacks.
2026-08-08 16:21:49 -07:00
Yang Yang
c03a6076ce test(xai): cover Responses tool usage attach helpers
Add unit tests for server_side_tool_usage_details extraction/attach and
streaming completed-event pass-through in XAIResponsesAPIConfig.
2026-08-08 16:21:49 -07:00
Yang Yang
7aaa9358aa fix(xai): read web_search per-call rate from model_info
Use search_context_cost_per_query from the model cost map (with $5/1k
fallback) so web search billing can change via pricing JSON updates.
2026-08-08 16:21:49 -07:00
Yang Yang
8687d7372a fix(xai): gate web search cost on server_side_tool_usage_details
Treat positive web_search_calls as a web-search signal in built-in tool
cost gating, and mirror counts onto prompt_tokens_details.web_search_requests
when attaching xAI tool usage details so charges are not skipped.
2026-08-08 16:21:49 -07:00
Yang Yang
25ad7dcb41 fix(xai): bill web_search from server_side_tool_usage_details
Use usage.server_side_tool_usage_details.web_search_calls at $5/1k calls
instead of legacy num_sources_used/web_search_requests. Preserve tool usage
details through Responses usage transform for accurate response cost.
2026-08-08 16:20:09 -07:00
mateo-berri
7bffbbd1f2 refactor(vertex_ai): drop unreachable post-path status checks in batches handler
HTTPHandler.post and AsyncHTTPHandler.post call raise_for_status before returning, so the status_code != 200 branches after the create and cancel POSTs could never run. Non-2xx already surfaces as httpx.HTTPStatusError from inside the client. The checks after GETs stay: the get helpers return without raising. Tests that faked a non-raising POST response are replaced by HTTPStatusError propagation coverage.
2026-08-07 23:00:50 -07:00
Devin AI
21df36ed09 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vertex_batch_create_error_propagation 2026-08-08 05:42:05 +00:00
tin-berri
e50a42051c
fix(websearch): restore snippet text in native web_search_tool_result blocks (LIT-5315) (#36228)
* fix(websearch): restore snippet text in native web_search_tool_result blocks (LIT-5315)

The build_web_search_tool_result_block method copied url/title/page_age but
hardcoded encrypted_content to empty string, never reading SearchResult.snippet.
This left every native block content-free, forcing clients to web_fetch each
result to recover evidence—the reported symptom.

The Anthropic spec carries page text only in encrypted_content (an opaque
server-issued blob we cannot mint), so snippet is emitted as an additive key
alongside the spec fields. encrypted_content stays empty rather than holding
plaintext, which would assert encryption semantics that don't hold.

The anthropic SDK's BaseModel sets extra='allow', so the additive snippet key
survives SDK parsing. litellm has no typed model for web_search_result at all,
so nothing drops it internally. Turn-2 replay behavior is unaffected: the
empty encrypted_content already exists today.

Tests:
- Updated test_shape_with_results to assert snippet present
- Added test_snippet_carried_for_every_result to cover multi-result ordering
- Added test_missing_snippet_degrades_to_empty_string for edge case
- Mutation check: reverting source-only yields 3 test failures, restored to 117 passed

Fixes: LIT-5315
Co-Authored-By: Claude <noreply@anthropic.com>

* fix(websearch): make synthesized web_search blocks replayable by native clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(websearch): flatten a resultless replayed search block so Bedrock accepts the next turn

The flatten added for LIT-5315 bails when the replayed web_search_tool_result
carries an empty content list, but that is exactly what the interceptor emits
when a search legitimately returns nothing and when a search raises. The block
survived into the outbound body, Bedrock rejected the tag, and the conversation
died on the following turn just as it did before the flatten existed.

An empty content list has no encrypted_content to respect and no evidence to
preserve, so it flattens safely, and its paired server_tool_use goes with it.
The rendered text now says so explicitly rather than emitting a bare header.

Adds the multi-turn replay coverage that existed nowhere: the outbound Bedrock
invoke body is asserted free of both block types, parametrized over the
results-present and resultless cases, and built from the interceptor's own
builder so the fixture cannot drift from what it emits.

Resolves LIT-5320

* test(websearch): pin flatten idempotency for the agentic-loop re-entry

The agentic loop re-enters the same /v1/messages entry point for its follow-up
call and hands it the original client history, so the flatten runs again over
already-flattened messages once per iteration. Bedrock always takes that path,
since its config reports web search as natively handled and the short-circuit
is skipped.

A pass that appended the rendered text instead of replacing the block would
duplicate the evidence on every iteration and re-ship the unsupported tag, and
no existing single-pass test sees it. Mutation checked: keeping the original
block alongside the rendered text fails this test on its own.

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-08-07 17:28:55 -07:00
Yassin Kortam
ae1d1cb05e
fix(http): stop pooled clients persisting cookies on the aiohttp jar too (#36149)
#35978 stopped the pooled A2A client replaying one upstream's Set-Cookie to
another by installing a blocking policy on that client's httpx cookie jar. That
covers only one of the two jars on the request path. AiohttpTransport is the
default transport unless it is explicitly disabled, and the aiohttp ClientSession
behind it keeps its own cookie jar which no httpx-level assertion can observe, so
the leak is still live on the default path: a live proxy on that commit still
delivers agent-alpha's session cookie to agent-beta's card fetch and JSON-RPC
call.

The reason it looked fixed is that aiohttp's default CookieJar is built with
unsafe=False and refuses to store cookies for IP hosts, so a proof addressed to
127.0.0.1 comes back clean whether or not that jar is blocked.

Cookie persistence is now blocked where the clients are built rather than at one
call site: blocked_cookie_jar() gives every httpx client, async and sync, a jar
whose DefaultCookiePolicy(allowed_domains=()) rejects every domain in both
directions, and both ClientSession constructions litellm owns, the transport's
session factory and the proxy's shared startup session, get a DummyCookieJar.
LiteLLM reads a response cookie nowhere, and an explicitly supplied Cookie header
still goes out, so passthrough forwarding and an agent's extra_headers are
unaffected. The A2A-scoped policy #35978 added is removed, since it is now dead.

The two suites that drive the aiohttp session factory synchronously mock
ClientSession because a real one needs a running event loop; DummyCookieJar has
the same requirement, so they mock it for the same reason.
2026-08-07 11:05:59 -07:00
mateo-berri
8ad75e5ae9 test(bedrock): cover the batch record classifier fallbacks and pin metadata handling 2026-08-06 22:35:56 -07:00
mateo-berri
9f8f9cd64d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_wt_35675 2026-08-06 22:34:09 -07:00
Mateo Wang
281e52ac49
Merge pull request #35314 from BerriAI/devin_ai_fix_lit5034_empty_choices
fix(anthropic adapter): stop indexing choices[0] on choiceless streaming chunks
2026-08-06 21:49:41 -07:00
mateo-berri
10e2395395 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_wt_35148_merge 2026-08-06 19:17:18 -07:00
Devin AI
4068b4a2f0 chore: merge litellm_internal_staging into devin_ai_fix_lit5034_empty_choices
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-07 02:16:26 +00:00
mateo-berri
28ff7f3f0b fix(guardrails): scan function-role results and dedupe returned tools
Under scan_only_tool_results, legacy OpenAI function-role messages now count as tool results, and duplicate names among guardrail-returned tools keep only the first occurrence. CustomGuardrail.structured_messages_cover_full_request lets CrowdStrike AIDR declare that its writeback already rebuilds the whole conversation, so handlers install it as-is instead of merging it into the full message list a second time and duplicating out-of-scope rows. Lint budget ceilings ratchet down to match the tree
2026-08-06 00:52:16 -07:00
mateo-berri
2cc65a94c3 merge: litellm_internal_staging into litellm_bedrock_batch_sse_kms
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-08-06 00:09:23 -07:00
mateo-berri
2b58053561 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vertex_batch_create_error_propagation
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-08-06 00:04:30 -07:00
mateo-berri
7c2b709727 Merge branch 'litellm_internal_staging' into litellm_bedrock_batch_non_chat_records
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-08-05 23:56:23 -07:00
mateo-berri
7d745521bf fix(guardrails): merge synthesized tools under scan_only_tool_results and reject role-filtered no-op combos at init 2026-08-05 23:46:40 -07:00
mateo-berri
c2998dea75 fix(guardrails): guard tools write-back under scan_only_tool_results and warn on role-filtered no-op scans 2026-08-05 20:49:15 -07:00
mateo-berri
3c808f9c8f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_scan_only_tool_results 2026-08-05 18:55:43 -07:00
mateo-berri
d70e10982a fix(guardrails): keep tool-results-only scans off function definitions and merge scoped write-backs
Gate the OpenAI handler's tools forwarding behind scan_only_tool_results,
matching the Anthropic handler, so a tool-results-only scan can no longer
evaluate or rewrite trusted function definitions.

When a guardrail returns a replacement structured_messages list, substitute
the returned messages back into the positions their scoped originals came
from instead of installing the scoped list as the whole conversation, so
out-of-scope messages (system prompt, prior turns) survive redaction on
both the OpenAI and Anthropic paths.
2026-08-05 18:21:58 -07:00
Mateo Wang
60d9e6012c
Merge pull request #35999 from BerriAI/litellm_guardrails_v1_messages_tool_traffic
fix(guardrails): scan /v1/messages tool traffic
2026-08-05 18:08:47 -07:00
mateo-berri
2ba4e91766 feat(guardrails): add scan_only_tool_results to scope unified guardrails to tool results 2026-08-05 15:38:08 -07:00
devin-ai-integration[bot]
f54cd287a2
fix(bedrock): grant bedrock:CountTokens in OIDC session policy (#33145)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-05 14:55:46 -07:00
Yassin Kortam
9d58923065
fix(langfuse): stop a collected httpx handler from closing a shared client (#35981)
A cached HTTPHandler hands its raw httpx.Client out to consumers that keep it
for the process lifetime. When the shared client cache expires the entry on its
TTL or evicts it under the 200-entry cap, nothing references the handler, so it
is collected and its finalizer closed the client those consumers still hold.
Langfuse ingestion then failed silently on the SDK's background flush thread
until the process restarted.

A finalizer running proves only that nothing references the handler; it proves
nothing about the client. Both handlers now close the client during finalization
only when they built it and are still its sole referrer, so an unshared client is
still released promptly and a handed-out one is left alone. That keeps the
pooled-socket reclamation the finalizer was providing, which measures identical
to base over 2000 handler create-and-drop cycles.

Explicit close() stays, now gated on _owns_client so the wrapper never closes a
caller-injected client, and __aexit__ routes through it.

LangFuseLogger also keeps a reference to the handler whose client it hands the
SDK. Previously that handler was a local that went out of scope immediately,
leaving the client reachable only from the SDK. It still shares the cached
client, so no extra clients are created per logger.
2026-08-05 14:54:31 -07:00
Yassin Kortam
309e96c27b
fix(jina_ai): resolve the documented JINA_API_KEY as a fallback (#35992)
The Jina key fallback chain read JINA_AI_API_KEY three times in a row
before falling through to JINA_AI_TOKEN, so two of the four slots were
dead. Jina's own documentation publishes JINA_API_KEY, and litellm's
rerank validate_environment already tells users to set that name, but
nothing ever read it: a user who set only JINA_API_KEY got no key
resolved and Jina answered AUTH_MISSING_API_KEY.

Replace one of the repeats with JINA_API_KEY and drop the other.
JINA_AI_API_KEY stays first so no install that resolves a key today
changes which key it picks.
2026-08-05 14:13:42 -07:00
mateo-berri
bee787b4b5 fix(guardrails): scan /v1/messages tool traffic
Guardrails silently skipped three surfaces on the Anthropic Messages
path, so an agent loop driven by /v1/messages ran unguarded:

- The Anthropic input translation never walked tool_result blocks, so
  content returned by a local tool (a curl, a file read, an MCP call)
  reached the model unscanned in both the string and list content
  shapes, images inside a tool_result included.
- tool_permission only understood ModelResponse, so an Anthropic
  non-streaming response or a raw SSE stream carrying tool_use blocks
  passed through with no rule ever evaluated.
- ContentFilterGuardrail scanned inputs["texts"] but never
  inputs["tool_calls"], so the arguments a model proposes for a tool
  call went unchecked.

Tool call arguments are parsed as JSON before filtering so a MASK
action rewrites the value and leaves the payload valid JSON; non-JSON
arguments fall back to scanning the raw string. Denied tool_use blocks
are dropped from the Anthropic content array and replaced with a text
block, and stop_reason resets to end_turn when nothing tool-shaped
survives.
2026-08-05 14:11:27 -07:00
Yassin Kortam
267adf709a
fix(bedrock): sign Bedrock managed-file S3 requests with S3SigV4Auth (#35983)
S3 rebuilds the canonical request from the wire path with single
percent-encoding, which botocore models as S3SigV4Auth. Generic SigV4Auth
quotes the already encoded path a second time, so an object key holding any
character that percent-encodes was signed over %2520 while the request
carried %20, and S3 answered 403 SignatureDoesNotMatch.

The S3 logger was corrected in #35726; the Bedrock managed-files upload and
retrieval paths copied that same pre-fix pattern and were left behind, so a
configured bucket prefix with a space 403s every file upload and every file
content read.
2026-08-05 13:43:30 -07:00