litellm/tests/integration/providers/test_websearch_interception_wire.py
devin-ai-integration[bot] 5c0b374f0a
test(integration): regression tests for August provider translation and streaming bugs (#42621)
* test(integration): Bedrock batch files upload completions and responses records as user messages (Pylon #6882)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): client Anthropic OAuth token never replaces Bedrock SigV4 authorization (Pylon #6888)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bridge /v1/messages and /v1/responses streams through empty-choices chunks (Pylon #6992)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prepend azure content-filter metadata chunk to the messages stream (Pylon #6992)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fireworks routers/ slug reaches the provider as accounts/fireworks/routers/<id> (Pylon #7030)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock hidden thinking tokens are not reported as text tokens (Pylon #7067)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): azure_ai FLUX.2-flex image generation targets the flex provider path with the BFL body (Pylon #7092)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep Databricks streaming usage and cache reads in the client stream and spend log (Pylon #7094)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): openai-compatible image edits forward provider-specific form fields to the backend (Pylon #7122)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): replayed intercepted web search turn reaches Bedrock as text through /v1/messages (Pylon #7181)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): streamed web search turn capped by max_agentic_loops ends the turn with snippets and ordered blocks (Pylon #7230)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock rerank keeps forwarded client headers out of the SigV4 signature (Pylon #7284)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): azure_ai rerank authenticates with an Entra token when no api key is set (Pylon #7303)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): perplexity stream with cost breakdown object completes and bills total_cost (Pylon #7331)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): azure_ai strips Anthropic message fields before the Foundry request (Pylon #7336)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): capped intercepted web search ends the turn without an internal tool_use block (Pylon #7378)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock passthrough converse-stream keeps event-stream content-type (Pylon #7482)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): v1/messages success exposes v3 priority rate limit headers (Pylon #7532)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): config deployment dropped by a stale boot cost map is restored after reload (Pylon #7564)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock Mantle project id reaches the provider as anthropic-workspace-id (Pylon #7583)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prefixed Opus 4.8 reasoning_effort reaches Bedrock as adaptive thinking (Pylon #7586)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): dashscope chat forwards reasoning_effort to the provider (Pylon #7606)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): failing stream logging callback still releases the max_parallel_requests slot (Pylon #7608)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): gen 5 Claude Bedrock Invoke tool search sends the Bedrock beta field (Pylon #7642)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): databricks ai gateway api_base requests OAuth token from workspace origin (Pylon #7724)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): deepseek vision image content list reaches the provider unchanged (Pylon #7729)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock Mantle context overflow surfaces as 400 prompt is too long (Pylon #7732)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Codex history items reach Bedrock Mantle as supported Responses input types (Pylon #7783)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): chat over responses deployment returns finish_reason length when output tokens run out (Pylon #7784)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock_mantle rewrites Codex history items before the Responses wire (Pylon #7812)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): advisor sub-call on /v1/messages uses the configured advisor deployment (Pylon #7828)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tencent thinking reaches the provider body instead of failing the request (Pylon #7834)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): xAI chat web search reaches /v1/responses with instructions and nested filters (Pylon #7835)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): send Bedrock Converse config blocks once at top level (Pylon #7839)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock converse sends gpt-5 reasoning_effort as reasoning.effort (Pylon #7850)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock cohere.embed-english-v3 embeddings accept encoding_format and dimensions (Pylon #7963)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): streamed chat completions emit SSE keepalive pings while the upstream is silent before its first token (Pylon #7987)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(integration): format the TTFT keepalive regression test (Pylon #7987)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): openai chat drops tool_choice when the request has no tools (Pylon #8022)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): stream whose first chunk has no choices falls back and bills the fallback (Pylon #8006)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "test(integration): Bedrock Mantle project id reaches the provider as anthropic-workspace-id (Pylon #7583)"

This reverts commit 864b65811f.

* test(integration): reconcile contracts manifest for bundled regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): isolate proxy config writes in bundled regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): address review feedback on keepalive, cost map reload and websearch order

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:52:00 -07:00

299 lines
13 KiB
Python

import json
from pathlib import Path
from typing import Final
import pytest
import yaml
from integration._support.client import Gateway
from integration._support.process import owned_proxy
from integration._support.wire import Reply, Request, wire_server
BEDROCK_MODEL: Final = "us.anthropic.claude-haiku-4-5-20251001-v1:0"
INVOKE_TARGET: Final = f"/model/{BEDROCK_MODEL}/invoke"
SEARCH_TARGET: Final = "/tavily/search"
SEARCH_RESULT: Final = {
"title": "Synthetic result",
"url": "https://example.test/result",
"content": "the snippet text",
}
def sse_events(text: str) -> tuple[tuple[str, dict[str, object]], ...]:
frames: Final = tuple(frame for frame in text.split("\n\n") if frame.strip())
return tuple(
(
next(line.removeprefix("event: ") for line in frame.splitlines() if line.startswith("event: ")),
json.loads(next(line.removeprefix("data: ") for line in frame.splitlines() if line.startswith("data: "))),
)
for frame in frames
)
@pytest.mark.covers("other.provider_wire.bedrock.websearch_interception_streamed_capped_turn_ends_with_native_results")
def test_streamed_web_search_turn_capped_by_max_agentic_loops_ends_turn_with_snippets_and_ordered_blocks(
gateway: Gateway, tmp_path: Path
) -> None:
def respond(request: Request) -> Reply:
assert request.method == "POST", request.target
body: Final = json.loads(request.body)
if request.target == SEARCH_TARGET:
assert request.headers["authorization"] == "Bearer synthetic-tavily-key"
assert body["query"] == "query-0", body
return Reply(body=json.dumps({"query": "query-0", "results": [SEARCH_RESULT]}).encode())
assert request.target == INVOKE_TARGET
assert request.headers["authorization"] == "Bearer synthetic-bedrock-token"
assert [tool["name"] for tool in body["tools"]] == ["litellm_web_search"], body["tools"]
assert "stream" not in body, body
depth: Final = sum(
1
for message in body["messages"]
if isinstance(message["content"], list)
for block in message["content"]
if block["type"] == "tool_result"
)
if depth == 1:
assert body["messages"][2]["content"] == [
{
"type": "tool_result",
"tool_use_id": "toolu_0",
"content": "Title: Synthetic result\nURL: https://example.test/result\nSnippet: the snippet text",
}
], body["messages"]
return Reply(
body=json.dumps(
{
"id": f"msg_{depth}",
"type": "message",
"role": "assistant",
"model": BEDROCK_MODEL,
"content": [
{"type": "text", "text": f"turn-{depth}"},
{
"type": "tool_use",
"id": f"toolu_{depth}",
"name": "litellm_web_search",
"input": {"query": f"query-{depth}"},
},
],
"stop_reason": "tool_use",
"stop_sequence": None,
"usage": {"input_tokens": 10, "output_tokens": 4},
}
).encode()
)
with wire_server(respond) as wire:
config: Final = yaml.safe_load(Path("tests/integration/proxy_config.yaml").read_text())
config["search_tools"] = [
{
"search_tool_name": "integration-search",
"litellm_params": {
"search_provider": "tavily",
"api_key": "synthetic-tavily-key",
"api_base": wire.url + "/tavily",
},
}
]
config["litellm_settings"].update(
{
"callbacks": ["websearch_interception"],
"websearch_interception_params": {
"enabled_providers": ["bedrock"],
"search_tool_name": "integration-search",
"max_agentic_loops": 1,
},
}
)
path: Final = tmp_path / "websearch.yaml"
path.write_text(yaml.safe_dump(config))
with owned_proxy(gateway, tmp_path, {}, config=path) as candidate, candidate.scenario() as scenario:
model: Final = scenario.model(
model=f"bedrock/{BEDROCK_MODEL}",
api_key="synthetic-bedrock-token",
api_base=wire.url,
aws_region_name="us-east-1",
aws_bedrock_runtime_endpoint=wire.url,
)
response: Final = candidate.request(
"POST",
"/v1/messages",
{
"model": model,
"max_tokens": 64,
"stream": True,
"messages": [{"role": "user", "content": "search control"}],
"tools": [{"type": "web_search_20250305", "name": "web_search"}],
},
)
assert response.status_code == 200, response.text
events: Final = sse_events(response.text)
assert [name for name, _ in events][:1] == ["message_start"], response.text
assert [name for name, _ in events][-2:] == ["message_delta", "message_stop"], response.text
for position, (name, event) in enumerate(events):
if name == "content_block_stop":
assert event["index"] in {
earlier_event["index"]
for earlier, earlier_event in events[:position]
if earlier == "content_block_start"
}, response.text
started: Final = tuple(event["content_block"] for name, event in events if name == "content_block_start")
search_ids: Final = tuple(block["id"] for block in started if block["type"] == "server_tool_use")
assert search_ids and all(search_id.startswith("srvtoolu_") for search_id in search_ids), response.text
assert started[-1] == {"type": "text", "text": ""}, response.text
assert started[:-1] == tuple(
block
for search_id in search_ids
for block in (
{"type": "server_tool_use", "id": search_id, "name": "web_search", "input": {"query": "query-0"}},
{
"type": "web_search_tool_result",
"tool_use_id": search_id,
"content": [
{
"type": "web_search_result",
"url": "https://example.test/result",
"title": "Synthetic result",
"page_age": None,
"encrypted_content": "",
"snippet": "the snippet text",
}
],
},
)
), response.text
assert (
"".join(event["delta"]["text"] for name, event in events if name == "content_block_delta") == "turn-1"
), response.text
assert [event["delta"]["stop_reason"] for name, event in events if name == "message_delta"] == [
"end_turn"
], response.text
assert "litellm_web_search" not in response.text, response.text
assert [request.target for request in wire.drain()] == [INVOKE_TARGET, SEARCH_TARGET, INVOKE_TARGET]
import threading
import uuid
from typing import Final
from urllib.parse import parse_qs, urlsplit
import httpx
import pytest
from integration._support.client import Gateway, eventually
_QUERY: Final = "integration capped search"
_TEXT_BLOCK: Final = {"type": "text", "text": "searching once more"}
_NOT_INTERCEPTED: Final = "native tool reached the provider"
_SEARCH_RESULT_BLOCK: Final = {
"type": "web_search_result",
"url": "https://owned.invalid/a",
"title": "Owned result",
"page_age": None,
"encrypted_content": "",
"snippet": "owned snippet",
}
def _search_tool_use(identity: str) -> dict[str, object]:
return {"type": "tool_use", "id": identity, "name": "litellm_web_search", "input": {"query": _QUERY}}
def _anthropic_reply(identity: str, content: list[dict[str, object]], stop_reason: str) -> Reply:
return Reply(
body=json.dumps(
{
"id": identity,
"type": "message",
"role": "assistant",
"model": "claude-sonnet-4-5-20250929",
"content": content,
"stop_reason": stop_reason,
"stop_sequence": None,
"usage": {"input_tokens": 10, "output_tokens": 4},
}
).encode()
)
@pytest.mark.covers(
"other.provider_wire.anthropic.websearch_interception_capped_loop_ends_turn_without_internal_tool_use"
)
def test_capped_websearch_interception_loop_ends_turn_instead_of_exposing_internal_tool_use(
gateway: Gateway, tmp_path: Path
) -> None:
identity: Final = "websearch-wire-" + uuid.uuid4().hex
searched: Final = threading.Event()
def respond(request: Request) -> Reply:
parts: Final = urlsplit(request.target)
if request.method == "GET" and parts.path == "/search":
assert parse_qs(parts.query)["q"] == [_QUERY], request.target
searched.set()
return Reply(
body=json.dumps(
{
"results": [
{"title": "Owned result", "url": "https://owned.invalid/a", "content": "owned snippet"}
]
}
).encode()
)
assert request.method == "POST" and parts.path == "/v1/messages", request.target
body: Final = json.loads(request.body)
if any(tool.get("type") == "web_search_20250305" for tool in body["tools"]):
return _anthropic_reply(identity, [{"type": "text", "text": _NOT_INTERCEPTED}], "end_turn")
assert [tool["name"] for tool in body["tools"]] == ["litellm_web_search"], body["tools"]
return _anthropic_reply(identity, [_TEXT_BLOCK, _search_tool_use(identity)], "tool_use")
def send(candidate: Gateway, model: str) -> httpx.Response:
return candidate.request(
"POST",
"/v1/messages",
{
"model": model,
"max_tokens": 64,
"messages": [{"role": "user", "content": identity + " attempt " + uuid.uuid4().hex}],
"tools": [{"type": "web_search_20250305", "name": "web_search", "max_uses": 3}],
},
)
def searched_through_proxy(response: httpx.Response) -> bool:
return searched.is_set() and _NOT_INTERCEPTED not in response.text
with wire_server(respond) as wire:
config: Final = yaml.safe_load(Path("tests/integration/proxy_config.yaml").read_text())
config["search_tools"] = [
{
"search_tool_name": "integration-searxng",
"litellm_params": {"search_provider": "searxng", "api_base": wire.url},
}
]
config["litellm_settings"].update(
{
"callbacks": ["websearch_interception"],
"websearch_interception_params": {
"enabled": True,
"enabled_providers": ["anthropic"],
"search_tool_name": "integration-searxng",
},
}
)
path: Final = tmp_path / "websearch.yaml"
path.write_text(yaml.safe_dump(config))
with owned_proxy(gateway, tmp_path, {}, config=path) as candidate, candidate.scenario() as scenario:
model: Final = scenario.model(
model="anthropic/claude-sonnet-4-5-20250929", api_base=wire.url, api_key="synthetic-anthropic-key"
)
response: Final = eventually(lambda: send(candidate, model), searched_through_proxy, seconds=40)
assert response.status_code == 200, response.text
body: Final = response.json()
assert body["stop_reason"] == "end_turn", response.text
content: Final = body["content"]
assert [block["type"] for block in content] == ["server_tool_use", "web_search_tool_result", "text"], (
response.text
)
assert content[0]["name"] == "web_search" and content[0]["input"] == {"query": _QUERY}, response.text
assert content[1]["tool_use_id"] == content[0]["id"], response.text
assert content[1]["content"] == [_SEARCH_RESULT_BLOCK], response.text
assert content[2] == _TEXT_BLOCK, response.text
targets: Final = tuple((request.method, urlsplit(request.target).path) for request in wire.drain())
assert targets[-3:] == (("POST", "/v1/messages"), ("GET", "/search"), ("POST", "/v1/messages")), targets