litellm/tests/integration/providers/test_responses_bridge_namespace_tools.py
devin-ai-integration[bot] a286ebf42e
test(integration): regression tests for July provider translation, routing and streaming bugs (#42693)
* test(integration): optional Anthropic tool properties stay optional on the OpenAI Responses wire (Pylon #6619)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock InvokeModel count-suffixed cache usage fields are reported and charged (Pylon #6708)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): anthropic messages honors the deployment request timeout (Pylon #6505)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop client_metadata before the Bedrock Converse body reaches the provider (Pylon #6645)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): repeat Bedrock requests under one session name assume the role once (Pylon #6681)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): messages stream keeps include_usage off the Responses wire with always_include_stream_usage (Pylon #6466)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): clamp sub-16 max_tokens to the Responses API floor instead of 400 (Pylon #6539)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): forwarded client x- headers reach the provider on /v1/responses (Pylon #6565)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Anthropic messages stop_sequences reach OpenAI-compatible providers as stop (Pylon #6536)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Codex namespace tools reach a chat upstream flattened and round-trip through /v1/responses (Pylon #6409)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep Claude 4.6 legacy thinking budget_tokens on /v1/messages (Pylon #6727)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): nvidia nim ranking keeps image passages and applies top_n without sending top_k (Pylon #6401)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tpm-only model rejects priority traffic once recorded tokens reach the model tpm (Pylon #6344)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): reasoning-only chunks open an Anthropic thinking block at index zero on /v1/messages streams (Pylon #6337)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): file content streams to the client before the upstream finishes sending (Pylon #6315)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): agent whose card lives only at agentCard/v1.0 is reached with bearer auth (Pylon #6249)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): vertex batch create returns a batch when outputInfo is null (Pylon #6374)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fireworks session id is sent as x-session-affinity and cached tokens land in spend log metadata (Pylon #6220)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): advisor sub-call failure does not cool down the executor deployment (Pylon #6212)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Gemini /v1/messages cache_control creates cachedContent with Anthropic ttl (Pylon #6221)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock_mantle max_output_tokens below 16 is clamped before reaching Mantle (Pylon #6262)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): missing thinking signature 400 on /v1/messages retries without thinking blocks (Pylon #6222)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): format Gemini messages cache_control wire test (Pylon #6221)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rebuilt shared aiohttp session keeps the configured keepalive timeout (Pylon #6387)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): sagemaker_chat signs the inference component header and sends hf_model_name as the body model (Pylon #6187)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock Converse DeepSeek drops Anthropic thinking and sends V3 reasoning_effort raw (Pylon #6149)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): concurrent team model TPM requests are reserved before the provider call (Pylon #6075)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): /v1/messages honors the configured timeout against a stalled upstream (Pylon #6025)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Codex additional_tools input items reach Bedrock Mantle as top-level tools (Pylon #6012)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): advisor api_base without api_key never sends the proxy Anthropic key to the caller host (Pylon #6226)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rerank responses carry call id, latency and cost headers (Pylon #5981)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): parse the outbound Anthropic body with the typed JSON adapter (Pylon #6025)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Bedrock Knowledge Base search forwards userContext to the Retrieve body (Pylon #5991)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): Marengo 3.0 text embeddings reach Bedrock nested under inputType (Pylon #5949)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): sub-16 max_tokens over a responses deployment reaches OpenAI as 16 (Pylon #6008)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): midturn system correction reaches the OpenAI Responses wire via /v1/messages (Pylon #6449)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): concurrent requests over a key tpm limit are rejected before reaching the provider (Pylon #5737)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): chat to responses bridge keeps deployment AWS credentials for Bedrock Mantle SigV4 (Pylon #5870)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): vertex gemini stream split across many fragments completes without stalling the proxy (Pylon #5838)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bedrock mantle /v1/messages stream keeps stream true and relays SSE events (Pylon #5596)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): large chat payloads are released from worker memory after the request ends (Pylon #5920)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): streaming success logs v3 rate limit remaining values for callbacks (Pylon #5767)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): a database created search tool backs Anthropic web search interception (Pylon #5669)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): register july provider regression contracts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): make july provider regression tests deterministic under cache and worker sharing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop order-fragile worker memory probe pending a real retention regression check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): apply ruff import sorting and formatting

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop stale contract entry and pass question to advisor executor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use tiktoken-backed executor model in advisor tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 09:51:06 -07:00

159 lines
6.1 KiB
Python

import json
import uuid
from typing import Final
import pytest
from integration._support.client import Gateway
from integration._support.wire import Reply, Request, wire_server
from pydantic import JsonValue, TypeAdapter
JSON_OBJECT: Final = TypeAdapter(dict[str, JsonValue])
JSON_LIST: Final = TypeAdapter(list[dict[str, JsonValue]])
NAMESPACE: Final = "mcp__everything"
TOOL_NAME: Final = "get_sum"
FLATTENED_NAME: Final = f"{NAMESPACE}__{TOOL_NAME}"
CALL_ID: Final = "call_synthetic_get_sum"
ARGUMENTS: Final = json.dumps({"a": 2, "b": 3})
PARAMETERS: Final[dict[str, JsonValue]] = {
"type": "object",
"required": ["a", "b"],
"properties": {"a": {"type": "number"}, "b": {"type": "number"}},
}
NAMESPACE_TOOL: Final[dict[str, JsonValue]] = {
"type": "namespace",
"name": NAMESPACE,
"description": "Tools exposed by the everything MCP server",
"tools": [
{
"type": "function",
"name": TOOL_NAME,
"description": "Adds two numbers",
"strict": False,
"parameters": PARAMETERS,
}
],
}
EXPECTED_CHAT_TOOLS: Final[list[JsonValue]] = [
{
"type": "function",
"function": {
"name": FLATTENED_NAME,
"description": "Tools exposed by the everything MCP server\n\nAdds two numbers",
"parameters": PARAMETERS,
"strict": False,
},
}
]
def tool_call_completion(marker: str) -> bytes:
return json.dumps(
{
"id": f"chatcmpl-{marker}",
"object": "chat.completion",
"created": 1789788253,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"finish_reason": "tool_calls",
"message": {
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": CALL_ID,
"type": "function",
"function": {"name": FLATTENED_NAME, "arguments": ARGUMENTS},
}
],
},
}
],
"usage": {"prompt_tokens": 30, "completion_tokens": 12, "total_tokens": 42},
}
).encode()
def text_completion(marker: str) -> bytes:
return json.dumps(
{
"id": f"chatcmpl-{marker}-final",
"object": "chat.completion",
"created": 1789788254,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"finish_reason": "stop",
"message": {"role": "assistant", "content": "The sum is 5"},
}
],
"usage": {"prompt_tokens": 40, "completion_tokens": 5, "total_tokens": 45},
}
).encode()
@pytest.mark.covers("other.provider_wire.responses_bridge.codex_namespace_tools_reach_chat_upstream_and_round_trip")
def test_codex_namespace_tool_is_flattened_for_chat_upstream_and_restored_in_responses_output(
gateway: Gateway,
) -> None:
marker: Final = uuid.uuid4().hex
prompt: Final = f"add 2 and 3 {marker}"
def chat_peer(request: Request) -> Reply:
assert request.method == "POST" and request.target == "/v1/chat/completions", request.target
body: Final = JSON_OBJECT.validate_json(request.body)
assert body["tools"] == EXPECTED_CHAT_TOOLS, body
messages: Final = JSON_LIST.validate_python(body["messages"])
if len(messages) == 1:
return Reply(body=tool_call_completion(marker))
assert messages[1]["role"] == "assistant", messages
history_calls: Final = JSON_LIST.validate_python(messages[1]["tool_calls"])
assert [(call["id"], call["function"]) for call in history_calls] == [
(CALL_ID, {"name": FLATTENED_NAME, "arguments": ARGUMENTS})
], messages
assert messages[2] == {"role": "tool", "tool_call_id": CALL_ID, "content": "5"}, messages
return Reply(body=text_completion(marker))
with wire_server(chat_peer) as wire, gateway.scenario() as scenario:
model: Final = scenario.model(model="deepseek/gpt-4o-mini", api_base=wire.url + "/v1")
first: Final = gateway.request(
"POST",
"/v1/responses",
{"model": model, "input": prompt, "tools": [NAMESPACE_TOOL], "store": False},
)
assert first.status_code == 200, first.text
first_output: Final = JSON_LIST.validate_python(JSON_OBJECT.validate_json(first.content)["output"])
calls: Final = tuple(item for item in first_output if item["type"] == "function_call")
assert len(calls) == 1, first.text
assert calls[0]["name"] == TOOL_NAME, first.text
assert calls[0]["namespace"] == NAMESPACE, first.text
assert calls[0]["call_id"] == CALL_ID, first.text
assert calls[0]["arguments"] == ARGUMENTS, first.text
second: Final = gateway.request(
"POST",
"/v1/responses",
{
"model": model,
"input": [
{"type": "message", "role": "user", "content": [{"type": "input_text", "text": prompt}]},
{
"type": "function_call",
"call_id": CALL_ID,
"name": TOOL_NAME,
"namespace": NAMESPACE,
"arguments": ARGUMENTS,
},
{"type": "function_call_output", "call_id": CALL_ID, "output": "5"},
],
"tools": [NAMESPACE_TOOL],
"store": False,
},
)
assert second.status_code == 200, second.text
second_output: Final = JSON_LIST.validate_python(JSON_OBJECT.validate_json(second.content)["output"])
assert [item["type"] for item in second_output] == ["message"], second.text
assert JSON_LIST.validate_python(second_output[0]["content"])[0]["text"] == "The sum is 5", second.text
assert len(wire.drain()) == 2