mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-05 08:07:05 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
131 lines
4.8 KiB
Python
131 lines
4.8 KiB
Python
"""web_search x Azure.
|
|
|
|
Drive the real `claude` CLI against a running LiteLLM proxy that routes
|
|
to Azure, allow the built-in `WebSearch` tool, ask a question that
|
|
requires fresh web data, and assert that the upstream emitted a
|
|
`tool_use` block calling `WebSearch` — proving the proxy preserves
|
|
Claude Code's tool definitions and the upstream's tool-use response
|
|
end-to-end.
|
|
|
|
Note: Claude Code's `WebSearch` is a *client-side* tool (the CLI
|
|
executes the search itself and feeds the result back as a `tool_result`
|
|
block), so the wire shape is `tool_use` with `name="WebSearch"` rather
|
|
than the Anthropic-managed `server_tool_use` / `web_search_tool_result`
|
|
blocks (which only appear when the request includes the
|
|
`web_search_20250305` server tool definition — something the CLI does
|
|
not currently inject). A regression where the proxy strips the
|
|
`WebSearch` tool from the request or drops the `tool_use` block from
|
|
the response will break this assertion.
|
|
|
|
The (feature, provider) for this cell is inferred from the file path by
|
|
`tests/e2e/claude_code/conftest.py`:
|
|
|
|
tests/e2e/claude_code/web_search/test_azure.py
|
|
^^^^^^^^^^ ^^^^^^^^^
|
|
feature_id provider
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Any, Mapping, Sequence
|
|
|
|
import pytest
|
|
|
|
from claude_code._env import require_proxy
|
|
from claude_code.cli_driver import (
|
|
ClaudeCLIError,
|
|
failure_diagnostic,
|
|
run_claude_models_parallel,
|
|
)
|
|
|
|
|
|
AZURE_MODELS = [
|
|
"claude-haiku-4-5-azure",
|
|
"claude-sonnet-4-5-azure",
|
|
"claude-opus-4-7-azure",
|
|
]
|
|
|
|
# A prompt the model cannot answer from training data alone — it forces
|
|
# the model to actually hit the web_search server tool rather than
|
|
# replying from memory. We pick "this week" as the freshness anchor
|
|
# because it's stable across long-running test schedules without
|
|
# pinning to a specific date that would go stale.
|
|
WEB_SEARCH_PROMPT = (
|
|
"Use web search to find a news headline published this week about "
|
|
"Anthropic. Reply with one sentence summarizing what you found."
|
|
)
|
|
# Allow only WebSearch so the model has no fallback path: if the proxy
|
|
# strips the server tool, the run will fail loudly rather than silently
|
|
# answering from training data via a different tool.
|
|
WEB_SEARCH_ARGS = ["--allowed-tools", "WebSearch"]
|
|
|
|
# The CLI tool name surfaced as `tool_use.name` when WebSearch fires.
|
|
WEB_SEARCH_TOOL_NAME = "WebSearch"
|
|
|
|
|
|
def _has_web_search_tool_use(events: Sequence[Mapping[str, Any]]) -> bool:
|
|
"""Walk the stream-json events and return True if any assistant
|
|
message included a `tool_use` block calling `WebSearch`."""
|
|
for event in events:
|
|
if event.get("type") != "assistant":
|
|
continue
|
|
message = event.get("message") or {}
|
|
content = message.get("content")
|
|
if not isinstance(content, list):
|
|
continue
|
|
for block in content:
|
|
if not isinstance(block, dict):
|
|
continue
|
|
if (
|
|
block.get("type") == "tool_use"
|
|
and block.get("name") == WEB_SEARCH_TOOL_NAME
|
|
):
|
|
return True
|
|
return False
|
|
|
|
|
|
@pytest.mark.covers("llm.messages.azure_foundry.web_search.nonstream.works")
|
|
def test_web_search_azure(compat_result):
|
|
"""Drive the `claude` CLI against the LiteLLM proxy and assert the
|
|
upstream emitted a `tool_use` block calling `WebSearch`, proving
|
|
the proxy preserved both the request-side tool definition and the
|
|
response-side tool_use block."""
|
|
base_url, api_key = require_proxy(compat_result)
|
|
|
|
outcomes = run_claude_models_parallel(
|
|
models=AZURE_MODELS,
|
|
prompt=WEB_SEARCH_PROMPT,
|
|
base_url=base_url,
|
|
api_key=api_key,
|
|
extra_args=WEB_SEARCH_ARGS,
|
|
)
|
|
|
|
failures = []
|
|
for model in AZURE_MODELS:
|
|
outcome = outcomes[model]
|
|
if isinstance(outcome, ClaudeCLIError):
|
|
error = f"[{model}] {outcome}"
|
|
compat_result.add({"status": "fail", "error": error})
|
|
failures.append(error)
|
|
continue
|
|
|
|
if outcome.exit_code != 0:
|
|
error = f"[{model}] claude CLI failed: {failure_diagnostic(outcome)}"
|
|
compat_result.add({"status": "fail", "error": error})
|
|
failures.append(error)
|
|
continue
|
|
|
|
if not _has_web_search_tool_use(outcome.events):
|
|
error = (
|
|
f"[{model}] no `tool_use` block with name=WebSearch observed; "
|
|
"the proxy may have stripped the WebSearch tool definition from "
|
|
"the request or the tool_use block from the response"
|
|
)
|
|
compat_result.add({"status": "fail", "error": error})
|
|
failures.append(error)
|
|
continue
|
|
|
|
compat_result.add({"status": "pass"})
|
|
|
|
if failures:
|
|
pytest.fail("; ".join(failures), pytrace=False)
|