mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
155 lines
6 KiB
Python
155 lines
6 KiB
Python
"""Shared body for the `basic_messaging_*` × <provider> compat cells.
|
||
|
||
Every basic_messaging cell follows the same skeleton:
|
||
|
||
1. Read the proxy base URL + API key from env, fail-early if missing.
|
||
2. Fan the three Claude tiers out via `run_claude_models_parallel`.
|
||
3. Inspect each model's outcome and report one `compat_result` row per
|
||
model — `ClaudeCLIError`, non-zero exit, and empty assistant text
|
||
are all per-model fails; everything else is a per-model pass.
|
||
4. Surface a joined failure message via `pytest.fail(...)` so the
|
||
pytest run also goes red.
|
||
|
||
The streaming variant additionally passes `verify_streaming=True`,
|
||
which adds the `--include-partial-messages` CLI flag and asserts that
|
||
the proxy actually streamed the response (see the helper docstring for
|
||
the wire-level rationale).
|
||
|
||
The conftest infers `(feature_id, provider)` purely from the test file
|
||
path, so each per-provider file just declares its model list and calls
|
||
`run_basic_messaging_cell(...)`. This keeps all cell logic in one place
|
||
— a future tweak to the env-missing guard or the failure-loop shape
|
||
now propagates to every cell automatically.
|
||
|
||
The leading underscore in the filename is what keeps pytest from
|
||
collecting this module as a test file.
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from typing import Any, Callable, Mapping, Sequence
|
||
|
||
import pytest
|
||
|
||
from claude_code._env import require_proxy
|
||
from claude_code.cli_driver import (
|
||
ClaudeCLIError,
|
||
DriverResult,
|
||
failure_diagnostic,
|
||
run_claude_models_parallel,
|
||
)
|
||
|
||
|
||
ClaudeRunner = Callable[..., Mapping[str, DriverResult | ClaudeCLIError]]
|
||
|
||
# Floor on the number of `stream_event` records (with delta payloads)
|
||
# we expect to see when the proxy actually streams. With
|
||
# `--include-partial-messages`, the CLI emits one `stream_event` per
|
||
# raw upstream SSE event — a fully-streamed response produces many
|
||
# (`message_start`, multiple `content_block_delta`s, `content_block_stop`,
|
||
# `message_delta`, `message_stop`); a proxy that buffers the upstream
|
||
# and returns a single non-streaming chunk produces 0 or 1. Floor of 2
|
||
# is safely above the buffered case for any non-trivial reply, which
|
||
# is why the streaming cells use a "count from 1 to 5" style prompt.
|
||
MIN_STREAM_DELTA_EVENTS = 2
|
||
|
||
|
||
def _count_stream_event_deltas(events: Sequence[Mapping[str, Any]]) -> int:
|
||
"""Count `stream_event` records that carry an SSE event payload.
|
||
|
||
With `--include-partial-messages`, Claude Code wraps every upstream
|
||
SSE event in a `{"type": "stream_event", "event": {...}}` record.
|
||
A buffering proxy collapses the upstream stream into a single
|
||
non-streaming response, so these records vanish. Counting them
|
||
(rather than just `len(events)`) is the wire-level signal that
|
||
"did the proxy preserve streaming?" — independent of the `system`
|
||
/`assistant`/`result` boilerplate records the CLI always emits.
|
||
"""
|
||
count = 0
|
||
for event in events:
|
||
if event.get("type") != "stream_event":
|
||
continue
|
||
if isinstance(event.get("event"), Mapping):
|
||
count += 1
|
||
return count
|
||
|
||
|
||
def run_basic_messaging_cell(
|
||
*,
|
||
compat_result,
|
||
models: Sequence[str],
|
||
prompt: str,
|
||
verify_streaming: bool = False,
|
||
env: Mapping[str, str] | None = None,
|
||
runner: ClaudeRunner = run_claude_models_parallel,
|
||
) -> None:
|
||
"""Run the shared `basic_messaging_*` × <provider> cell body.
|
||
|
||
When ``verify_streaming=True``, the cell additionally asserts that
|
||
the proxy streamed the response end-to-end. The check works by
|
||
passing ``--include-partial-messages`` to the `claude` CLI, which
|
||
causes it to emit one ``stream_event`` record per raw upstream SSE
|
||
event (``message_start``, ``content_block_delta``, ``message_stop``,
|
||
etc.). A proxy that buffers the upstream stream and returns a
|
||
single non-streaming response collapses those records to zero —
|
||
so a floor of ``MIN_STREAM_DELTA_EVENTS`` ``stream_event`` records
|
||
catches the buffering regression without needing a streaming-aware
|
||
driver.
|
||
|
||
This is the same shape of check as ``tool_use_streaming`` uses,
|
||
just keyed off the explicit partial-message flag so it works for
|
||
plain assistant replies (where the CLI would otherwise collapse a
|
||
streamed reply to a single ``assistant`` event in
|
||
``--print --output-format stream-json`` mode).
|
||
"""
|
||
base_url, api_key = require_proxy(compat_result, env=env)
|
||
|
||
extra_args: Sequence[str] = (
|
||
("--include-partial-messages",) if verify_streaming else ()
|
||
)
|
||
|
||
outcomes = runner(
|
||
models=models,
|
||
prompt=prompt,
|
||
base_url=base_url,
|
||
api_key=api_key,
|
||
extra_args=extra_args,
|
||
)
|
||
|
||
failures = []
|
||
for model in models:
|
||
outcome = outcomes[model]
|
||
if isinstance(outcome, ClaudeCLIError):
|
||
error = f"[{model}] {outcome}"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if outcome.exit_code != 0:
|
||
error = f"[{model}] claude CLI failed: {failure_diagnostic(outcome)}"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if not outcome.text.strip():
|
||
error = f"[{model}] claude returned empty assistant text"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if verify_streaming:
|
||
stream_event_count = _count_stream_event_deltas(outcome.events)
|
||
if stream_event_count < MIN_STREAM_DELTA_EVENTS:
|
||
error = (
|
||
f"[{model}] only {stream_event_count} stream_event records "
|
||
f"observed (< {MIN_STREAM_DELTA_EVENTS}); proxy likely "
|
||
f"buffered the upstream response instead of streaming it"
|
||
)
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
compat_result.add({"status": "pass"})
|
||
|
||
if failures:
|
||
pytest.fail("; ".join(failures), pytrace=False)
|