mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
198 lines
7.5 KiB
Python
198 lines
7.5 KiB
Python
"""Matrix JSON Builder.
|
|
|
|
Pure-function module that consumes the pytest-produced `compat-results.json`,
|
|
the manifest, and run metadata, and emits the final `compatibility-matrix.json`
|
|
conforming to the schema published in the PRD.
|
|
|
|
This module is deliberately free of subprocess, network, or filesystem side
|
|
effects in its public API — the public entry points take pre-loaded inputs
|
|
and return data structures, so they can be exercised by golden-file tests
|
|
without I/O. A small `build_from_paths()` convenience wrapper does the I/O
|
|
for callers that need it (the daily-cron publisher).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
from pathlib import Path
|
|
from typing import Any, Dict, List, Mapping, Optional, Sequence
|
|
|
|
import yaml
|
|
|
|
SCHEMA_VERSION = "1"
|
|
VALID_STATUSES = {"pass", "fail", "not_applicable", "not_tested"}
|
|
|
|
|
|
class ManifestError(ValueError):
|
|
"""Raised when `manifest.yaml` is malformed."""
|
|
|
|
|
|
class ResultsError(ValueError):
|
|
"""Raised when the pytest results artifact is malformed."""
|
|
|
|
|
|
def load_manifest(path: Path) -> Dict[str, Any]:
|
|
"""Load and validate `manifest.yaml`.
|
|
|
|
Returns a dict with keys: schema_version, providers, features. Raises
|
|
ManifestError on missing fields or schema mismatch.
|
|
"""
|
|
raw = yaml.safe_load(path.read_text())
|
|
if not isinstance(raw, dict):
|
|
raise ManifestError(f"manifest at {path} is not a mapping")
|
|
schema_version = str(raw.get("schema_version", ""))
|
|
if schema_version != SCHEMA_VERSION:
|
|
raise ManifestError(
|
|
f"manifest schema_version {schema_version!r} does not match "
|
|
f"builder version {SCHEMA_VERSION!r}"
|
|
)
|
|
providers = raw.get("providers")
|
|
if not isinstance(providers, list) or not providers:
|
|
raise ManifestError("manifest.providers must be a non-empty list")
|
|
features = raw.get("features")
|
|
if not isinstance(features, list) or not features:
|
|
raise ManifestError("manifest.features must be a non-empty list")
|
|
for feature in features:
|
|
if not isinstance(feature, dict):
|
|
raise ManifestError("each feature must be a mapping")
|
|
if not feature.get("id") or not feature.get("name"):
|
|
raise ManifestError("each feature must have id and name")
|
|
return raw
|
|
|
|
|
|
def load_results(path: Path) -> List[Dict[str, Any]]:
|
|
"""Load the pytest results artifact and return its `results` list."""
|
|
raw = json.loads(path.read_text())
|
|
if not isinstance(raw, dict) or not isinstance(raw.get("results"), list):
|
|
raise ResultsError(f"results artifact at {path} has no `results` list")
|
|
return raw["results"]
|
|
|
|
|
|
def build_matrix(
|
|
*,
|
|
manifest: Mapping[str, Any],
|
|
results: Sequence[Mapping[str, Any]],
|
|
litellm_version: str,
|
|
claude_code_version: str,
|
|
generated_at: str,
|
|
) -> Dict[str, Any]:
|
|
"""Build the published matrix JSON from pre-loaded inputs.
|
|
|
|
Empty cells (no test ran for a (feature, provider) and no
|
|
`not_applicable` was declared) are filled in with `not_tested`. If
|
|
multiple results report on the same cell — e.g. a per-feature test
|
|
file containing one parametrize per Claude model — the cell aggregates
|
|
to `pass` only if every model passed; otherwise `fail` with the first
|
|
breaking model surfaced in the error.
|
|
"""
|
|
providers: List[str] = list(manifest["providers"])
|
|
feature_specs: List[Dict[str, Any]] = list(manifest["features"])
|
|
|
|
grouped: Dict[tuple, List[Dict[str, Any]]] = {}
|
|
for entry in results:
|
|
if not isinstance(entry, Mapping):
|
|
continue
|
|
feature_id = entry.get("feature_id")
|
|
provider = entry.get("provider")
|
|
result = entry.get("result")
|
|
if not feature_id or not provider or not isinstance(result, Mapping):
|
|
continue
|
|
if result.get("status") not in VALID_STATUSES:
|
|
continue
|
|
grouped.setdefault((feature_id, provider), []).append(dict(result))
|
|
|
|
features_out: List[Dict[str, Any]] = []
|
|
for spec in feature_specs:
|
|
feature_id = spec["id"]
|
|
cells: Dict[str, Dict[str, Any]] = {}
|
|
for provider in providers:
|
|
cell_results = grouped.get((feature_id, provider), [])
|
|
cells[provider] = _aggregate_cell(cell_results)
|
|
features_out.append(
|
|
{
|
|
"id": feature_id,
|
|
"name": spec["name"],
|
|
"providers": cells,
|
|
}
|
|
)
|
|
|
|
return {
|
|
"schema_version": SCHEMA_VERSION,
|
|
"generated_at": generated_at,
|
|
"litellm_version": litellm_version,
|
|
"claude_code_version": claude_code_version,
|
|
"providers": providers,
|
|
"features": features_out,
|
|
}
|
|
|
|
|
|
def _aggregate_cell(results: Sequence[Mapping[str, Any]]) -> Dict[str, Any]:
|
|
"""Aggregate a list of per-model results into a single cell status.
|
|
|
|
Order of precedence (most informative wins):
|
|
- Any `fail` → cell is `fail` with every failing model's error
|
|
joined by `"; "` so a multi-tier breakage doesn't silently hide
|
|
all but the first error from the published matrix.
|
|
- Any `pass` → cell is `pass`. A mix of (pass, not_applicable) —
|
|
e.g. a tier where the feature isn't supported alongside tiers
|
|
where it works — surfaces as `pass` so the published cell
|
|
reflects that the feature *does* work on this provider rather
|
|
than silently demoting it to `not_applicable` and discarding
|
|
the passing tiers.
|
|
- All `not_applicable` → cell is `not_applicable` with the first
|
|
row's reason.
|
|
- empty / nothing recognized → `not_tested`.
|
|
|
|
`not_tested` rows are treated as absent data: they're dropped before
|
|
aggregation so a mix of (pass, not_tested) — e.g. from a partial
|
|
crash or a test that explicitly recorded "this tier didn't run" —
|
|
still surfaces the passing tiers rather than silently demoting the
|
|
whole cell to `not_tested`. A cell is only `not_tested` when *every*
|
|
row is `not_tested` (or there are no rows at all).
|
|
"""
|
|
if not results:
|
|
return {"status": "not_tested"}
|
|
|
|
observed = [r for r in results if r.get("status") != "not_tested"]
|
|
if not observed:
|
|
return {"status": "not_tested"}
|
|
|
|
failures = [r for r in observed if r.get("status") == "fail"]
|
|
if failures:
|
|
errors = [str(r.get("error", "test failed")) for r in failures]
|
|
return {"status": "fail", "error": "; ".join(errors)}
|
|
|
|
if any(r.get("status") == "pass" for r in observed):
|
|
return {"status": "pass"}
|
|
|
|
if all(r.get("status") == "not_applicable" for r in observed):
|
|
return {
|
|
"status": "not_applicable",
|
|
"reason": str(observed[0].get("reason", "not applicable")),
|
|
}
|
|
|
|
return {"status": "not_tested"}
|
|
|
|
|
|
def build_from_paths(
|
|
*,
|
|
manifest_path: Path,
|
|
results_path: Path,
|
|
litellm_version: str,
|
|
claude_code_version: str,
|
|
generated_at: str,
|
|
output_path: Optional[Path] = None,
|
|
) -> Dict[str, Any]:
|
|
"""I/O wrapper around ``build_matrix``: reads the manifest and per-test results from disk, calls ``build_matrix``, and (optionally) writes the compat-matrix JSON to ``output_path``. Whatever orchestrator publishes the matrix (currently the ECR image) invokes this."""
|
|
manifest = load_manifest(manifest_path)
|
|
results = load_results(results_path)
|
|
matrix = build_matrix(
|
|
manifest=manifest,
|
|
results=results,
|
|
litellm_version=litellm_version,
|
|
claude_code_version=claude_code_version,
|
|
generated_at=generated_at,
|
|
)
|
|
if output_path is not None:
|
|
output_path.write_text(json.dumps(matrix, indent=2, sort_keys=False) + "\n")
|
|
return matrix
|