mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-17 23:51:30 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
108 lines
4.1 KiB
Bash
Executable file
108 lines
4.1 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Run the Claude Code compat matrix end-to-end against a live LiteLLM
|
|
# proxy, with per-provider rate limits applied via the cross-process
|
|
# token bucket in `tests/e2e/claude_code/rate_limiter.py`.
|
|
#
|
|
# Designed for binary-searching the ideal X / Y / Z req/s per provider:
|
|
# 1. Pick an initial rate (e.g. 5/s for everyone).
|
|
# 2. Run this script.
|
|
# 3. Read `compat-rate-limit-summary.json` to see whether any provider
|
|
# hit a 429-shaped error during the run.
|
|
# 4. If a provider has `rate_limited > 0`, halve its rate; else, double it.
|
|
# 5. Repeat until the highest no-429 rate is found.
|
|
#
|
|
# Required env (proxy connection), same names as the rest of tests/e2e:
|
|
# LITELLM_PROXY_URL e.g. http://localhost:4000
|
|
# LITELLM_MASTER_KEY e.g. sk-1234
|
|
#
|
|
# Optional env (rate limits, all default to 5 req/s; 0 disables a column):
|
|
# LITELLM_COMPAT_RATE_ANTHROPIC
|
|
# LITELLM_COMPAT_RATE_AZURE
|
|
# LITELLM_COMPAT_RATE_VERTEX_AI
|
|
# LITELLM_COMPAT_RATE_BEDROCK_CONVERSE
|
|
# LITELLM_COMPAT_RATE_BEDROCK_INVOKE
|
|
# LITELLM_COMPAT_RATE_BURST override per-bucket burst
|
|
#
|
|
# Optional env (parallelism):
|
|
# COMPAT_XDIST_WORKERS passed to `pytest -n` (default: auto)
|
|
#
|
|
# Optional env (artifacts):
|
|
# COMPAT_RESULTS_PATH default: compat-results.json
|
|
# COMPAT_RATE_LIMIT_SUMMARY_PATH default: compat-rate-limit-summary.json
|
|
|
|
set -euo pipefail
|
|
|
|
if [[ -z "${LITELLM_PROXY_URL:-}" || -z "${LITELLM_MASTER_KEY:-}" ]]; then
|
|
echo "error: LITELLM_PROXY_URL and LITELLM_MASTER_KEY must be set" >&2
|
|
exit 64
|
|
fi
|
|
|
|
# Reset the cross-process rate-limiter state from any prior run. Stale
|
|
# token-bucket files would let a previous run's accumulated budget bleed
|
|
# into the new one, which subtly biases the binary search.
|
|
state_dir="${LITELLM_COMPAT_RATE_STATE_DIR:-${TMPDIR:-/tmp}/litellm-claude-compat-ratelimit}"
|
|
if [[ -d "$state_dir" ]]; then
|
|
rm -rf "$state_dir"
|
|
fi
|
|
|
|
# Worker count. `auto` picks one worker per CPU; the rate limiter
|
|
# enforces aggregate provider rates regardless of worker count, so
|
|
# this is a "go as fast as the limiter allows" knob, not a tuning knob.
|
|
workers="${COMPAT_XDIST_WORKERS:-auto}"
|
|
|
|
# Where the artifacts land. We resolve them now so the summary file is
|
|
# always at a known path the caller can grep, even if they didn't set
|
|
# the env explicitly.
|
|
results_path="${COMPAT_RESULTS_PATH:-compat-results.json}"
|
|
summary_path="${COMPAT_RATE_LIMIT_SUMMARY_PATH:-compat-rate-limit-summary.json}"
|
|
|
|
echo "[run_compat] rates:"
|
|
for provider in ANTHROPIC AZURE VERTEX_AI BEDROCK_CONVERSE BEDROCK_INVOKE; do
|
|
var="LITELLM_COMPAT_RATE_${provider}"
|
|
echo " ${provider}=${!var:-default(5/s)}"
|
|
done
|
|
echo " BURST=${LITELLM_COMPAT_RATE_BURST:-default(=rate)}"
|
|
echo "[run_compat] xdist workers: ${workers}"
|
|
echo "[run_compat] results: ${results_path}"
|
|
echo "[run_compat] summary: ${summary_path}"
|
|
|
|
# Run only the per-feature live tests; skip the unit-test directories
|
|
# (they're under directories starting with `_`). The dist=loadfile
|
|
# scheduler keeps each test file pinned to a single worker, which is
|
|
# what we want — every test in a file shares a single ThreadPoolExecutor
|
|
# fanout, and we don't gain anything by splitting it across workers.
|
|
start=$(date +%s)
|
|
set +e
|
|
COMPAT_RESULTS_PATH="${results_path}" \
|
|
COMPAT_RATE_LIMIT_SUMMARY_PATH="${summary_path}" \
|
|
PATH="$HOME/.local/bin:$PATH" \
|
|
uv run pytest \
|
|
tests/e2e/claude_code/basic_messaging_non_streaming \
|
|
tests/e2e/claude_code/basic_messaging_streaming \
|
|
tests/e2e/claude_code/thinking \
|
|
tests/e2e/claude_code/tool_use \
|
|
tests/e2e/claude_code/vision \
|
|
tests/e2e/claude_code/prompt_caching_5m \
|
|
-n "${workers}" \
|
|
--dist=loadfile \
|
|
-q \
|
|
"$@"
|
|
|
|
exit_code=$?
|
|
set -e
|
|
end=$(date +%s)
|
|
echo "[run_compat] wall time: $((end - start))s"
|
|
|
|
# Surface the rate-limit summary inline so a human reader doesn't have
|
|
# to `cat` the JSON file. The full file is still on disk for the binary
|
|
# search loop.
|
|
if [[ -f "${summary_path}" ]]; then
|
|
echo "[run_compat] summary: ${summary_path}"
|
|
if command -v jq >/dev/null 2>&1; then
|
|
jq '.totals, .per_provider' "${summary_path}"
|
|
else
|
|
cat "${summary_path}"
|
|
fi
|
|
fi
|
|
|
|
exit "${exit_code}"
|