mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-14 23:21:35 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
216 lines
7.2 KiB
YAML
216 lines
7.2 KiB
YAML
# local setup to run e2e tests
|
|
configs:
|
|
dd_sink_script:
|
|
content: |
|
|
# Minimal DataDog logs-intake sink for the logging suite: records every
|
|
# POST (gunzipping the compressed batches the integration sends) and
|
|
# replays them as JSON on GET /requests so tests can assert delivery.
|
|
import gzip, json
|
|
from http.server import BaseHTTPRequestHandler, HTTPServer
|
|
|
|
REQUESTS = []
|
|
|
|
class Handler(BaseHTTPRequestHandler):
|
|
def do_POST(self):
|
|
body = self.rfile.read(int(self.headers.get("Content-Length", 0)))
|
|
if self.headers.get("Content-Encoding") == "gzip":
|
|
body = gzip.decompress(body)
|
|
REQUESTS.append({"path": self.path, "body": body.decode("utf-8", "replace")})
|
|
self.send_response(202)
|
|
self.end_headers()
|
|
self.wfile.write(b"{}")
|
|
|
|
def do_GET(self):
|
|
self.send_response(200)
|
|
if self.path == "/health":
|
|
self.send_header("Content-Type", "text/plain")
|
|
self.end_headers()
|
|
self.wfile.write(b"ok")
|
|
return
|
|
self.send_header("Content-Type", "application/json")
|
|
self.end_headers()
|
|
self.wfile.write(json.dumps({"requests": REQUESTS}).encode())
|
|
|
|
def log_message(self, *args):
|
|
pass
|
|
|
|
HTTPServer(("0.0.0.0", 8080), Handler).serve_forever()
|
|
|
|
litellm_config:
|
|
content: |
|
|
general_settings:
|
|
master_key: os.environ/LITELLM_MASTER_KEY
|
|
database_url: os.environ/DATABASE_URL
|
|
store_prompts_in_spend_logs: true
|
|
proxy_budget_rescheduler_min_time: 5
|
|
proxy_budget_rescheduler_max_time: 10
|
|
|
|
litellm_settings:
|
|
drop_params: true
|
|
num_retries: 3
|
|
request_timeout: 600
|
|
cache: true
|
|
cache_params:
|
|
type: redis
|
|
host: redis
|
|
port: 6379
|
|
# OTEL v2 trace destination for the logging suite's trace-completeness
|
|
# tests: the arize_phoenix preset is OTLP with a configurable endpoint
|
|
# (PHOENIX_COLLECTOR_HTTP_ENDPOINT below points it at the jaeger service),
|
|
# so gen-AI spans export through a preset-owned provider - the code path
|
|
# where trace splits actually happen - with no cloud credentials needed.
|
|
callbacks: ["arize_phoenix", "datadog"]
|
|
|
|
router_settings:
|
|
routing_strategy: simple-shuffle
|
|
num_retries: 3
|
|
allowed_fails: 5
|
|
cooldown_time: 30
|
|
fallbacks:
|
|
- gemini-2.5-flash: ["gpt-5.5", "claude-haiku-4-5"]
|
|
|
|
finetune_settings:
|
|
- custom_llm_provider: openai
|
|
api_key: os.environ/OPENAI_API_KEY
|
|
|
|
files_settings:
|
|
- custom_llm_provider: openai
|
|
api_key: os.environ/OPENAI_API_KEY
|
|
- custom_llm_provider: azure
|
|
api_base: os.environ/AZURE_API_BASE
|
|
api_key: os.environ/AZURE_API_KEY
|
|
api_version: "2024-05-01-preview"
|
|
|
|
model_list:
|
|
- model_name: gpt-5.5
|
|
litellm_params:
|
|
model: openai/gpt-5.5
|
|
api_key: os.environ/OPENAI_API_KEY
|
|
|
|
- model_name: claude-haiku-4-5
|
|
litellm_params:
|
|
model: anthropic/claude-haiku-4-5
|
|
api_key: os.environ/ANTHROPIC_API_KEY
|
|
|
|
- model_name: gemini-2.5-flash
|
|
litellm_params:
|
|
model: gemini/gemini-2.5-flash
|
|
api_key: os.environ/GEMINI_API_KEY
|
|
|
|
- model_name: openai-text-embedding-3-small
|
|
litellm_params:
|
|
model: openai/text-embedding-3-small
|
|
api_key: os.environ/OPENAI_API_KEY
|
|
|
|
# v2 auto-router with the LLM complexity classifier. SIMPLE stays on the
|
|
# openai backend; every higher tier routes to the anthropic backend, so the
|
|
# served deployment (read back from the spend log's model) reveals whether
|
|
# the LLM classifier actually ran or silently fell back to heuristic scoring.
|
|
- model_name: complexity-smart-router
|
|
litellm_params:
|
|
model: auto_router/complexity_router
|
|
complexity_router_config:
|
|
classifier_type: llm
|
|
classifier_llm_config:
|
|
model: gpt-5.5
|
|
tiers:
|
|
SIMPLE: gpt-5.5
|
|
MEDIUM: claude-haiku-4-5
|
|
COMPLEX: claude-haiku-4-5
|
|
REASONING: claude-haiku-4-5
|
|
|
|
services:
|
|
litellm:
|
|
image: ghcr.io/berriai/litellm:main-latest
|
|
depends_on:
|
|
db:
|
|
condition: service_healthy
|
|
redis:
|
|
condition: service_healthy
|
|
jaeger:
|
|
condition: service_healthy
|
|
dd-sink:
|
|
condition: service_healthy
|
|
env_file: .env
|
|
environment:
|
|
LITELLM_MASTER_KEY: sk-1234
|
|
STORE_MODEL_IN_DB: "True"
|
|
DD_API_KEY: local-sink-noauth
|
|
DD_SITE: datadoghq.com
|
|
DD_BASE_URL: http://dd-sink:8080
|
|
LITELLM_OTEL_V2: "true"
|
|
PHOENIX_COLLECTOR_HTTP_ENDPOINT: http://jaeger:4318/v1/traces
|
|
PHOENIX_API_KEY: local-jaeger-noauth
|
|
DATABASE_URL: postgresql://litellm:litellm@db:5432/litellm
|
|
UI_USERNAME: admin
|
|
UI_PASSWORD: sk-1234
|
|
AWS_S3_BUCKET_NAME: ${AWS_S3_BUCKET_NAME:-${AWS_BATCH_S3_BUCKET:-}}
|
|
AWS_BATCH_S3_BUCKET: ${AWS_BATCH_S3_BUCKET:-${AWS_S3_BUCKET_NAME:-}}
|
|
AWS_BATCH_ROLE_ARN: ${AWS_BATCH_ROLE_ARN:-}
|
|
AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID:-}
|
|
AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY:-}
|
|
AWS_REGION: ${AWS_REGION:-us-east-1}
|
|
GCS_BUCKET_NAME: ${GCS_BUCKET_NAME:-}
|
|
VERTEXAI_PROJECT: ${VERTEXAI_PROJECT:-}
|
|
VERTEXAI_CREDENTIALS: ${VERTEXAI_CREDENTIALS:-}
|
|
GOOGLE_APPLICATION_CREDENTIALS: ${GOOGLE_APPLICATION_CREDENTIALS:-}
|
|
MISTRAL_API_KEY: ${MISTRAL_API_KEY:-}
|
|
AZURE_API_BASE: ${AZURE_API_BASE:-}
|
|
AZURE_API_KEY: ${AZURE_API_KEY:-}
|
|
AZURE_AI_API_BASE: ${AZURE_AI_API_BASE:-}
|
|
AZURE_AI_API_KEY: ${AZURE_AI_API_KEY:-}
|
|
ports:
|
|
- "4000:4000"
|
|
configs:
|
|
- source: litellm_config
|
|
target: /app/config.yaml
|
|
command: ["--config", "/app/config.yaml", "--port", "4000"]
|
|
|
|
# throwaway db
|
|
db:
|
|
image: postgres:16
|
|
environment:
|
|
POSTGRES_USER: litellm
|
|
POSTGRES_PASSWORD: litellm
|
|
POSTGRES_DB: litellm
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U litellm"]
|
|
interval: 3s
|
|
timeout: 3s
|
|
retries: 20
|
|
|
|
redis:
|
|
image: redis:7
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 3s
|
|
timeout: 3s
|
|
retries: 20
|
|
|
|
# throwaway OTEL trace destination (OTLP ingest on 4318 inside the network,
|
|
# query API on host 16686 for test read-back; see E2E_OTEL_QUERY_URL)
|
|
jaeger:
|
|
image: jaegertracing/all-in-one:1.62.0
|
|
ports:
|
|
- "16686:16686"
|
|
healthcheck:
|
|
test: ["CMD", "wget", "-qO-", "http://localhost:14269/"]
|
|
interval: 3s
|
|
timeout: 3s
|
|
retries: 20
|
|
|
|
# throwaway DataDog logs-intake sink (records POSTs, replays on GET /requests;
|
|
# see E2E_DD_SINK_URL)
|
|
dd-sink:
|
|
image: python:3.12-alpine
|
|
command: ["python", "/sink.py"]
|
|
configs:
|
|
- source: dd_sink_script
|
|
target: /sink.py
|
|
ports:
|
|
- "9915:8080"
|
|
healthcheck:
|
|
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:8080/health"]
|
|
interval: 3s
|
|
timeout: 3s
|
|
retries: 20
|