mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-14 23:21:35 +00:00
* refactor(e2e/claude_code): align proxy env names with the rest of tests/e2e
Every claude_code compat cell used to read its own `LITELLM_PROXY_BASE_URL` and `LITELLM_PROXY_API_KEY` and duplicate the same 12-line "missing env, hard fail" block. The rest of `tests/e2e/` reads `LITELLM_PROXY_URL` and `LITELLM_MASTER_KEY` from `e2e_config.py`, so anyone standing up a live proxy for one suite had to export a second spelling for claude_code, and every cell repeated the same boilerplate.
Centralize the resolution in `claude_code/_env.py`. `resolve_proxy()` prefers the suite-wide `LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY` names and falls back to the legacy pair so existing CI wiring on stage keeps working during the roll-out. `require_proxy(compat_result)` is the one-liner cells call to bind `(base_url, api_key)` or hard-fail with a message that names both spellings.
55 cell files, `_basic_messaging.py`, and the driver's own unit-test fixture now go through the helper. `run_compat.sh` accepts either spelling and normalizes to the primary names before invoking pytest. `cron_vm/run_daily.sh` exports the primary names when launching pytest.
`_pr_gate_unit_tests/test_env_resolution.py` pins the resolution rules so a future edit cannot silently reintroduce the drift: primary names win on tie, legacy names still resolve when primary is unset, mixed URL-primary key-legacy still resolves, empty-string exports are treated as unset, `require_proxy` names both spellings in its error message.
Net diff: 71 files, +370/-1240.
* fix(e2e): anchor claude_code Bash pin at parents[1] so container run collects
`test_bash_tool_restrictions.py` derived `REPO_ROOT = Path(__file__).resolve().parents[4]` and then joined `tests/e2e/claude_code/<feature>`. That works locally, but the stage container mounts tests/e2e/ at /app/e2e/, so parents[4] resolves to filesystem root and the `_bash_cells()` assertion looks for `/tests/e2e/claude_code/tool_use` — a path that doesn't exist. Collection interrupts before any test runs, so the entire e2e suite appears broken.
Fix: `CLAUDE_CODE_DIR = Path(__file__).resolve().parents[1]` resolves to the sibling `claude_code/` dir in either layout, and the `relative_to(REPO_ROOT)` calls become `relative_to(CLAUDE_CODE_DIR)` so test IDs and error messages read the same.
Adds `test_claude_code_dir_anchor_is_layout_independent` as a regression pin: it checks the anchor lands on a directory named `claude_code` that contains this test file, which would fail under the old parents[4] anchor when run from /app/e2e/.
* feat(e2e/claude_code): register compat deployments via /model/new from a session fixture
Every compat cell hardcodes a virtual model name like `claude-sonnet-4-6` or `claude-sonnet-4-6-bedrock-invoke` and hits the proxy expecting it to be routable. On stage those live in the deployed model_list; locally the `docker-config.yaml` under tests/e2e/ only declares one of them, so anything past haiku 400s with `Invalid model name`.
`claude_code/test_config.yaml` is the ground-truth compat matrix config the deployment already uses. `_compat_models.py` loads it, normalizes the yaml keys pydantic would silently drop (vertex_ai_* → vertex_*), and selects the subset whose provider credentials are present in the environment. An autouse session fixture in `conftest.py` POSTs each selected deployment to `/model/new`, blocks until it is servable on the data plane, and tears them all down on session exit. Skips silently when the proxy env is unset so pure-unit runs stay hermetic.
`test_compat_models.py` pins the invariants that keep this safe. Every cell-referenced name must have a yaml entry (drift check catches a cell probing a name the fixture never registered); the yaml has no unused declarations; the fixture registers exactly 15 deployments (3 tiers × 5 provider surfaces); vertex_ai_* yaml keys populate the pydantic body's vertex_* fields (they got silently dropped historically); Azure needs both AZURE_FOUNDRY_* env vars; Bedrock lifts creds from the ambient AWS chain; Vertex needs both the yaml refs AND ambient GCP credentials.
* refactor(e2e/claude_code): inject env + runner instead of monkeypatching
`require_proxy` and `_basic_messaging.run_basic_messaging_cell` now take the env mapping (and the CLI runner) as constructor-style arguments with `os.environ` and `run_claude_models_parallel` as defaults. Tests exercise the branching by passing dicts and callables directly, so `monkeypatch.setenv` and `monkeypatch.setattr(_basic_messaging, "run_claude_models_parallel", ...)` are gone from every unit test in this refactor's blast radius.
`test_env_resolution.py` drops the `monkeypatch.setenv`/`delenv` fixtures and passes `env={...}` dicts to `require_proxy`. Added a new pinned check that a successful resolution leaves `compat_result` untouched, and split the "unset env" test into three explicit shapes (empty, primary-only, legacy-only) so a regression that swaps the precedence rule can no longer hide behind a single monkeypatched fixture.
`test_basic_messaging.py` (driver) replaces the `_install_fake_runner(monkeypatch, ...)` helper with `_make_fake_runner(...)` that returns a `(callable, captured_dict)` pair the test passes in via the helper's new `runner=` kwarg. Also drops the autouse `_proxy_env` fixture in favor of a module-level `_PROXY_ENV` dict each test wires through the helper's new `env=` kwarg. Added a regression pin that a missing-env call hard-fails without ever invoking the runner (so the guard order stays correct).
`test_run_daily_pytest_scrubs_env.py` updates its pin to assert the new suite-wide env spellings (`LITELLM_PROXY_URL` / `LITELLM_MASTER_KEY`) instead of the legacy `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` that `run_daily.sh` used to export.
* handwrote rules
202 lines
4.3 KiB
Python
202 lines
4.3 KiB
Python
"""Registry row schema: the contract every denominator cell validates against.
|
|
|
|
A cell is one customer-noticeable behavior a single e2e test can assert pass/fail
|
|
on. `module` is the id's segment-1 prefix (eight of them); dashboard rollups can
|
|
split or merge those prefixes. The union is discriminated on `module`, so an LLM
|
|
row cannot carry a guardrail field and vice versa.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from enum import Enum
|
|
from typing import Annotated, Literal
|
|
|
|
from pydantic import BaseModel, ConfigDict, Field, TypeAdapter
|
|
|
|
|
|
class Tier(str, Enum):
|
|
P0 = "P0"
|
|
P1 = "P1"
|
|
P2 = "P2"
|
|
|
|
|
|
class FailBeforeFix(str, Enum):
|
|
proven = "proven"
|
|
unproven = "unproven"
|
|
|
|
|
|
LlmEndpoint = Literal[
|
|
"chat_completions",
|
|
"messages",
|
|
"responses",
|
|
"embeddings",
|
|
"batches",
|
|
"files",
|
|
"rerank",
|
|
"images_generations",
|
|
"audio_speech",
|
|
"audio_transcriptions",
|
|
"moderations",
|
|
"realtime",
|
|
]
|
|
|
|
LlmRoute = Literal[
|
|
"anthropic",
|
|
"azure_foundry",
|
|
"azure_openai",
|
|
"bedrock_converse",
|
|
"bedrock_invoke",
|
|
"cohere",
|
|
"openai",
|
|
"together_ai",
|
|
"vertex",
|
|
]
|
|
|
|
LlmCapability = Literal[
|
|
"basic",
|
|
"count_tokens",
|
|
"long_context_1m",
|
|
"mid_conversation_system",
|
|
"pdf_input",
|
|
"prompt_cache_1h",
|
|
"prompt_cache_5m",
|
|
"service_tier",
|
|
"structured_output",
|
|
"thinking",
|
|
"thinking_with_tool_use",
|
|
"tool_search",
|
|
"tool_use",
|
|
"vision",
|
|
"web_search",
|
|
]
|
|
|
|
|
|
class _Base(BaseModel):
|
|
model_config = ConfigDict(frozen=True, extra="forbid")
|
|
|
|
id: str
|
|
tier: Tier
|
|
assertions: tuple[str, ...]
|
|
source: str
|
|
rationale: str = ""
|
|
fail_before_fix: FailBeforeFix = FailBeforeFix.unproven
|
|
supported: bool = True
|
|
|
|
|
|
class LlmCell(_Base):
|
|
module: Literal["llm"]
|
|
subject_endpoint: LlmEndpoint
|
|
route: LlmRoute
|
|
capability: LlmCapability
|
|
streaming: Literal["stream", "nonstream", "na"]
|
|
|
|
|
|
class MgmtCell(_Base):
|
|
module: Literal["mgmt"]
|
|
surface: Literal["api", "ui"]
|
|
|
|
|
|
class McpCell(_Base):
|
|
module: Literal["mcp"]
|
|
operation: str
|
|
auth_family: Literal["none", "api_key", "bearer", "oauth"]
|
|
|
|
|
|
class ReliabilityCell(_Base):
|
|
module: Literal["reliability"]
|
|
behavior: str
|
|
variant: str
|
|
exercised_on: tuple[str, ...]
|
|
|
|
|
|
class QuotaCell(_Base):
|
|
module: Literal["quota_management"]
|
|
behavior: Literal["ratelimit", "budget", "spend_tracking"]
|
|
variant: str
|
|
exercised_on: tuple[str, ...]
|
|
|
|
|
|
class LoggingCell(_Base):
|
|
module: Literal["logging"]
|
|
event: str
|
|
exercised_on: tuple[str, ...]
|
|
|
|
|
|
class GuardrailCell(_Base):
|
|
module: Literal["guardrail"]
|
|
hook_point: str
|
|
exercised_on: tuple[str, ...]
|
|
|
|
|
|
class OtherCell(_Base):
|
|
module: Literal["other"]
|
|
area: str
|
|
|
|
|
|
Cell = Annotated[
|
|
LlmCell
|
|
| MgmtCell
|
|
| McpCell
|
|
| ReliabilityCell
|
|
| QuotaCell
|
|
| LoggingCell
|
|
| GuardrailCell
|
|
| OtherCell,
|
|
Field(discriminator="module"),
|
|
]
|
|
|
|
CELL_ADAPTER: TypeAdapter[Cell] = TypeAdapter(Cell)
|
|
|
|
CORE_LLM_ENDPOINTS: frozenset[str] = frozenset(
|
|
{
|
|
"chat_completions",
|
|
"messages",
|
|
"responses",
|
|
}
|
|
)
|
|
|
|
PREFIX_ROLLUP: dict[str, str] = {
|
|
"mcp": "MCPs",
|
|
"mgmt": "Management/UI",
|
|
"reliability": "Reliability & Performance",
|
|
"quota_management": "Quota Management",
|
|
"logging": "Logging & Guardrails",
|
|
"guardrail": "Logging & Guardrails",
|
|
"other": "Other",
|
|
}
|
|
|
|
MODULE_ORDER: tuple[str, ...] = (
|
|
"Core LLMs",
|
|
"Non-Core LLMs",
|
|
"MCPs",
|
|
"Management/UI",
|
|
"Reliability & Performance",
|
|
"Quota Management",
|
|
"Logging & Guardrails",
|
|
"Other",
|
|
)
|
|
|
|
LOKI_MODULE_LABELS: dict[str, str] = {
|
|
"Core LLMs": "core_llms",
|
|
"Non-Core LLMs": "non_core_llms",
|
|
"MCPs": "mcp",
|
|
"Management/UI": "management_ui",
|
|
"Reliability & Performance": "reliability_performance",
|
|
"Quota Management": "quota_management",
|
|
"Logging & Guardrails": "logging_guardrails",
|
|
"Other": "other",
|
|
}
|
|
|
|
|
|
def dashboard_module(cell: Cell) -> str:
|
|
"""Return the Grafana/reporting module for a registry cell."""
|
|
if isinstance(cell, LlmCell):
|
|
if cell.subject_endpoint in CORE_LLM_ENDPOINTS:
|
|
return "Core LLMs"
|
|
return "Non-Core LLMs"
|
|
return PREFIX_ROLLUP[cell.module]
|
|
|
|
|
|
def loki_module_label(module: str) -> str:
|
|
"""Return the log-safe Loki label for a dashboard module."""
|
|
return LOKI_MODULE_LABELS[module]
|