* test(e2e): budget reset diagonal for team, org, user, and #32005 team-member keys Adds E2E-7/8/10/11 from the budget-level x key-kind coverage matrix: each budget level serves traffic again after its budget_duration window elapses, walking the same ladder as the enforcement diagonal. New registry rows and tests cover the team, organization, and internal-user reset rungs, plus the #32005 interplay where a team-member key frozen by its owner's user budget comes back when the user's window renews; the bare-key and per-team-member rungs already had coverage Each case isolates the cap to one entity, drives spend to a budget_exceeded block, then polls past the window until a call succeeds, holding every refusal as a budget block so a reset that no-ops (stays blocked forever) or crashes (leaks a 5xx) fails the test. budget_duration becomes an optional param on the budget_client create_team / create_user / create_org helpers * test(e2e): fold the reset diagonal into test_budget_reset_e2e.py and address greptile nits Move the team / org / user / #32005 reset cases out of the standalone test_budget_reset_diagonal_e2e.py and into test_budget_reset_e2e.py, absorbing the pre-existing bare-key reset into the same TestBudgetResetDiagonal spec class so the whole reset ladder reads as one file (mirroring how the enforcement diagonal lives in test_budget_enforcement_e2e.py) and the drive/poll helpers are defined once instead of duplicated across reset files. Greptile nits: bound the drive phase to under one window (12 attempts x 2s < 30s) so a block is observed before the reset job can fire, and replace the bare assert in the poll loop with a pytest.fail that prints the HTTP status, so a provider 429 or a crashed reset path is distinguishable from a budget block at a glance. * test(e2e): trim reset diagonal docstrings back to the file's original style * test(e2e): inline single-use drive-loop bounds * test(e2e): cut the reset module docstring to one line * test(e2e): make the org reset test wait for a scheduled window (bugbot) /organization/new stores budget_duration without scheduling budget_reset_at, so the reset job's NULL catch-up branch zeroes org spend on its first 5-10s tick; the org reset test could pass off that catch-up instead of a real window roll (tracked as LIT-4570). The test now reads the org's budget_id and polls /budget/info until budget_reset_at is scheduled before driving spend, so the recovery it observes can only come from a genuine window expiry. Verified live: the org case now runs ~33s (a full window) instead of beating the rescheduler |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| collector.py | ||
| guardrail.yaml | ||
| llm_claude_code_compat.yaml | ||
| llm_conversational.yaml | ||
| llm_nonconversational.yaml | ||
| logging.yaml | ||
| mcp.yaml | ||
| mgmt.yaml | ||
| other.yaml | ||
| quota_management.yaml | ||
| README.md | ||
| registry.py | ||
| reliability.yaml | ||
| schema.py | ||
| test_collector.py | ||
e2e coverage registry
This directory is the denominator for e2e test coverage: the set of behaviors we
want covered, one row per behavior, checked into the repo so coverage is a number we
can track instead of a guess. It implements the plan in the "E2E Coverage Tracking"
note; the naming grammar lives in tests/e2e/CLAUDE.md.
The model
A cell is one customer-noticeable behavior a single e2e test can assert pass/fail
on, for example llm.chat_completions.bedrock_converse.tool_use.stream.works. Cells are
grouped module > feature > test, with LLM cells split into Core LLMs and
Non-Core LLMs for dashboarding. Each cell carries a tier (P0/P1/P2), a source, and a
fail_before_fix flag.
The rows live in per-prefix YAML files (llm_*.yaml, mgmt.yaml, mcp.yaml,
reliability.yaml, quota_management.yaml, logging.yaml, guardrail.yaml,
other.yaml) and validate against
the discriminated union in schema.py, so an LLM row cannot carry a guardrail field and
vice versa. llm rows with subject_endpoint of chat_completions, messages, or
responses roll up to Core LLMs; all other LLM endpoints roll up to Non-Core LLMs.
LLM endpoint, route, and capability values are typed in schema.py, so new taxonomy
values require an explicit schema change. logging and guardrail are two id-prefixes
that roll up into the single Logging & Guardrails dashboard module.
A test declares what it covers with a marker:
@pytest.mark.covers("llm.chat_completions.openai.tool_use.stream.works")
def test_openai_streaming_tool_calls(self) -> None:
...
The number
collector.py diffs the registry against those markers and reports coverage per module.
It is static: a collect-only pass reads the markers, so it runs no test and needs no live
proxy. Whether a covered cell currently passes or fails is a separate, live concern.
cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector
Use --format loki after the e2e pytest run in the same Kubernetes job/pod to print
structured stdout lines for Loki:
cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --format loki --strict
This emits exactly one COVERAGE_TOTAL line and one COVERAGE_MODULE line per module
in MODULE_ORDER, in that order. Loki uses log-safe module= labels from
LOKI_MODULE_LABELS (core_llms, management_ui, etc.) so existing JSON and
Prometheus consumers keep their human-readable module names unchanged.
The headline is overall coverage. The collector also lists markers that point at ids not in the registry, so a typo or an unenumerated behavior surfaces instead of being silently dropped.
Use strict mode in CI once existing draft markers are reconciled:
cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --strict
Strict mode exits non-zero on @pytest.mark.covers(...) ids that are not checked into
the registry. Add --fail-on-collection-errors when the job should also fail on pytest
collection errors.
Status: this is a draft for review
The cells were enumerated from the codebase and the tiers are a first proposal. Known things to settle before treating the set as final:
- tiers are proposed, not signed off; 125 P0 is a lot to prove fail-before-fix, so P0 may want tightening
- a few cells need a support check or a prune (for example
llm.embeddings.anthropic.*andreliability.perf.throughput.under_slo) - auth is covered in two places (
other.auth.*and the mgmt authz assertions); the boundary needs a decision, and the auth cluster may deserve promotion to its own module - the P2 "niche" cells each stand in for a large tail of integrations/providers by design, so the denominator is deliberately P0-weighted rather than a full inventory