litellm/tests/e2e/ui/coverage.yaml
Yassin Kortam b5028b81c6 fix(e2e): make coverage collection independent of the runner
The coverage number is documented as static, a property of the source tree, but
the same tree reported 312, 314, or 315 covered cells depending on the machine.
Markers were harvested from session.items in pytest_collection_finish, which
runs after every pytest_collection_modifyitems hook, so anything the runner
dropped vanished from the numerator: the weekly load test is deselected unless
E2E_WEEKLY_ANOMALY is set, and the MCP OAuth module sits behind
pytest.importorskip("mcp") / importorskip("playwright.async_api").

Markers are now read straight off the source with ast, so a cell counts when a
test declaring it exists, whatever the runner does with it. The collect-only
pytest pass stays for the two things the source text cannot give: markers built
at import time (pytest.mark.covers(*fn(...)) inside a pytest.param) and the
nodeids that failed to import, whose cells are genuinely unknowable and still
drive --fail-on-collection-errors. Files that cannot be parsed are reported the
same way instead of being dropped.

The TypeScript Playwright suite emits no pytest markers, so the two
surface: ui mgmt rows were structurally uncoverable. It now declares its cells
in tests/e2e/ui/coverage.yaml, which the collector unions in; an id there that
is not in the registry surfaces as an orphan marker exactly like a mistyped
pytest marker. Each row also names the spec and the Playwright test title behind
it, and both are resolved against the tree on every run, so renaming, deleting
or commenting out that test drops the cell out of the numerator and fails
--strict instead of leaving it counted forever.

Title matching is per declaration, never per file. Comments are stripped before
titles are read, by a string-aware pass that keeps literals intact so a "//"
inside a title is never mistaken for a comment opener. An interpolated title is
matched against its literal segments, so it is checked as far as it can be, and
a dynamic title elsewhere in the spec never exempts a row naming a literal one.
A title with no literal text, or one assembled from variables, backs no row, so
an unmatchable shape fails loudly rather than waving a whole file through. Only
mgmt.key.update.happy_path is claimed, proven by the "Update key TPM and RPM
limits" spec. mgmt.key.generate.happy_path is left unclaimed and documented:
that row scopes itself to SSO-driven key gen and every role in the suite logs in
with username/password.

Pruned llm.embeddings.anthropic.basic.nonstream.works. Anthropic ships no
embeddings API, the anthropic handler's embedding() is a pass stub, no anthropic
dispatch exists in embedding(), and no anthropic row in
model_prices_and_context_window.json carries mode: embedding, so the row could
never pass.

Also corrected the suite-folder list in tests/e2e/CLAUDE.md: gateway/,
embeddings/, security/, and a top-level realtime/ do not exist (realtime lives
under llm_translation/), and guardrails/ was undocumented.

Same tree now reports 316/430 with and without E2E_WEEKLY_ANOMALY and with or
without mcp/playwright installed.
2026-07-27 17:15:30 -07:00

40 lines
2 KiB
YAML

# Coverage-registry cells this Playwright suite covers.
#
# Every other e2e suite is Python and declares coverage with
# @pytest.mark.covers("<cell id>"). This suite is TypeScript, so it declares the
# same thing here and tests/e2e/coverage_registry/collector.py unions these ids
# into the covered set.
#
# To add a row: write the UI test first, then name the registry cell it proves,
# the spec file it lives in, and the test title. The id must exist in
# tests/e2e/coverage_registry/*.yaml; a typo lands in the collector's orphan
# marker list exactly like a typo in a pytest marker, and `--strict` fails on it.
# Claim a cell only when the spec actually asserts that behavior.
#
# `spec` is relative to this file and `test` is the Playwright test title, and
# the collector resolves both against the tree on every run. Rename or delete
# that test and the row stops counting and fails `--strict`, so a claim here
# cannot outlive the test that backs it.
#
# Commented-out code does not count as a test, so a row whose test is commented
# out fails the same way a deleted one does.
#
# An interpolated title such as test(`${role} sidebar`) is checked against its
# literal segments, so write out the title as it renders (e.g. "admin sidebar").
# The check is per row, never per file: an interpolated title elsewhere in the
# spec does not exempt your row. A title with no literal text at all, or one
# assembled from variables, cannot back a row; give that test a title with
# something literal in it.
covers:
- id: mgmt.key.update.happy_path
spec: tests/proxy-admin/keys.spec.ts
test: Update key TPM and RPM limits
# Not claimed yet:
#
# mgmt.key.generate.happy_path - the registry row scopes this to SSO-driven key
# generation ("SSO-driven key gen (UI path)", source ui_sso.py). This suite
# creates keys through the dashboard, but every role logs in with
# username/password (globalSetup.ts), so no test drives the SSO path. Claim it
# once a spec exercises SSO login, or retarget the registry row at plain
# dashboard key creation and claim it then.