mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-30 01:52:18 +00:00
A test's JUnit report says what happened to it, never what it was about.
`@pytest.mark.covers("cell.id")` is a registry key, not a description: it
cannot answer "which tests drive /v1/responses on Anthropic", and nothing
in the report says where a failing test actually died.
Two halves, deliberately separated, both riding out as JUnit <property>
entries the downstream emitter already knows how to read.
DECLARED - `@meta(Subject(...))` from the new tests/e2e/e2e_metadata.py.
One frozen dataclass, every field a closed enum (domain, route, provider,
model, capabilities, mode), so a typo is a basedpyright error at the call
site rather than a property that silently never appears. Serialization is
one pass over `dataclasses.asdict`, so a new scalar field needs no
serializer edit; `capabilities` is deduped and sorted at declaration so
committed run artifacts diff cleanly whatever order a test spelled it in.
Empty fields emit nothing - the suite does not pad every testcase with
five empty entries.
RECORDED - `@step("POST /chat/completions")` on harness helpers, never on
tests. Each call appends its label to the running test's user_properties
in call order, so the list IS the test's user story and cannot drift from
what the test did. The label is recorded BEFORE the wrapped call, so a
helper that raises still leaves its own label last: a failing test's last
step is where it died. Consecutive duplicates collapse and the log caps at
50, so a poll loop is one beat of the story rather than fifty.
Steps cannot be attached where the other properties are -
`pytest_collection_modifyitems` runs before any test body, so the recorder
is empty there. They attach from the existing `pytest_runtest_makereport`
wrapper on the call phase, which is what puts them on failures too, and an
autouse fixture empties the log at setup. The attach drops any prior step
entries first, because the suite runs `--reruns 1` and a retry would
otherwise stack a second copy of the story behind the first.
`covers` is untouched: the marker is separate because a dataclass passed
to `covers` would be dropped silently by `dedupe_covers` and would hard-
fail collection in tests/integration/conftest.py. The fixed
package/covers/source prefix stays byte-identical, and `@meta` goes BELOW
`@covers` so `Item.location` still anchors at the first decorator and
every `source` deep link keeps pointing where it pointed.
`Provider` mirrors litellm's `LlmProviders` values instead of importing
them, so nothing here - the module or its call sites - imports litellm.
tests/e2e is a black-box HTTP suite that is copied to the runner image on
its own, so a `from litellm...` at the top of a test module would make the
package a COLLECTION-time dependency: where it is absent, every test in
the suite errors out before running rather than importing slowly. The
mirror cannot drift silently - `TestProviderMirrorsLitellm` asserts every
value is a real `LlmProviders` value wherever litellm is importable, and
skips where it is not, which is the property it is guarding.
A declared `model` names the constant the test drives, never a copy of its
value: `CHEAP_ANTHROPIC_MODEL` and `CHEAP_OPENAI_MODEL` are env-overridable
(`E2E_CHEAP_ANTHROPIC_MODEL`, `E2E_CHEAP_OPENAI_MODEL`), so a hardcoded
default would have reported a model the run never touched. The same holds
for a file's own `BACKEND`/`MODEL` constant, where the copy was merely
waiting to drift.
Pilot: tests/e2e/quota_management, all 85 tests annotated and its three
clients plus cost_rows @step-decorated, to prove the API against real
tests rather than a toy. The rest of the suite is a later backfill.
Verified: 42 harness unit tests in test_junit_properties.py (19 new), 550
harness unit tests green, basedpyright over tests/e2e at the same 29
pre-existing errors as origin/main (all in the untouched mcp/
oauth_chat_client.py), ruff clean, the coverage-registry collector
byte-identical before and after (473/586), and the test-quality gate OK
against origin/main. Collection was also run with the litellm package
blocked at the import hook: 1339/1350 collected either way, the one error
being the pre-existing missing `httpx2` in mcp/. The live e2e tests need a
deployed proxy and real provider keys and were not run.
134 lines
6 KiB
Python
134 lines
6 KiB
Python
"""Custom per-test signals for the standard JUnit reporter.
|
|
|
|
The e2e suite ships results to Loki/Grafana from a standard pytest JUnit report
|
|
(`--junitxml=e2e-report.xml`), not a bespoke log line. JUnit already records
|
|
outcome, duration, and node id for every `<testcase>`; the signals it cannot
|
|
derive on its own are the normalized suite package, the coverage-registry cell
|
|
ids a test covers, and where the test's source lives. Those ride along as JUnit
|
|
`<property>` entries via each item's `user_properties`, attached in
|
|
`conftest.py::pytest_collection_modifyitems`.
|
|
|
|
`source` is a property rather than the `file=` / `line=` attributes pytest used
|
|
to write, because the `xunit2` family this suite runs on drops those, and
|
|
switching families would change the XML for every consumer of it -- the
|
|
Buildkite Test Engine upload and the Loki pipeline included.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from collections.abc import Iterable
|
|
|
|
import pytest
|
|
from coverage_registry.management_cases import case_properties
|
|
from e2e_metadata import step_properties, subject_properties
|
|
|
|
# Hardcoded because the runner image copies tests/e2e/ to /app/e2e, so nothing
|
|
# at runtime names this suite's place in the repo. test_junit_properties.py
|
|
# fails from a checkout if it moves.
|
|
SUITE_ROOT = "tests/e2e"
|
|
|
|
|
|
def suite_parts(path_part: str) -> tuple[str, ...]:
|
|
"""Path components of a suite file relative to tests/e2e, however it ran.
|
|
|
|
Pytest paths are rootdir-relative, and rootdir moves with the invocation: a
|
|
repo-root run gives `tests/e2e/logging/test_x.py`, a suite-cwd run (the
|
|
runner image) gives `logging/test_x.py`. Both collapse to the same tuple.
|
|
"""
|
|
raw = tuple(p for p in path_part.replace("\\", "/").split("/") if p and p != ".")
|
|
return raw[2:] if len(raw) >= 3 and raw[0] == "tests" and raw[1] == "e2e" else raw
|
|
|
|
|
|
def package_from_nodeid(nodeid: str) -> str:
|
|
"""Top-level suite package under tests/e2e/, or 'root' for top-level files."""
|
|
parts = suite_parts(nodeid.split("::", 1)[0])
|
|
if len(parts) <= 1:
|
|
return "root"
|
|
return parts[0]
|
|
|
|
|
|
def source_from_location(path: str, lineno: int | None) -> str:
|
|
"""Repo-relative `path:line` for a test, or '' when nothing is linkable.
|
|
|
|
`pytest.Item.location` gives a rootdir-relative path and a ZERO-based line.
|
|
The path is re-rooted at SUITE_ROOT so consumers need not know how pytest was
|
|
started, and the line is emitted ONE-based to match editors, tracebacks and
|
|
code hosts. A decorated test anchors at its first decorator, which is where
|
|
pytest reports it.
|
|
|
|
Empty rather than a guess for anything unlinkable: no line, a path reaching
|
|
upward, or a path carrying a colon, which is both how an absolute Windows
|
|
path arrives and a character `path:line` has no way to represent.
|
|
"""
|
|
if lineno is None:
|
|
return ""
|
|
normalized = path.replace("\\", "/")
|
|
if normalized.startswith("/") or ":" in normalized or ".." in normalized.split("/"):
|
|
return ""
|
|
parts = suite_parts(normalized)
|
|
if not parts:
|
|
return ""
|
|
return f"{'/'.join((SUITE_ROOT, *parts))}:{lineno + 1}"
|
|
|
|
|
|
def source_from_item(item: pytest.Item) -> str:
|
|
"""Read the repo-relative `path:line` off a pytest Item's reported location."""
|
|
path, lineno, _ = item.location
|
|
return source_from_location(path, lineno)
|
|
|
|
|
|
def dedupe_covers(marker_args: Iterable[tuple[object, ...]]) -> tuple[str, ...]:
|
|
"""Flatten @pytest.mark.covers arg lists into unique, order-preserving cell
|
|
ids, dropping anything that is not a non-empty string."""
|
|
return tuple(dict.fromkeys(arg for args in marker_args for arg in args if isinstance(arg, str) and arg))
|
|
|
|
|
|
def covers_from_item(item: pytest.Item) -> tuple[str, ...]:
|
|
"""Read @pytest.mark.covers cell ids off a pytest Item, order-preserving."""
|
|
return dedupe_covers(marker.args for marker in item.iter_markers(name="covers"))
|
|
|
|
|
|
def result_properties(item: pytest.Item) -> tuple[tuple[str, str], ...]:
|
|
"""The custom signals a standard reporter cannot derive: the normalized suite
|
|
package, the comma-joined coverage-registry cell ids this test covers, the
|
|
repo-relative `path:line` its source sits at, and the typed `@meta(Subject(...))`
|
|
fields.
|
|
|
|
The fixed three-tuple prefix is load-bearing and stays byte-identical: Loki,
|
|
Grafana, the status page and tests/integration/conftest.py all read
|
|
`package`/`covers`/`source` today. `subject_properties` only ever appends, and
|
|
appends nothing at all for a test with no `meta` marker -- which is every test
|
|
in the suite until the backfill lands.
|
|
"""
|
|
fixed = (
|
|
("package", package_from_nodeid(item.nodeid)),
|
|
("covers", ",".join(covers_from_item(item))),
|
|
("source", source_from_item(item)),
|
|
)
|
|
return fixed + case_properties(item.nodeid) + subject_properties(item)
|
|
|
|
|
|
def attach_result_properties(item: pytest.Item) -> None:
|
|
"""Attach result_properties to an item's user_properties, idempotently: a
|
|
second call is a no-op, so a collection that runs the hook more than once
|
|
never emits duplicate <property> entries."""
|
|
if any(name == "package" for name, _ in item.user_properties):
|
|
return
|
|
item.user_properties.extend(result_properties(item))
|
|
|
|
|
|
def attach_step_properties(item: pytest.Item) -> None:
|
|
"""Attach the runtime-recorded steps after the call phase.
|
|
|
|
Separate from `attach_result_properties` because it cannot share its home:
|
|
that one runs in `pytest_collection_modifyitems`, before any test body has
|
|
executed, so the recorder is necessarily empty there.
|
|
|
|
Any `step` entries already on the item are dropped first. The suite runs with
|
|
`--reruns 1`, so a flaky test's second attempt would otherwise append a second
|
|
copy of the story behind the first, and the report would read as one very long
|
|
test that did everything twice. Last attempt wins, which is the attempt whose
|
|
outcome JUnit records.
|
|
"""
|
|
item.user_properties[:] = [entry for entry in item.user_properties if entry[0] != "step"]
|
|
item.user_properties.extend(step_properties())
|