* test(e2e): move the harness self-tests out of tests/e2e
The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead
* fix(ci): point the providers replay controls at tests/e2e_harness
The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension
* style(tests): mark the raw-HTTP check's bindings Final
* feat(e2e): give e2e tests typed metadata for what they drive
@meta(Subject(domain, route, providers, models, capabilities, mode)) declares
what a test is about with closed enums, and each field lands in the JUnit report
as a property. The quota_management suites are the first to declare it.
* docs(e2e): say e2e_metadata avoids litellm, not that it is stdlib-only
It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import
* test(e2e): declare models through the constant each test drives
43 @meta declarations in quota_management typed the model name out again, so changing the call would leave the coverage report naming the old model. Each file now has one constant used by both, and a guard fails on any model written as a string literal in @meta
* refactor(e2e): set route only when the endpoint is what the test checks
A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting
* refactor(e2e): build the declared properties without mutating a list
subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping
* fix(e2e): tag each spend-route probe with the endpoint it checks
The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
* feat(e2e): record each e2e test's steps, starting with ProxyClient
@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.
* docs(e2e): rewrite the recorded test steps guide in plain language
* fix(e2e): keep logging callback credentials out of recorded steps
* fix(e2e): mask the run's credentials in every recorded step
* fix(e2e): attach steps before the oauth failure snapshot
The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach
* fix(e2e): name the saved credential in its recorded step
The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
`path:line` cannot represent a path that itself contains a colon, and the
one way pytest produces one is a Windows absolute location: separator
normalization turns `C:\app\e2e\a2a\test_x.py` into `C:/app/...`, which
slipped past the leading-slash check and composed the nonsense repo path
`tests/e2e/C:/app/e2e/a2a/test_x.py`.
Reject the colon itself rather than special-casing a drive letter: it is
the character the format reserves, so no path containing one was ever
linkable.
Claude-Session: https://claude.ai/code/session_017dTKXwJkzhtVLzDhePHsKG
The JUnit report is the only thing that leaves the e2e run, and it says
where a test's results came from but never where its code lives. A reader
looking at `test_cell_claimed_only_by_a_skipped_test_is_uncovered` on the
status page has a name and nothing else -- no file, no line, no way to
reach the source short of grepping the repo by hand.
Pytest knows the location; the report format loses it. The `xunit1` family
wrote `file=` and `line=` onto every `<testcase>`, and the `xunit2` default
this suite runs on drops both. Switching families back would change the
document for every consumer of the same XML -- the Buildkite Test Engine
upload and the Loki pipeline included -- so add the location the way this
suite already adds `package` and `covers`: as a `<property>`, which is
purely additive.
`source` is repo-relative and one-based (`tests/e2e/a2a/test_x.py:41`), so
a consumer can build a link without knowing how pytest was started. That
takes normalizing the two launch shapes -- the runner image runs from its
own copy at /app/e2e, a developer runs from the repo root -- which is the
same normalization `package_from_nodeid` was already doing in reverse, now
factored into `suite_parts` so the two cannot drift apart. Paths that
escape the suite, and tests pytest reports no line for, emit an empty
string: a test with no link beats a link that 404s.
Claude-Session: https://claude.ai/code/session_017dTKXwJkzhtVLzDhePHsKG
* refactor(e2e): replace bespoke result reporter with standard JUnit report
tests/e2e/e2e_result_reporter.py hand-rolled a per-test logfmt emitter that
reimplemented outcome mapping, logfmt escaping, and node-id parsing to print one
E2E_RESULT line per finished test. Outcome, duration, and node id are all things
a standard pytest reporter already produces, so the only genuinely custom data is
the covers marker ids and the normalized package label
Delete the module and emit a standard pytest JUnit XML report (--junitxml)
instead, carrying the two custom signals as user_properties (JUnit <property>
entries) attached at collection time in pytest_collection_modifyitems, so they
land on every test on every outcome including skips and setup errors. The small
package/covers extraction lives in junit_properties.py and is unit tested plus
checked end to end against a real JUnit artifact in test_junit_properties.py
Shipping the JUnit report to Loki is a thin infra-side transform, documented in
grafana/status_history_panels.md
* chore(e2e): remove grafana status history panels doc and junit properties e2e test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>