From dc83a9c97909f18b7f85f5dd3e2191ba1c066b09 Mon Sep 17 00:00:00 2001
From: "devin-ai-integration[bot]"
<158243242+devin-ai-integration[bot]@users.noreply.github.com>
Date: Thu, 24 Sep 2026 16:43:34 -0700
Subject: [PATCH] docs(e2e): carve harness tests out of the no-unit-tests hard
rule (#43076)
* docs(e2e): carve harness tests out of the no-unit-tests hard rule
The Hard Rules bullet in tests/e2e/AGENTS.md banned unit tests of any
kind, the harness's own included, while the same file's claude_code/
and load/ entries, CONTRIBUTING.md, and the harness conftest all
describe markerless harness tests that run without a proxy. Reword the
rule to keep the product-feature and no-mocks bans, name the harness
trees as the one exception with the standard they are judged by, and
put the harness sentence back in the marker paragraph so the two docs
agree
* docs(e2e): name env vars set through monkeypatch as inputs, not patches
The rule banned monkeypatching anywhere under tests/e2e while its harness
carve-out named root-level test_*.py files that set env vars through
pytest's monkeypatch fixture. Say the ban is about patching code and that
an env var set that way is an input, so the examples and the rule agree
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
---
tests/e2e/AGENTS.md | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/tests/e2e/AGENTS.md b/tests/e2e/AGENTS.md
index 94035cfe849..b6abcdb6ba2 100644
--- a/tests/e2e/AGENTS.md
+++ b/tests/e2e/AGENTS.md
@@ -98,7 +98,7 @@ Each suite provides its own `client` fixture (see `llm_translation/passthrough_c
Request and response bodies are typed pydantic models in `models.py`; only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a `Result[R]` tagged union (`Success`, `NetworkError`, `UnauthorizedError`, `RateLimitedError`, `ValidationError`, `UnknownApiError`). Handle them with `match`, or call `unwrap(...)` when a non-success should fail the test. The harness hard-fails and never skips: a test marked `e2e` fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
-Mark live tests with `@pytest.mark.e2e` (on the class or the module). Add `@pytest.mark.quiet_stack` to a test that measures the proxy itself (RSS, latency): the shared stack lock in `stack_lock.py` then runs it while no other test on the host is hitting the stack, marked or not, so the reading depends only on the test's own traffic. Use `scoped_key` for a fresh all-models key that auto-deletes, `resources` when you need to create and tear down more than a key, and `unique_marker()` from `e2e_config` to keep prompts, tags, and customer ids from colliding across concurrent runs and the shared response cache
+Mark live tests with `@pytest.mark.e2e` (on the class or the module). Coverage of the harness itself carries no marker and runs whether or not a proxy is up. Add `@pytest.mark.quiet_stack` to a test that measures the proxy itself (RSS, latency): the shared stack lock in `stack_lock.py` then runs it while no other test on the host is hitting the stack, marked or not, so the reading depends only on the test's own traffic. Use `scoped_key` for a fresh all-models key that auto-deletes, `resources` when you need to create and tear down more than a key, and `unique_marker()` from `e2e_config` to keep prompts, tags, and customer ids from colliding across concurrent runs and the shared response cache
## Record and replay fixtures
@@ -251,7 +251,7 @@ other...
```
## Hard Rules
-- no unit tests of any kind under `tests/e2e`. a product feature is proven end to end against a live proxy, never with a unit test, and the harness itself is not unit-tested here either. no monkeypatching or mock tests. if a contributor asks you to write an end to end test, do NOT stage a unit test with it; if you find a product gap, call it out in the PR description
+- no unit tests of a product feature under `tests/e2e`, and no mock tests or monkeypatching of code anywhere in it: a product feature is proven end to end against a live proxy, never with a unit test. if a contributor asks you to write an end to end test, do NOT stage a unit test with it; if you find a product gap, call it out in the PR description. the harness's own plumbing is the one exception: the markerless tests in the root-level `test_*.py` files, `coverage_registry/test_collector.py`, `guardrails/test_guardrails_client.py`, the `claude_code/_*_unit_tests/` trees, and the `load/` aggregation tests carry no `e2e` marker, run without a proxy, and take their inputs as arguments or env vars (setting an env var through pytest's `monkeypatch` fixture is fine, patching a function, class, or module is not), and no coverage-registry or compat-matrix cell rests on them. judge a change inside one of them by that standard, not as a misplaced product test
- use model management endpoints to create new models for a test. this could be in a conftest / inline for each test. ask the user what they want.