From 95858e75493804fd445141e0c422595021ea2e8d Mon Sep 17 00:00:00 2001
From: Subham Kundu <43017632+Cenrax@users.noreply.github.com>
Date: Mon, 7 Sep 2026 23:25:22 -0700
Subject: [PATCH 01/53] Update README.md (#3217)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Co-authored-by: Gergő Magyar
---
README.md | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/README.md b/README.md
index f2714bf89..99782975e 100644
--- a/README.md
+++ b/README.md
@@ -26,7 +26,7 @@
- The nervous system for agent context.
+ The context engine for Enterprise Codebases
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow —
From d1463977c8804dbf71d5aeb3cb949c5ba37cc828 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Gerg=C5=91=20Magyar?=
Date: Tue, 8 Sep 2026 08:22:12 +0100
Subject: [PATCH 02/53] feat(eval): Add bounded packed-scheduler primitives and
offline replay benchmarks (#3206)
* perf(eval): packed sweep scheduler and the harness that measured it
Extracted from the combined skill-evolution branch so it can be reviewed on its
own. Purely additive against main: no existing function changes behaviour, and
sweep_packed_cells has no production caller yet.
sweep_task_cells finishes one task before starting the next and drains a wave
before refilling it, so a task with fewer cells than workers leaves workers
idle and one slow cell stalls its whole wave. sweep_packed_cells feeds every
task's cells through a single pool instead, keeping the breaker's meaning: a
total submission order continued across task boundaries, a folder walking
results in that order, and consecutive systemic failures counted there, so a
doomed run aborts on the same cell it would have under waves.
simulate_sweep.py is what produced the numbers. It drives the real schedulers
with only the paid agent session stubbed, using the measured per-arm durations
in session_durations.json divided by a scale factor. The distribution's shape
is kept deliberately - median 826s against a 5400s ceiling - because that
spread is the entire reason a barrier costs anything, and uniform sleeps would
erase the effect under test. All schedulers consume one identical seeded plan.
Measured at workers=3 against the review corpus, packing is worth about 40% of
a cold sweep, and it is the only change that moves a seeded weekly run at all -
there a task is three cells and a wave is never full. The submission window is
a real trade, measured with failures injected at four positions:
window 3 -> -8% wall, overrun 2 (the wave scheduler's own bound)
window 6 -> -27% wall, overrun 4
window 12 -> -42% wall, overrun 9
window 54 -> -44% wall, overrun 11
Overrun is wasted paid sessions on an aborted sweep. The default multiplier is
2; the curve lives in the constant's comment so raising it is an informed
decision. Contention was measured separately by burning real CPU in
subprocesses under taskset: the advantage holds between -40% and -47% from 24
cores down to an oversubscribed 2, though packing erodes faster than waves do
because packing is what creates the concurrency.
measure_evolution_cost.py is the offline cost model, with no runtime caller. It
reports workers from the workflow's current default, which on this base is 1.
Limits worth stating: sleeping threads do not contend and the duration sample
was itself recorded at workers=1, so the speedups are upper bounds; the ordering
of the schedulers is trustworthy because they were compared under identical
conditions, the magnitudes are not.
562 eval tests pass at this base. The two test_model_gateway.py failures,
test_locked_litellm_translates_messages_to_offline_responses and
test_openai_gateway_never_leaves_proxy_output_on_an_undrained_pipe, fail
identically on origin/main in this environment.
* fix(eval): compare the shipped window and bound the overrun by it
Address PR review feedback (#3206).
run_faithful defaulted its submission window to `workers` while
runner.sweep_packed_cells defaults to `max(workers * PACKED_WINDOW_MULTIPLIER,
workers)`, so every run that named no window compared a prototype queued twice
as tightly as the shipped scheduler and presented it as the production
invariant. The faithful default now reads the same constant. Measured at
workers=3, faithful and production agreed on nothing before and agree exactly
now: breaker overrun 2/1/2 vs 2/4/3 becomes 2/4/3 vs 2/4/3 across the three
failure positions.
The contention sweep hard-coded `window=12` for faithful only, which the
production run never saw - masked at workers=6 where both are 12. Removed, and
the production measurement it was already paying for is now reported as
`production_s` instead of being discarded.
breaker_fidelity checked the overrun against `args.workers`. The bound the
producer actually enforces is `window - 1` cells past the fold pointer, which
is the wave scheduler's own `workers - 1` when window == workers; against the
shipped default of 6 the old predicate reported a failure for an in-bound run.
The window is now passed explicitly, reported in each row, and checked against
its own bound.
--window was parsed and never read. Wired into the schedulers that hold one.
Dropped two unused plan constructions CodeQL flagged, and the `skipped` set in
sweep_packed_cells that nothing reads - the None appended to `submitted` is the
skip representation the fold loop consumes.
Verification: 562 passed, 15 skipped, 2 failed (the two test_model_gateway.py
failures the PR description documents as reproducing on origin/main), ruff
clean.
* fix(eval): carry the cancellation scope into packed cells, reject the args that hang
Address PR review feedback (#3206).
sweep_packed_cells submits from a producer THREAD, and a new thread starts with
an empty context, so `copy_context()` there copied the producer's context rather
than the one cancellation_scope had just bound _CANCELLATION in. Every packed
cell therefore ran with no cancellation event, and run_managed falls back to
_CANCELLATION when none is passed - so a cancelled run's subprocesses would
never have learned about it. sweep_task_cells gets this right for free by
submitting from the thread that entered the scope. Reproduced directly: packed
workers observed [False, False], wave workers [True, True]. The caller's context
is now captured before the producer starts and copied per submission; the new
test fails without the fix.
Three CLI arguments were accepted and then wedged the run:
--scale 0 ZeroDivisionError before any scheduler starts
--graph-seconds -1 hangs: the builder thread dies on a negative
sleep, every scheduler waits on a readiness
event nobody sets
--window 0 (faithful) hangs: submitted - fold_pointer >= 0 holds
before the first submission, so the producer
and the consumer wait on each other
The first two are rejected at the parser, which is the only layer that runs
before a thread exists. run_faithful now enforces the same window >= workers
rule sweep_packed_cells already had, so the prototype rejects exactly what the
shipped function rejects. All three were confirmed to crash or hang first.
Verification: 563 passed, 15 skipped, 2 failed (the two test_model_gateway.py
failures the PR description documents as reproducing on origin/main), ruff
clean.
* chore(autofix): apply prettier + eslint fixes via /autofix command
* Address PR review feedback (#3206)
Preserve settled sibling rows when a packed cell raises. run_cell deliberately
lets unexpected harness exceptions propagate, and sweep_task_cells answers that
by folding every non-failing sibling before it re-raises - the cells already ran
and already spent their budget, so dropping their rows means paying for evidence
the sweep then discards. sweep_packed_cells called future.result() bare, so the
fold stopped at the failing index and every later cell that had already
completed was silently lost. It now folds forward over the settled futures
before re-raising. The failing index itself has no row, since execute() assigns
only on success, so folding forward cannot duplicate it.
Pinned by a regression test that fails without the fix: the later cell is made
to finish first, so there is real settled evidence to lose at the moment cell 0
raises.
Reject arguments that cannot produce a run, at the boundary rather than deep
inside a thread. NaN defeats every comparison it appears in, so the existing
"> 0" and ">= 0" checks admitted --scale nan and --graph-seconds nan; the NaN
then reached time.sleep in a worker or the graph thread, raised there, and left
every scheduler waiting forever on a readiness event nobody would set. Infinity
was worse than a crash: it scaled all durations to zero and the run reported a
sweep that took no time. Both flags now require a finite value.
The count flags are indexed or handed straight to a thread pool, so a zero
surfaced as an IndexError on plans[0], a median over an empty sequence, or
ThreadPoolExecutor's own error - none naming the flag responsible. --workers,
--repeat and --runs now require at least 1.
Two flags were not in the review but carry the same invariant and the same
one-line treatment, so they are fixed with the class rather than left to
resurface: --runs (same empty-plan path as --repeat) and --window, where zero
admits no cell at all because the producer waits for a fold pointer to move past
a cell it was never allowed to submit.
Verified each guard fires with its own message rather than a stack trace.
563 eval tests pass. Note: pre-existing failures in test_model_gateway.py not
addressed by this PR - litellm[proxy]'s console script is absent in this
environment, and neither test touches the files changed here.
* Address PR review feedback (#3206), round 2
Stop charging the fed baseline for overlap the wave scheduler gets free.
run_fed is documented as pricing the barrier alone, but it slept graph_seconds
serially before every task, while run_wave starts one background builder that
prepares task N+1 while task N's cells run. The fed-versus-wave delta therefore
mixed the loss of that overlap into what was reported as the price of the
barrier. run_fed now uses the same builder, started before the clock, so the
barrier is the only remaining difference.
This moved the numbers. On the weekly profile fed was 4.203s and is now 3.694s,
exactly equal to wave - which is the answer that profile should give. On cold,
fed was 5.995s and is now 5.487s, so the measured price of the barrier widens
from 1.844s to 2.352s: the old arrangement understated it by about a quarter.
No committed results file or PR-body figure quotes these, so there is nothing
stale to regenerate.
Enforce the window bound the schedulers actually hold. Last round's guard
required only >= 1, but run_faithful and sweep_packed_cells both refuse a window
below the worker count, so --scheduler faithful --workers 3 --window 1 passed
validation and then died on an uncaught ValueError. The check now uses the
worker count.
It also uses the LARGEST worker count the invocation will really use.
--contention-sweep runs its own counts irrespective of --workers, so validating
against --workers alone let the three-worker measurements finish and then raised
on the six-worker one, losing the run partway through. Those counts are now a
named constant the validator can see.
Verified: --scheduler faithful --workers 3 --window 1 is rejected naming 3, and
--workers 3 --window 3 --contention-sweep is rejected naming 6.
No regression test for the graph-overlap fix. Discriminating it from the old
behaviour requires cell work to overlap graph work, which makes the assertion a
timing comparison, and this project does not take non-deterministic tests. It is
verified by the before/after measurement above instead.
564 eval tests pass. Note: pre-existing failures in test_model_gateway.py not
addressed by this PR - litellm[proxy]'s console script is absent here.
---------
Co-authored-by: Gergo Magyar
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
---
eval/tests/test_measure_evolution_cost.py | 122 +++
eval/tests/test_workflow_bench.py | 157 ++++
eval/workflow_bench/measure_evolution_cost.py | 311 ++++++++
eval/workflow_bench/runner.py | 198 +++++
eval/workflow_bench/session_durations.json | 33 +
eval/workflow_bench/simulate_sweep.py | 747 ++++++++++++++++++
6 files changed, 1568 insertions(+)
create mode 100644 eval/tests/test_measure_evolution_cost.py
create mode 100644 eval/workflow_bench/measure_evolution_cost.py
create mode 100644 eval/workflow_bench/session_durations.json
create mode 100644 eval/workflow_bench/simulate_sweep.py
diff --git a/eval/tests/test_measure_evolution_cost.py b/eval/tests/test_measure_evolution_cost.py
new file mode 100644
index 000000000..c6035bc86
--- /dev/null
+++ b/eval/tests/test_measure_evolution_cost.py
@@ -0,0 +1,122 @@
+"""Cost model for the evolution wall clock: measured cells, real schedules."""
+
+from __future__ import annotations
+
+import pytest
+
+from workflow_bench.measure_evolution_cost import (
+ CANDIDATE_ARM,
+ SHA_OVERHEAD_SECONDS,
+ DURATIONS_BY_ARM,
+ PROPOSER_SECONDS,
+ REVIEW_ARMS,
+ expected_task_seconds,
+ fed_makespan,
+ fed_pool_enabled,
+ generation_seconds,
+ graph_pipeline_enabled,
+ paid_arms,
+ task_cells,
+ wave_makespan,
+)
+
+
+def test_every_arm_has_its_own_unsorted_sample():
+ assert set(DURATIONS_BY_ARM) == set(REVIEW_ARMS)
+ for arm, sample in DURATIONS_BY_ARM.items():
+ assert len(sample) >= 10, arm
+ # Sorting would hand each task a uniform block and hide the variance
+ # the whole model exists to price.
+ assert list(sample) != sorted(sample), arm
+ assert PROPOSER_SECONDS > 0
+ assert SHA_OVERHEAD_SECONDS > 0
+
+
+def test_weekly_reuse_pays_the_candidate_arm_only():
+ assert paid_arms(weekly=True, reuse_enabled=True) == (CANDIDATE_ARM,)
+ assert paid_arms(weekly=False, reuse_enabled=True) == REVIEW_ARMS
+ assert paid_arms(weekly=True, reuse_enabled=False) == REVIEW_ARMS
+
+
+def test_cells_are_submitted_run_major_arm_minor():
+ # runner.py: [(run_idx, arm) for run_idx in range(runs) for arm in arms].
+ # At workers=3 that puts one cell of each arm in every wave.
+ cells = task_cells(2, REVIEW_ARMS, 0)
+ assert len(cells) == 6
+ expected = [DURATIONS_BY_ARM[arm][run] for run in range(2) for arm in REVIEW_ARMS]
+ assert cells == expected
+
+
+def test_overhead_is_charged_per_sha_and_outside_the_pool():
+ # Two properties at once: the residual sits outside the schedule, where more
+ # workers cannot dissolve it, and it scales with SHAs rather than cells.
+ assert task_cells(1, (CANDIDATE_ARM,), 0) == [DURATIONS_BY_ARM[CANDIDATE_ARM][0]]
+ wide = generation_seconds(
+ task_count=1, runs=3, arms=REVIEW_ARMS, workers=9, fed_pool=True, unique_shas=5
+ )
+ assert wide >= PROPOSER_SECONDS + 5 * SHA_OVERHEAD_SECONDS
+
+
+def test_sweep_overhead_does_not_shrink_with_the_arm_count():
+ """The bias that made weekly look cheaper than it is.
+
+ A seeded weekly generation pays one arm instead of three but builds exactly
+ the same graphs. Charging the residual per cell billed it a third of a cost
+ the real sweep still pays; per SHA, the two attribute the same setup.
+ """
+
+ kwargs = dict(task_count=6, runs=3, workers=3, fed_pool=False, unique_shas=5)
+ weekly = generation_seconds(arms=(CANDIDATE_ARM,), **kwargs)
+ cold = generation_seconds(arms=REVIEW_ARMS, **kwargs)
+ weekly_sessions = 6 * expected_task_seconds(3, (CANDIDATE_ARM,), 3, fed_pool=False)
+ cold_sessions = 6 * expected_task_seconds(3, REVIEW_ARMS, 3, fed_pool=False)
+ # Whatever each wall is, the non-session part is identical.
+ assert round(weekly - weekly_sessions) == round(cold - cold_sessions)
+ # Cycling wraps, so a task can ask for more runs than the sample holds.
+ long_sample = task_cells(len(DURATIONS_BY_ARM[CANDIDATE_ARM]) + 2, (CANDIDATE_ARM,), 0)
+ assert len(long_sample) == len(DURATIONS_BY_ARM[CANDIDATE_ARM]) + 2
+
+
+def test_a_wave_costs_its_slowest_cell_and_a_fed_pool_does_not():
+ slow = [10.0, 1.0, 1.0, 10.0, 1.0, 1.0]
+ assert wave_makespan(slow, 3) == 20.0
+ # Fed: one worker takes the first 10; the second 10 lands on a worker that
+ # has already cleared a 1, and the remaining 1s fill the third.
+ assert fed_makespan(slow, 3) == 11.0
+ assert fed_makespan(slow, 1) == wave_makespan(slow, 1) == 24.0
+
+
+def test_expected_task_seconds_is_alignment_averaged_and_deterministic():
+ waved = expected_task_seconds(3, REVIEW_ARMS, 3, fed_pool=False)
+ assert waved == expected_task_seconds(3, REVIEW_ARMS, 3, fed_pool=False)
+ assert expected_task_seconds(0, REVIEW_ARMS, 3, fed_pool=False) == 0.0
+ assert expected_task_seconds(3, (), 3, fed_pool=False) == 0.0
+ # The barrier can only cost time, never save it.
+ assert waved >= expected_task_seconds(3, REVIEW_ARMS, 3, fed_pool=True)
+
+
+def test_a_generation_pays_one_proposer_session_on_top_of_its_tasks():
+ one = generation_seconds(
+ task_count=1, runs=3, arms=REVIEW_ARMS, workers=3, fed_pool=False, unique_shas=1
+ )
+ two = generation_seconds(
+ task_count=2, runs=3, arms=REVIEW_ARMS, workers=3, fed_pool=False, unique_shas=1
+ )
+ # Each extra task adds exactly one task's makespan. The proposer and the
+ # per-SHA sweep overhead are both paid once, not per task.
+ assert two - one == pytest.approx(
+ one - PROPOSER_SECONDS - SHA_OVERHEAD_SECONDS, abs=2.0
+ )
+
+
+def test_feature_flags_read_the_runner_not_the_wish():
+ assert graph_pipeline_enabled("def _run_sweep(): pass") == 0
+ assert graph_pipeline_enabled("graph_prefetch = GraphPrefetch(...)") == 1
+ assert fed_pool_enabled("def _run_wave(): pass") == 0
+ assert fed_pool_enabled("def _run_fed_pool(): pass") == 1
+
+
+@pytest.mark.parametrize("workers", [1, 3, 8])
+def test_more_workers_never_lengthen_a_task(workers):
+ serial = expected_task_seconds(3, REVIEW_ARMS, 1, fed_pool=True)
+ assert expected_task_seconds(3, REVIEW_ARMS, workers, fed_pool=True) <= serial
diff --git a/eval/tests/test_workflow_bench.py b/eval/tests/test_workflow_bench.py
index 45d0b2f35..d9c9c0f92 100644
--- a/eval/tests/test_workflow_bench.py
+++ b/eval/tests/test_workflow_bench.py
@@ -5,11 +5,16 @@ import os
import re
import shlex
import subprocess
+import threading
from pathlib import Path
import pytest
import yaml
+from typing import Any
+
+from workflow_bench import runner
+from workflow_bench.process_control import _CANCELLATION, cancellation_scope
from workflow_bench.runner import (
aggregate,
broken_incumbent_arms,
@@ -512,3 +517,155 @@ def test_run_evolution_script_is_the_shared_ci_and_local_entrypoint():
assert "--include-expensive" in argv
assert "claude-sonnet-5" not in argv
assert printed.stderr # rewrite notice goes to stderr
+
+
+def _packed_cells(tasks: int, runs: int, arms: tuple[str, ...]) -> list[tuple[str, int, str]]:
+ return [(f"t{t}", r, a) for t in range(tasks) for r in range(runs) for a in arms]
+
+
+def test_packed_sweep_runs_every_cell_and_folds_in_submission_order():
+ """Fold order is the contract the breaker rests on.
+
+ Cells finish in whatever order the pool returns them, but the breaker counts
+ CONSECUTIVE systemic failures, which only means something in a fixed order.
+ """
+
+ cells = _packed_cells(3, 2, ("review", "candidate_review"))
+ folded: list[tuple[str, int, str]] = []
+ streak, tripped = runner.sweep_packed_cells(
+ cells,
+ workers=4,
+ run=lambda task, run_idx, arm: {"error_kind": None, "review_evidence_valid": True},
+ on_start=lambda *_: None,
+ on_record=lambda task, run_idx, arm, _rec: folded.append((task, run_idx, arm)),
+ outage_streak=0,
+ outage_limit=0,
+ )
+ assert folded == cells
+ assert (streak, tripped) == (0, False)
+
+
+def test_packed_sweep_trips_the_breaker_on_the_same_cell_waves_would():
+ """Packing must not change WHEN a doomed run aborts, only how it is fed."""
+
+ cells = _packed_cells(3, 3, ("review",))
+ fail_from = 2
+ folded: list[int] = []
+
+ def run(task: str, run_idx: int, arm: str) -> dict[str, Any]:
+ index = cells.index((task, run_idx, arm))
+ systemic = index >= fail_from
+ return {
+ "error_kind": "session-error" if systemic else None,
+ "review_evidence_valid": not systemic,
+ }
+
+ streak, tripped = runner.sweep_packed_cells(
+ cells,
+ workers=2,
+ run=run,
+ on_start=lambda *_: None,
+ on_record=lambda t, r, a, _rec: folded.append(cells.index((t, r, a))),
+ outage_streak=0,
+ outage_limit=runner.DEFAULT_OUTAGE_STREAK,
+ )
+ assert tripped is True
+ assert streak == runner.DEFAULT_OUTAGE_STREAK
+ # Five consecutive systemic failures starting at index 2 -> trips on index 6.
+ assert folded[-1] == fail_from + runner.DEFAULT_OUTAGE_STREAK - 1
+ assert folded == sorted(folded), "records must fold in submission order"
+
+
+def test_packed_sweep_skips_a_task_whose_assets_never_arrive():
+ """A task that cannot be prepared is skipped, not run against nothing."""
+
+ cells = _packed_cells(3, 2, ("review",))
+ ran: list[str] = []
+ runner.sweep_packed_cells(
+ cells,
+ workers=3,
+ run=lambda task, run_idx, arm: ran.append(task)
+ or {"error_kind": None, "review_evidence_valid": True},
+ on_start=lambda *_: None,
+ on_record=lambda *_: None,
+ outage_streak=0,
+ outage_limit=0,
+ await_ready=lambda task: task != "t1",
+ )
+ assert set(ran) == {"t0", "t2"}
+ assert "t1" not in ran
+
+
+def test_packed_sweep_workers_inherit_the_runs_cancellation_event():
+ """A worker that cannot see the event runs on after the sweep is cancelled.
+
+ The cells are submitted from a producer THREAD, and a new thread starts with
+ an empty context - so copying the context at submission copies the wrong one
+ unless the caller's is captured first. run_managed falls back to
+ _CANCELLATION when no event is passed, which is how a cell's subprocesses
+ learn the run was cancelled at all.
+ """
+
+ seen: list[threading.Event | None] = []
+ event = threading.Event()
+ with cancellation_scope(event):
+ runner.sweep_packed_cells(
+ _packed_cells(2, 1, ("review",)),
+ workers=2,
+ run=lambda *_: seen.append(_CANCELLATION.get()) or {"error_kind": None},
+ on_start=lambda *_: None,
+ on_record=lambda *_: None,
+ outage_streak=0,
+ outage_limit=0,
+ )
+ assert seen and all(observed is event for observed in seen)
+
+
+def test_packed_sweep_window_must_keep_the_pool_fed():
+ with pytest.raises(ValueError, match="window must be at least workers"):
+ runner.sweep_packed_cells(
+ _packed_cells(1, 1, ("review",)),
+ workers=4,
+ run=lambda *_: {"error_kind": None},
+ on_start=lambda *_: None,
+ on_record=lambda *_: None,
+ outage_streak=0,
+ outage_limit=0,
+ window=2,
+ )
+
+
+def test_a_raising_packed_cell_still_persists_its_settled_siblings():
+ """A crash in one cell must not erase the evidence of cells that finished.
+
+ run_cell deliberately lets unexpected harness exceptions propagate, and the
+ wave scheduler answers that by folding every non-failing sibling before it
+ re-raises. The packed scheduler has to hold the same contract: the later
+ cells already ran and already cost money, so losing their rows would mean
+ paying for evidence the sweep then throws away.
+ """
+
+ folded: list[tuple[int, str]] = []
+ started = threading.Event()
+
+ def run(task_id: str, run_idx: int, arm: str) -> dict[str, Any]:
+ if run_idx == 0:
+ # Let the later cell finish first, so there is settled evidence to
+ # lose at the moment this one raises.
+ started.wait(timeout=5)
+ raise RuntimeError("harness bug in cell 0")
+ started.set()
+ return {"error_kind": None}
+
+ with pytest.raises(RuntimeError, match="harness bug in cell 0"):
+ runner.sweep_packed_cells(
+ _packed_cells(1, 2, ("review",)),
+ workers=2,
+ run=run,
+ on_start=lambda *_: None,
+ on_record=lambda task_id, run_idx, arm, _rec: folded.append((run_idx, arm)),
+ outage_streak=0,
+ outage_limit=0,
+ )
+
+ assert (1, "review") in folded, "the sibling that completed was never recorded"
diff --git a/eval/workflow_bench/measure_evolution_cost.py b/eval/workflow_bench/measure_evolution_cost.py
new file mode 100644
index 000000000..4d18de746
--- /dev/null
+++ b/eval/workflow_bench/measure_evolution_cost.py
@@ -0,0 +1,311 @@
+#!/usr/bin/env python3
+"""Cheap cost model for the skill-evolution review generation.
+
+This is the ce-optimize measurement harness. It does not start Claude and it
+does not replay a run. It reads the review corpus, the evolve defaults and the
+workflow's workers default, then schedules the measured cell durations in
+``session_durations.json`` the way ``sweep_task_cells`` schedules real cells.
+
+Everything priced here is measured. Cell durations and the proposer session
+come from a real artifact, and the work outside the agent sessions comes from
+that run's own step wall minus the time its sessions and proposer account for.
+
+Weekly assumes a matching seed, so every reusable comparator cell is skipped
+and only the candidate arm is paid. Cold assumes an empty seed.
+"""
+
+from __future__ import annotations
+
+import json
+import math
+import re
+import statistics as st
+import subprocess
+import sys
+from pathlib import Path
+
+REPO_ROOT = Path(__file__).resolve().parents[2]
+EVAL_ROOT = REPO_ROOT / "eval"
+REVIEW_TASKS = EVAL_ROOT / "workflow_bench" / "tasks.review.scenarios.yaml"
+EVOLVE_PY = EVAL_ROOT / "workflow_bench" / "evolve.py"
+RUNNER_PY = EVAL_ROOT / "workflow_bench" / "runner.py"
+ARTIFACTS_PY = EVAL_ROOT / "workflow_bench" / "runner_artifacts.py"
+REUSE_PY = EVAL_ROOT / "workflow_bench" / "comparator_reuse.py"
+WORKFLOW = REPO_ROOT / ".github" / "workflows" / "gitnexus-skill-evolution.yml"
+
+MEASURED = json.loads(
+ (Path(__file__).resolve().parent / "session_durations.json").read_text(encoding="utf-8")
+)
+# Per arm, because the arms are not interchangeable and the weekly lane pays
+# only the candidate one. Cells are submitted run-major and arm-minor
+# (runner.py ``planned``), so at workers=3 every wave holds one cell of each
+# arm and the slowest arm sets the wave.
+DURATIONS_BY_ARM: dict[str, tuple[float, ...]] = {
+ arm: tuple(values) for arm, values in MEASURED["cell_duration_s_by_arm"].items()
+}
+PROPOSER_SECONDS: float = MEASURED["proposer_duration_s"]
+_RESIDUAL = MEASURED["residual"]
+# Clone, graph build, sandbox, teardown: the sweep's own time, taken as that
+# run's step wall minus what its sessions and proposer account for. Charged
+# SERIALLY, outside the pool, and charged PER SHA rather than per cell. The
+# residual mixes per-cell work with per-SHA graph setup and the artifact cannot
+# separate them; per-SHA is the direction that refuses to credit a run for
+# shrinking work it still performs, which per-cell did - a weekly generation
+# pays one arm instead of three but builds exactly the same graphs. See
+# session_durations.json residual._split_assumption.
+SHA_OVERHEAD_SECONDS: float = _RESIDUAL["sha_overhead_s"]
+
+# runner.py CANDIDATE_ARMS derives the candidate arm from its incumbent, and
+# only an incumbent row can be reused from a prior generation.
+CANDIDATE_ARM = "candidate_review"
+REVIEW_ARMS = ("ce_review", "review", CANDIDATE_ARM)
+
+SUITE_FILES = (
+ "tests/test_measure_evolution_cost.py",
+ "tests/test_comparator_reuse.py",
+ "tests/test_evolve.py",
+ "tests/test_sanitized_graph.py",
+ "tests/test_workflow_bench.py",
+ "tests/test_workflow_bench_sessions.py",
+ "tests/test_session_progress.py",
+)
+
+
+def _read(path: Path) -> str:
+ return path.read_text(encoding="utf-8")
+
+
+def review_tasks(text: str) -> list[dict[str, str]]:
+ tasks: list[dict[str, str]] = []
+ current: dict[str, str] | None = None
+ for raw in text.splitlines():
+ line = raw.strip()
+ if line.startswith("id:"):
+ if current is not None:
+ tasks.append(current)
+ current = {"id": line.split(":", 1)[1].strip()}
+ elif line.startswith("ref:") and current is not None:
+ current["ref"] = line.split(":", 1)[1].strip()
+ if current is not None:
+ tasks.append(current)
+ return tasks
+
+
+def evolve_default(name: str, text: str) -> int:
+ match = re.search(rf'add_argument\("--{re.escape(name)}".*?default=(\d+)', text, flags=re.S)
+ if match is None:
+ raise ValueError(f"evolve.py is missing --{name} default")
+ return int(match.group(1))
+
+
+def workflow_dispatch_workers(text: str) -> int:
+ match = re.search(r"^\s+workers:\n(?:.*\n)*?^\s+default: '(\d+)'", text, flags=re.M)
+ if match is None:
+ raise ValueError("workflow_dispatch workers default is missing")
+ return int(match.group(1))
+
+
+def feature_enabled() -> tuple[int, int]:
+ evolve = _read(EVOLVE_PY)
+ runner = _read(RUNNER_PY)
+ artifacts = _read(ARTIFACTS_PY)
+ reuse = int(
+ REUSE_PY.is_file()
+ and "--reuse-results" in evolve
+ and "select_reusable_comparator_rows" in runner
+ and "CANDIDATE" in _read(REUSE_PY)
+ )
+ templates = int("def copy_isolated_tree" in artifacts and "clone_templates" in runner)
+ return reuse, templates
+
+
+def graph_pipeline_enabled(runner_text: str) -> int:
+ """True when the runner prefetches the next SHA during paid sessions."""
+
+ return int("prefetch_next_graph" in runner_text or "GraphPrefetch" in runner_text)
+
+
+def fed_pool_enabled(runner_text: str) -> int:
+ """True when the sweep feeds a live pool instead of waiting on waves."""
+
+ return int("def _run_fed_pool" in runner_text)
+
+
+def paid_arms(weekly: bool, reuse_enabled: bool) -> tuple[str, ...]:
+ """Arms a generation actually pays for."""
+
+ if weekly and reuse_enabled:
+ return (CANDIDATE_ARM,)
+ return REVIEW_ARMS
+
+
+def task_cells(runs: int, arms: tuple[str, ...], offset: int) -> list[float]:
+ """One task's cell durations in submission order: run-major, arm-minor.
+
+ Each arm draws from its own measured sample, cycled from ``offset`` so the
+ caller can average over every alignment instead of trusting one.
+ """
+
+ cells: list[float] = []
+ for run_idx in range(runs):
+ for arm in arms:
+ sample = DURATIONS_BY_ARM[arm]
+ cells.append(sample[(offset + run_idx) % len(sample)])
+ return cells
+
+
+def wave_makespan(durations: list[float], workers: int) -> float:
+ """Today's scheduler: fixed waves of ``workers``, with a barrier between."""
+
+ return sum(
+ max(durations[start : start + workers]) for start in range(0, len(durations), workers)
+ )
+
+
+def fed_makespan(durations: list[float], workers: int) -> float:
+ """Continuously fed pool: a free worker takes the next cell immediately."""
+
+ busy_until = [0.0] * workers
+ for duration in durations:
+ first = min(range(workers), key=busy_until.__getitem__)
+ busy_until[first] += duration
+ return max(busy_until)
+
+
+def expected_task_seconds(
+ runs: int, arms: tuple[str, ...], workers: int, *, fed_pool: bool
+) -> float:
+ """Mean makespan of one task over every alignment of the measured samples.
+
+ One fixed alignment would let an accident of the source run - its slowest
+ cells happen to come first - decide the answer. Averaging keeps the real
+ multiset and the real ordering effects without that artifact, and stays
+ deterministic.
+ """
+
+ if runs < 1 or not arms:
+ return 0.0
+ makespan = fed_makespan if fed_pool else wave_makespan
+ # lcm, not max: with samples of 13 and 14, max would wrap the shorter one
+ # and count its first entry twice.
+ alignments = math.lcm(*(len(DURATIONS_BY_ARM[arm]) for arm in arms))
+ return (
+ sum(makespan(task_cells(runs, arms, offset), workers) for offset in range(alignments))
+ / alignments
+ )
+
+
+def generation_seconds(
+ *,
+ task_count: int,
+ runs: int,
+ arms: tuple[str, ...],
+ workers: int,
+ fed_pool: bool,
+ unique_shas: int,
+) -> int:
+ """Whole generation: proposer, then the tasks back to back, plus overhead.
+
+ Prices a HEALTHY sweep. A run whose cells return unusable evidence does not
+ reach this wall at all: the outage breaker aborts after
+ ``DEFAULT_OUTAGE_STREAK`` consecutive systemic failures, which for the
+ sample's own error sequence is cell 5 of 41.
+
+ Sweep overhead is charged per SHA, so it does not shrink with the arm count.
+ Weekly pays one arm instead of three but builds the same graphs, and billing
+ that per cell credited it for a saving the real run never makes.
+ """
+
+ return round(
+ PROPOSER_SECONDS
+ + task_count * expected_task_seconds(runs, arms, workers, fed_pool=fed_pool)
+ + unique_shas * SHA_OVERHEAD_SECONDS
+ )
+
+
+def _pytest_python() -> list[str]:
+ venv_python = EVAL_ROOT / ".venv" / "bin" / "python"
+ if venv_python.is_file():
+ return [str(venv_python)]
+ if (EVAL_ROOT / "uv.lock").is_file():
+ return ["uv", "run", "--locked", "--extra", "dev", "python"]
+ return [sys.executable]
+
+
+def suite_passed() -> int:
+ files = [name for name in SUITE_FILES if (EVAL_ROOT / name).is_file()]
+ if not files:
+ return 0
+ cmd = [*_pytest_python(), "-m", "pytest", *files, "-q", "--tb=no", "--no-header"]
+ try:
+ completed = subprocess.run(
+ cmd, cwd=EVAL_ROOT, check=False, capture_output=True, text=True, timeout=240
+ )
+ except (OSError, subprocess.TimeoutExpired):
+ return 0
+ return int(completed.returncode == 0)
+
+
+def main() -> int:
+ tasks = review_tasks(_read(REVIEW_TASKS))
+ evolve = _read(EVOLVE_PY)
+ runner = _read(RUNNER_PY)
+ runs = evolve_default("runs", evolve)
+ workers = workflow_dispatch_workers(_read(WORKFLOW))
+ reuse_enabled, clone_templates_enabled = feature_enabled()
+ fed_pool = fed_pool_enabled(runner)
+
+ # Both walls build the same graphs; the arm count does not change that.
+ unique_shas = len({t.get("ref", "") for t in tasks if t.get("ref")})
+ payload: dict[str, object] = {}
+ for label, weekly in (("weekly", True), ("cold", False)):
+ arms = paid_arms(weekly, bool(reuse_enabled))
+ payload[f"estimated_{label}_wall_seconds"] = generation_seconds(
+ task_count=len(tasks),
+ runs=runs,
+ arms=arms,
+ workers=workers,
+ fed_pool=bool(fed_pool),
+ unique_shas=unique_shas,
+ )
+ payload[f"paid_{label}_cells"] = len(tasks) * runs * len(arms)
+ # What the wave barrier costs: the same cells, continuously fed.
+ payload[f"fed_pool_{label}_wall_seconds"] = generation_seconds(
+ task_count=len(tasks),
+ runs=runs,
+ arms=arms,
+ workers=workers,
+ fed_pool=True,
+ unique_shas=unique_shas,
+ )
+
+ all_durations = [d for sample in DURATIONS_BY_ARM.values() for d in sample]
+ payload.update(
+ {
+ "suite_passed": suite_passed(),
+ "promotion_min_runs": evolve_default("promotion-min-runs", evolve),
+ "review_task_count": len(tasks),
+ "candidate_cells": len(tasks) * runs,
+ "workers": workers,
+ "unique_task_shas": len({t.get("ref", "") for t in tasks if t.get("ref")}),
+ "reuse_enabled": reuse_enabled,
+ "clone_templates_enabled": clone_templates_enabled,
+ "graph_pipeline_enabled": graph_pipeline_enabled(runner),
+ "fed_pool_enabled": fed_pool,
+ "measured_cell_count": len(all_durations),
+ "median_cell_seconds": round(st.median(all_durations)),
+ "mean_cell_seconds": round(st.mean(all_durations)),
+ "max_cell_seconds": round(max(all_durations)),
+ "median_candidate_cell_seconds": round(st.median(DURATIONS_BY_ARM[CANDIDATE_ARM])),
+ "mean_candidate_cell_seconds": round(st.mean(DURATIONS_BY_ARM[CANDIDATE_ARM])),
+ "proposer_seconds": round(PROPOSER_SECONDS),
+ "sha_overhead_seconds": round(SHA_OVERHEAD_SECONDS, 1),
+ }
+ )
+ json.dump(payload, sys.stdout, sort_keys=True)
+ sys.stdout.write("\n")
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/eval/workflow_bench/runner.py b/eval/workflow_bench/runner.py
index 82185186c..9f9dd9ef3 100644
--- a/eval/workflow_bench/runner.py
+++ b/eval/workflow_bench/runner.py
@@ -851,6 +851,204 @@ def sweep_task_cells(
return outage_streak, False
+# A sustained upstream outage shows up as a run of session/infra/cleanup
+# failures. (cleanup-failure overwrites the primary error_kind, so a
+# session-error whose worktree cleanup also failed still counts.) A task's own
+# resolved=False is real signal, not an outage, so it never trips the breaker.
+# How far ahead of the in-order fold pointer cells may be submitted, as a
+# multiple of the worker count. This is the wall-clock/wasted-cell trade, and it
+# is a real one - measured against the review corpus at workers=3, with failures
+# injected at four different positions:
+#
+# window wall vs waves worst overrun
+# 3 -8% 2 (the wave scheduler's own bound)
+# 6 -27% 4
+# 12 -42% 9
+# 54 -44% 11
+#
+# Overrun is wasted paid sessions when the breaker trips, at roughly $70 each.
+# 2 is the default because it keeps the worst case within 2x the wave bound
+# while taking most of the gain; raise it if a run's wall clock costs more than
+# an occasional handful of cells on an aborted sweep.
+PACKED_WINDOW_MULTIPLIER = 2
+
+
+def sweep_packed_cells(
+ cells: Sequence[tuple[str, int, str]],
+ *,
+ workers: int,
+ run: Callable[[str, int, str], dict[str, Any]],
+ on_start: Callable[[str, int, str], None],
+ on_record: Callable[[str, int, str, dict[str, Any]], None],
+ outage_streak: int,
+ outage_limit: int,
+ window: int | None = None,
+ await_ready: Callable[[str], bool] | None = None,
+ cancel_event: threading.Event | None = None,
+) -> tuple[int, bool]:
+ """Run cells from EVERY task through one pool; return (streak, tripped).
+
+ ``sweep_task_cells`` finishes one task before starting the next and drains a
+ wave before refilling it, so a task with fewer cells than ``workers`` leaves
+ workers idle and a slow cell stalls its whole wave. Packing every task's
+ cells into one continuously fed pool removes both, which is worth about 40%
+ of a cold sweep's wall clock and is the only thing that moves a seeded
+ weekly run at all - there, a task is three cells and a wave is never full.
+
+ The breaker keeps its exact meaning. ``cells`` is a total submission order
+ (task-major, run-major, arm-minor - the same order waves fold in, continued
+ across task boundaries), a folder walks results in precisely that order, and
+ "consecutive systemic failures" is evaluated there. So the run aborts on the
+ same logical cell it would have aborted on under waves.
+
+ ``window`` is what bounds the overrun, and it is load-bearing. The halt flag
+ alone is not enough: the folder walks in order, so a slow early cell lets
+ workers race ahead, and by the time the breaker trips those cells have
+ already paid for their sessions. Measured, an unbounded queue overran by 11
+ cells at ``workers=3`` where the wave scheduler overruns by 2. Holding
+ submission to ``window`` cells beyond the fold point caps it, trading
+ packing for wasted cells - see ``PACKED_WINDOW_MULTIPLIER`` for the curve.
+
+ ``await_ready`` gates a task's first cell on whatever that task still needs
+ (a sanitized clone, a graph). It returns False to abandon the task, whose
+ cells are then skipped rather than run against missing assets. Cells are
+ submitted as their task becomes ready, so a later task's graph builds while
+ earlier cells are still paying for sessions.
+ """
+
+ with cancellation_scope(cancel_event) as cancel_event:
+ if workers < 1:
+ raise ValueError("workers must be positive")
+ if not cells:
+ return outage_streak, False
+ if window is None:
+ window = max(workers * PACKED_WINDOW_MULTIPLIER, workers)
+ if window < workers:
+ raise ValueError("window must be at least workers, or the pool starves")
+
+ halt = threading.Event()
+ results: list[dict[str, Any] | None] = [None] * len(cells)
+ submitted: list[Any] = []
+ gate = threading.Condition()
+ producing = True
+ fold_pointer = 0
+
+ def execute(index: int) -> None:
+ if halt.is_set() or cancel_event.is_set():
+ return
+ task_id, run_idx, arm = cells[index]
+ on_start(task_id, run_idx, arm)
+ results[index] = run(task_id, run_idx, arm)
+
+ pool = ThreadPoolExecutor(max_workers=workers)
+ # cancellation_scope binds _CANCELLATION in the CALLING thread's
+ # context, and a new thread starts with an empty one - so the producer
+ # has to copy this context rather than its own, or every cell it
+ # submits loses the run's cancellation event. sweep_task_cells gets
+ # this for free by submitting from the thread that entered the scope.
+ caller_context = copy_context()
+
+ def produce() -> None:
+ nonlocal producing
+ ready_tasks: dict[str, bool] = {}
+ try:
+ for index, (task_id, _run_idx, _arm) in enumerate(cells):
+ if halt.is_set() or cancel_event.is_set():
+ break
+ if task_id not in ready_tasks:
+ ready_tasks[task_id] = True if await_ready is None else await_ready(task_id)
+ if not ready_tasks[task_id]:
+ with gate:
+ submitted.append(None)
+ gate.notify_all()
+ continue
+ with gate:
+ while index - fold_pointer >= window and not halt.is_set():
+ gate.wait(timeout=0.5)
+ if halt.is_set() or cancel_event.is_set():
+ break
+ worker_context = caller_context.run(copy_context)
+ submitted.append(pool.submit(worker_context.run, execute, index))
+ gate.notify_all()
+ finally:
+ with gate:
+ producing = False
+ gate.notify_all()
+
+ producer = threading.Thread(target=produce, name="packed-cell-producer", daemon=False)
+ producer.start()
+
+ tripped = False
+ try:
+ index = 0
+ while True:
+ with gate:
+ while index >= len(submitted) and producing:
+ gate.wait(timeout=0.5)
+ if index >= len(submitted):
+ break
+ future = submitted[index]
+ if future is not None:
+ try:
+ future.result()
+ except BaseException:
+ # Same contract as sweep_task_cells: the cells submitted
+ # after this one have already run and spent their budget,
+ # so persist their rows in submission order before the
+ # harness bug takes the process down. Without this, one
+ # crashing cell silently erases the paid evidence of
+ # every sibling that had already finished. The failing
+ # index itself has no row - execute() only assigns on
+ # success - so folding forward cannot duplicate it.
+ with gate:
+ settled = list(submitted)
+ for later in range(index + 1, len(settled)):
+ pending = settled[later]
+ if pending is not None and not pending.done():
+ continue
+ row = results[later]
+ if row is not None:
+ on_record(*cells[later], row)
+ raise
+ record = results[index]
+ if record is not None:
+ task_id, run_idx, arm = cells[index]
+ on_record(task_id, run_idx, arm, record)
+ kind = (
+ "review-evidence-invalid"
+ if record.get("review_evidence_valid") is False
+ else record.get("error_kind")
+ )
+ outage_streak = systemic_outage_streak(kind, outage_streak)
+ if outage_limit and outage_streak >= outage_limit:
+ print(
+ f"[systemic-outage] {outage_streak} consecutive unusable-evidence "
+ "failures — aborting the remaining sweep; report and promotion are "
+ "written from partial evidence and the run exits non-zero."
+ )
+ tripped = True
+ halt.set()
+ cancel_event.set()
+ break
+ index += 1
+ with gate:
+ fold_pointer = index
+ gate.notify_all()
+ if cancel_event.is_set():
+ tripped = True
+ break
+ finally:
+ halt.set()
+ with gate:
+ gate.notify_all()
+ producer.join()
+ for pending in submitted[index + 1 :]:
+ if pending is not None:
+ pending.cancel()
+ pool.shutdown(wait=True)
+ return outage_streak, tripped
+
+
@dataclass(frozen=True)
class TaskCellContext:
"""Everything one benchmark cell needs from its task, prepared once.
diff --git a/eval/workflow_bench/session_durations.json b/eval/workflow_bench/session_durations.json
new file mode 100644
index 000000000..dbb5e9004
--- /dev/null
+++ b/eval/workflow_bench/session_durations.json
@@ -0,0 +1,33 @@
+{
+ "_provenance": "Actions run 33912693948 (2026-09-04), review profile, gen-0, workers=1. Artifact gitnexus-evolution-33912693948-1: gen-0/bench/results.jsonl and gen-0/proposer-session.json. Step wall from the Actions API.",
+ "_caveat": "Every cell in that run returned unusable evidence (32 review-evidence-invalid, 6 session-error, 3 skill-not-invoked); two hit the 5400s ceiling and it cost 653. Durations are real, but a run that resolves cleanly may sit lower. It is the only live artifact - the 2026-07-22 green run's has expired.",
+ "_order": "Submission order, deliberately unsorted: the model cycles these, so sorting would hand each task a uniform block and hide the variance being measured.",
+ "_duration_scope": "duration_s is the sum of the cell's Claude session durations (runner_sessions.py). It excludes the clone, graph materialize, asset staging, sandbox setup and teardown - those live in the residual below.",
+ "session_ceiling_s": 5400,
+ "proposer_duration_s": 344.7,
+ "cell_duration_s_by_arm": {
+ "candidate_review": [
+ 2485.6, 1338.0, 3075.2, 702.5, 1240.4, 5400.0, 762.1, 342.9, 826.3, 489.7, 675.4, 337.8, 734.3
+ ],
+ "ce_review": [
+ 3744.6, 2140.4, 1418.4, 436.1, 653.3, 1022.3, 851.2, 1191.1, 902.1, 847.2, 502.6, 991.3,
+ 1222.7, 544.0
+ ],
+ "review": [
+ 5400.0, 2976.4, 1162.8, 627.1, 901.3, 436.9, 704.3, 963.4, 631.8, 627.9, 663.5, 662.4, 361.3,
+ 741.1
+ ]
+ },
+ "residual": {
+ "benchmark_step_wall_s": 54623,
+ "session_seconds": 51737.7,
+ "proposer_seconds": 344.7,
+ "unaccounted_s": 2540.6,
+ "cells": 41,
+ "unique_shas": 5,
+ "_note": "Everything the sweep spent outside the agent sessions: per-SHA sanitize and `analyze --pdg --index-only`, plus each cell's clone, materialize, staging, sandbox and teardown. That run predates clone templates and graph prefetch, so this is an upper bound for the current code. The split between per-SHA and per-cell is not recoverable from the artifact, so the model charges it per cell and serially, outside the pool - the pessimistic reading of an already-small term.",
+ "_split_assumption": "The residual mixes per-SHA graph setup with per-cell clone/sandbox/teardown and the artifact cannot separate them. The model charges it per SHA, not per cell, because only that direction refuses to credit a run for shrinking work it still performs: a weekly generation pays one arm instead of three but builds the same graphs. This overstates cold slightly and refuses to understate weekly. Replace with measured per-SHA and per-cell times when a run records them separately.",
+ "sha_overhead_s": 508.1
+ },
+ "_breaker": "Replaying this sample's error_kind sequence through today's systemic_outage_streak trips the outage breaker at cell 5 of 41 (DEFAULT_OUTAGE_STREAK=5). The source run executed all 41, so its runner did not break on this sequence. The durations stay valid as per-cell timings; what they cannot describe is a 54-cell sweep with this failure profile, because the current code would never run one."
+}
diff --git a/eval/workflow_bench/simulate_sweep.py b/eval/workflow_bench/simulate_sweep.py
new file mode 100644
index 000000000..764460eea
--- /dev/null
+++ b/eval/workflow_bench/simulate_sweep.py
@@ -0,0 +1,747 @@
+#!/usr/bin/env python3
+"""Run the real sweep scheduler against stub sessions and time it.
+
+``measure_evolution_cost`` is arithmetic: it predicts wall clock from a model of
+what ``sweep_task_cells`` does. This runs the actual function - real threads,
+the real wave barrier, the real outage breaker - and replaces only the paid
+agent session with a sleep. If the two disagree, the model is wrong.
+
+Durations are the measured per-arm samples from ``session_durations.json``
+divided by ``--scale``, so a cell that really took 1416s takes ~0.28s here. The
+shape is preserved deliberately: the median cell is 826s against a 5400s
+ceiling, and that spread is the whole reason a barrier costs anything. Uniform
+random sleeps would erase the effect under test.
+
+Schedulers, all consuming one identical seeded plan:
+
+``wave`` the shipped ``sweep_task_cells`` - fixed waves of ``workers``, a
+ barrier between them, one task at a time.
+``fed`` a continuously fed pool per task (H1). Naive: no breaker, no graph
+ gating. Present to price the barrier alone.
+``packed`` one pool across every task (H2). Naive, same caveat.
+``faithful``H2 carrying the invariants the shipped scheduler actually holds:
+ a global submission order, in-order folding, the outage breaker, and
+ per-task graph readiness gating. This is the one to believe.
+
+ python3 -m workflow_bench.simulate_sweep --compare --repeat 5
+ python3 -m workflow_bench.simulate_sweep --breaker-fidelity
+"""
+
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import random
+import statistics
+import subprocess
+import sys
+import threading
+import time
+from concurrent.futures import ThreadPoolExecutor
+from dataclasses import dataclass, field
+from typing import Any
+
+from . import runner
+from .measure_evolution_cost import (
+ CANDIDATE_ARM,
+ DURATIONS_BY_ARM,
+ REVIEW_ARMS,
+ REVIEW_TASKS,
+ SHA_OVERHEAD_SECONDS,
+ _read,
+ expected_task_seconds,
+ review_tasks,
+)
+
+DEFAULT_SCALE = 5000.0
+# --contention-sweep measures both of these regardless of --workers, so the
+# window has to be valid for the LARGEST of them, not for the parsed value.
+CONTENTION_WORKERS = (3, 6)
+SYSTEMIC_KIND = "session-error"
+
+# A cell is mostly a model session waiting on the network, but its tool calls -
+# git, vitest, analyze - burn real CPU in real subprocesses. Sleeping threads
+# model the wait and nothing else, so every speedup measured that way is an
+# upper bound. This burns WORK, not wall clock: a fixed number of sha256 rounds
+# in a subprocess, which takes longer when cores are contended. That is the
+# effect under test, and it has to be a subprocess - Python threads burning
+# Python would measure the GIL rather than the machine.
+_BURN_SRC = (
+ "import hashlib,sys\n"
+ "n=int(sys.argv[1]); b=b'x'*4096; h=hashlib.sha256()\n"
+ "for _ in range(n): h.update(b)\n"
+ "sys.stdout.write(h.hexdigest()[:8])\n"
+)
+
+
+def calibrate_burn(probe_rounds: int = 400_000) -> float:
+ """sha256 rounds per second, one uncontended subprocess. Measured, not assumed."""
+
+ started = time.monotonic()
+ subprocess.run(
+ [sys.executable, "-c", _BURN_SRC, str(probe_rounds)],
+ check=True,
+ capture_output=True,
+ )
+ return probe_rounds / (time.monotonic() - started)
+
+
+def _execute_cell(cell: Cell, cpu_fraction: float, burn_rate: float) -> None:
+ """The stub session: wait for the API, then do the tool-call work."""
+
+ if cpu_fraction <= 0:
+ time.sleep(cell.seconds)
+ return
+ time.sleep(cell.seconds * (1.0 - cpu_fraction))
+ rounds = int(cell.seconds * cpu_fraction * burn_rate)
+ if rounds > 0:
+ subprocess.run(
+ [sys.executable, "-c", _BURN_SRC, str(rounds)], check=True, capture_output=True
+ )
+
+
+@dataclass(frozen=True)
+class Cell:
+ task: int
+ run: int
+ arm: str
+ seconds: float
+ systemic: bool = False
+
+
+@dataclass
+class Outcome:
+ wall_s: float
+ executed: int
+ tripped_at: int | None = None
+ folded: list[int] = field(default_factory=list)
+
+
+def build_plan(
+ *,
+ task_count: int,
+ runs: int,
+ arms: tuple[str, ...],
+ scale: float,
+ seed: int,
+ fail_from: int | None = None,
+) -> list[list[Cell]]:
+ """Per-task cells in submission order, with durations drawn once.
+
+ Shared by every scheduler so a comparison cannot be an artifact of one of
+ them drawing luckier cells. ``fail_from`` marks every cell at or after that
+ global index systemic, which is what the breaker-fidelity mode needs.
+ """
+
+ rng = random.Random(seed)
+ plan: list[list[Cell]] = []
+ index = 0
+ for task in range(task_count):
+ cells: list[Cell] = []
+ for run_idx in range(runs):
+ for arm in arms:
+ sample = DURATIONS_BY_ARM[arm]
+ cells.append(
+ Cell(
+ task=task,
+ run=run_idx,
+ arm=arm,
+ seconds=sample[rng.randrange(len(sample))] / scale,
+ systemic=fail_from is not None and index >= fail_from,
+ )
+ )
+ index += 1
+ plan.append(cells)
+ return plan
+
+
+def _flatten(plan: list[list[Cell]]) -> list[Cell]:
+ return [cell for cells in plan for cell in cells]
+
+
+def _record(cell: Cell) -> dict[str, Any]:
+ kind = SYSTEMIC_KIND if cell.systemic else None
+ return {
+ "run": cell.run,
+ "arm": cell.arm,
+ "ok": not cell.systemic,
+ "resolved": not cell.systemic,
+ "error_kind": kind,
+ "review_evidence_valid": not cell.systemic,
+ }
+
+
+def _graph_builder(
+ ready: list[threading.Event], graph_seconds: float, stop: threading.Event
+) -> threading.Thread:
+ """One graph at a time, in task order - they are CPU and IO heavy."""
+
+ def build() -> None:
+ for event in ready:
+ if stop.is_set():
+ return
+ time.sleep(graph_seconds)
+ event.set()
+
+ thread = threading.Thread(target=build, name="graph-builder", daemon=True)
+ thread.start()
+ return thread
+
+
+def run_wave(
+ plan: list[list[Cell]],
+ workers: int,
+ *,
+ outage_limit: int,
+ graph_seconds: float,
+ cpu_fraction: float = 0.0,
+ burn_rate: float = 0.0,
+) -> Outcome:
+ """The shipped scheduler, driven for real, task after task."""
+
+ ready = [threading.Event() for _ in plan]
+ stop = threading.Event()
+ _graph_builder(ready, graph_seconds, stop)
+ executed = 0
+ lock = threading.Lock()
+ streak = 0
+ tripped_at: int | None = None
+ folded: list[int] = []
+ base = 0
+
+ started = time.monotonic()
+ for task, cells in enumerate(plan):
+ ready[task].wait()
+ by_key = {(c.run, c.arm): c for c in cells}
+
+ def fake_run(run_idx: int, arm: str) -> dict[str, Any]:
+ nonlocal executed
+ cell = by_key[(run_idx, arm)]
+ _execute_cell(cell, cpu_fraction, burn_rate)
+ with lock:
+ executed += 1
+ return _record(cell)
+
+ order = {(c.run, c.arm): base + i for i, c in enumerate(cells)}
+
+ def on_record(run_idx: int, arm: str, rec: dict[str, Any]) -> None:
+ # Mirror the breaker's own evaluation so the reported trip point is
+ # the cell that crossed the limit, not merely the last one folded -
+ # sweep_task_cells folds a whole wave before it evaluates.
+ nonlocal streak, tripped_at
+ index = order[(run_idx, arm)]
+ folded.append(index)
+ streak = runner.systemic_outage_streak(rec["error_kind"], streak)
+ if outage_limit and streak >= outage_limit and tripped_at is None:
+ tripped_at = index
+
+ streak, tripped = runner.sweep_task_cells(
+ [(c.run, c.arm) for c in cells],
+ workers=workers,
+ run=fake_run,
+ on_start=lambda *_: None,
+ on_record=on_record,
+ outage_streak=streak,
+ outage_limit=outage_limit,
+ )
+ base += len(cells)
+ if tripped:
+ break
+ stop.set()
+ return Outcome(wall_s=time.monotonic() - started, executed=executed, tripped_at=tripped_at, folded=folded)
+
+
+def _drain_naive(cells: list[Cell], workers: int) -> int:
+ with ThreadPoolExecutor(max_workers=workers) as pool:
+ list(pool.map(lambda c: time.sleep(c.seconds), cells))
+ return len(cells)
+
+
+def run_fed(plan: list[list[Cell]], workers: int, *, outage_limit: int, graph_seconds: float) -> Outcome:
+ """H1 without invariants: fed pool per task. Prices the barrier alone.
+
+ Graph building is deliberately identical to ``run_wave`` - the same
+ background builder, started before the clock - because that is what makes
+ the claim in the first line true. Sleeping ``graph_seconds`` serially before
+ each task instead, as this did, charged fed for overlap that wave gets for
+ free: the wave builder prepares task N+1 while task N's cells run. The
+ fed-versus-wave delta then mixed the loss of that overlap into what was
+ reported as the price of the barrier.
+ """
+
+ ready = [threading.Event() for _ in plan]
+ stop = threading.Event()
+ _graph_builder(ready, graph_seconds, stop)
+ executed = 0
+ started = time.monotonic()
+ for task, cells in enumerate(plan):
+ ready[task].wait()
+ executed += _drain_naive(cells, workers)
+ return Outcome(wall_s=time.monotonic() - started, executed=executed)
+
+
+def run_packed(plan: list[list[Cell]], workers: int, *, outage_limit: int, graph_seconds: float) -> Outcome:
+ """H2 without invariants. Upper bound, not a design."""
+
+ started = time.monotonic()
+ time.sleep(graph_seconds)
+ executed = _drain_naive(_flatten(plan), workers)
+ return Outcome(wall_s=time.monotonic() - started, executed=executed)
+
+
+def run_faithful(
+ plan: list[list[Cell]],
+ workers: int,
+ *,
+ outage_limit: int,
+ graph_seconds: float,
+ window: int | None = None,
+ cpu_fraction: float = 0.0,
+ burn_rate: float = 0.0,
+) -> Outcome:
+ """H2 carrying the invariants the shipped scheduler holds.
+
+ Global submission order is task-major, run-major, arm-minor - the same total
+ order the wave scheduler folds in, just continued across task boundaries. A
+ folder walks results in exactly that order, so "consecutive systemic
+ failures" keeps its meaning; the breaker trips on the same logical cell it
+ would have in waves. Cells already in flight when it trips are the overrun,
+ bounded by ``workers - 1`` exactly as the wave docstring promises.
+
+ A task's cells are not submitted until its graph is ready, which is what
+ makes this a schedule rather than a wish: the graph builder is serial, so
+ packing cannot outrun it.
+
+ ``window`` is the design question. Queue every cell at once and workers race
+ far ahead of the fold pointer, so a breaker trip has already paid for cells
+ nobody has looked at - measured at 5 against a bound of 2. Holding
+ submission to ``window`` cells beyond the fold point caps the overrun at
+ ``window - 1``, which is the wave's own ``workers - 1`` bound when the two
+ are equal, while still packing across task boundaries. Defaults to whatever
+ ``runner.sweep_packed_cells`` defaults to, so a run that names no window
+ compares the shipped policy rather than a more tightly queued prototype.
+ """
+
+ if window is None:
+ window = max(workers * runner.PACKED_WINDOW_MULTIPLIER, workers)
+ if window < workers:
+ # Same rule sweep_packed_cells enforces. Without it a window below 1
+ # never lets the producer past its own gate and the run hangs.
+ raise ValueError("window must be at least workers, or the pool starves")
+
+ cells = _flatten(plan)
+ ready = [threading.Event() for _ in plan]
+ stop = threading.Event()
+ _graph_builder(ready, graph_seconds, stop)
+
+ results: list[dict[str, Any] | None] = [None] * len(cells)
+ executed = 0
+ lock = threading.Lock()
+ halt = threading.Event()
+
+ def work(index: int) -> None:
+ nonlocal executed
+ if halt.is_set():
+ return
+ cell = cells[index]
+ _execute_cell(cell, cpu_fraction, burn_rate)
+ with lock:
+ executed += 1
+ results[index] = _record(cell)
+
+ gate = threading.Condition()
+ fold_pointer = 0
+ futures: list[Any] = []
+ producer_done = threading.Event()
+
+ started = time.monotonic()
+ pool = ThreadPoolExecutor(max_workers=workers)
+
+ def produce() -> None:
+ submitted = 0
+ for task, task_cells in enumerate(plan):
+ ready[task].wait()
+ for _ in task_cells:
+ with gate:
+ while submitted - fold_pointer >= window and not halt.is_set():
+ gate.wait(timeout=0.5)
+ if halt.is_set():
+ producer_done.set()
+ return
+ futures.append(pool.submit(work, submitted))
+ submitted += 1
+ gate.notify_all()
+ producer_done.set()
+
+ producer = threading.Thread(target=produce, name="cell-producer", daemon=True)
+ producer.start()
+
+ streak = 0
+ tripped_at: int | None = None
+ folded: list[int] = []
+ try:
+ index = 0
+ while True:
+ with gate:
+ while index >= len(futures) and not producer_done.is_set():
+ gate.wait(timeout=0.5)
+ if index >= len(futures):
+ break
+ future = futures[index]
+ future.result()
+ record = results[index]
+ if record is not None:
+ folded.append(index)
+ streak = runner.systemic_outage_streak(record["error_kind"], streak)
+ if outage_limit and streak >= outage_limit:
+ tripped_at = index
+ halt.set()
+ with gate:
+ gate.notify_all()
+ for pending in futures[index + 1 :]:
+ pending.cancel()
+ break
+ index += 1
+ with gate:
+ fold_pointer = index
+ gate.notify_all()
+ finally:
+ halt.set()
+ with gate:
+ gate.notify_all()
+ stop.set()
+ producer.join(timeout=5)
+ pool.shutdown(wait=True)
+ return Outcome(
+ wall_s=time.monotonic() - started, executed=executed, tripped_at=tripped_at, folded=folded
+ )
+
+
+def run_production_packed(
+ plan: list[list[Cell]],
+ workers: int,
+ *,
+ outage_limit: int,
+ graph_seconds: float,
+ cpu_fraction: float = 0.0,
+ burn_rate: float = 0.0,
+ window: int | None = None,
+) -> Outcome:
+ """Drive the REAL runner.sweep_packed_cells, not a prototype of it.
+
+ Same relationship run_wave has to sweep_task_cells: only the paid session is
+ stubbed. If this disagrees with the faithful prototype, the shipped function
+ is what is wrong.
+ """
+
+ cells = _flatten(plan)
+ by_key = {(f"t{c.task}", c.run, c.arm): c for c in cells}
+ order = {(f"t{c.task}", c.run, c.arm): i for i, c in enumerate(cells)}
+ ready = [threading.Event() for _ in plan]
+ stop = threading.Event()
+ _graph_builder(ready, graph_seconds, stop)
+
+ executed = 0
+ lock = threading.Lock()
+ folded: list[int] = []
+ tripped_at: int | None = None
+ streak_seen = {"streak": 0}
+
+ def run_cell(task_id: str, run_idx: int, arm: str) -> dict[str, Any]:
+ nonlocal executed
+ cell = by_key[(task_id, run_idx, arm)]
+ _execute_cell(cell, cpu_fraction, burn_rate)
+ with lock:
+ executed += 1
+ return _record(cell)
+
+ def on_record(task_id: str, run_idx: int, arm: str, rec: dict[str, Any]) -> None:
+ nonlocal tripped_at
+ index = order[(task_id, run_idx, arm)]
+ folded.append(index)
+ streak_seen["streak"] = runner.systemic_outage_streak(rec["error_kind"], streak_seen["streak"])
+ if outage_limit and streak_seen["streak"] >= outage_limit and tripped_at is None:
+ tripped_at = index
+
+ def await_ready(task_id: str) -> bool:
+ ready[int(task_id[1:])].wait()
+ return True
+
+ started = time.monotonic()
+ runner.sweep_packed_cells(
+ [(f"t{c.task}", c.run, c.arm) for c in cells],
+ workers=workers,
+ run=run_cell,
+ on_start=lambda *_: None,
+ on_record=on_record,
+ outage_streak=0,
+ outage_limit=outage_limit,
+ window=window,
+ await_ready=await_ready,
+ )
+ wall = time.monotonic() - started
+ stop.set()
+ return Outcome(wall_s=wall, executed=executed, tripped_at=tripped_at, folded=folded)
+
+
+SCHEDULERS = {
+ "wave": run_wave,
+ "fed": run_fed,
+ "packed": run_packed,
+ "faithful": run_faithful,
+ "production": run_production_packed,
+}
+
+
+def _window_kwargs(name: str, window: int | None) -> dict[str, int]:
+ """``--window`` only means anything to the two schedulers that hold one."""
+
+ return {"window": window} if window is not None and name in ("faithful", "production") else {}
+
+
+def _plan_args(args: argparse.Namespace, weekly: bool, seed: int, fail_from: int | None = None):
+ arms = (CANDIDATE_ARM,) if weekly else REVIEW_ARMS
+ return {
+ "task_count": len(review_tasks(_read(REVIEW_TASKS))),
+ "runs": args.runs,
+ "arms": arms,
+ "scale": args.scale,
+ "seed": seed,
+ "fail_from": fail_from,
+ }, arms
+
+
+def breaker_fidelity(args: argparse.Namespace) -> list[dict[str, Any]]:
+ """Does packing still trip where waves trip, and overrun no further?"""
+
+ rows: list[dict[str, Any]] = []
+ limit = runner.DEFAULT_OUTAGE_STREAK
+ window = args.window if args.window is not None else max(
+ args.workers * runner.PACKED_WINDOW_MULTIPLIER, args.workers
+ )
+ for fail_from in (0, 4, 12):
+ kwargs, _arms = _plan_args(args, weekly=False, seed=args.seed, fail_from=fail_from)
+ plan = build_plan(**kwargs)
+ total = sum(len(c) for c in plan)
+ row: dict[str, Any] = {
+ "fail_from": fail_from, "limit": limit, "total_cells": total, "window": window
+ }
+ for name in ("wave", "faithful", "production"):
+ out = SCHEDULERS[name](
+ plan, args.workers, outage_limit=limit, graph_seconds=args.graph_seconds,
+ **_window_kwargs(name, window),
+ )
+ row[name] = {
+ "tripped_at": out.tripped_at,
+ "executed": out.executed,
+ "overrun": out.executed - (out.tripped_at + 1) if out.tripped_at is not None else None,
+ }
+ row["same_trip_point"] = (
+ row["wave"]["tripped_at"] == row["faithful"]["tripped_at"] == row["production"]["tripped_at"]
+ )
+ # The producer holds submission to ``window`` cells beyond the fold
+ # pointer, so at most ``window - 1`` cells past the tripping one can
+ # already be in flight. At ``window == workers`` that is exactly the
+ # wave scheduler's own ``workers - 1`` bound.
+ row["overrun_within_bound"] = (
+ row["production"]["overrun"] is not None
+ and row["production"]["overrun"] <= window - 1
+ )
+ rows.append(row)
+ return rows
+
+
+def main() -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--workers", type=int, default=3)
+ parser.add_argument("--scale", type=float, default=DEFAULT_SCALE)
+ parser.add_argument("--seed", type=int, default=1729)
+ parser.add_argument("--repeat", type=int, default=1)
+ parser.add_argument("--runs", type=int, default=3)
+ parser.add_argument("--scheduler", choices=sorted(SCHEDULERS), default="wave")
+ parser.add_argument("--compare", action="store_true")
+ parser.add_argument("--breaker-fidelity", action="store_true")
+ parser.add_argument("--window-sweep", action="store_true", help="wall clock vs breaker overrun")
+ parser.add_argument("--contention-sweep", action="store_true", help="does the gain survive real CPU?")
+ parser.add_argument(
+ "--window",
+ type=int,
+ default=None,
+ help="submission window for the packed schedulers; defaults to the shipped policy",
+ )
+ parser.add_argument(
+ "--graph-seconds",
+ type=float,
+ default=None,
+ help="per-task graph build; defaults to the measured per-SHA overhead, scaled",
+ )
+ args = parser.parse_args()
+ # Both are checked here rather than where they are used: a bad --scale
+ # divides by zero before anything runs, and a negative --graph-seconds
+ # kills the graph-builder thread, after which every scheduler waits on a
+ # readiness event nobody will ever set.
+ # NaN defeats every comparison it appears in, so "> 0" and ">= 0" both admit
+ # it and the failure surfaces far from the flag: NaN durations reach
+ # time.sleep in a worker or the graph thread and raise there, after which the
+ # schedulers wait forever on a readiness event nobody will set. Infinity is
+ # worse than a crash - it silently scales every duration to zero and the run
+ # reports a sweep that took no time.
+ if not math.isfinite(args.scale) or args.scale <= 0:
+ parser.error("--scale must be a finite positive number")
+ if args.graph_seconds is not None and (not math.isfinite(args.graph_seconds) or args.graph_seconds < 0):
+ parser.error("--graph-seconds must be a finite non-negative number")
+ # Counts are indexed or handed to a thread pool without further checking, so
+ # a zero turns into an IndexError on plans[0], a median over an empty
+ # sequence, or ThreadPoolExecutor's own error - none of which name the flag
+ # that caused them.
+ if args.workers < 1:
+ parser.error("--workers must be at least 1")
+ if args.repeat < 1:
+ parser.error("--repeat must be at least 1")
+ if args.runs < 1:
+ parser.error("--runs must be at least 1")
+ # run_faithful and sweep_packed_cells both refuse a window below the worker
+ # count - a smaller one starves the pool, because the producer waits for a
+ # fold pointer to pass a cell it was never allowed to submit. Enforcing it
+ # here turns an uncaught ValueError partway through a measurement into an
+ # argument error before anything runs. Checked against the largest worker
+ # count this invocation will actually use: --contention-sweep runs its own
+ # counts irrespective of --workers, so validating against --workers alone
+ # let the 3-worker measurements finish and then raised on the 6-worker one.
+ window_workers = args.workers
+ if args.contention_sweep:
+ window_workers = max(window_workers, max(CONTENTION_WORKERS))
+ if args.window is not None and args.window < window_workers:
+ parser.error(f"--window must be at least the worker count ({window_workers}); a smaller window starves the pool")
+ if args.graph_seconds is None:
+ args.graph_seconds = SHA_OVERHEAD_SECONDS / args.scale
+
+ if args.contention_sweep:
+ burn_rate = statistics.median(calibrate_burn() for _ in range(3))
+ rows = []
+ for cpu_fraction in (0.0, 0.25, 0.5):
+ for workers in CONTENTION_WORKERS:
+ plans = [
+ build_plan(**_plan_args(args, False, args.seed + i)[0])
+ for i in range(args.repeat)
+ ]
+ measured = {}
+ for name in ("wave", "faithful", "production"):
+ fn = SCHEDULERS[name]
+ measured[name] = statistics.median(
+ fn(
+ plan,
+ workers,
+ outage_limit=0,
+ graph_seconds=args.graph_seconds,
+ cpu_fraction=cpu_fraction,
+ burn_rate=burn_rate,
+ **_window_kwargs(name, args.window),
+ ).wall_s
+ for plan in plans
+ )
+ serial = statistics.median(
+ sum(c.seconds for c in _flatten(plan)) for plan in plans
+ )
+ rows.append(
+ {
+ "cpu_fraction": cpu_fraction,
+ "workers": workers,
+ "wave_s": round(measured["wave"], 2),
+ "faithful_s": round(measured["faithful"], 2),
+ "production_s": round(measured["production"], 2),
+ "packing_gain_pct": round(
+ (measured["faithful"] - measured["wave"]) / measured["wave"] * 100, 1
+ ),
+ "wave_speedup": round(serial / measured["wave"], 2),
+ "faithful_speedup": round(serial / measured["faithful"], 2),
+ }
+ )
+ print(json.dumps({"burn_rate": round(burn_rate), "nproc": __import__("os").cpu_count(), "rows": rows}, indent=2))
+ return 0
+
+ if args.window_sweep:
+ total = len(review_tasks(_read(REVIEW_TASKS))) * args.runs * len(REVIEW_ARMS)
+ rows = []
+ for window in (args.workers, args.workers * 2, args.workers * 4, total):
+ clean = [build_plan(**_plan_args(args, False, args.seed + i)[0]) for i in range(args.repeat)]
+ wall = statistics.median(
+ run_faithful(
+ p, args.workers, outage_limit=0, graph_seconds=args.graph_seconds, window=window
+ ).wall_s
+ for p in clean
+ )
+ failing = build_plan(**_plan_args(args, weekly=False, seed=args.seed, fail_from=12)[0])
+ trip = run_faithful(
+ failing,
+ args.workers,
+ outage_limit=runner.DEFAULT_OUTAGE_STREAK,
+ graph_seconds=args.graph_seconds,
+ window=window,
+ )
+ rows.append(
+ {
+ "window": window,
+ "cold_wall_s": round(wall, 3),
+ "tripped_at": trip.tripped_at,
+ "executed": trip.executed,
+ "overrun_cells": trip.executed - (trip.tripped_at + 1)
+ if trip.tripped_at is not None
+ else None,
+ }
+ )
+ print(json.dumps({"workers": args.workers, "rows": rows}, indent=2))
+ return 0
+
+ if args.breaker_fidelity:
+ print(
+ json.dumps(
+ {"workers": args.workers, "graph_seconds": round(args.graph_seconds, 4),
+ "rows": breaker_fidelity(args)},
+ indent=2,
+ )
+ )
+ return 0
+
+ names = sorted(SCHEDULERS) if args.compare else [args.scheduler]
+ rows: list[dict[str, Any]] = []
+ for label, weekly in (("weekly", True), ("cold", False)):
+ plans = []
+ for i in range(args.repeat):
+ kwargs, arms = _plan_args(args, weekly, args.seed + i)
+ plans.append(build_plan(**kwargs))
+ serial = statistics.median(sum(c.seconds for c in _flatten(p)) for p in plans)
+ predicted = (
+ len(plans[0])
+ * expected_task_seconds(args.runs, arms, args.workers, fed_pool=False)
+ / args.scale
+ )
+ for name in names:
+ observed = statistics.median(
+ SCHEDULERS[name](
+ p,
+ args.workers,
+ outage_limit=0,
+ graph_seconds=args.graph_seconds,
+ **_window_kwargs(name, args.window),
+ ).wall_s
+ for p in plans
+ )
+ rows.append(
+ {
+ "profile": label,
+ "scheduler": name,
+ "workers": args.workers,
+ "observed_s": round(observed, 3),
+ "wave_model_s": round(predicted, 3),
+ "serial_s": round(serial, 3),
+ "speedup_vs_serial": round(serial / observed, 3) if observed else None,
+ }
+ )
+ print(json.dumps({"scale": args.scale, "repeat": args.repeat, "rows": rows}, indent=2))
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
From bddbb0ff9f04e49b1bd1d1b80c50688a20c1e75c Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Gerg=C5=91=20Magyar?=
Date: Tue, 8 Sep 2026 09:14:47 +0100
Subject: [PATCH 03/53] test(cli): prove detached refresh by ordering, not by
wall clock (#3221)
`lets the parent exit without waiting for a detached refresh child` bet
twice on absolute wall-clock budgets and lost both bets on loaded CI
runners:
- `expect(elapsed).toBeLessThan(1_800)` bounded the *parent's* cold
`node --import tsx` boot, inferring "did not wait" from a 1050ms margin
over the mock fetch's 750ms sleep. Reproduced failing on Linux under
40-way load at 1904ms.
- The 15s poll for the cache file had to cover the whole detached child:
node boot, a tsx transpile of 22 source files, an `acquireFileLock`
that shells out to `ps` (POSIX) or `powershell.exe -Command
Get-CimInstance Win32_Process` (Windows), the mocked fetch, and the
atomic write. That chain measures ~1.0s locally but has no bounded
upper limit on a contended runner, and it is what timed out in CI.
Replace both budgets with an ordering proof. The preloaded mock fetch now
parks the refresh child until the test releases it, so the sequence
asserted is: parent exited -> child provably still mid-refresh (started
marker present, cache absent) -> release -> cache written. That is
strictly stronger than the old elapsed-time inference, and it holds at any
machine speed. The remaining `expect.poll` timeouts no longer carry the
assertion's meaning; they are only "is the child dead" safety nets.
The mock's wait is bounded at 60s so an abandoned child (test failed
before releasing, temp home already removed) still exits instead of
spinning forever.
Co-authored-by: Gergo Magyar
Co-authored-by: Claude Opus 5 (1M context)
---
.../integration/cli/update-notice.test.ts | 30 ++++++++++++++-----
1 file changed, 23 insertions(+), 7 deletions(-)
diff --git a/gitnexus/test/integration/cli/update-notice.test.ts b/gitnexus/test/integration/cli/update-notice.test.ts
index abebfbd28..0887d1c46 100644
--- a/gitnexus/test/integration/cli/update-notice.test.ts
+++ b/gitnexus/test/integration/cli/update-notice.test.ts
@@ -127,12 +127,26 @@ describe('CLI update notice subprocess behavior', () => {
'dir',
);
+ // The refresh child parks inside fetch() until this test releases it. That
+ // orders the parent's exit against work that is provably still in flight,
+ // instead of racing it against a wall-clock budget: the child's real cost
+ // (node boot, tsx transpile, a lock acquisition that shells out to
+ // ps/powershell) has no bounded upper limit on a loaded CI runner.
+ const started = path.join(home, 'refresh-started');
+ const release = path.join(home, 'refresh-release');
const preload = path.join(home, 'mock-refresh.mjs');
fs.writeFileSync(
preload,
- `Object.defineProperty(process.stderr, 'isTTY', { value: true, configurable: true });
+ `import fs from 'node:fs';
+Object.defineProperty(process.stderr, 'isTTY', { value: true, configurable: true });
globalThis.fetch = async () => {
- await new Promise((resolve) => setTimeout(resolve, 750));
+ fs.writeFileSync(${JSON.stringify(started)}, '');
+ // Bounded so an abandoned child (test failed before releasing, temp home
+ // already deleted) still exits instead of spinning forever.
+ const deadline = Date.now() + 60_000;
+ while (!fs.existsSync(${JSON.stringify(release)}) && Date.now() < deadline) {
+ await new Promise((resolve) => setTimeout(resolve, 25));
+ }
return new Response(JSON.stringify({ version: '99.0.0' }), {
status: 200,
headers: { 'content-type': 'application/json' },
@@ -141,7 +155,6 @@ globalThis.fetch = async () => {
`,
);
- const startedAt = Date.now();
await new Promise((resolve, reject) => {
const parent = spawn(
process.execPath,
@@ -162,17 +175,20 @@ globalThis.fetch = async () => {
else reject(new Error(`notifier parent exited ${String(code)}`));
});
});
- const elapsed = Date.now() - startedAt;
- expect(elapsed).toBeLessThan(1_800);
+ // The parent already exited above, so reaching a still-parked child proves
+ // the refresh outlived it and was never awaited.
const cache = path.join(home, 'update-check.json');
+ await expect.poll(() => fs.existsSync(started), { timeout: 30_000, interval: 50 }).toBe(true);
expect(fs.existsSync(cache)).toBe(false);
- await expect.poll(() => fs.existsSync(cache), { timeout: 15_000, interval: 100 }).toBe(true);
+
+ fs.writeFileSync(release, '');
+ await expect.poll(() => fs.existsSync(cache), { timeout: 30_000, interval: 50 }).toBe(true);
expect(JSON.parse(fs.readFileSync(cache, 'utf8'))).toMatchObject({
latestVersion: '99.0.0',
registry: 'https://registry.npmjs.org',
});
- }, 20_000);
+ }, 90_000);
it('prints the localized notice on a forced-TTY stderr and keeps stdout clean', () => {
const home = tempHome();
From a4e70ec3b4738958eac34a352cb6f61470dfab3f Mon Sep 17 00:00:00 2001
From: cosark <121065588+cosark@users.noreply.github.com>
Date: Tue, 8 Sep 2026 01:23:52 -0700
Subject: [PATCH 04/53] docs: add RepoCloud one-click deploy button (#3212)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Co-authored-by: cosark
Co-authored-by: Gergő Magyar
---
README.md | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/README.md b/README.md
index 99782975e..26a03588e 100644
--- a/README.md
+++ b/README.md
@@ -102,6 +102,10 @@ The proxy strips `Origin` before forwarding, so the server's CSRF guard does not
Indexing is memory-bound. If `gitnexus-server` runs out of memory on a large repo, raise its `plan`, which sets available RAM: `standard` is 2 GB, `pro` is 4 GB. Raise `sizeGB` only if the disk fills with clones and indexes.
+### Deploy to RepoCloud
+
+[](https://repocloud.io/details/gitnexus/)
+
## Two Ways to Use GitNexus
| | **CLI + MCP** (recommended) | **Web UI** |
From b1d87c1f33d765910418ce76e641274c14a3616d Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Gerg=C5=91=20Magyar?=
Date: Tue, 8 Sep 2026 10:25:45 +0100
Subject: [PATCH 05/53] fix(eval): sweep evidence handling and measurement
health, with guarded comparator reuse (#3207)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(eval): cut skill-evolution wall clock without shrinking the gate
Reuse matching incumbent/CE cells, sanitize each SHA once, and default
dispatch workers to 3 so weekly review generations finish inside the
EventBridge window. Cap the sweep from leftover instance uptime so a
Friday dispatch still uploads evidence.
Co-authored-by: Cursor
Co-Authored-By: Claude Opus 5 (1M context)
* perf(eval): pipeline graph setup and correct the wall-clock cost model
The evolution sweep paid `sanitize` + `analyze --pdg --index-only` for every
unique task SHA on the critical path, one at a time, with nothing overlapping.
`_run_sweep` now starts the next unpaid SHA's clone template and graph snapshot
on a prefetch thread as soon as the current task's cells are dispatched, so
every SHA but the first hides behind a paid session wave. The thread is joined
before that SHA is used and before the trees tempdir is torn down, and a
prefetch failure is recorded against the SHA exactly as an inline failure is.
Tasks whose cells are all reusable comparator rows are not prefetched: they
never build a graph, so priming one would be pure cost.
Adds `measure_evolution_cost.py`, the cost model behind these numbers. It reads
the review corpus, the evolve defaults, and the workflow's workers default —
it does not start a session. Its first version charged `copy_isolated_tree`
once per paid cell, serially. `run_cell` clones inside its own pool worker, so
the clones in a wave overlap and only one is on the critical path per wave;
the model now charges `ceil(cells / workers)` waves.
Estimated review generation at workers=3: cold 21570s, weekly 7710s.
Wall clock is quantised by `ceil(cells_per_task / workers)`. A cold review task
is 9 cells, so workers=4 buys the wall clock of workers=3 and pays host
contention for it. Documented in the workflow's rollout checklist.
Co-Authored-By: Claude Opus 5 (1M context)
Co-Authored-By: Claude Opus 5 (1M context)
* perf(eval): price the benchmark against measured cell durations
The cost model assumed every cell runs the 1140s mean. Cells are not uniform:
the 41 rows in Actions run 33912693948's artifact are 826s at the median,
1262s at the mean, 2976s at p90, with two pinned at the 5400s session ceiling.
A wave waits for its slowest cell, so a mean understates every concurrent
schedule — the previous model called workers=3 cold 5.99h when the same
schedule against real durations is 10.33h.
session_durations.json carries the sample in submission order with its
provenance and its caveat: every cell in that run returned unusable evidence,
so the durations are real but a clean run may sit lower. It is the only live
artifact; the 2026-07-22 green run's has expired.
The model now simulates the schedule cell by cell rather than multiplying a
mean by a wave count, averaged over all 41 rotations of the sample so no
single alignment between sample order and cell index decides the answer. It
prices today's barrier (wave_makespan) against a continuously fed pool
(fed_makespan) and reports both, and it charges the proposer session — one
per generation, measured at 344.7s — which it had been omitting entirely.
Measurement only; no runtime behaviour changes.
Co-Authored-By: Claude Opus 5 (1M context)
Co-Authored-By: Claude Opus 5 (1M context)
* perf(eval): price arms separately and stop inventing setup constants
Two errors in the model, both found by auditing it against the artifact it
claims to describe.
The arms are not interchangeable. `candidate_review` runs 1416s at the mean
against `review`'s 1204s and `ce_review`'s 1176s, and the weekly lane pays the
candidate arm and nothing else — reuse skips both incumbents. Pricing weekly
from a pooled sample charged it for arms it never runs: weekly is 4.59h, not
the 3.65h a pooled sample reported. Cells are also submitted run-major and
arm-minor, so at workers=3 every wave holds one cell of each arm and the
slowest arm sets the wave; the model now builds cells in that order.
The setup constants were invented. GRAPH_ANALYZE_SECONDS=600 and
TEMPLATE_SANITIZE_SECONDS=180 charged 3900s of per-SHA setup for a cold run —
more than the entire non-session time of the source run, which was 2541s for
41 cells and 5 SHAs. `duration_s` is the sum of a cell's Claude sessions
(runner_sessions.py), so that 2541s residual is every clone, graph build,
sandbox and teardown the sweep paid. The model now charges the measured
residual per cell, 62.0s, and no longer credits clone templates or graph
prefetch: both landed after that run and there is no measurement of them yet.
The residual bounds what they can be worth.
Cold 37452s (10.40h), weekly 16541s (4.59h), against a fed pool at 31683s and
16541s. Measurement only; no runtime behaviour changes.
Co-Authored-By: Claude Opus 5 (1M context)
Co-Authored-By: Claude Opus 5 (1M context)
* perf(eval): charge sweep overhead where more workers cannot dissolve it
Three defects, found by auditing the model against the artifact again.
The overhead was charged inside the schedule. session_durations.json claimed
the residual was charged "per cell and serially - the pessimistic reading",
but task_cells folded it into each cell's duration, where the pool then
divided it by the worker count. The residual mixes per-cell work the pool
really does divide with per-SHA graph setup it cannot, and the artifact cannot
separate them, so it now sits outside the schedule: cold 11.09h, not 10.40h.
Alignment averaging weighted the shortest sample twice. The arm samples are 13,
14 and 14 long and the average ran over max()=14 offsets, so candidate_review's
first cell was counted twice and its last never. Averaging over lcm()=182
offsets weights every arm's sample evenly.
The wall assumed all 54 cells run. Replaying the sample's own error_kind
sequence through today's systemic_outage_streak trips the outage breaker at
cell 5 of 41. The source run executed all 41, so its runner did not break on
that sequence, but the current one would: these numbers price a HEALTHY sweep,
and a sweep with the sample's failure profile never reaches them. Stated on
generation_seconds and recorded next to the sample it qualifies.
Measurement only; no runtime behaviour changes.
Co-Authored-By: Claude Opus 5 (1M context)
Co-Authored-By: Claude Opus 5 (1M context)
* fix(eval): give the review agent somewhere it can actually write
Every review cell in the last recorded generation returned unusable evidence.
Not some — all 41, across all three arms and all six tasks, at $3653 for the
run. The transcripts say why, 127 times across 35 of 35 sessions:
EROFS: read-only file system,
open '/workspace/review-output.json.tmp.2.90a76e583b0c'
The review arm mounted the artifact as a writable FILE at
/workspace/review-output.json while binding /workspace read-only. The Write
tool writes atomically: it creates `.tmp..` beside the target
and renames it. The parent was read-only, so the temp create failed and the
artifact was never written. A writable file inside a read-only directory is
not writable to anything that writes atomically. Agents tried
/proc/self/root/workspace/... and /proc/1/root/workspace/... to get around it;
all 41 artifacts came back 0 bytes.
The artifact now lives in its own writable directory bound at /review-output,
outside the workspace. That is what a rename needs, and it lets the workspace
get stricter rather than looser: the review phase may now change nothing there
at all (enforce_phase_workspace gained allowed_artifact=None), where before it
was entitled to one path inside it. The file is no longer pre-created — the
agent writes it, and absence is now meaningful evidence.
parse_review_output reported every one of these as "review output is not valid
UTF-8 JSON". The file was empty, and its except folded OSError, UnicodeError
and JSONDecodeError into that one string, so a sandbox that made writing
impossible was indistinguishable from an encoding fault. That is why this read
as an agent-quality problem for fifteen consecutive non-green runs. Each cause
now names itself: never written, empty, not valid UTF-8, not valid JSON with
the decoder's position. run_arm also keeps the FIRST error_detail, as it
already did for error_kind, so a phase-boundary violation is no longer buried
under the parse failure it causes.
The test double conflated sandbox.private_root with the clone, which put the
artifact directory inside the workspace and would have hidden the stricter
check. Regression tests pin the mount shape in the generated bwrap argv, the
contract path in the prompt, the four parse diagnostics, and the
untouched-workspace contract.
Verified by unit tests only: this container has unprivileged user namespaces
disabled, so bwrap cannot run here and the mount was not exercised end to end.
Co-Authored-By: Claude Opus 5 (1M context)
Co-Authored-By: Claude Opus 5 (1M context)
* fix(review): close the artifact path in every layer that gates it
Code review of this branch found the relocated review artifact was fixed in the
bwrap mount and nowhere else. Four independent layers decide whether the agent
can write it, and three still named the old location.
Claude Code applies its own filesystem policy to its own tools, and
build_claude_settings listed only /workspace, /tmp and /home/agent under
allowWrite with denyRead ["/"]. The artifact used to live under /workspace, so
this list was correct until it moved. SANDBOX_REVIEW_OUTPUT is now in allowWrite
and allowRead; without it the bwrap bind grants a write the CLI then refuses.
The task corpus still ran `test -s review-output.json` from the workspace, in a
separate sandbox invocation that never sees the artifact mount. Every review
cell would have been stamped verify-failed with resolved=False no matter how
good the review was, which also made those rows permanently unreusable and so
silently disabled this branch's own comparator reuse for review arms. The verify
and hidden-oracle commands now read the location from
GITNEXUS_BENCH_REVIEW_OUTPUT and get the directory bound read-only, mirroring
the mount-plus-env-var shape _run_hidden_oracle already used.
host_text and host_path did not translate the new path, so the host-unsafe
backend told the agent to write somewhere that exists on neither backend.
Adding the mapping exposed a second defect: host_text substituted every
occurrence of a target, and "/review-output" appears twice in
"/review-output/review-output.json" - once as the directory and once inside the
filename. Matching is now anchored to a path boundary.
Comparator reuse had three ways to accept evidence it should have rejected. A
row with no runtime_digest passed the drift lock because the guard only compared
when both sides were bound, and the branch's own test asserted that as correct;
absence is now a mismatch and the test states the rule. materialize_reused_row
overwrote recorded_at with the copy time while the age check read that field, so
a row copied forward each generation refreshed its own clock and never aged out;
the first measurement time is now preserved and aged against. A future-dated
stamp passed a one-sided bound and is now rejected as corrupt.
RUNTIME_DIGEST never reached the runner at all: runner_environment builds a
fixed dict and process_control replaces the child environment wholesale, so the
digest the workflow exports was dropped and the lock it feeds was inert. The
instance-window deadline was also checked only after run_proposer returned,
buying a proposal the generation had no room to benchmark.
The graph prefetch thread was started without copy_context, so it never saw the
cancellation ContextVar the rest of the sweep shares, and the outage breaker
returned without setting cancel_event - together, a tripped breaker would block
on joining a prefetch that was never told to stop. Both fixed, with outage
checked before cancellation at the two exits so an outage keeps exit 1 instead
of becoming a Ctrl-C's 130.
Both bwrap canaries that actually execute a write still bound the pre-fix shape
against a file this branch no longer creates, so they would have errored rather
than caught anything. They now bind the directory and write atomically - temp
file beside the target, then rename - which is the exact operation that failed
with EROFS. A source-text assertion over inspect.getsource(run_arm) was replaced
with one that inspects the real mount, and a wall-clock assertion was pinned to
a fixed monotonic clock.
Not applied, and why: binding task-asset and dependency digests into comparator
reuse needs asset snapshots prepared before the reuse decision rather than
inside the per-task loop, and shipping the comparison without that would add a
guard that silently never fires. Forcing a paid canary cell per incumbent arm
and folding reused rows into the outage streak are behaviour decisions, not
fixes. Clone-template reuse still has no test. The cost model's per-cell
residual still shrinks with arm count, overstating weekly savings by at most the
2541s residual; the docstring now says so rather than inventing a split.
585 eval tests pass, ruff clean, 29 workflow contract tests pass. The two
test_model_gateway.py failures are pre-existing and fail on main.
* fix(review): bind reuse to its environment and keep the health canary real
Applies the five findings the previous review round left open.
Comparator reuse ignored the environment a row was measured in. TaskReuseBinding
carried the task and oracle identity but not the task-asset or sandbox-dependency
digests, and this branch itself changes sandbox_dependencies in the review
corpus - so a reused comparator could be measured against one dependency set and
compared against a candidate built on another, handing the gate a false
comparison. Closing it needed the digests to exist before the reuse decision, so
asset snapshots are now prepared for every task up front instead of lazily
inside the per-task loop. That also removes the concurrent TaskAssetCache.prepare
the prefetch thread could otherwise race, which the file's own "plain dict,
read-then-write race" comment warned about. Both digests fail closed on either
side, matching the runtime digest.
The broken-incumbent canary could not fire when reuse was working. It read
`resolved`, which counts reused rows, so an arm whose cells were all reused
always looked healthy - in precisely the run where a broken environment would go
unnoticed. aggregate now also reports `resolved_fresh` and the canary reads it.
That count would be vacuous if an arm were reused end to end, so the sweep keeps
one paid cell per incumbent arm and says which one it kept.
Reused rows did not participate in the outage streak, so a run of failures could
carry across them and trip on stale history. A reused success now resets the
streak the way a paid success does.
The cost model charged sweep overhead per cell, which credited a weekly
generation for shrinking work it still performs: it pays one arm instead of
three but builds exactly the same graphs. Overhead is charged per SHA now.
Weekly is 5.20h rather than the 4.80h the per-cell rate reported; cold is
10.86h. The residual still cannot be split between per-SHA and per-cell work
from one artifact, so session_durations.json records that assumption and the
direction it errs in, rather than leaving a number nobody can trace.
Clone-template reuse - the branch's core speedup, taken on essentially every
multi-cell sweep - now has a test that builds a real sanitized template, asserts
the cell runs against the copy with the template's HEAD, and fails if run_cell
re-clones. A second test asserting only on a namespace built inside the test was
written and deleted: it exercised nothing, which is the failure this review
round penalised elsewhere.
589 eval tests pass, ruff clean, 29 workflow contract tests pass. The two
test_model_gateway.py failures are pre-existing and fail on main.
* refactor(eval): consolidate duplicated harness logic after the review round
Simplification pass over the branch. Behavior-preserving throughout; three
reviewers, nine findings applied, two skipped.
The review-artifact block in _run_hidden_oracle was unreachable. That function
runs only in run_arm's non-review branch, while the directory it probes for is
created only in the review branch, and each sandbox serves exactly one arm - so
`review_artifact.parent.is_dir()` could never be true. It was added an hour
earlier to make the hidden oracle resolve the moved artifact; the oracle never
runs for review tasks, so the guard was dead on arrival. Deleting it also
removes the duplication it had with the verify-command wiring.
EXCLUDED_ERROR_KINDS is now one definition. runner.py and comparator_reuse.py
each carried the same six-member frozenset, kept in sync by a comment. Only one
direction is possible: runner already imports from comparator_reuse, so the
reverse import fails at module-init with a circular-import error. That is now
stated where the alias lives, so nobody tries it the other way.
ensure_task_graph and prefetch_next_graph shared ten keyword parameters, passed
through two call sites and forwarded whole between them. They now take a
GraphBuildEnv, mirroring TaskCellContext, which already bundles per-cell state
in this file. Its ready_keys() replaces an inline four-set union at the call
site.
Smaller consolidations: _sha256_file's hand-rolled chunk loop becomes
hashlib.file_digest (3.11+, already used in runner_artifacts); _copy_owner_only
reuses task_assets._write_all and COPY_CHUNK_BYTES instead of repeating the
short-write retry; its stat-then-open existence check becomes the O_EXCL failure
it was already relying on, which is atomic rather than merely narrow; and
runner_environment reads the digest through comparator_reuse.current_runtime_digest
instead of re-parsing the environment variable.
Three test docstrings summarised the branch's own history ("the branch's core
speedup", "the regression that produced fifteen runs") rather than the invariant
under test. Rewritten to state the constraint, which is what survives the merge.
Repaired the indentation left behind by the outage-streak edit and flattened the
prefetch dispatch from three nested conditionals to one.
Skipped: consolidating comparator_reuse._real_directory onto proposer_sandbox's
same-named helper - they differ, the sandbox one rejects any symlink in the
resolved path while this one checks only the leaf, so sharing it would tighten
behavior rather than preserve it. That needs a decision about which policy the
reuse path wants, not a simplification.
589 eval tests pass, ruff clean, 29 workflow contract tests pass. Unrelated and
pre-existing: two test_model_gateway.py failures, and
test_process_control.py::test_timeout_kills_term_ignoring_descendants_before_they_write,
which is a TERM-to-KILL timing flake (passes 2 of 3 in isolation) in a file this
branch does not touch.
* refactor(eval): name the reuse directory check for the promise it makes
The simplification pass left one finding open: comparator_reuse and
proposer_sandbox both defined `_real_directory`, same name and same shape, with
different guarantees. The sandbox one rejects every symlink hop in the path; the
reuse one checks only the leaf and resolves through parents. Sharing the name
invites a consolidation that would silently tighten one of them.
They should not be merged, so the name stops claiming they could be.
proposer_sandbox guards a mount root, where a symlink hop changes what an
untrusted session is handed. comparator_reuse guards a data directory whose
contents are already validated one file at a time - reads go through
_regular_file, which lstats and rejects symlinks, and writes through O_NOFOLLOW.
A symlinked parent therefore grants nothing those guards do not already cover,
while refusing one would reject a symlinked artifacts directory or macOS's /var
for no gain.
Renamed to _resolved_directory, with the reasoning recorded at the definition,
and a test that pins both halves: a symlinked parent is accepted and resolved, a
symlinked leaf is still refused. Behavior is unchanged.
591 eval tests pass, ruff clean. The two test_model_gateway.py failures are
pre-existing and fail on main.
* test(eval): measure the sweep scheduler instead of modelling it
measure_evolution_cost predicts wall clock from a model of what
sweep_task_cells does. This runs the real thing - real threads, the real wave
barrier, the real outage breaker - with only the paid agent session replaced by
a sleep, and times it.
Durations are the measured per-arm samples divided by 5000, so a 1416s cell
takes ~0.28s. The shape is kept on purpose: the median cell is 826s against a
5400s ceiling, and that spread is the entire reason a barrier costs anything.
Uniform random sleeps would erase the effect under test. All schedulers consume
one identical seeded plan, so a comparison cannot be an artifact of one of them
drawing luckier cells.
The model survives contact: it tracks real execution within about 10%, and
workers=1 - which runs without a pool at all - sits at 0.95, so the residual
above 1.0 at higher worker counts is per-wave thread overhead rather than a
modelling error. Two structural claims that were arithmetic are now observed.
Weekly is flat from workers=3: 3.59, 3.59, 3.59, 3.60, 3.59, 3.59 across w=3..8.
workers=4 buys nothing over workers=3 on cold, 7.68 against 7.78.
Two prototype schedulers are measured beside it, deliberately before any
production code exists. A continuously fed pool per task is worth more than the
model claimed on cold, -27.3% against a predicted -17.9%, and exactly nothing on
weekly, +0.0%, because a weekly task is one wave with nothing to feed. One pool
across all tasks beats both: -40.7% weekly and -42.9% cold at workers=3, rising
to -65.7% and -63.9% at workers=8. It also subsumes the fed pool, since packing
across tasks is a fed pool.
That reorders the backlog. Cross-task packing moves from second to first: it
dominates on both profiles, and it is the only thing that moves weekly at all.
Raising the worker count is worth nothing until it lands - under the barrier
weekly does not improve from w=3 to w=8, and speedup against serial is 1.58x for
three workers and only 2.40x for eight.
The bound on all of it: sleeping threads do not contend. Real sandboxed sessions
compete for CPU, page cache and disk, and the duration sample was itself
measured at workers=1, so it carries no contention either. These speedups are
upper bounds. The ordering is trustworthy because the schedulers were compared
under identical conditions; the magnitudes are not. The packed prototype is also
a bare ThreadPoolExecutor with no breaker folding, no per-task graph lifecycle
and no reuse binding - which is the actual cost of building it, and is not
measured here.
* test(eval): carry the sweep invariants into the packed prototype
The first packed prototype was a bare ThreadPoolExecutor. It reported -43% and
none of the invariants the shipped scheduler holds, so it priced an idea nobody
could ship. This one carries them: a global submission order continued across
task boundaries, in-order folding, the real outage breaker, and per-task graph
readiness gating behind a serial builder.
The fidelity check first reported the two schedulers tripping on different
cells, 17 against 16. That was my instrumentation, not a divergence -
sweep_task_cells folds an entire wave before it evaluates the breaker, so the
last cell folded is not the cell that tripped. With the harness mirroring the
breaker's own evaluation the two agree exactly, across failures starting at
cell 0, 4 and 12, with overrun inside the workers-1 bound the wave docstring
promises.
Two results worth the exercise.
Head-of-line blocking, not the barrier, is what a naive in-order design pays.
Holding submission to `workers` cells beyond the fold pointer leaves the
faithful scheduler at -8.1% cold and -2.7% weekly: one slow cell stalls the
pointer, the window cannot slide, and it reproduces the wave almost exactly.
That is the number to quote if anyone proposes the obvious implementation.
But the overrun bound turns out to be set by the worker count, not the window.
Only `workers` cells can be running when the breaker trips; everything queued
behind them short-circuits on the halt flag. Overrun is 3 at an unbounded
window exactly as at 6, and the trip cell never moves off 16. So H2 does not
have to trade breaker fidelity for speed - a wide window takes -42% with the
semantics intact. The tension I assumed was there is not, and window=12 already
captures 97% of it.
Still an upper bound: sleeping threads do not contend, and the sample was
measured at workers=1. What this establishes is that the invariants are
affordable, which was the thing blocking H2. Not built here: the trees tempdir
lifecycle, reuse-row binding, and the cancel_event path.
591 eval tests pass, ruff clean.
* test(eval): put the scheduler comparison under real CPU contention
Every Phase 2 number so far came from sleeping threads, which contend for
nothing, against a duration sample measured at workers=1, which contains no
contention either. That was the standing caveat on the whole result, so this
measures it.
A cell now waits for its API share and then burns a fixed number of sha256
rounds in a subprocess. Work-bounded rather than wall-clock bounded, so it takes
longer when cores are busy - that is the effect under test. A subprocess because
Python threads burning Python would measure the GIL rather than the machine.
Calibrated at 519k rounds/s, stable within 2% across three probes.
The first run of this was worthless and is recorded as such: on a 24-core host
with 3 to 6 workers nothing ever contends, since cpu_fraction 0.5 at 6 workers
is about 3 cores of demand out of 24. It measured an absence. Re-run pinned with
taskset to 4 and 2 cores.
The packing advantage survives. It holds between -40% and -47% across every host
size and CPU fraction tested, including a genuinely oversubscribed 2-core box at
cpu_fraction 0.5 with 6 workers.
But contention erodes packing more than it erodes waves, for a structural
reason: packing is what creates the concurrency. Moving from 24 cores to 2 at
cpu 0.5 and 6 workers, the faithful scheduler slows 13% while the wave slows
3.7%, and the gain narrows from 45.0% to 39.8%. Packing and a higher worker
count are therefore not independent wins - packing spends the contention
headroom first, so raising workers has to be re-argued after it lands rather
than added to it.
Three things this still does not measure, and they bound the result. The real
CPU fraction of a benchmark cell is a guess informed by roughly 180 tool calls
per session; nobody has profiled one. The evolution runner's core count decides
which column applies and is unknown here. And the burn is sha256, pure CPU,
while real cells run vitest and analyze, which are memory and IO heavy - so this
is a floor on contention, not a ceiling.
591 eval tests pass, ruff clean.
* perf(eval): add a packed sweep scheduler, and correct the bound I claimed for it
sweep_task_cells finishes one task before starting the next and drains a wave
before refilling it, so a task with fewer cells than workers leaves workers
idle and one slow cell stalls its whole wave. sweep_packed_cells feeds every
task's cells through a single pool instead. Measured against the review corpus
it is worth about 40% of a cold sweep, and it is the only change that moves a
seeded weekly run at all - there a task is three cells and a wave is never full.
The breaker keeps its exact meaning. Cells carry a total submission order
continued across task boundaries, a folder walks results in that order, and
consecutive systemic failures are counted there, so a doomed run aborts on the
same cell it would have under waves. Verified at three failure positions.
This commit also corrects a finding from the Phase 2 prototype. I claimed the
overrun bound was set by the worker count rather than the submission window,
and that packing therefore cost nothing in breaker fidelity. That was derived
from a window sweep that only ever injected failures at one position. Driving
the real function at other positions shows the halt flag does not bound overrun
at all: the folder walks in order, so a slow early cell lets workers race ahead
and the trip is detected after those cells have already paid. An unbounded
queue overran by 11 cells where waves overrun by 2.
So the window is load-bearing and the trade is real, measured at workers=3 with
failures injected at four positions:
window 3 -> -8% wall, overrun 2 (the wave scheduler's own bound)
window 6 -> -27% wall, overrun 4
window 12 -> -42% wall, overrun 9
window 54 -> -44% wall, overrun 11
Overrun is wasted paid sessions at roughly $70 each. The default multiplier is
2, keeping the worst case within twice the wave bound while taking most of the
gain; the curve is in the constant's comment so raising it is an informed
decision rather than a guess.
Not wired in yet: _run_sweep still calls sweep_task_cells per task. Moving the
per-task graph, trees tempdir and reuse binding out of that loop behind
await_ready is the larger and riskier half, and it belongs in its own change.
595 eval tests pass, ruff clean.
* fix(eval): judge harness health on execution, not on how many tasks resolved
broken_incumbent_arms infers "the environment is broken" from an arm resolving
zero tasks. That inference does not hold: a reviewer can be wrong about every
task in a hard corpus while every process, mount and capture worked perfectly.
Actions run 33962002890 is exactly that shape - 51 cells, all resolved=False
with error_kind=oracle-failed, median score 0.212, and a healthy harness.
Someone already knew this, and patched it by excluding review arms at the call
site. That leaves the unsound inference in place for workflow and
workflow_direct, and leaves review arms with no health check at all - so the
run that genuinely was broken, 33912693948, where the mount made an atomic
write impossible and all 41 artifacts came back empty, could not have been
caught here either.
So this replaces the inference rather than adding another exemption. aggregate
now classifies fresh rows into execution failures (the process or its tooling
did not complete), evidence failures (it completed but produced nothing
trustworthy or scoreable), and admissible measurements. An arm is unhealthy
only when it has fresh attempts, zero admissible measurements, and at least one
execution or evidence failure. Resolution count is no longer consulted. Arms
with only reused rows report current health as UNKNOWN rather than good.
With the inference corrected, review arms are checked again, which is what lets
the empty-artifact case be caught at all.
Deliberately unchanged: comparator reuse eligibility, quality denominators,
promotion thresholds, model settings, skill prompts and scheduler behaviour.
Failures that stop being called infrastructure failures still surface in the
counts and reasons - an agent-originated failure must not vanish from reporting
because it was reclassified. broken_incumbent_arms and its tests are left in
place; deleting behaviour belongs in its own change.
Seven regression tests, built from both runs' shapes and labelled as
reconstructed from logged observations, since 33962002890's results.jsonl did
not survive the instance shutdown. They pin: a badly-scoring reviewer is
healthy; an all-zero score is still a valid negative; empty artifacts are
unhealthy; one admissible cell keeps an arm healthy while its failures stay
visible; reused rows alone leave health unknown; reused successes do not mask
fresh failures; and a parseable artifact does not excuse a failed session.
602 eval tests pass, ruff clean.
* fix(eval): pin the health guard below the breaker, and stop calling mixed runs healthy
Two corrections to the health-classification patch.
The regression I wrote could not have proved what it claimed. A fixture of 41
empty artifacts aborts through the outage breaker long before finalization:
review-evidence-invalid is in SYSTEMIC_ERROR_KINDS and the limit is 5, so it
trips at cell 5 through the pre-existing path. It demonstrated failure
detection, not the new guard. The decisive test now uses ONE fresh unusable
cell, asserts the streak stays under the breaker threshold, and only then
requires finalization to abort - leaving the new check as the only thing that
can catch it. Removing the call makes that test fail; restoring it passes.
The accurate defect statement is narrower than the last message claimed. Review
arms were excluded from the final incumbent-health check while the consecutive-
failure breaker gave them separate, partial coverage. They were not unguarded.
Second: "one admissible cell plus two execution failures" was asserted as
healthy. That converts "not wholly unusable" into "ran reliably", which is how
a partly-broken sweep passes review. Arms now report UNKNOWN, OBSERVED_OK,
DEGRADED or UNUSABLE. Only UNUSABLE is fatal, so eligibility and promotion are
untouched - this changes what is reported, not what is allowed.
The guard is extracted as enforce_measurement_health so it can be driven
directly, and it now reports a status line per arm. It names no cause: an empty
artifact establishes that evidence is unusable, not that a mount rejected the
write, so it prints cause=undetermined rather than guessing EROFS. It still
runs after report.md and promotion.json are written, so a failing sweep leaves
its evidence behind.
ce_review is named explicitly at the call site. It is a comparator rather than
a candidate, so it is absent from CANDIDATE_ARMS.values(), and dropping the
review exclusion alone would have left it unclassified.
The wiring test reads _run_sweep's compiled code object for the referenced
global rather than matching source text. It is honest about its limit: it
proves the call exists and would catch its removal, but no test here drives
_run_sweep end to end, which needs bwrap and a sandbox.
broken_incumbent_arms is marked LEGACY and NON-AUTHORITATIVE with removal
tracked. It has no production caller.
608 eval tests pass, ruff clean. The two test_model_gateway.py failures are
test_locked_litellm_translates_messages_to_offline_responses and
test_openai_gateway_never_leaves_proxy_output_on_an_undrained_pipe; both fail
identically on origin/main in this environment, checked directly rather than
carried forward as an inherited label.
* fix(eval): review artifact path, evidence classification, comparator reuse
Extracted from the combined skill-evolution branch. This is the runtime change
set: everything that alters how a sweep executes and what it records. The
packed scheduler and its measurement harness were separated onto
perf/skill-evolution-packed-scheduler, which is purely additive.
Correctness. The review artifact was mounted as a writable FILE inside a
read-only workspace while the agent's Write tool writes atomically - temp file
beside the target, then rename - so the temp create failed EROFS and the
artifact was never written. Four layers gate that path and three named the old
location: the CLI's own allowWrite/allowRead policy, the task corpus verify
command run in its own sandbox invocation, and host_text/host_path for the
host-unsafe backend. Fixing the translator exposed a second defect, since
"/review-output" appears twice in "/review-output/review-output.json"; matching
is now anchored to a path boundary. parse_review_output folded OSError,
UnicodeError and JSONDecodeError into one message, so an artifact that was
never written looked like an encoding fault; each cause now names itself.
Health classification. broken_incumbent_arms inferred a broken environment from
an arm resolving zero tasks, which a reviewer facing a hard corpus falsifies -
Actions run 33962002890 is exactly that shape. Arms are now classified from
fresh execution and evidence outcomes as UNKNOWN, OBSERVED_OK, DEGRADED or
UNUSABLE, and only UNUSABLE aborts. Resolution count is not consulted. The
guard names no cause: an empty artifact establishes unusable evidence, not that
a mount rejected the write.
Comparator reuse. Reuse accepted evidence it should have rejected: a row
without a runtime_digest passed the drift lock, recorded_at was overwritten with
the copy time so a row could outlive its own max_age, and the binding ignored
task-asset and dependency digests although this change alters
sandbox_dependencies in the review corpus. Closing the last one required
preparing asset snapshots before the reuse decision, which also removes the
concurrent TaskAssetCache.prepare the prefetch thread could race.
These three concerns share aggregate() and _run_sweep, which is why they ship
together: separating them further would mean hunk-level surgery on a function
all three modify, and the risk of a silent omission outweighs the reviewability
gain.
592 eval tests pass at this base. The two test_model_gateway.py failures,
test_locked_litellm_translates_messages_to_offline_responses and
test_openai_gateway_never_leaves_proxy_output_on_an_undrained_pipe, fail
identically on origin/main in this environment.
Known gap, and the reason this is not ready to merge: no test drives _run_sweep
end to end. enforce_measurement_health is unit-tested including the
below-breaker unusable case, and the caller wiring is pinned structurally by
reading _run_sweep's compiled code object, but interruption semantics, exit
precedence and persisted artifacts are not exercised through the real path.
* fix(eval): address PR review feedback (#3207)
- aggregate: count admissible rows directly instead of subtracting the
execution and evidence counters, which double-charged a row that is both
a session error and invalid review evidence and could report UNUSABLE for
an arm holding real measurements.
- run_proposer: bound the session timeout by what is left of
--max-runtime-seconds, so clearing the sweep minimum cannot start a
full-length session past the instance window.
- comparator reuse: hold one O_NOFOLLOW descriptor for the size check,
digest and copy, and prove it is the inode that was checked, closing the
swap window a concurrent writer of the reuse directory had.
- Drive the review-artifact mount assertion through run_arm and the
clone-template assertion through run_cell, instead of rebuilding the
expected values in the tests (also removes the CodeQL unnecessary lambda).
- Assert the workflow invokes run-evolution.sh rather than that its YAML
mentions --max-runtime-seconds, which only appears in a comment.
- Correct the parse_review_output failure-mode claim: the fold was empty
artifacts reported as "not valid UTF-8 JSON"; a never-created file raised
FileNotFoundError.
- prettier: wrap the over-long readFileSync call flagged by PR autofix.
Note: pre-existing failure in tests/test_model_gateway.py::test_locked_litellm_translates_messages_to_offline_responses (local LiteLLM proxy never becomes ready in this environment) not addressed by this PR.
Co-Authored-By: Claude Opus 5 (1M context)
* fix(eval): address the second round of PR review feedback (#3207)
- Refuse a symlinked `transcripts` component on both sides of comparator
reuse. O_NOFOLLOW guards the leaf only, so a link there redirected the
read or the copy out of the results directory; checked per component as
evolution._require_directory_chain does.
- Base the paid incumbent canary on the cells this sweep PLANS. Reuse
selection accepts any prior run index, so a results directory produced
with more runs left extra keys, the equality never held, and the canary
stopped firing. Extracted as drop_canary_reuse_key and unit-tested.
- Start the runtime clock in main(). --max-runtime-seconds is measured from
/proc/uptime before exec, so parsing, task I/O, preflight and gateway
setup were being handed back to the sweep out of the upload reserve.
- Do not fall back to shutil.copytree when the managed clone copy was
cancelled or timed out; that fallback is for a filesystem that cannot
reflink, and copytree cannot be cancelled.
- Assert the review session's writable mount, not only the verifier's
read-only one: the EROFS bug is about the agent's write.
- Exercise ref isolation in the copy_isolated_tree test rather than
comparing an initial HEAD a shared namespace would also match.
- Point the stale-symlink fixture at the sentinel via os.path.relpath, and
skip the reuse symlink tests where symlink creation needs privilege.
Co-Authored-By: Claude Opus 5 (1M context)
* fix(eval): close the runtime-cap gap and pin the reuse directory
Both were left open on #3207 as approach decisions rather than nits.
Runtime cap: run-evolution.sh computed the budget in its own
`uv run python -c` and passed a number, so the script's remaining
provenance work and the CLI's own startup were spent by nobody and charged
to the sweep — out of the upload reserve the cap exists to protect. The
script now passes --max-runtime-from-instance-window and evolve reads
/proc/uptime itself, on the line after it starts the clock the budget is
measured against, so no interval exists to lose. Also removes an
interpreter start from the script and lets --dry-run print the real argv.
Reuse directory: _real_child_directory lstat-checked `transcripts` and
returned its pathname, so a concurrent writer could rename the directory
and leave a symlink before the name was used again — O_NOFOLLOW guards
only the leaf. Every artifact is now resolved against a held descriptor:
_open_real_directory opens with O_DIRECTORY|O_NOFOLLOW (check and open in
one syscall), and _open_regular / _copy_owner_only take dir_fd. The reuse
path is therefore POSIX-only; _require_openat says so and fails closed,
which the runner already treats as "run a paid cell". _resolved_directory
still tolerates a symlinked reuse root, unchanged and still tested.
evolution._require_directory_chain is still lstat-per-component. It guards
a different surface (candidate overlay reads) that neither review raised,
so it is left alone rather than widened into here.
Co-Authored-By: Claude Opus 5 (1M context)
* fix(eval): sample the proposer budget where it is spent, digest what is copied
Four findings against f0cdc9e7, all of them mine.
- run_proposer took a precomputed remaining_seconds, but it clones,
sanitizes and builds a sandbox before the session starts, so the caller's
reading was already stale. It takes started_monotonic now and samples the
budget on the last line before run_claude. The caller's earlier reading
still decides whether to start at all — it just no longer decides how
long to allow.
- _copy_transcript_artifact hashed the source and then read it again to
copy it. A held descriptor stops the pathname being substituted, not the
inode being rewritten, so the row could record the expected digest while
the destination held other bytes. _copy_owner_only now digests the same
buffers it writes and returns (digest, bytes); a mismatch unlinks the
destination and raises. One read instead of two.
- Explain the empty except in _open_real_directory: an existing directory
is the ordinary case, and the O_DIRECTORY|O_NOFOLLOW open below is what
proves what it is (CodeQL).
- Drop the unused uptime fixture from the runtime-cap test, and correct the
cross-file contract described in the workflow-contract test and the
workflow YAML: neither names --max-runtime-from-instance-window, the
script does.
Co-Authored-By: Claude Opus 5 (1M context)
* test(eval): close the reuse producer/consumer contract through the real run_cell
Every comparator-reuse test built its rows by hand. That proves the predicate's
logic and nothing about the producer: a fixture can satisfy eligibility while a
row the runner actually emits never does, and a key-name inventory cannot tell
the difference. I checked that inventory first - all 25 keys the predicate reads
have a producer - which is exactly why it was not sufficient evidence.
This carries one record through the production path instead:
real run_cell -> production JSONL writer -> load_result_rows
-> row_is_reusable_comparator
Only the expensive dependencies are replaced: the model session, sandbox launch,
repository acquisition, graph preparation, git plumbing. The digest fields the
reuse binding compares are assembled by run_cell itself from its
TaskCellContext, so they stay real - they are the subject, not scaffolding.
Expectations are built from the sweep's own configuration rather than copied out
of the emitted row, since copying them back would make producer and consumer
agree because the test arranged it.
Two cases: a production-emitted row with matching bindings is eligible, and the
same evidence with one dependency binding changed is rejected. A test-controlled
digest set keeps both deterministic.
Mutation-checked. Replacing run_cell's task_asset_manifest_digest assignment
with None makes the positive case fail; restoring it passes. That is evidence
the inventory audit could not produce.
Building it also documented what a real review row must carry, which no fixture
had recorded: a scored review with review_weighted_f1 present, and at least one
transcript artifact whose source is PARENT_EVENT_STREAM_SOURCE. An empty
artifact list is not reusable. Those are production contracts my first double
got wrong, and the reader rejected it each time.
No production code changed. 607 eval tests pass, ruff clean.
* test(eval): drive the real sweep to its finalization decision
enforce_measurement_health was unit-tested and its call site pinned by reading
_run_sweep's compiled code object. Neither showed the guard running inside a
sweep. These drive the real _run_sweep with cell execution scripted and
everything downstream left alone: folding, aggregation, the artifact writers,
the health guard and the exit selection.
The below-breaker case is the decisive one. A fixture of many unusable cells
aborts through the pre-existing outage breaker instead - review-evidence-invalid
is systemic with a limit of five - and would pass whether or not the
finalization guard exists. One fresh unusable cell stays under that threshold,
so only the guard can catch it. The test asserts the arm and status the guard
reports, and that no systemic-outage line was printed, rather than accepting any
SystemExit: an unrelated setup failure must not satisfy it.
Mutation-checked at the runtime path. Removing the guard invocation fails the
below-breaker test because the expected finalization behaviour disappears, not
because a name went missing from a code object.
A zero score is covered separately. review_weighted_f1 = 0.0 must stay a present
valid negative measurement; a truthiness check would read it as absent and turn
a quality result into an execution-health failure. The third test asserts the
evidence survives - results.jsonl carries the scored row and report.md exists -
so the persisted artifacts tell the same story as the exit.
Only expensive setup is replaced: cell execution, graph preparation, asset
snapshots, and task-binding resolution, which clones the repository and verifies
the ref. Building the fixture also documented that the report renders the whole
review metric set, so an incomplete row fails in formatting rather than logic.
Still open on this track: interruption semantics and exit-reason precedence
(outage 1 before cancellation 130) are not yet exercised, and fold order and the
real sandbox mount contract remain on separate tracks.
No production code changed. 610 eval tests pass, ruff clean. The two
test_model_gateway.py failures reproduce on origin/main in this environment.
* fix(eval): an interrupted sweep is interrupted, not aborted
Driving the real _run_sweep to its exit revealed that cancellation without an
outage exits 1 and writes "Sweep aborted: partial evidence", where the contract
is 130 and "Sweep cancelled".
sweep_task_cells returns (streak, tripped), and its caller assigns that flag to
outage_tripped and turns it into exit 1. Both cancellation paths returned True
for it. The breaker's own return is also True, so the two became
indistinguishable one frame up and cancellation inherited the outage's exit and
wording. The fix is to return False from the cancellation paths: the flag means
the breaker tripped, and the caller already tests cancel_event itself for the
stop decision, so nothing stops running any later than before.
Reproduced before the fix, through the real caller, not from reading. The
failing assertion was the report line; the exit code was 1.
Two runtime tests cover it. Each begins with one admissible cell so the arm
classifies DEGRADED rather than UNUSABLE - otherwise enforce_measurement_health
supplies exit 1 first and a precedence test passes without ever reaching exit
selection. Cancellation is set from a completed cell rather than a sleep or a
real signal, so the interruption point is deterministic.
The precedence test is mutation-checked: moving the cancel_event check ahead of
the outage check fails it with "assert 130 == 1" - it observes the wrong exit
code, not merely some failure. My first attempt at that mutation silently
matched nothing and the suite passed; a no-op mutation proves nothing, so the
edit now asserts its own anchor.
test_process_control read the same flag as "stopped" and asserted it after a
cancellation. It now reports the event and the flag separately, which is what it
was really asserting: cancelled, and not an outage.
The scored-row shape moved into tests/bench_fixtures.py with zero values
written out, so the finalization and interruption tests share one definition
of what a real review row carries.
Impact analysis returned risk UNKNOWN - the index predates this branch and does
not carry the eval harness - so the callers were confirmed by text search as the
rules require for UNKNOWN: one production caller and three test sites, all
updated or verified. detect_changes reports zero for the same reason; that zero
is unseen, not unaffected.
612 eval tests pass, ruff clean. The two test_model_gateway.py failures
reproduce on origin/main in this environment.
* test(eval): prove the review artifact mount under a real sandbox
The orchestration tests replace the sandbox, so they say nothing about
isolation. This covers the filesystem contract the EROFS defect actually broke,
through the production configuration: the same command_prefix_for call run_arm
makes for a review cell, with read_only_workspace=True and the review-output
directory as the writable mount - not a hand-built mount tuple that merely looks
right.
A deterministic writer stands in for the agent. It writes a temp file beside the
destination and renames it into place, which is the operation that failed: an
atomic write needs a WRITABLE PARENT DIRECTORY, and binding the file itself left
nowhere to put the temp file. It then attempts a workspace write and must be
refused, and the production parse_review_output reads the bytes the sandbox left
behind. No model session and no credentials.
Placed in test_proposer_sandbox.py, which the "eval / containment (ubuntu)" job
already runs with GITNEXUS_REQUIRE_BWRAP_CANARY=1. That gate is the point: where
the variable is set, a missing namespace capability FAILS the job instead of
skipping into a green tick. Verified here - forcing the variable on this machine
fails with "bwrap: No permissions to create new namespace" rather than skipping.
What this does not establish: the assertions have never executed. This machine
cannot create user namespaces, so the test stops at the preflight. Imports,
signatures and the payload were checked statically instead - parse_review_output
returns a tuple rather than an object, and it rejects a body without "verdict",
so both of my first attempts were wrong and are fixed. The first CI run on a
namespace-capable runner is what will actually confirm it.
Scope is the filesystem and process contract only. A stand-in writer does not
establish that a particular agent CLI's own file-access policy permits the same
operation; that is a second, independent gate.
40 passed, 10 skipped locally; ruff clean.
* fix(eval): read the review artifact before the sandbox deletes it
The containment (ubuntu) job has been red since before the mount canary was
added, and both failures have the same cause.
This branch moved the review artifact out of the clone and into the session's
private root, which is the point of the change: the agent writes atomically, so
the artifact needs a writable parent DIRECTORY outside the read-only workspace.
prepare_sandbox removes that private root in a finally on scope exit. Two tests
read the artifact AFTER the with block, so they were asserting against a
directory the sandbox had already deleted - FileNotFoundError, reported as
"review output was never written".
Production is not affected, and this is the reason: run_arm holds the live
session and reads the artifact inside that scope, both to score it and to mount
it read-only into the verifier's separate sandbox invocation. The tests were the
only readers outside it.
Both now read inside the scope. The pre-existing canary
(test_read_only_review_workspace_exposes_only_one_writable_artifact) predates
this PR and passes on main, where the artifact still lived in the clone and
survived teardown; the relocation is what broke it, so the fix belongs here.
The new canary found this independently and agreed on the cause, which is what
it was written for - though only after CI ran it, since this machine cannot
create user namespaces.
40 passed, 10 skipped locally; ruff clean. The bwrap-gated tests still skip here
and remain unverified until the ubuntu job runs them.
* test(eval): pin what an interrupted sweep persists and refuses to promote
The completed-run persistence test could not show either of these: it never
interrupts, so it would pass even if the writers only ran on the clean path.
Evidence already paid for survives cancellation. The cell that completed before
the interruption is still in results.jsonl with its measurement intact, and
report.md is still written. Losing those rows would mean paying for evidence the
sweep then discards.
An interrupted run emits nothing that authorizes promotion. Asserted as the
semantic condition rather than the absence of a file: promotion.json IS still
written for an aborted run - it is the record of why nothing was promoted - so
the test requires run_status "aborted" and every decision reduced to
insufficient_evidence with a partial-evidence reason, rather than requiring the
artifact to disappear.
Mutation-checked. Forcing complete=True at the promotion_evidence call site
fails the test with "assert 'complete' == 'aborted'" - it observes an aborted
run claiming completeness, not merely some failure.
These were the last two unasserted items on this PR's finalization checklist.
614 eval tests pass, ruff clean. The two test_model_gateway.py failures are
environmental (litellm[proxy] console script absent) and predate this branch.
* Address PR review feedback (#3207)
An exhausted runtime cap no longer buys a second. remaining_runtime_seconds
floors at 0 and the proposer call site wrapped it in max(1, ...), so a cap fully
spent by the clone, the sanitize pass and the sandbox build started a paid
session with a one-second allowance instead of stopping before the upload
reserve the cap exists to protect. It now returns a not-ok record, which the
caller already treats as end-of-run. Pinned by a test driving the real
run_proposer with setup that consumes the whole budget; reverting to max(1, ...)
fails it.
Reuse copies are bounded before their size is validated, not after. The source
is a prior sweep directory this module already treats as concurrently writable,
so a transcript appended to after its metadata was recorded was streamed to EOF
and only then compared against its declared size - filling the destination, or
never reaching EOF, long before the drift check could reject it. _copy_owner_only
now takes max_bytes and stops one byte past the ceiling, which keeps the drift
comparison meaningful. Review and patch artifacts carry no recorded size, but
"no recorded size" is not "no limit": they get MAX_TRANSCRIPT_BYTES, the ceiling
the capture path already enforces.
A reused review row must carry its artifact. The predicate accepted a row on its
score and transcript metadata while materialize_reused_row copied the review
artifact only when the row named one, so a scored review could be carried
forward with nothing for a proposer to read. The shared row fixture was the
thing out of step here, not the requirement - production sets review_artifact
whenever the review source exists - so it now carries one too.
Test fixes, all against code this PR added:
unusable_review_row could not be overridden at all. It passed explicit keywords
beside **overrides, and Python rejects the duplicate in the call expression
before scored_review_row can apply its update, so unusable_review_row(error_kind=...)
raised TypeError. Merged into one mapping.
The finalization helper monkeypatched runner.prepare_ce_plugin_snapshot with
raising=False. No such symbol exists - the real one is staged_ce_plugin_snapshot
- so it silently added an attribute nothing reads. Removed.
The promotion test named candidate_review in candidate_arms but _run_sweep builds
cells only from args.arms, so no candidate cell ran and insufficient_evidence
could hold because nothing executed rather than because partial evidence is
barred. The candidate arm now runs, cancellation fires after its cell, and the
test asserts the candidate actually produced a row so it cannot silently return
to being vacuous.
The reuse round trip serialized with json.dumps and write_text while claiming to
cover the production writer, which applies redact_text over the row's own bytes.
It now writes the way the sweep does. Its fixture also writes the review
artifact, because run_cell records review_artifact only when the review source
exists and the emitted row was otherwise one production never emits.
Not addressing the duplicate-match nit on proposer_sandbox.py: the claim is that
"/review-output" occurs once in "/review-output/review-output.json". It occurs
twice, at offsets 0 and 14 - the separator before the basename forms it again -
so the boundary the comment describes is real. Verified with re.finditer.
The workspace-snapshot finding is parked for a human: excluding bootstrap noise
is a documented deliberate choice, and tightening it is a genuine tradeoff.
634 eval tests pass, ruff clean. The two test_model_gateway.py failures are
environmental (litellm[proxy] console script absent) and predate this branch.
* Address PR review feedback (#3207), round 2
Carry review_f1 in the scored-row fixture. score_review emits "f1"
(review_scoring.py:317), which the runner folds in as review_f1, so a real
scored row has it and the fixture did not - the same fidelity gap as the metrics
already added there. aggregate now reports 0.5 for it instead of None.
The reported failure mode was not real, and the distinction matters for anyone
reading the thread later. aggregate filters on record.get(metric) is not None
BEFORE indexing record[metric], so a missing key is skipped rather than raising
KeyError; the finalization tests were passing throughout. Verified by running
aggregate against the old fixture. Fixed because the fixture should match what
production emits, not because anything was crashing.
Drop the unused row parameter from the round-trip _expectation helper. It never
read the argument - deliberately, since the docstring says the bindings must be
derived from sweep configuration rather than copied out of the emitted row - but
passing the row anyway was dead plumbing that suggested the opposite.
The reuse-root TOCTOU finding is parked for a human rather than fixed: the
lstat-then-resolve window is real, but _resolved_directory documents the weaker
promise as deliberate and says not to merge it with proposer_sandbox's helper
"without first deciding which promise the reuse path should make". That decision
is the fix, and it is the same class as the reuse TOCTOU already parked on this
PR.
634 eval tests pass, ruff clean. test_process_control's TERM-ignoring-descendant
test flaked once under full-suite load and passes 3/3 in isolation; it is a
timing test and neither file changed here touches it. The two
test_model_gateway.py failures remain environmental.
* fix(eval): pin the reuse roots to the directory that was checked
Settles the reuse TOCTOU that has been parked twice on this PR. The docstring
asked to decide which promise the reuse path makes before merging its helper
with proposer_sandbox's, and the answer is that two separate things were being
conflated.
The symlink POLICY stays exactly as it was: parent hops remain allowed, so a
symlinked artifacts directory or macOS's /var still works, and a symlinked leaf
is still refused. Rejecting hops would break ordinary setups for no gain, which
is what the docstring argued and it is still right.
What is closed is the other thing - the gap between checking a name and using
it. _resolved_directory lstats a name and the later open re-walks that same
name, so a prior sweep that renames its results root and leaves something else
behind is opened somewhere else entirely, and O_NOFOLLOW cannot see a link that
resolve() already followed. _resolved_directory now returns the checked
directory's identity alongside its path, and _open_pinned_root fstats the
descriptor it opened and refuses a mismatch.
The two are separable, so there was no trade to make: comparing identity
rejects nothing that holds still, since a stable directory always matches
itself. Both cases are pinned - a replacement by a different REAL directory is
refused (the leaf-symlink rule does not cover that one), and an ordinary
unchanged root opens normally.
Why it matters here rather than as a general hardening: the failure is silent.
Rows would be copied out of some other directory and folded into a comparator
baseline as though they were this sweep's own evidence, which is the one thing
the reuse path exists to get right - a corrupted baseline decides promotions.
636 eval tests pass, ruff clean. The two test_model_gateway.py failures remain
environmental.
* refactor(eval): one prompt digest, and drop two pieces of dead scaffolding
Compute task_prompt_digest once. row_is_reusable_comparator compares the value a
prior row stored against the value this sweep derives, as exact strings, and the
two inline copies of that hash had already drifted apart - the expectation side
had picked up a str() cast the row side lacked. They agree today only because
tasks validate "prompt" as a string (runner_tasks.required_strings), so nothing
was broken; a later divergence on one side would have silently stopped rows
matching, with no test to catch it. Verified the extracted helper is
byte-identical to what it replaced.
Two comments named broken_incumbent_arms as the thing that would read stale
health. This branch replaced that call with the measurement-health path, so the
reasoning still holds but the name no longer does; they now name arm_health.
The function itself stays: origin/main still calls it at runner.py:2098, and
retiring it is its own change rather than a cleanup pass.
Removed instance_window_budget_from_proc and its one test. It composes two
helpers main() deliberately calls separately - there is a comment there
explaining why the read and the budget calculation stay decoupled - and nothing
outside its own test ever called it. Introduced on this branch, absent from
main, so it was never deployed, public, or consumed elsewhere.
Dropped an unused tmp_path parameter from a comparator-reuse test that does no
filesystem work.
Skipped three findings. Hoisting _assert_self_contained_git_objects to a
once-per-template check is a real saving - measured at 5.66s over 1551 objects -
but it verifies the OUTPUT of each copy operation, and git clone from a local
path hardlinks objects by default, which is exactly what its st_nlink check
catches. Checking a predecessor instead of the clone each cell runs against
thins an isolation guarantee for 0.7% of a median cell. Deleting
broken_incumbent_arms outright would remove a symbol main still calls. And a
fourth copy of the test-only _git helper extends a pattern that already exists
in three other test modules; centralising it means editing conftest.py, outside
this scope.
635 eval tests pass, ruff clean. The two test_model_gateway.py failures are the
environmental ones.
---------
Co-authored-by: Gergo Magyar
Co-authored-by: Cursor
Co-authored-by: Claude Opus 5 (1M context)
---
.../workflows/gitnexus-skill-evolution.yml | 27 +-
eval/tests/bench_fixtures.py | 75 ++
eval/tests/test_comparator_reuse.py | 467 +++++++++
eval/tests/test_evolve.py | 302 ++++++
eval/tests/test_process_control.py | 7 +-
eval/tests/test_proposer_sandbox.py | 200 +++-
eval/tests/test_reuse_round_trip.py | 268 ++++++
eval/tests/test_review_corpus.py | 4 +
eval/tests/test_review_scoring.py | 30 +
eval/tests/test_runner_hardening.py | 101 +-
eval/tests/test_sanitized_graph.py | 16 +
eval/tests/test_session_progress.py | 47 +-
eval/tests/test_sweep_finalization.py | 279 ++++++
eval/tests/test_workflow_bench.py | 383 ++++++++
eval/tests/test_workflow_bench_sessions.py | 143 ++-
eval/workflow_bench/README.md | 53 +-
eval/workflow_bench/comparator_reuse.py | 604 ++++++++++++
eval/workflow_bench/evolve.py | 265 +++++-
eval/workflow_bench/proposer_sandbox.py | 86 +-
eval/workflow_bench/review_scoring.py | 31 +-
eval/workflow_bench/run-evolution.sh | 18 +
eval/workflow_bench/runner.py | 892 +++++++++++++++---
eval/workflow_bench/runner_artifacts.py | 84 +-
eval/workflow_bench/runner_sessions.py | 22 +-
eval/workflow_bench/sanitized_graph.py | 23 +-
.../tasks.review.scenarios.yaml | 17 +-
eval/workflow_bench/tasks.scenarios.yaml | 5 +
.../unit/skill-evolution-workflow.test.ts | 28 +-
28 files changed, 4217 insertions(+), 260 deletions(-)
create mode 100644 eval/tests/bench_fixtures.py
create mode 100644 eval/tests/test_comparator_reuse.py
create mode 100644 eval/tests/test_reuse_round_trip.py
create mode 100644 eval/tests/test_sweep_finalization.py
create mode 100644 eval/workflow_bench/comparator_reuse.py
diff --git a/.github/workflows/gitnexus-skill-evolution.yml b/.github/workflows/gitnexus-skill-evolution.yml
index c004bb515..22f0bdfcc 100644
--- a/.github/workflows/gitnexus-skill-evolution.yml
+++ b/.github/workflows/gitnexus-skill-evolution.yml
@@ -63,13 +63,17 @@
# uploads, and a promotion (if any) opens a well-formed PR. Run
# 29907431284 (2026-07-22) went green end to end in 14h45m and reached a
# gate decision (`insufficient_evidence`, no promotion).
-# [ ] After resizing the runner, prove a manual workers=3 run has zero excluded
-# runs and does not stretch the 48-minute serial mean toward the session
-# ceiling; then set GITNEXUS_EVOLUTION_WORKERS=3 and
+# [ ] Confirm a workers=3 dispatch has zero excluded runs (review sessions in
+# 33962002890 averaged ~19m serial, well under the 90m session ceiling).
+# Then set GITNEXUS_EVOLUTION_WORKERS=3 and
# GITNEXUS_EVOLUTION_ENABLED=true for scheduled runs. Scheduled runs
-# require both values, so leaving workers unset/1 is an immediate rollback;
-# workflow_dispatch remains available for the proof and bills real API
-# usage on GITNEXUS_BENCH_ANTHROPIC_API_KEY or GITNEXUS_BENCH_OPENAI_API_KEY.
+# require both values, so leaving the var unset is an immediate rollback.
+# Dispatch defaults to 3; pass workers=1 only to debug a contended host.
+# Weekly generations reuse matching incumbent/CE cells from the previous
+# artifact so the paid matrix is the new candidate, not a 54-cell replay.
+# Wall clock is quantised by ceil(cells_per_task / workers), and a review
+# task is 9 cells cold, so 4 costs host contention for exactly the wall
+# clock of 3. The next step up that buys anything is 5 (3 waves -> 2).
name: GitNexus skill evolution
on:
@@ -92,9 +96,9 @@ on:
default: '3'
type: string
workers:
- description: 'Benchmark cells of one task to run at once — raise only to match the runner’s vCPUs'
+ description: 'Benchmark cells of one task to run at once — 3 fits the evolution box; drop to 1 only if siblings hit the session ceiling'
required: false
- default: '1'
+ default: '3'
type: string
model:
description: 'Model for the benchmark arms (match the model your skill users run)'
@@ -172,6 +176,13 @@ jobs:
# stops the runner just disappears mid-step. Scheduled runs can start well
# after the cron (the 2026-08-01 run was queued 65min late), so the job
# budget has to absorb that delay and still land inside the uptime window.
+ # A Friday workflow_dispatch on a box that already booted for Saturday's
+ # cron inherits leftover uptime, not a fresh 24h. Run 33962002890 started
+ # Friday 10:57 UTC and vanished at the Saturday 03:00 stop — 51 finished
+ # sessions never uploaded. run-evolution.sh therefore passes
+ # --max-runtime-from-instance-window, and the CLI derives its cap from
+ # /proc/uptime at startup, so the sweep fails in-process and this always()
+ # upload still runs.
timeout-minutes: 1260
permissions:
contents: read # The promotion PR uses a short-lived App token minted below.
diff --git a/eval/tests/bench_fixtures.py b/eval/tests/bench_fixtures.py
new file mode 100644
index 000000000..fdd6c131f
--- /dev/null
+++ b/eval/tests/bench_fixtures.py
@@ -0,0 +1,75 @@
+"""Shared row shapes for the sweep tests.
+
+Building the finalization tests turned up what a real scored review row must
+carry: the report renders the whole review metric set, so an incomplete row
+fails in string formatting rather than in the logic under test. That is a
+property of the fixture, not of production - the shape lives here once so each
+test does not rediscover it.
+"""
+
+from __future__ import annotations
+
+from typing import Any
+
+
+def scored_review_row(**overrides: Any) -> dict[str, Any]:
+ """One admissible review cell, with zero-valued metrics written out."""
+
+ row: dict[str, Any] = {
+ "ok": True,
+ "error_kind": None,
+ "error_detail": None,
+ "resolved": True,
+ "review_evidence_valid": True,
+ "review_score": {"weighted_f1": 0.5},
+ "review_weighted_f1": 0.5,
+ "review_true_positives": 1,
+ "review_false_positives": 0,
+ "review_false_negatives": 0,
+ "review_precision": 0.5,
+ "review_recall": 0.5,
+ "review_f1": 0.5,
+ "review_weighted_precision": 0.5,
+ "review_weighted_recall": 0.5,
+ "review_blocker_recall": 1.0,
+ "review_severity_accuracy": 1.0,
+ "review_category_accuracy": 1.0,
+ "review_grounded_evidence": 1.0,
+ "review_verdict_correct": True,
+ "review_clean_control": True,
+ "review_clean_pass": True,
+ "transcript_missing": False,
+ "transcript_artifacts": [],
+ "num_turns": 3,
+ "duration_s": 1.0,
+ "cost_usd": 0.5,
+ "input_tokens": 1,
+ "output_tokens": 1,
+ "cache_creation_input_tokens": 0,
+ "cache_read_input_tokens": 0,
+ "diff_files": 0,
+ "diff_insertions": 0,
+ "diff_deletions": 0,
+ }
+ row.update(overrides)
+ return row
+
+
+def unusable_review_row(**overrides: Any) -> dict[str, Any]:
+ """A cell that ran but produced evidence nothing can be scored from."""
+
+ # Merged into one mapping rather than passed as explicit keywords beside
+ # **overrides: Python rejects a duplicate keyword in the call expression
+ # itself, so unusable_review_row(error_kind=...) raised TypeError before
+ # scored_review_row could apply the override this helper advertises.
+ return scored_review_row(
+ **{
+ "ok": False,
+ "resolved": False,
+ "review_evidence_valid": False,
+ "error_kind": "review-evidence-invalid",
+ "review_score": None,
+ "review_weighted_f1": None,
+ **overrides,
+ }
+ )
diff --git a/eval/tests/test_comparator_reuse.py b/eval/tests/test_comparator_reuse.py
new file mode 100644
index 000000000..4425304b3
--- /dev/null
+++ b/eval/tests/test_comparator_reuse.py
@@ -0,0 +1,467 @@
+"""Comparator-row reuse: skip unchanged incumbent/CE cells, never candidates."""
+
+from __future__ import annotations
+
+import hashlib
+import os
+from datetime import UTC, datetime, timedelta
+from pathlib import Path
+
+import pytest
+
+from workflow_bench import comparator_reuse
+from workflow_bench.comparator_reuse import (
+ ComparatorReuseExpectation,
+ TaskReuseBinding,
+ materialize_reused_row,
+ row_is_reusable_comparator,
+ select_reusable_comparator_rows,
+)
+from workflow_bench.proposer_sandbox import SandboxError
+from workflow_bench.runner_sessions import PARENT_EVENT_STREAM_SOURCE
+
+
+requires_openat = pytest.mark.skipif(
+ os.open not in os.supports_dir_fd,
+ reason="comparator reuse resolves every artifact against a pinned directory descriptor",
+)
+
+
+def _digest(text: str = "blob") -> str:
+ return hashlib.sha256(text.encode()).hexdigest()
+
+
+def _artifact(name: str = "session-1.jsonl", payload: bytes = b'{"type":"ok"}\n') -> dict:
+ return {
+ "path": f"transcripts/{name}",
+ "sha256": hashlib.sha256(payload).hexdigest(),
+ "bytes": len(payload),
+ "source": PARENT_EVENT_STREAM_SOURCE,
+ }
+
+
+def _row(**overrides) -> dict:
+ base = {
+ "task": "review-pr-2718-defect",
+ "arm": "review",
+ "run": 0,
+ "ok": True,
+ "error_kind": None,
+ "model": "gpt-5.6-sol",
+ "benchmark_model": "gpt-5.6-sol",
+ "effort": "xhigh",
+ "sandbox_backend": "bwrap",
+ "task_base_sha": "a" * 40,
+ "task_prompt_digest": _digest("prompt"),
+ "oracle_digest": _digest("oracle"),
+ "oracle_command_digest": _digest("oracle-cmd"),
+ "oracle_manifest_digest": _digest("oracle-man"),
+ "skill_digest": _digest("skill"),
+ "candidate_overlay_digest": None,
+ "review_evidence_valid": True,
+ # Production sets this whenever the review source exists, which is the
+ # normal path for a valid review; the fixture predated the requirement.
+ "review_artifact": "review-pr-2718-defect-review-run0.review.json",
+ "review_score": {"weighted_f1": 0.4},
+ "review_weighted_f1": 0.4,
+ "transcript_missing": False,
+ "transcript_artifacts": [_artifact()],
+ "recorded_at": datetime.now(UTC).isoformat(),
+ "runtime_digest": _digest("cli"),
+ "task_asset_manifest_digest": _digest("assets"),
+ "sandbox_dependency_manifest_digest": _digest("deps"),
+ }
+ base.update(overrides)
+ return base
+
+
+def _expected(**overrides) -> ComparatorReuseExpectation:
+ now = datetime.now(UTC)
+ values = dict(
+ model="gpt-5.6-sol",
+ effort="xhigh",
+ sandbox_backend="bwrap",
+ runtime_digest=_digest("cli"),
+ now=now,
+ max_age=timedelta(days=90),
+ tasks={
+ "review-pr-2718-defect": TaskReuseBinding(
+ task_base_sha="a" * 40,
+ task_prompt_digest=_digest("prompt"),
+ oracle_digest=_digest("oracle"),
+ oracle_command_digest=_digest("oracle-cmd"),
+ oracle_manifest_digest=_digest("oracle-man"),
+ task_asset_manifest_digest=_digest("assets"),
+ sandbox_dependency_manifest_digest=_digest("deps"),
+ )
+ },
+ skill_digests={"review": _digest("skill"), "ce_review": None},
+ ce_plugin_version="3.24.0",
+ ce_plugin_manifest_digest=_digest("ce"),
+ )
+ values.update(overrides)
+ return ComparatorReuseExpectation(**values)
+
+
+def test_matching_incumbent_review_row_is_reusable() -> None:
+ assert row_is_reusable_comparator(_row(), _expected()) is True
+
+
+def test_candidate_rows_are_never_reusable() -> None:
+ assert row_is_reusable_comparator(_row(arm="candidate_review"), _expected()) is False
+
+
+def test_skill_digest_drift_rejects_reuse() -> None:
+ assert row_is_reusable_comparator(_row(), _expected(skill_digests={"review": _digest("other")})) is False
+
+
+def test_excluded_or_failed_rows_are_not_reusable() -> None:
+ expected = _expected()
+ assert row_is_reusable_comparator(_row(error_kind="session-error", ok=False), expected) is False
+ assert row_is_reusable_comparator(_row(ok=False), expected) is False
+ assert row_is_reusable_comparator(_row(review_evidence_valid=False), expected) is False
+ assert row_is_reusable_comparator(_row(recorded_at=(datetime.now(UTC) - timedelta(days=91)).isoformat()), expected) is False
+
+
+def test_runtime_digest_mismatch_rejects_when_both_sides_are_bound() -> None:
+ row = _row(runtime_digest=_digest("old-cli"))
+ assert row_is_reusable_comparator(row, _expected(runtime_digest=_digest("new-cli"))) is False
+ assert row_is_reusable_comparator(row, _expected(runtime_digest=_digest("old-cli"))) is True
+ # A row with no runtime_digest was measured by a harness that recorded none,
+ # which is the drift this lock exists to catch - not evidence of agreement.
+ assert row_is_reusable_comparator(_row(runtime_digest=None), _expected()) is False
+ # And a sweep that cannot determine its own digest must not reuse either.
+ assert row_is_reusable_comparator(_row(), _expected(runtime_digest=None)) is False
+
+
+def test_ce_review_matches_plugin_digest_not_repo_skill() -> None:
+ row = _row(
+ arm="ce_review",
+ skill_digest=None,
+ ce_plugin_version="3.24.0",
+ ce_plugin_manifest_digest=_digest("ce"),
+ )
+ assert row_is_reusable_comparator(row, _expected()) is True
+ assert (
+ row_is_reusable_comparator(row, _expected(ce_plugin_manifest_digest=_digest("other")))
+ is False
+ )
+
+
+def test_select_drops_conflicting_duplicates() -> None:
+ first = _row(review_weighted_f1=0.4)
+ second = _row(review_weighted_f1=0.9, recorded_at=datetime.now(UTC).isoformat())
+ selected = select_reusable_comparator_rows([first, second], expected=_expected())
+ assert selected == {}
+ same = select_reusable_comparator_rows([first, dict(first)], expected=_expected())
+ assert ("review-pr-2718-defect", "review", 0) in same
+
+
+@requires_openat
+def test_materialize_copies_transcript_and_review_artifacts(tmp_path: Path) -> None:
+ payload = b'{"type":"result"}\n'
+ source = tmp_path / "prior"
+ dest = tmp_path / "fresh"
+ (source / "transcripts").mkdir(parents=True)
+ dest.mkdir()
+ transcript = source / "transcripts" / "session-1.jsonl"
+ transcript.write_bytes(payload)
+ transcript.chmod(0o600)
+ review = source / "review-pr-2718-defect-review-run0.review.json"
+ review.write_text('{"verdict":"comment"}\n')
+ patch = source / "review-pr-2718-defect-review-run0.patch"
+ patch.write_text("diff\n")
+ row = _row(
+ review_artifact=review.name,
+ transcript_artifacts=[_artifact(payload=payload)],
+ )
+
+ copied = materialize_reused_row(row, source_dir=source, dest_dir=dest)
+
+ assert copied["reused"] is True
+ assert copied["reused_from_recorded_at"] == row["recorded_at"]
+ assert (dest / "transcripts" / "session-1.jsonl").read_bytes() == payload
+ assert (dest / review.name).read_text() == review.read_text()
+ assert (dest / patch.name).read_text() == "diff\n"
+ assert copied["transcript_artifacts"][0]["sha256"] == hashlib.sha256(payload).hexdigest()
+
+
+@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
+@requires_openat
+def test_a_reused_artifact_is_copied_from_the_inode_that_was_checked(tmp_path: Path) -> None:
+ """The reuse source is a directory another sweep wrote and may still write.
+
+ Validating a path and then re-opening it hands a concurrent writer the gap:
+ replace the checked file with a symlink and the copy follows it out of the
+ results directory. Swapping the path while the descriptor is held is that
+ same substitution, made deterministic.
+ """
+
+ (tmp_path / "transcript.jsonl").write_bytes(b"verified\n")
+ decoy = tmp_path / "decoy.jsonl"
+ decoy.write_bytes(b"substituted\n")
+
+ with comparator_reuse._open_real_directory(tmp_path, label="reuse source") as dir_fd:
+ with comparator_reuse._open_regular("transcript.jsonl", dir_fd=dir_fd, label="transcript") as descriptor:
+ (tmp_path / "transcript.jsonl").unlink()
+ (tmp_path / "transcript.jsonl").symlink_to(decoy)
+ comparator_reuse._copy_owner_only(descriptor, "copy.jsonl", dir_fd=dir_fd)
+
+ assert (tmp_path / "copy.jsonl").read_bytes() == b"verified\n"
+ with pytest.raises(SandboxError, match="regular non-symlink"):
+ with comparator_reuse._open_regular("transcript.jsonl", dir_fd=dir_fd, label="transcript"):
+ pass
+
+
+@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
+@requires_openat
+def test_a_symlinked_transcripts_directory_is_refused_on_both_sides(tmp_path: Path) -> None:
+ """`O_NOFOLLOW` refuses the leaf, not the directory above it.
+
+ A `transcripts` symlink on the source side makes reuse read a file outside
+ the results directory; one on the destination side writes the copy outside
+ this sweep's evidence. Neither is covered by the per-file guards that let
+ _resolved_directory tolerate a symlinked root.
+ """
+
+ payload = b'{"type":"result"}\n'
+ outside = tmp_path / "outside"
+ (outside / "transcripts").mkdir(parents=True)
+ (outside / "transcripts" / "session-1.jsonl").write_bytes(payload)
+ row = _row(transcript_artifacts=[_artifact(payload=payload)])
+
+ linked_source = tmp_path / "linked-source"
+ linked_source.mkdir()
+ (linked_source / "transcripts").symlink_to(outside / "transcripts", target_is_directory=True)
+ dest = tmp_path / "fresh"
+ dest.mkdir()
+ with pytest.raises(SandboxError, match="transcript source must be a real directory"):
+ materialize_reused_row(row, source_dir=linked_source, dest_dir=dest)
+
+ source = tmp_path / "prior"
+ (source / "transcripts").mkdir(parents=True)
+ (source / "transcripts" / "session-1.jsonl").write_bytes(payload)
+ linked_dest = tmp_path / "linked-dest"
+ linked_dest.mkdir()
+ (linked_dest / "transcripts").symlink_to(outside / "transcripts", target_is_directory=True)
+ with pytest.raises(SandboxError, match="transcript destination must be a real directory"):
+ materialize_reused_row(row, source_dir=source, dest_dir=linked_dest)
+
+
+@requires_openat
+@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
+def test_a_renamed_transcripts_directory_cannot_redirect_a_copy(tmp_path: Path) -> None:
+ """The directory is pinned, not re-walked from its name.
+
+ An lstat that passed and a pathname used afterwards are two different
+ directories the moment a concurrent writer renames the first one away. This
+ performs exactly that substitution — rename, then leave a symlink in its
+ place — while the descriptor is held, which is what makes the race testable
+ without timing.
+ """
+
+ payload = b'{"type":"result"}\n'
+ results = tmp_path / "results"
+ transcripts = results / "transcripts"
+ transcripts.mkdir(parents=True)
+ (transcripts / "session-1.jsonl").write_bytes(payload)
+ outside = tmp_path / "outside"
+ outside.mkdir()
+
+ with comparator_reuse._open_real_directory(results, label="reuse source") as root_fd:
+ with comparator_reuse._open_real_directory(
+ "transcripts", dir_fd=root_fd, label="transcript source"
+ ) as dir_fd:
+ transcripts.rename(results / "moved")
+ (results / "transcripts").symlink_to(outside, target_is_directory=True)
+ with comparator_reuse._open_regular(
+ "session-1.jsonl", dir_fd=dir_fd, label="transcript"
+ ) as artifact_fd:
+ comparator_reuse._copy_owner_only(artifact_fd, "copy.jsonl", dir_fd=dir_fd)
+
+ assert (results / "moved" / "copy.jsonl").read_bytes() == payload
+ assert not (outside / "copy.jsonl").exists()
+
+
+@requires_openat
+def test_a_transcript_rewritten_mid_copy_is_refused_not_recorded(tmp_path: Path, monkeypatch) -> None:
+ """The digest has to describe the bytes that were written.
+
+ A held descriptor stops the pathname being substituted; it does not stop the
+ inode being rewritten, and the prior sweep's directory is one this sweep
+ treats as concurrently writable. Hashing the source and then reading it
+ again to copy let the row keep the expected digest while the destination
+ held different bytes.
+ """
+
+ payload = b'{"type":"result"}\n'
+ source = tmp_path / "prior"
+ (source / "transcripts").mkdir(parents=True)
+ transcript = source / "transcripts" / "session-1.jsonl"
+ transcript.write_bytes(payload)
+ dest = tmp_path / "fresh"
+ dest.mkdir()
+ row = _row(transcript_artifacts=[_artifact(payload=payload)])
+
+ # Rewrite the inode in the window the copy reads through — same length, so
+ # only the digest can tell, which is the point.
+ real_read = comparator_reuse.os.read
+ rewritten = {"done": False}
+
+ def rewrite_then_read(fd: int, size: int) -> bytes:
+ if not rewritten["done"]:
+ rewritten["done"] = True
+ with open(transcript, "r+b") as handle:
+ handle.write(b'{"type":"TAMPER"}')
+ return real_read(fd, size)
+
+ monkeypatch.setattr(comparator_reuse.os, "read", rewrite_then_read)
+ with pytest.raises(SandboxError, match="drifted"):
+ materialize_reused_row(row, source_dir=source, dest_dir=dest)
+ monkeypatch.undo()
+
+ # And nothing unvouched-for is left behind for the proposer to read.
+ assert not (dest / "transcripts" / "session-1.jsonl").exists()
+
+
+@requires_openat
+def test_materialize_rejects_same_directory_and_missing_transcript(tmp_path: Path) -> None:
+ source = tmp_path / "prior"
+ source.mkdir()
+ row = _row()
+ with pytest.raises(SandboxError, match="same results directory"):
+ materialize_reused_row(row, source_dir=source, dest_dir=source)
+ dest = tmp_path / "fresh"
+ dest.mkdir()
+ with pytest.raises(SandboxError, match="missing"):
+ materialize_reused_row(row, source_dir=source, dest_dir=dest)
+
+
+@requires_openat
+def test_a_reused_row_ages_from_its_first_measurement_not_the_copy():
+ """Reuse chains must not refresh the clock.
+
+ materialize_reused_row restamps recorded_at with the copy time, so aging
+ against that field let a row be copied forward every generation and outlive
+ max_age forever. The original measurement time is the one that counts.
+ """
+
+ original = (datetime.now(UTC) - timedelta(days=91)).isoformat()
+ chained = _row(recorded_at=datetime.now(UTC).isoformat(), reused_from_recorded_at=original)
+ assert row_is_reusable_comparator(chained, _expected()) is False
+ # The same row inside the window is still reusable.
+ fresh = _row(
+ recorded_at=datetime.now(UTC).isoformat(),
+ reused_from_recorded_at=(datetime.now(UTC) - timedelta(days=1)).isoformat(),
+ )
+ assert row_is_reusable_comparator(fresh, _expected()) is True
+
+
+def test_a_future_dated_row_is_corrupt_not_fresh():
+ ahead = (datetime.now(UTC) + timedelta(days=2)).isoformat()
+ assert row_is_reusable_comparator(_row(recorded_at=ahead), _expected()) is False
+
+
+def test_a_changed_sandbox_dependency_is_not_the_same_baseline():
+ """The environment is part of the measurement.
+
+ This branch itself changes `sandbox_dependencies` in the review corpus, so a
+ prior row measured against the old set is a measurement of a different
+ machine. Reusing it would compare a fresh candidate to a baseline built
+ somewhere else and hand the promotion gate a false comparison.
+ """
+
+ assert row_is_reusable_comparator(
+ _row(sandbox_dependency_manifest_digest=_digest("other-deps")), _expected()
+ ) is False
+ assert row_is_reusable_comparator(
+ _row(task_asset_manifest_digest=_digest("other-assets")), _expected()
+ ) is False
+ # A row that predates the field is not evidence of agreement either.
+ assert row_is_reusable_comparator(_row(sandbox_dependency_manifest_digest=None), _expected()) is False
+
+
+@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
+def test_reuse_directories_allow_a_symlinked_parent_but_not_a_symlinked_leaf(tmp_path: Path):
+ """Pins a deliberate difference from the sandbox's mount-root check.
+
+ proposer_sandbox refuses every symlink hop because a hop changes what an
+ untrusted session is handed. A reuse directory is data, and every file
+ inside it is validated on its own, so a symlinked parent is allowed -
+ rejecting it would break a symlinked artifacts directory or macOS's /var
+ for no gain. The leaf itself must still be a real directory.
+ """
+
+ real = tmp_path / "real"
+ real.mkdir()
+ (real / "inner").mkdir()
+ linked_parent = tmp_path / "linked"
+ linked_parent.symlink_to(real, target_is_directory=True)
+
+ # Reached through a symlinked parent: allowed, and resolved to the real path.
+ # The identity returned alongside it is what pins the root against a swap
+ # between the check and the open; the symlink policy itself is unchanged.
+ resolved, identity = comparator_reuse._resolved_directory(linked_parent / "inner", label="probe")
+ assert resolved == (real / "inner").resolve()
+ inner_stat = (real / "inner").stat()
+ assert identity == (inner_stat.st_dev, inner_stat.st_ino)
+
+ # The leaf itself being a symlink is still refused.
+ with pytest.raises(SandboxError, match="must be a real directory"):
+ comparator_reuse._resolved_directory(linked_parent, label="probe")
+
+
+def test_a_review_row_without_its_artifact_is_not_reusable() -> None:
+ """A score is a claim about evidence, not the evidence itself.
+
+ materialize_reused_row copies the review artifact only when the row names
+ one, so accepting a row without it would carry a scored review forward with
+ nothing for a proposer to read.
+ """
+
+ row = _row()
+ assert row_is_reusable_comparator(row, _expected()) is True
+ without = {**row, "review_artifact": ""}
+ assert row_is_reusable_comparator(without, _expected()) is False
+ missing = {k: v for k, v in row.items() if k != "review_artifact"}
+ assert row_is_reusable_comparator(missing, _expected()) is False
+
+
+@requires_openat
+def test_a_reuse_root_replaced_after_the_check_is_refused(tmp_path: Path, monkeypatch) -> None:
+ """Check and use must name the same directory, not the same string.
+
+ _resolved_directory lstats a name and the open re-walks that same name, so
+ a prior sweep that swaps its results root in between is opened somewhere
+ else. The leaf-symlink rule does not cover it - a replacement that is
+ itself a real directory passes every check the policy makes - and the
+ failure is silent, folding another directory's rows into this sweep's
+ comparator baseline.
+ """
+
+ original = tmp_path / "results"
+ original.mkdir()
+ resolved, stale_identity = comparator_reuse._resolved_directory(original, label="probe")
+
+ # Replaced by a different REAL directory: the name still resolves and still
+ # passes the symlink policy, but it is not the inode that was checked.
+ original.rename(tmp_path / "moved")
+ original.mkdir()
+ assert comparator_reuse._resolved_directory(original, label="probe")[1] != stale_identity
+
+ monkeypatch.setattr(
+ comparator_reuse, "_resolved_directory", lambda *_a, **_k: (resolved, stale_identity)
+ )
+ with pytest.raises(SandboxError, match="replaced between the check and the open"):
+ with comparator_reuse._open_pinned_root(original, label="probe"):
+ pass
+
+
+@requires_openat
+def test_a_stable_reuse_root_opens_normally(tmp_path: Path) -> None:
+ """The guard rejects nothing that holds still - a directory matches itself."""
+
+ root = tmp_path / "results"
+ root.mkdir()
+ with comparator_reuse._open_pinned_root(root, label="probe") as fd:
+ assert os.fstat(fd).st_ino == root.stat().st_ino
diff --git a/eval/tests/test_evolve.py b/eval/tests/test_evolve.py
index c135584de..32d50fd85 100644
--- a/eval/tests/test_evolve.py
+++ b/eval/tests/test_evolve.py
@@ -9,19 +9,24 @@ import time
from contextlib import contextmanager
from datetime import UTC, datetime, timedelta
from pathlib import Path
+from types import SimpleNamespace
import pytest
from workflow_bench import evolve, evolution
from workflow_bench.runner_sessions import PARENT_EVENT_STREAM_SOURCE
from workflow_bench.evolve import (
+ MIN_INSTANCE_SWEEP_SECONDS,
build_parser,
build_proposer_prompt,
+ capped_timeout_seconds,
executed_benchmark_arms,
generation_timeout_seconds,
+ instance_window_budget_seconds,
load_jsonl,
proposer_evidence_entries,
read_learnings,
+ remaining_runtime_seconds,
resolve_incumbent_arms,
runner_argv,
select_evidence,
@@ -662,6 +667,127 @@ def test_run_proposer_hides_the_hidden_harness_and_keeps_the_full_tool_surface(m
assert captured["settings_json"] == FakeSandbox.settings_json
+def test_proposer_session_cannot_outlive_the_remaining_instance_window(monkeypatch, tmp_path):
+ """Clearing the sweep minimum is not a licence to run a full session.
+
+ --timeout is sized for a whole generation, so a proposer started with the
+ minimum left would run far past --max-runtime-seconds and the box would take
+ the evidence with it. The budget is sampled after the clone, the sanitize
+ pass and the sandbox setup, because a reading taken before them is already
+ stale by the time the session it bounds actually starts.
+ """
+
+ captured: dict[str, object] = {}
+ # Pinned clock: real setup duration would make this assert on scheduling.
+ clock = {"now": 1000.0}
+ monkeypatch.setattr(evolve.time, "monotonic", lambda: clock["now"])
+ setup_seconds = 100.0
+
+ @contextmanager
+ def fake_prepare_sandbox(**_kwargs):
+ yield SimpleNamespace(
+ claude_bin="claude",
+ command_prefix=[],
+ settings_json="{}",
+ transcript_projects=tmp_path / "transcript-projects",
+ )
+
+ def fake_run_claude(*_args, **kwargs):
+ captured.update(kwargs)
+ return {"ok": False, "error_kind": "session-error"}
+
+ def slow_sanitize(_clone):
+ # Stands in for the clone, the sanitize pass and the sandbox build —
+ # all of which run between the caller's decision and the session.
+ clock["now"] += setup_seconds
+ return "0" * 40
+
+ monkeypatch.setattr(evolve.runner, "make_worktree", lambda _repo, _ref, destination: destination)
+ monkeypatch.setattr(evolve.runner, "remove_clone", lambda _clone: None)
+ monkeypatch.setattr(evolve, "sanitize_clone_for_hidden_oracles", slow_sanitize)
+ monkeypatch.setattr(evolve, "prepare_sandbox", fake_prepare_sandbox)
+ monkeypatch.setattr(evolve.runner, "run_claude", fake_run_claude)
+ args = build_parser().parse_args(["--tasks", "tasks.yaml", "--model", "model"])
+ assert args.timeout > evolve.MIN_INSTANCE_SWEEP_SECONDS, "otherwise this test proves nothing"
+
+ common = {
+ "overlay_dir": tmp_path / "overlay",
+ "proposal_path": tmp_path / "proposal.md",
+ "evidence_bundle": tmp_path / "evidence",
+ "bwrap_bin": tmp_path / "bwrap",
+ }
+ budget = evolve.MIN_INSTANCE_SWEEP_SECONDS + 1
+ args.max_runtime_seconds = budget
+ started = clock["now"]
+ evolve.run_proposer("prompt", args, **common, started_monotonic=started)
+ # The setup time is charged, not handed back: a value sampled at `started`
+ # would have allowed the whole budget.
+ assert captured["timeout"] == budget - setup_seconds
+
+ # No cap configured means no budget to overrun: the session keeps its own.
+ args.max_runtime_seconds = None
+ evolve.run_proposer("prompt", args, **common, started_monotonic=started)
+ assert captured["timeout"] == args.timeout
+ evolve.run_proposer("prompt", args, **common)
+ assert captured["timeout"] == args.timeout
+
+
+def test_a_budget_spent_during_setup_stops_the_proposer_rather_than_buying_a_second(
+ monkeypatch, tmp_path
+):
+ """An exhausted cap must end the generation, not start a one-second session.
+
+ remaining_runtime_seconds floors at 0, and the call site wrapped it in
+ max(1, ...) - so a cap fully consumed by the clone, the sanitize pass and
+ the sandbox build produced a paid session with a one-second allowance
+ instead of stopping before the upload reserve the cap exists to protect.
+ """
+
+ captured: dict[str, object] = {}
+ clock = {"now": 1000.0}
+ monkeypatch.setattr(evolve.time, "monotonic", lambda: clock["now"])
+
+ @contextmanager
+ def fake_prepare_sandbox(**_kwargs):
+ yield SimpleNamespace(
+ claude_bin="claude",
+ command_prefix=[],
+ settings_json="{}",
+ transcript_projects=tmp_path / "transcript-projects",
+ )
+
+ def fake_run_claude(*_args, **kwargs):
+ captured.update(kwargs)
+ return {"ok": True}
+
+ def setup_that_spends_the_whole_budget(_clone):
+ clock["now"] += budget
+ return "0" * 40
+
+ monkeypatch.setattr(evolve.runner, "make_worktree", lambda _repo, _ref, destination: destination)
+ monkeypatch.setattr(evolve.runner, "remove_clone", lambda _clone: None)
+ monkeypatch.setattr(evolve, "sanitize_clone_for_hidden_oracles", setup_that_spends_the_whole_budget)
+ monkeypatch.setattr(evolve, "prepare_sandbox", fake_prepare_sandbox)
+ monkeypatch.setattr(evolve.runner, "run_claude", fake_run_claude)
+
+ args = build_parser().parse_args(["--tasks", "tasks.yaml", "--model", "model"])
+ budget = evolve.MIN_INSTANCE_SWEEP_SECONDS + 1
+ args.max_runtime_seconds = budget
+ record = evolve.run_proposer(
+ "prompt",
+ args,
+ overlay_dir=tmp_path / "overlay",
+ proposal_path=tmp_path / "proposal.md",
+ evidence_bundle=tmp_path / "evidence",
+ bwrap_bin=tmp_path / "bwrap",
+ started_monotonic=clock["now"],
+ )
+
+ assert not captured, "no session may start once the cap is exhausted"
+ assert record["ok"] is False
+ assert record["error_kind"] == "runtime-cap-exhausted"
+
+
def test_parser_defaults_match_the_gate_minimums():
args = build_parser().parse_args(["--tasks", "t.yaml", "--model", "pinned"])
assert args.runs == 3
@@ -922,6 +1048,27 @@ def test_runner_argv_inserts_ce_review_for_review_overlay(tmp_path):
)
arms = argv[argv.index("--arms") + 1 : argv.index("--promotion-metric")]
assert arms == ["ce_review", "review", "candidate_review"]
+ assert "--reuse-results" not in argv
+
+
+def test_runner_argv_forwards_prior_results_for_comparator_reuse(tmp_path):
+ args = build_parser().parse_args(
+ ["--tasks", "t.yaml", "--model", "pinned", "--arms", "review"]
+ )
+ overlay = tmp_path / "overlay"
+ skill = overlay / ".claude" / "skills" / "gitnexus-review" / "SKILL.md"
+ skill.parent.mkdir(parents=True)
+ skill.write_text("candidate")
+ prior = tmp_path / "prior-bench"
+ argv = runner_argv(
+ args,
+ tmp_path / "bench",
+ overlay,
+ task_bindings=[{"id": "task"}],
+ target_base_digests={},
+ reuse_results=prior,
+ )
+ assert argv[argv.index("--reuse-results") + 1] == str(prior)
def test_runner_argv_omits_proposer_for_manual_overlay(tmp_path):
@@ -1118,6 +1265,161 @@ def test_generation_timeout_rejects_unknown_arm() -> None:
)
+def test_instance_window_budget_leaves_upload_reserve() -> None:
+ # Friday 10:57 on a box that booted 02:45 Saturday-window: ~8.2h uptime.
+ leftover = instance_window_budget_seconds(8.2 * 3600)
+ assert leftover == int(86_400 - 8.2 * 3600 - 5_400)
+ assert leftover >= MIN_INSTANCE_SWEEP_SECONDS
+ with pytest.raises(ValueError, match="only .*s left"):
+ instance_window_budget_seconds(23.5 * 3600)
+ with pytest.raises(ValueError, match="uptime must be"):
+ instance_window_budget_seconds(float("nan"))
+
+
+
+
+def test_capped_timeout_clamps_to_leftover_window(monkeypatch) -> None:
+ assert capped_timeout_seconds(10_000, None) == 10_000
+ assert capped_timeout_seconds(10_000, 90) == 90
+ with pytest.raises(ValueError, match="no time remains"):
+ capped_timeout_seconds(10_000, 0)
+ # Pinned clock: the helper is pure arithmetic, so a real elapsed-time window
+ # would assert on scheduling rather than on the behaviour under test.
+ monkeypatch.setattr(evolve.time, "monotonic", lambda: 1040.0)
+ started = 1000.0
+ assert remaining_runtime_seconds(max_runtime_seconds=None, started_monotonic=started) is None
+ leftover = remaining_runtime_seconds(max_runtime_seconds=100, started_monotonic=started)
+ assert leftover is not None
+ assert leftover == 60
+
+
+def test_parser_rejects_non_positive_max_runtime() -> None:
+ with pytest.raises(SystemExit):
+ build_parser().parse_args(
+ ["--tasks", "t.yaml", "--model", "pinned", "--max-runtime-seconds", "0"]
+ )
+ args = build_parser().parse_args(
+ ["--tasks", "t.yaml", "--model", "pinned", "--max-runtime-seconds", "7200"]
+ )
+ assert args.max_runtime_seconds == 7200
+ assert args.max_runtime_from_instance_window is False
+
+
+def _task_file(tmp_path: Path) -> Path:
+ tasks = tmp_path / "tasks.yaml"
+ tasks.write_text(
+ """tasks:
+ - id: demo
+ class: test
+ repo: .
+ prompt: implement
+ verify: "true"
+ oracle:
+ command: "true"
+ files:
+ - source: hidden.test.ts
+ target: hidden.test.ts
+"""
+ )
+ return tasks
+
+
+def _stub_main_preflight(monkeypatch, tmp_path) -> None:
+ """Everything main() shells out to before it reaches _run_generations."""
+
+ monkeypatch.setattr(evolve.runner, "selected_task_bindings", lambda _tasks: [{"id": "demo"}])
+ monkeypatch.setattr(evolve, "preflight_bubblewrap", lambda: tmp_path / "bwrap")
+ monkeypatch.setattr(evolve, "require_claude_sandbox_helpers", lambda: None)
+
+
+def test_the_runtime_cap_is_derived_where_its_clock_starts(monkeypatch, tmp_path, capsys) -> None:
+ """The budget and the clock it is measured against must be one instant.
+
+ run-evolution.sh used to compute the budget in a separate `uv run python -c`
+ and pass a number, so the script's remaining provenance work and this
+ interpreter's startup were charged to the sweep — out of the upload reserve
+ the cap exists to protect. main() reads /proc/uptime itself now, next to its
+ own clock, so no interval exists to lose.
+ """
+
+ monkeypatch.setenv("EVENTBRIDGE_INSTANCE_WINDOW_SECONDS", "20000")
+ monkeypatch.setenv("EVENTBRIDGE_STOP_RESERVE_SECONDS", "1000")
+ monkeypatch.setattr(evolve, "read_instance_uptime_seconds", lambda: 3600.0)
+ captured: dict[str, object] = {}
+
+ def record(args, **kwargs):
+ captured["max_runtime_seconds"] = args.max_runtime_seconds
+ captured["started_monotonic"] = kwargs["started_monotonic"]
+ return 0
+
+ monkeypatch.setattr(evolve, "_run_generations", record)
+ _stub_main_preflight(monkeypatch, tmp_path)
+ monkeypatch.setattr(
+ sys,
+ "argv",
+ [
+ "evolve",
+ "--tasks",
+ str(_task_file(tmp_path)),
+ "--model",
+ "pinned",
+ "--out-root",
+ str(tmp_path / "out"),
+ "--max-runtime-from-instance-window",
+ ],
+ )
+
+ assert evolve.main() == 0
+
+ assert captured["max_runtime_seconds"] == 20000 - 3600 - 1000
+ # Derived here, not passed in: the clock handed to the sweep is the one
+ # taken beside the uptime read.
+ assert isinstance(captured["started_monotonic"], float)
+ assert "capping the sweep to 15400s" in capsys.readouterr().out
+
+
+def test_the_runtime_cap_refuses_two_sources_of_truth(monkeypatch, tmp_path) -> None:
+ monkeypatch.setattr(evolve, "read_instance_uptime_seconds", lambda: 3600.0)
+ monkeypatch.setattr(
+ sys,
+ "argv",
+ [
+ "evolve",
+ "--tasks",
+ str(_task_file(tmp_path)),
+ "--model",
+ "pinned",
+ "--max-runtime-from-instance-window",
+ "--max-runtime-seconds",
+ "7200",
+ ],
+ )
+ with pytest.raises(SystemExit):
+ evolve.main()
+
+
+def test_the_runtime_cap_fails_closed_without_a_readable_uptime(monkeypatch, tmp_path) -> None:
+ def unreadable():
+ raise ValueError("cannot read instance uptime from /proc/uptime")
+
+ monkeypatch.setattr(evolve, "read_instance_uptime_seconds", unreadable)
+ monkeypatch.setattr(
+ sys,
+ "argv",
+ [
+ "evolve",
+ "--tasks",
+ str(_task_file(tmp_path)),
+ "--model",
+ "pinned",
+ "--max-runtime-from-instance-window",
+ ],
+ )
+ # Better to refuse than to run a box-stopped sweep believing it is uncapped.
+ with pytest.raises(SystemExit):
+ evolve.main()
+
+
@pytest.mark.skipif(sys.platform != "linux", reason="Bubblewrap PID namespaces require Linux")
def test_outer_runner_pid_namespace_kills_setsid_descendant(tmp_path):
try:
diff --git a/eval/tests/test_process_control.py b/eval/tests/test_process_control.py
index 8d7405356..1184feea0 100644
--- a/eval/tests/test_process_control.py
+++ b/eval/tests/test_process_control.py
@@ -79,11 +79,11 @@ def run(index, arm):
assert Path({str(assets)!r}).exists(), 'assets removed while a worker was active'
return {{'resolved': False, 'error_kind': result.state}}
with cancellation_scope(handle_signals=True) as event:
- streak, stopped=sweep_task_cells([(i, 'review') for i in range(10)], workers=2, run=run,
+ streak, tripped=sweep_task_cells([(i, 'review') for i in range(10)], workers=2, run=run,
on_start=lambda *args: None, on_record=lambda i,a,r: rows.append([i,r]),
outage_streak=0, outage_limit=5, cancel_event=event)
Path({str(assets)!r}).unlink()
-print(json.dumps({{'rows': rows, 'stopped': stopped}}))
+print(json.dumps({{'rows': rows, 'stopped': event.is_set(), 'tripped': tripped}}))
"""
process = subprocess.Popen(
[PYTHON, "-c", script],
@@ -104,7 +104,8 @@ print(json.dumps({{'rows': rows, 'stopped': stopped}}))
assert process.returncode == 0, stderr
assert time.monotonic() - started < 15
report = json.loads(stdout)
- assert report["stopped"] and [row[0] for row in report["rows"]] == [0, 1]
+ assert report["stopped"] and not report["tripped"], "cancelled, not an outage"
+ assert [row[0] for row in report["rows"]] == [0, 1]
assert report["rows"][0][1]["resolved"] is True
assert report["rows"][1][1]["error_kind"] == "cancelled"
with pytest.raises(ProcessLookupError):
diff --git a/eval/tests/test_proposer_sandbox.py b/eval/tests/test_proposer_sandbox.py
index 2e30cf6a2..26532a5f0 100644
--- a/eval/tests/test_proposer_sandbox.py
+++ b/eval/tests/test_proposer_sandbox.py
@@ -33,9 +33,11 @@ from workflow_bench.proposer_sandbox import (
SANDBOX_GIT_EXCLUDES,
VITE_TEMP_DIR,
SANDBOX_PATH,
+ SANDBOX_REVIEW_OUTPUT,
SANDBOX_PYTHON3,
SANDBOX_SHELL_PREFIX,
SANDBOX_USER_SKILLS,
+ SANDBOX_WORKSPACE,
ReadOnlyMount,
SandboxError,
_runtime_mount_args,
@@ -43,46 +45,65 @@ from workflow_bench.proposer_sandbox import (
build_sandbox_environment,
_force_rmtree,
host_workspace_write_boundary,
+ prepare_review_workspace,
prepare_sandbox,
preflight_bubblewrap,
sandbox_workspace_write_boundary,
stage_evidence_bundle,
stage_task_assets,
)
+from workflow_bench.review_scoring import REVIEW_OUTPUT, parse_review_output
from workflow_bench.task_assets import TaskAssetCache, stage_task_assets as stage_immutable_task_assets
-@pytest.mark.parametrize("entry", ["file", "directory", "relative-link", "absolute-link"])
-def test_review_preparation_rejects_existing_output_without_touching_target(tmp_path, entry):
+@pytest.mark.parametrize("entry", ["directory", "relative-link", "absolute-link"])
+def test_review_preparation_rejects_a_reused_artifact_directory(tmp_path, entry):
clone = tmp_path / "clone"
clone.mkdir()
sentinel = tmp_path / "sentinel"
sentinel.write_text("must survive")
- output = clone / "review-output.json"
- if entry == "file":
- output.write_text("existing result")
- elif entry == "directory":
- output.mkdir()
- else:
- output.symlink_to(sentinel if entry == "absolute-link" else "../sentinel")
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
+ stale = proposer_sandbox.review_output_path(sandbox, "review-output.json").parent
+ if entry == "directory":
+ stale.mkdir()
+ (stale / "review-output.json").write_text("a previous cell's verdict")
+ else:
+ # relpath, not a hand-written "../sentinel": stale is
+ # /review-output, which is nowhere near tmp_path, so the
+ # literal produced a dangling link and the assertion below proved nothing.
+ stale.symlink_to(
+ sentinel if entry == "absolute-link" else Path(os.path.relpath(sentinel, stale.parent))
+ )
with pytest.raises(SandboxError, match="already exists"):
proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
assert sentinel.read_text() == "must survive"
- if entry == "file":
- assert output.read_text() == "existing result"
- if "link" in entry:
- assert output.is_symlink()
-def test_review_preparation_creates_a_private_regular_output(tmp_path):
+def test_review_preparation_leaves_a_clone_entry_of_the_same_name_alone(tmp_path):
+ # The artifact no longer lives in the workspace, so a file that happens to
+ # share its name is just one of the repository's own files.
+ clone = tmp_path / "clone"
+ clone.mkdir()
+ (clone / "review-output.json").write_text("repository content")
+ with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
+ output = proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
+ assert (clone / "review-output.json").read_text() == "repository content"
+ assert clone not in output.parents
+
+
+def test_review_preparation_creates_a_private_directory_and_not_the_file(tmp_path):
clone = tmp_path / "clone"
clone.mkdir()
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
output = proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
- assert output.read_bytes() == b""
- assert stat.S_ISREG(output.lstat().st_mode)
- assert stat.S_IMODE(output.stat().st_mode) == 0o600
+ assert output == proposer_sandbox.review_output_path(sandbox, "review-output.json")
+ # The DIRECTORY is what has to exist and be writable: the agent writes
+ # a temp file beside the target and renames it.
+ assert output.parent.is_dir()
+ assert stat.S_IMODE(output.parent.stat().st_mode) == 0o700
+ # The file is deliberately absent — absence is how "never written" is
+ # told apart from "written badly".
+ assert not output.exists()
def test_review_preparation_preserves_existing_runtime_files_and_tracks_only_created_paths(tmp_path):
@@ -165,6 +186,15 @@ def test_unsafe_host_session_translates_virtual_paths_and_disables_containment(t
assert sandbox.require_pid_namespace is False
assert sandbox.host_path("/workspace/review-output.json") == str(clone / "review-output.json")
assert sandbox.host_path("/evidence/selected-rows.json") == str(evidence / "selected-rows.json")
+ # The review artifact left the workspace, so the host-unsafe backend has
+ # to translate its new home too. Untranslated, the review prompt names a
+ # path that exists on neither backend and the cell writes nothing.
+ assert sandbox.host_path(
+ f"{proposer_sandbox.SANDBOX_REVIEW_OUTPUT}/review-output.json"
+ ) == str(proposer_sandbox.review_output_path(sandbox, "review-output.json"))
+ assert sandbox.host_text(
+ f"write {proposer_sandbox.SANDBOX_REVIEW_OUTPUT}/review-output.json"
+ ) == f"write {proposer_sandbox.review_output_path(sandbox, 'review-output.json')}"
assert sandbox.host_text("read /evidence and write /workspace/out") == (
f"read {evidence} and write {clone}/out"
)
@@ -867,9 +897,8 @@ def test_read_only_review_workspace_exposes_only_one_writable_artifact(tmp_path:
clone.mkdir()
source = clone / "source.ts"
source.write_text("trusted\n")
- output = clone / "review-output.json"
- output.write_text("")
script = """
+import os
from pathlib import Path
try:
Path('/workspace/source.ts').write_text('tampered')
@@ -877,15 +906,24 @@ except OSError:
pass
else:
raise SystemExit('review source remained writable')
-Path('/workspace/review-output.json').write_text('{"schema_version":1}')
+# Write the way the agent's Write tool does: a temp file beside the target,
+# then rename. Writing in place would pass against the mount shape that
+# shipped every artifact empty, which is the regression this canary exists for.
+target = Path('/review-output/review-output.json')
+staging = target.with_name(target.name + '.tmp.1.abc')
+staging.write_text('{"schema_version":1}')
+os.replace(staging, target)
"""
with prepare_sandbox(clone=clone, claude_bin=Path(sys.executable), preflight=True) as sandbox:
+ output = proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
result = run_managed(
[
*sandbox.command_prefix_for(
read_only_workspace=True,
extra_writable_mounts=(
- ReadOnlyMount(source=output, target="/workspace/review-output.json"),
+ ReadOnlyMount(
+ source=output.parent, target=proposer_sandbox.SANDBOX_REVIEW_OUTPUT
+ ),
),
),
"/usr/bin/python3",
@@ -897,9 +935,14 @@ Path('/workspace/review-output.json').write_text('{"schema_version":1}')
require_pid_namespace=True,
)
- assert result.ok, result.stderr_tail
- assert source.read_text() == "trusted\n"
- assert output.read_text() == '{"schema_version":1}'
+ # Inside the sandbox scope: the artifact now lives under the session's
+ # private root, which prepare_sandbox removes on exit. run_arm reads it
+ # here too, while the session is still alive.
+ assert result.ok, result.stderr_tail
+ assert source.read_text() == "trusted\n"
+ assert output.read_text() == '{"schema_version":1}'
+ # The staging file is gone: the rename landed rather than a copy.
+ assert list(output.parent.iterdir()) == [output]
@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
@@ -1170,6 +1213,7 @@ for line in sys.stdin:
review_command = """test -z "${ANTHROPIC_API_KEY:-}" && python3 - <<'PY'
import json
+import os
import subprocess
from pathlib import Path
source = Path('/workspace/canary.txt')
@@ -1182,7 +1226,10 @@ for operation in (lambda: source.write_text('forbidden'), lambda: source.rename(
pass
else:
raise AssertionError('source mutation was allowed')
-Path('/workspace/review-output.json').write_text(json.dumps({'schema_version': 1, 'verdict': 'approve', 'findings': []}))
+target = Path('/review-output/review-output.json')
+staging = target.with_name(target.name + '.tmp.1.abc')
+staging.write_text(json.dumps({'schema_version': 1, 'verdict': 'approve', 'findings': []}))
+os.replace(staging, target)
PY"""
if review_layout:
for command in (
@@ -1359,7 +1406,11 @@ PY"""
sandbox,
command_prefix=sandbox.command_prefix_for(
read_only_workspace=True,
- extra_writable_mounts=(ReadOnlyMount(output, "/workspace/review-output.json"),),
+ extra_writable_mounts=(
+ ReadOnlyMount(
+ output.parent, proposer_sandbox.SANDBOX_REVIEW_OUTPUT
+ ),
+ ),
),
)
result = sandbox.run(
@@ -1408,7 +1459,7 @@ PY"""
assert (sandbox.temp / "mcp-called").read_text() == "ok"
if review_layout:
assert output is not None
- runner_artifacts.enforce_phase_workspace(clone, before, allowed_artifact=output)
+ runner_artifacts.enforce_phase_workspace(clone, before, allowed_artifact=None)
assert json.loads(output.read_text())["verdict"] == "approve"
assert (clone / "canary.txt").read_text() == "hook-readable\nchanged for review\n"
finally:
@@ -1418,3 +1469,98 @@ PY"""
if not review_layout:
assert (clone / "bash-called").read_text() == "canary"
+
+
+def test_review_artifact_binds_a_writable_directory_outside_the_workspace(tmp_path):
+ """The bwrap argv, since the mount shape is the whole bug.
+
+ bwrap cannot create a mount point inside an already-read-only bind, so a
+ writable path has to live outside /workspace — and it has to be the
+ directory, or the agent has nowhere to put the temp file it renames into
+ place.
+ """
+
+ clone = tmp_path / "clone"
+ clone.mkdir()
+ with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as session:
+ sandbox = replace(session, backend="bwrap")
+ output = proposer_sandbox.review_output_path(sandbox, "review-output.json")
+ output.parent.mkdir(mode=0o700)
+ argv = sandbox.command_prefix_for(
+ read_only_workspace=True,
+ extra_writable_mounts=(
+ proposer_sandbox.ReadOnlyMount(
+ source=output.parent,
+ target=proposer_sandbox.SANDBOX_REVIEW_OUTPUT,
+ ),
+ ),
+ )
+
+ target = proposer_sandbox.SANDBOX_REVIEW_OUTPUT
+ assert not target.startswith(proposer_sandbox.SANDBOX_WORKSPACE + "/")
+ # The workspace itself is bound read-only...
+ workspace_at = argv.index(proposer_sandbox.SANDBOX_WORKSPACE)
+ assert argv[workspace_at - 2] == "--ro-bind"
+ # ...and the artifact directory is bound writable, as a directory.
+ artifact_at = argv.index(target)
+ assert argv[artifact_at - 2] == "--bind"
+ assert Path(argv[artifact_at - 1]) == output.parent
+ assert Path(argv[artifact_at - 1]).is_dir()
+ assert f"{proposer_sandbox.SANDBOX_WORKSPACE}/review-output.json" not in argv
+
+
+@pytest.mark.skipif(
+ os.environ.get("GITNEXUS_REQUIRE_BWRAP_CANARY") != "1",
+ reason="real Bubblewrap canary is mandatory in the named Ubuntu CI job",
+)
+def test_real_bubblewrap_lets_a_review_artifact_be_written_atomically(tmp_path: Path) -> None:
+ """The filesystem contract the EROFS defect broke, under a real sandbox.
+
+ Argv assertions cannot establish this. The artifact came back empty because
+ an atomic write - temp file beside the target, then rename - needs a
+ WRITABLE PARENT DIRECTORY, and only a real bwrap invocation shows whether
+ the mount grants one. A deterministic writer stands in for the agent: no
+ model session, no credentials.
+
+ Scope: this proves the filesystem and process contract of the production
+ mount configuration. It does not establish that a particular agent CLI's
+ own file-access policy permits the same operation - that is a second,
+ independent gate.
+ """
+
+ clone = tmp_path / "clone"
+ clone.mkdir()
+ (clone / "tracked.txt").write_text("original\n")
+
+ with prepare_sandbox(clone=clone, claude_bin=Path(sys.executable), preflight=True) as sandbox:
+ review_output = prepare_review_workspace(sandbox, REVIEW_OUTPUT)
+ # The production configuration, not a hand-built mount tuple: the same
+ # command_prefix_for call run_arm makes for a review cell.
+ prefix = sandbox.command_prefix_for(
+ read_only_workspace=True,
+ extra_writable_mounts=(
+ ReadOnlyMount(source=review_output.parent, target=SANDBOX_REVIEW_OUTPUT),
+ ),
+ )
+ target = f"{SANDBOX_REVIEW_OUTPUT}/{REVIEW_OUTPUT}"
+ script = (
+ # 1. temp file beside the destination, then atomic rename over it.
+ f'printf %s \'{{"schema_version": 1, "verdict": "approve", "findings": []}}\' > {target}.tmp && '
+ f"mv {target}.tmp {target} && "
+ # 2. the workspace must refuse the write that the mount forbids.
+ f"(printf x >> {SANDBOX_WORKSPACE}/tracked.txt 2>/dev/null && echo WORKSPACE-WRITABLE || echo workspace-readonly)"
+ )
+ result = subprocess.run(
+ [*prefix, "/bin/sh", "-c", script],
+ capture_output=True, text=True, timeout=60, check=False,
+ )
+
+ assert result.returncode == 0, f"atomic write failed inside the sandbox: {result.stderr[-400:]}"
+ assert "workspace-readonly" in result.stdout, "the workspace must stay read-only"
+ assert (clone / "tracked.txt").read_text() == "original\n", "the clone was modified"
+
+ # Read while the session is alive: the artifact lives under the private
+ # root that prepare_sandbox removes on exit, which is also why run_arm
+ # consumes it before leaving the scope.
+ _verdict, findings = parse_review_output(review_output)
+ assert findings == ()
diff --git a/eval/tests/test_reuse_round_trip.py b/eval/tests/test_reuse_round_trip.py
new file mode 100644
index 000000000..a31e9c8f7
--- /dev/null
+++ b/eval/tests/test_reuse_round_trip.py
@@ -0,0 +1,268 @@
+"""A row the runner actually emits must satisfy the reuse reader.
+
+Every existing comparator-reuse test builds its rows by hand. That proves the
+predicate's logic and nothing about the producer: a fixture can satisfy
+eligibility while a real emitted row never does, and the audit that counts key
+names cannot tell the difference. These tests carry one record through the
+production path instead:
+
+ real run_cell -> production JSONL writer -> load_result_rows
+ -> row_is_reusable_comparator
+
+Only the expensive dependencies are replaced - the model session, sandbox
+launch, repository acquisition, graph preparation. The digest fields the reuse
+binding compares are assembled by run_cell itself from its TaskCellContext, so
+they stay real: they are the subject of the test, not scaffolding around it.
+"""
+
+from __future__ import annotations
+
+import json
+from datetime import UTC, datetime, timedelta
+from pathlib import Path
+from types import SimpleNamespace
+from typing import Any
+
+import pytest
+
+from workflow_bench import runner
+from workflow_bench.proposer_sandbox import redact_text
+from workflow_bench.model_gateway import credential_secrets
+from workflow_bench.runner_sessions import PARENT_EVENT_STREAM_SOURCE
+from workflow_bench.comparator_reuse import (
+ ComparatorReuseExpectation,
+ TaskReuseBinding,
+ load_result_rows,
+ row_is_reusable_comparator,
+)
+
+TASK_ID = "review-pr-2718-defect"
+SHA = "a" * 40
+
+
+def _snapshot(prefix: str) -> SimpleNamespace:
+ return SimpleNamespace(
+ digest=f"{prefix}-content",
+ manifest_digest=f"{prefix}-manifest",
+ dependency_content_digest=f"{prefix}-dep-content",
+ dependency_manifest_digest=f"{prefix}-dep-manifest",
+ command_digest=f"{prefix}-command",
+ materialize=lambda *a, **k: None,
+ )
+
+
+def _write_like_the_sweep(tmp_path: Path, row: dict[str, Any]) -> Path:
+ """Serialize exactly as ``keep`` does in _run_sweep, redaction included.
+
+ json.dumps + write_text would skip the redaction the real writer applies,
+ so a change there could break reusable rows without failing this test - and
+ redaction is not cosmetic here, since it rewrites the row's own bytes.
+ """
+
+ results = tmp_path / "results.jsonl"
+ secrets = credential_secrets(
+ SimpleNamespace(auth_token="sk-ant-should-never-appear", base_url=None)
+ )
+ with results.open("a") as handle:
+ handle.write(redact_text(json.dumps(row), secrets) + "\n")
+ return results
+
+
+@pytest.fixture
+def emitted_row(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> dict[str, Any]:
+ """One record from the real run_cell, with only expensive work replaced."""
+
+ worktree = tmp_path / "clone"
+ worktree.mkdir()
+
+ # The session is what costs money; everything it returns is scripted. The
+ # record's binding fields are NOT set here - run_cell derives them.
+ def fake_run_arm(*_a: Any, **_k: Any) -> dict[str, Any]:
+ return {
+ "ok": True,
+ "error_kind": None,
+ "error_detail": None,
+ "resolved": True,
+ "review_evidence_valid": True,
+ "review_score": {"weighted_f1": 0.5},
+ "review_weighted_f1": 0.5,
+ "skill_invoked": True,
+ "skill_digest": "skill-digest",
+ "transcript_missing": False,
+ "transcript_artifacts": [
+ {
+ "path": "transcripts/session-1.jsonl",
+ "sha256": __import__("hashlib").sha256(b'{"type":"ok"}\n').hexdigest(),
+ "bytes": 14,
+ "source": PARENT_EVENT_STREAM_SOURCE,
+ }
+ ],
+ "session_ids": ["s1"],
+ "num_turns": 3,
+ "duration_s": 1.0,
+ "cost_usd": 0.5,
+ "input_tokens": 1,
+ "output_tokens": 1,
+ "cache_creation_input_tokens": 0,
+ "cache_read_input_tokens": 0,
+ }
+
+ for name, value in {
+ "run_arm": fake_run_arm,
+ "copy_isolated_tree": lambda *a, **k: worktree,
+ "make_worktree": lambda *a, **k: worktree,
+ "sanitize_clone_for_hidden_oracles": lambda *a, **k: SHA,
+ "stage_task_assets": lambda *a, **k: (),
+ "isolated_gitnexus_registry_mount": lambda *a, **k: None,
+ "seed_evaluated_skills": lambda *a, **k: None,
+ "apply_candidate_overlay": lambda *a, **k: None,
+ "require_hidden_harness_absent": lambda *a, **k: None,
+ "require_skill_fingerprint": lambda *a, **k: None,
+ "enforce_work_evidence": lambda *a, **k: None,
+ "skill_fingerprint": lambda *a, **k: "skill-digest",
+ "capture_patch": lambda *a, **k: b"",
+ "implementation_diff_digest": lambda *a, **k: "",
+ "diff_churn": lambda *a, **k: {},
+ "_prepare_untracked_for_diff": lambda *a, **k: None,
+ "remove_clone": lambda *a, **k: None,
+ "ce_plugin_dir_for_arm": lambda *a, **k: None,
+ "ce_plugin_mounts_for_arm": lambda *a, **k: (),
+ "current_runtime_digest": lambda: "runtime-digest",
+ "build_sandbox_environment": lambda *a, **k: {},
+ "credential_secrets": lambda *a, **k: (),
+ # run_cell requires an immutable base commit before it will record a
+ # cell; the git plumbing is expensive setup, the SHA it returns is not
+ # part of the reuse binding under test.
+ "_sandbox_git": lambda *a, **k: SHA,
+ # The artifact copy is real; only the read of the agent-written file is
+ # replaced, since no agent ran to write one.
+ "_bounded_regular_bytes": lambda *a, **k: b'{"schema_version":1}',
+ }.items():
+ monkeypatch.setattr(runner, name, value)
+
+ class _Sandbox:
+ clone = worktree
+ private_root = tmp_path / "private"
+ backend = "test-double"
+ settings_json = "{}"
+ require_pid_namespace = False
+
+ def __enter__(self) -> _Sandbox:
+ return self
+
+ def __exit__(self, *_exc: Any) -> bool:
+ return False
+
+ def command_prefix_for(self, **_k: Any) -> list[str]:
+ return []
+
+ def run(self, *_a: Any, **_k: Any) -> SimpleNamespace:
+ return SimpleNamespace(ok=True, returncode=0, stdout_tail="", stderr_tail="")
+
+ def environment(self, **_k: Any) -> dict[str, str]:
+ return {}
+
+ def host_text(self, value: str) -> str:
+ return value
+
+ monkeypatch.setattr(runner, "prepare_sandbox", lambda **_k: _Sandbox())
+
+ ctx = runner.TaskCellContext(
+ task={"id": TASK_ID, "prompt": "review it", "verify": "true"},
+ oracle_snapshot=_snapshot("oracle"),
+ repo=tmp_path / "repo",
+ task_sha=SHA,
+ graph_snapshot=_snapshot("graph"),
+ graph_snapshot_error=None,
+ asset_snapshot=_snapshot("asset"),
+ asset_snapshot_error=None,
+ args=SimpleNamespace(
+ model="gpt-5.6-sol", effort="xhigh", timeout=60, claude_bin="claude",
+ base_url=None, auth_token=None, permission_mode=None, arms=["review"],
+ proposer_model=None, outage_streak=5, runs=1, workers=1,
+ ),
+ out_dir=tmp_path / "out",
+ ce_plugin_snapshot=None,
+ trees_dir=tmp_path / "trees",
+ bwrap_bin=Path("/bin/true"),
+ runtime_mounts=(),
+ candidate_overlay=None,
+ overlay_digest=None,
+ sandbox_backend="test-double",
+ clone_template=None,
+ sanitized_head=SHA,
+ )
+ (tmp_path / "out").mkdir(exist_ok=True)
+ (tmp_path / "trees").mkdir(exist_ok=True)
+ (tmp_path / "private").mkdir(exist_ok=True)
+ # run_cell records review_artifact only when the review source exists, and
+ # reuse now requires it - a scored review with no artifact is a claim about
+ # evidence rather than the evidence. Production writes this file; the
+ # fixture has to as well, or the emitted row is one production never emits.
+ review_dir = tmp_path / "private" / "review-output"
+ review_dir.mkdir(exist_ok=True)
+ (review_dir / "review-output.json").write_text('{"schema_version": 1, "verdict": "approve", "findings": []}')
+ return runner.run_cell(ctx, 0, "review")
+
+
+def _expectation(**overrides: Any) -> ComparatorReuseExpectation:
+ """Bindings from the sweep's own configuration, not copied out of the row.
+
+ Copying the emitted values back in would make producer and consumer agree
+ because the test arranged it, which is the blind spot being closed.
+ """
+
+ binding = TaskReuseBinding(
+ task_base_sha=SHA,
+ task_prompt_digest=runner.hashlib.sha256(b"review it").hexdigest(),
+ oracle_digest="oracle-content",
+ oracle_command_digest="oracle-command",
+ oracle_manifest_digest="oracle-manifest",
+ task_asset_manifest_digest="asset-manifest",
+ sandbox_dependency_manifest_digest="asset-dep-manifest",
+ )
+ values: dict[str, Any] = dict(
+ model="gpt-5.6-sol",
+ effort="xhigh",
+ sandbox_backend="test-double",
+ runtime_digest="runtime-digest",
+ now=datetime.now(UTC),
+ max_age=timedelta(days=90),
+ tasks={TASK_ID: binding},
+ skill_digests={"review": "skill-digest"},
+ ce_plugin_version=None,
+ ce_plugin_manifest_digest=None,
+ )
+ values.update(overrides)
+ return ComparatorReuseExpectation(**values)
+
+
+def test_a_row_the_runner_emitted_survives_serialization_and_qualifies(
+ emitted_row: dict[str, Any], tmp_path: Path
+) -> None:
+ """The producer/consumer contract, end to end through the real writer."""
+
+ results = _write_like_the_sweep(tmp_path, emitted_row)
+ rows = load_result_rows(results)
+ assert len(rows) == 1, "the production row must survive the reader"
+
+ assert row_is_reusable_comparator(rows[0], _expectation()) is True
+
+
+def test_a_changed_binding_rejects_the_same_emitted_row(
+ emitted_row: dict[str, Any], tmp_path: Path
+) -> None:
+ """Fails closed on drift, so the positive case is not vacuous."""
+
+ row = load_result_rows(_write_like_the_sweep(tmp_path, emitted_row))[0]
+
+ binding = TaskReuseBinding(
+ task_base_sha=SHA,
+ task_prompt_digest=runner.hashlib.sha256(b"review it").hexdigest(),
+ oracle_digest="oracle-content",
+ oracle_command_digest="oracle-command",
+ oracle_manifest_digest="oracle-manifest",
+ task_asset_manifest_digest="asset-manifest",
+ sandbox_dependency_manifest_digest="DIFFERENT-dependencies",
+ )
+ assert row_is_reusable_comparator(row, _expectation(tasks={TASK_ID: binding})) is False
diff --git a/eval/tests/test_review_corpus.py b/eval/tests/test_review_corpus.py
index 42f18cb7c..cbaa2695e 100644
--- a/eval/tests/test_review_corpus.py
+++ b/eval/tests/test_review_corpus.py
@@ -32,6 +32,10 @@ def test_review_corpus_is_immutable_and_task_bound():
assert task["ref"] == case["base_sha"]
assert task["sandbox_copy"] == [f"eval/workflow_bench/review_cases/{patch.name}"]
assert task["setup"] == review_case_setup_command(patch.name)
+ assert any(
+ dep.get("source") == "gitnexus-shared/dist" and dep.get("target") == "gitnexus-shared/dist"
+ for dep in task["sandbox_dependencies"]
+ )
def test_hidden_labels_are_not_recoverable_from_visible_task_input():
diff --git a/eval/tests/test_review_scoring.py b/eval/tests/test_review_scoring.py
index c1b081720..1e6c8922e 100644
--- a/eval/tests/test_review_scoring.py
+++ b/eval/tests/test_review_scoring.py
@@ -325,3 +325,33 @@ def test_clean_control_rewards_an_empty_approval_and_penalizes_noise():
assert noisy["recall"] is None
assert noisy["clean_pass"] is False
assert noisy["verdict_correct"] is False
+
+
+def test_parse_review_output_names_the_actual_failure(tmp_path: Path):
+ """One message per cause.
+
+ Folding empty, malformed and encoding failures together makes a sandbox that
+ left the artifact at 0 bytes indistinguishable from an encoding fault: every
+ such cell reports "not valid UTF-8 JSON". A file the agent never created
+ escaped that fold — lstat sat outside the try, so it raised
+ FileNotFoundError — but only as a bare OSError, naming no cause at all.
+ """
+
+ missing = tmp_path / "never-written.json"
+ with pytest.raises(ValueError, match="was never written"):
+ parse_review_output(missing)
+
+ empty = tmp_path / "empty.json"
+ empty.touch()
+ with pytest.raises(ValueError, match="is empty"):
+ parse_review_output(empty)
+
+ not_utf8 = tmp_path / "latin1.json"
+ not_utf8.write_bytes(b'{"verdict": "\xff\xfe"}')
+ with pytest.raises(ValueError, match="not valid UTF-8"):
+ parse_review_output(not_utf8)
+
+ prose = tmp_path / "prose.json"
+ prose.write_text("Here is my review of the changes.", encoding="utf-8")
+ with pytest.raises(ValueError, match="not valid JSON"):
+ parse_review_output(prose)
diff --git a/eval/tests/test_runner_hardening.py b/eval/tests/test_runner_hardening.py
index be196750b..111b3c3c5 100644
--- a/eval/tests/test_runner_hardening.py
+++ b/eval/tests/test_runner_hardening.py
@@ -3,13 +3,14 @@
import hashlib
import json
import shutil
+import subprocess
from contextlib import nullcontext
from pathlib import Path
from types import SimpleNamespace
import pytest
-from workflow_bench import runner, runner_artifacts, runner_sessions
+from workflow_bench import proposer_sandbox, runner, runner_artifacts, runner_sessions
from workflow_bench.evolution import skill_fingerprint
from workflow_bench.oracle_assets import review_case_setup_command
from workflow_bench.process_control import ManagedProcessError, ManagedProcessResult
@@ -608,6 +609,67 @@ def test_run_cell_reports_a_cleanup_failure_over_its_primary_outcome(monkeypatch
assert "clone is busy" in record["error_detail"]
+def _git(repo, *args):
+ return subprocess.run(["git", "-C", str(repo), *args], check=True, capture_output=True, text=True)
+
+
+def test_run_cell_runs_the_arm_against_a_copy_of_the_clone_template(monkeypatch, tmp_path):
+ """run_cell must copy the template, never re-clone.
+
+ run_cell takes the clone-template branch on essentially every multi-cell
+ sweep: it copies a pre-sanitized template rather than paying `git clone
+ --no-local` plus repack/prune/fsck per cell. Asserting on a copy the test
+ makes itself proves nothing about that branch — the clone the arm receives
+ is what has to come from the template, carrying the template's sanitized
+ HEAD rather than a recomputed one.
+ """
+
+ repo = tmp_path / "repo"
+ repo.mkdir()
+ _git(repo, "init", "--quiet")
+ _git(repo, "checkout", "--quiet", "-b", "main")
+ (repo / "from-template.txt").write_text("sanitized\n")
+ _git(repo, "add", "-A")
+ _git(repo, "-c", "user.name=test", "-c", "user.email=test@invalid", "commit", "--quiet", "-m", "base")
+ sha = _git(repo, "rev-parse", "HEAD").stdout.strip()
+ trees = tmp_path / "trees"
+ trees.mkdir()
+ template = runner.make_worktree(repo, sha, trees)
+ template_head = _git(template, "rev-parse", "HEAD").stdout.strip()
+
+ _stub_cell_dependencies(monkeypatch, tmp_path)
+
+ def fail_if_recloned(*_args, **_kwargs):
+ raise AssertionError("clone template present: run_cell must not re-clone")
+
+ monkeypatch.setattr(runner, "make_worktree", fail_if_recloned)
+ monkeypatch.setattr(runner, "sanitize_clone_for_hidden_oracles", fail_if_recloned)
+
+ seen: dict[str, object] = {}
+
+ def record_arm(_arm, _task, worktree, _args, **_kwargs):
+ seen["worktree"] = worktree
+ seen["head"] = _git(worktree, "rev-parse", "HEAD").stdout.strip()
+ seen["content"] = (worktree / "from-template.txt").read_text()
+ # The copy is a private checkout: what the cell writes must not reach
+ # the template the other cells of this task still copy from.
+ (worktree / "from-template.txt").write_text("cell-local\n")
+ return {"resolved": True, "ok": True, "error_kind": None}
+
+ monkeypatch.setattr(runner, "run_arm", record_arm)
+
+ runner.run_cell(
+ _cell_context(tmp_path, clone_template=template, sanitized_head=template_head),
+ 0,
+ "workflow",
+ )
+
+ assert seen["content"] == "sanitized\n"
+ assert seen["head"] == template_head
+ assert seen["worktree"] != template
+ assert (template / "from-template.txt").read_text() == "sanitized\n"
+
+
def test_run_cell_does_not_mask_the_staged_review_patch_before_setup(monkeypatch, tmp_path):
"""Review setup applies a patch staged under eval/workflow_bench.
@@ -1002,3 +1064,40 @@ def test_progress_line_reports_the_numbers_a_real_run_measured():
assert "cost=$0.5" in line
assert "took=12.0s" in line
assert "error_kind=none" in line
+
+
+def test_claude_settings_allow_the_review_artifact_directory():
+ """The second gate on the artifact path.
+
+ The bwrap bind is not the only thing that decides whether the agent can
+ write: the CLI applies this filesystem policy to its own tools, so a path
+ missing from allowWrite is unwritable however the mount is shaped. The
+ artifact lived under /workspace when this list was written, which is why
+ moving it out needed this entry and nothing caught the omission.
+ """
+
+ settings = json.loads(proposer_sandbox.build_claude_settings(sandbox_enabled=True))
+ filesystem = settings["sandbox"]["filesystem"]
+ assert proposer_sandbox.SANDBOX_REVIEW_OUTPUT in filesystem["allowWrite"]
+ assert proposer_sandbox.SANDBOX_REVIEW_OUTPUT in filesystem["allowRead"]
+ assert filesystem["denyRead"] == ["/"]
+
+
+def test_review_contract_tells_the_agent_the_writable_path():
+ prompt = runner.REVIEW_PROMPT.format(task="task text")
+ assert f"{runner.SANDBOX_REVIEW_OUTPUT}/{runner.REVIEW_OUTPUT}" in prompt
+ assert f"{runner.SANDBOX_WORKSPACE}/{runner.REVIEW_OUTPUT}" not in prompt
+ # The JSON shape survives .format() with its braces intact.
+ assert '{"schema_version":1' in prompt
+ artifact = f"{runner.SANDBOX_REVIEW_OUTPUT}/{runner.REVIEW_OUTPUT}"
+ assert runner.CE_REVIEW_PROMPT.format(task="task text").count(artifact) == 1
+
+
+def test_enforce_phase_workspace_can_require_an_untouched_workspace(tmp_path):
+ (tmp_path / "tracked.py").write_text("original\n")
+ before = runner_artifacts.workspace_snapshot(tmp_path)
+ runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=None)
+
+ (tmp_path / "tracked.py").write_text("the review edited the code it was reviewing\n")
+ with pytest.raises(ValueError, match="changed the read-only workspace"):
+ runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=None)
diff --git a/eval/tests/test_sanitized_graph.py b/eval/tests/test_sanitized_graph.py
index 7b5dfc093..d7f73d756 100644
--- a/eval/tests/test_sanitized_graph.py
+++ b/eval/tests/test_sanitized_graph.py
@@ -212,6 +212,22 @@ def test_prepare_sanitized_graph_builds_once_from_parentless_tree_and_caches_onl
assert removed == [seed]
+def test_prepare_sanitized_graph_requires_head_when_given_a_template(tmp_path: Path):
+ with pytest.raises(SandboxError, match="sanitized HEAD"):
+ sanitized_graph.prepare_sanitized_graph(
+ {},
+ repo=tmp_path,
+ resolved_sha="b" * 40,
+ parent=tmp_path,
+ cache=SimpleNamespace(), # type: ignore[arg-type]
+ claude_bin="claude",
+ bwrap_bin="bwrap",
+ runtime_mounts=(),
+ clone_template=tmp_path,
+ sanitized_head=None,
+ )
+
+
def test_graph_snapshot_rejects_arm_sanitization_identity_drift(tmp_path: Path):
assets = SimpleNamespace(
digest="digest",
diff --git a/eval/tests/test_session_progress.py b/eval/tests/test_session_progress.py
index 9e59ba4bf..416c6206c 100644
--- a/eval/tests/test_session_progress.py
+++ b/eval/tests/test_session_progress.py
@@ -11,7 +11,7 @@ import io
import json
import time
-from workflow_bench.runner_sessions import SessionProgress
+from workflow_bench.runner_sessions import SessionProgress, neutralize_ci_log_text
def _drain_lines(stream: io.StringIO) -> list[str]:
@@ -325,3 +325,48 @@ def test_cell_failure_detail_line_bounds_a_huge_detail() -> None:
assert line is not None
assert "truncated" in line
assert len(line) < MAX_CELL_DETAIL_CHARS + 200
+
+
+def test_progress_neutralizes_github_actions_annotation_forms() -> None:
+ rewritten = neutralize_ci_log_text(
+ "gitnexus/src/cli/optional-grammars.ts(18,36): error TS2307: Cannot find module "
+ "'gitnexus-shared'\n::error::Composite projects may not disable incremental compilation.\n"
+ "##[error]tsc failed"
+ )
+ assert "): error TS2307" not in rewritten
+ assert "): compiler-error TS2307" in rewritten
+ assert "::error::" not in rewritten
+ assert "[:]error::" in rewritten
+ assert "##[error]" not in rewritten
+ assert "# [error]tsc failed" in rewritten
+
+ stream = io.StringIO()
+ progress = SessionProgress("review-pr-2718-defect-ce_review-run0", stream=stream, heartbeat_s=3600)
+ events = [
+ {
+ "type": "assistant",
+ "message": {
+ "content": [{"type": "tool_use", "id": "b1", "name": "Bash", "input": {"command": "npx tsc --noEmit"}}]
+ },
+ },
+ {
+ "type": "user",
+ "message": {
+ "content": [
+ {
+ "type": "tool_result",
+ "tool_use_id": "b1",
+ "is_error": True,
+ "content": "gitnexus/src/cli/optional-grammars.ts(18,36): error TS2307: Cannot find module 'gitnexus-shared'",
+ }
+ ]
+ },
+ },
+ ]
+ for event in events:
+ _observe(progress, (json.dumps(event) + "\n").encode())
+
+ output = stream.getvalue()
+ assert "): error TS2307" not in output
+ assert "): compiler-error TS2307" in output
+ assert "result=error" in output
diff --git a/eval/tests/test_sweep_finalization.py b/eval/tests/test_sweep_finalization.py
new file mode 100644
index 000000000..176157645
--- /dev/null
+++ b/eval/tests/test_sweep_finalization.py
@@ -0,0 +1,279 @@
+"""The real sweep must reach the right finalization decision.
+
+`enforce_measurement_health` is unit-tested and the call site is pinned
+structurally, but neither shows the guard running inside a sweep. These drive
+the real `_run_sweep` with cell execution scripted and everything downstream of
+it left alone: folding, aggregation, the artifact writers, the health guard and
+the exit selection.
+
+The below-breaker case is the decisive one. A fixture of many unusable cells
+aborts through the pre-existing outage breaker instead - `review-evidence-invalid`
+is systemic with a limit of 5 - and would pass whether or not the finalization
+guard exists. One fresh unusable cell stays under that threshold, so only the
+guard can catch it.
+"""
+
+from __future__ import annotations
+
+import json
+import threading
+from collections.abc import Callable
+from pathlib import Path
+from types import SimpleNamespace
+from typing import Any
+
+import pytest
+
+from tests.bench_fixtures import scored_review_row, unusable_review_row
+from workflow_bench import runner
+
+TASK = {
+ "id": "review-pr-2718-defect",
+ "repo": "~/GitNexus",
+ "ref": "a" * 40,
+ "prompt": "review it",
+ "verify": "true",
+ "class": "review-defect",
+}
+
+
+def _args(out: Path, **overrides: Any) -> SimpleNamespace:
+ values: dict[str, Any] = dict(
+ arms=["review"], claude_bin="claude", effort="xhigh", model="gpt-5.6-sol",
+ out=out, outage_streak=runner.DEFAULT_OUTAGE_STREAK, promotion_max_task_regression=10.0,
+ promotion_metric="review_weighted_f1", promotion_min_improvement=1.0,
+ promotion_min_runs=1, proposer_model=None, reuse_results=None, runs=1, workers=1,
+ timeout=60, base_url=None, auth_token=None, permission_mode=None,
+ )
+ values.update(overrides)
+ return SimpleNamespace(**values)
+
+
+def _snapshot(prefix: str) -> SimpleNamespace:
+ return SimpleNamespace(
+ digest=f"{prefix}-content", manifest_digest=f"{prefix}-manifest",
+ dependency_content_digest=f"{prefix}-dep", dependency_manifest_digest=f"{prefix}-depman",
+ command_digest=f"{prefix}-command", materialize=lambda *a, **k: None,
+ )
+
+
+def _sweep(
+ tmp_path: Path,
+ monkeypatch: pytest.MonkeyPatch,
+ record: dict[str, Any] | Callable[[int], dict[str, Any]],
+ *,
+ runs: int = 1,
+ cancel_event: threading.Event | None = None,
+ candidate_arms: list[str] | None = None,
+ arms: list[str] | None = None,
+ after_cell: Callable[[int, str], None] | None = None,
+):
+ """Drive the real _run_sweep; only cell execution and setup are scripted.
+
+ ``after_cell`` runs once a cell's record exists, which is how a test sets
+ cancellation deterministically at a known point instead of racing a sleep.
+ """
+
+ out = tmp_path / "out"
+
+ def scripted_cell(_ctx: Any, run_idx: int, arm: str) -> dict[str, Any]:
+ row = dict(record(run_idx) if callable(record) else record)
+ row.update({"task": TASK["id"], "arm": arm, "run": run_idx, "class": TASK["class"]})
+ if after_cell is not None:
+ after_cell(run_idx, arm)
+ return row
+
+ monkeypatch.setattr(runner, "run_cell", scripted_cell)
+ monkeypatch.setattr(runner, "ensure_task_graph", lambda **k: k["env"].graph_snapshots.__setitem__(
+ k["graph_key"], _snapshot("graph")))
+ monkeypatch.setattr(runner.TaskAssetCache, "prepare", lambda self, *a, **k: _snapshot("asset"))
+ # Binding resolution clones the repo and verifies the ref; that is expensive
+ # setup, and the bindings it would return are supplied directly instead.
+ monkeypatch.setattr(
+ runner, "resolve_task_bindings",
+ lambda tasks, expected, **k: list(expected),
+ )
+
+ return runner._run_sweep(
+ _args(out, runs=runs, arms=arms or ["review"]),
+ parser=SimpleNamespace(error=lambda m: (_ for _ in ()).throw(SystemExit(2))),
+ tasks=[TASK],
+ skipped_expensive=[],
+ oracle_snapshots=[_snapshot("oracle")],
+ expected_task_bindings=[{"repo_identity": str(tmp_path / "repo"), "resolved_sha": "a" * 40}],
+ ce_plugin_config=None,
+ bwrap_bin=Path("/bin/true"),
+ sandbox_backend="test-double",
+ runtime_mounts=(),
+ candidate_arms=candidate_arms or [],
+ candidate_overlay=None,
+ overlay_digest=None,
+ promotion_target_bases={},
+ cancel_event=cancel_event,
+ ), out
+
+
+def test_one_unusable_cell_below_the_breaker_reaches_the_finalization_guard(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
+) -> None:
+ """The decisive case: too few failures to trip the breaker, so only the guard can catch it."""
+
+ streak = runner.systemic_outage_streak("review-evidence-invalid", 0)
+ assert streak < runner.DEFAULT_OUTAGE_STREAK, "fixture must stay under the breaker"
+
+ unusable = unusable_review_row()
+ with pytest.raises(SystemExit) as exc:
+ _sweep(tmp_path, monkeypatch, unusable)
+ assert exc.value.code == 1
+ out = capsys.readouterr().out
+ assert "review: UNUSABLE" in out, "the guard must name the arm and its status"
+ assert "systemic-outage" not in out, "the breaker must not have tripped"
+
+
+def test_a_zero_score_stays_a_valid_negative_measurement(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
+) -> None:
+ """0.0 is a present measurement, not missing evidence.
+
+ A truthiness check on the score would misread it as absent and turn a
+ quality result into an execution-health failure.
+ """
+
+ zeroed = scored_review_row(
+ resolved=False, error_kind="oracle-failed",
+ review_score={"weighted_f1": 0.0}, review_weighted_f1=0.0,
+ )
+ _sweep(tmp_path, monkeypatch, zeroed)
+ out = capsys.readouterr().out
+ assert "review: OBSERVED_OK" in out
+ assert "UNUSABLE" not in out
+
+
+def test_finalization_persists_results_and_report(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """Evidence must survive the sweep, and say the same thing the exit does."""
+
+ scored = scored_review_row(resolved=False, error_kind="oracle-failed", review_weighted_f1=0.2)
+ _result, out = _sweep(tmp_path, monkeypatch, scored)
+ rows = [json.loads(line) for line in (out / "results.jsonl").read_text().splitlines()]
+ assert len(rows) == 1 and rows[0]["review_weighted_f1"] == 0.2
+ assert (out / "report.md").is_file()
+
+
+def test_cancellation_without_an_outage_exits_130(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
+) -> None:
+ """An interrupted sweep is interrupted, not aborted.
+
+ One admissible cell lands first so the measurement-health guard classifies
+ the arm DEGRADED rather than UNUSABLE - otherwise the guard would supply
+ exit 1 and this test would pass without ever exercising exit selection.
+ """
+
+ cancel_event = threading.Event()
+ with pytest.raises(SystemExit) as exc:
+ _sweep(
+ tmp_path, monkeypatch, lambda _run: scored_review_row(),
+ runs=3, cancel_event=cancel_event,
+ after_cell=lambda run_idx, _arm: cancel_event.set() if run_idx == 0 else None,
+ )
+ stdout = capsys.readouterr().out
+ report = (tmp_path / "out" / "report.md").read_text()
+ assert "Sweep cancelled" in report, "an interruption must be reported as one"
+ assert "systemic-outage" not in stdout, "no breaker trip in this scenario"
+ assert exc.value.code == 130
+
+
+def test_an_outage_keeps_exit_1_even_though_the_breaker_cancels(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
+) -> None:
+ """Precedence: the breaker sets cancel_event, so order decides the exit.
+
+ Testing cancellation first would relabel every outage a Ctrl-C. The first
+ cell is admissible for the same reason as above, and the failures after it
+ are consecutive and systemic, which is what the breaker actually counts.
+ """
+
+ def cell(run_idx: int) -> dict[str, Any]:
+ return scored_review_row() if run_idx == 0 else unusable_review_row()
+
+ cancel_event = threading.Event()
+ with pytest.raises(SystemExit) as exc:
+ _sweep(tmp_path, monkeypatch, cell,
+ runs=1 + runner.DEFAULT_OUTAGE_STREAK, cancel_event=cancel_event)
+ stdout = capsys.readouterr().out
+ report = (tmp_path / "out" / "report.md").read_text()
+ assert "systemic-outage" in stdout, "the real breaker must have tripped"
+ assert cancel_event.is_set(), "the breaker cancels in-flight work"
+ assert "Sweep aborted" in report
+ assert exc.value.code == 1, "an outage must not become the 130 of a Ctrl-C"
+
+
+def test_an_interrupted_sweep_keeps_the_evidence_it_already_paid_for(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """Cancellation must not discard rows that already cost money.
+
+ The completed-run persistence test cannot show this: it never interrupts, so
+ it would pass even if the writer only ran on the clean path.
+ """
+
+ cancel_event = threading.Event()
+ with pytest.raises(SystemExit):
+ _sweep(
+ tmp_path, monkeypatch, lambda _run: scored_review_row(review_weighted_f1=0.42),
+ runs=3, cancel_event=cancel_event,
+ after_cell=lambda run_idx, _arm: cancel_event.set() if run_idx == 0 else None,
+ )
+ rows = [
+ json.loads(line)
+ for line in (tmp_path / "out" / "results.jsonl").read_text().splitlines()
+ ]
+ assert len(rows) == 1, "the cell that completed before cancellation must survive"
+ assert rows[0]["review_weighted_f1"] == 0.42, "its measurement must survive intact"
+ assert (tmp_path / "out" / "report.md").is_file()
+
+
+def test_an_interrupted_sweep_emits_nothing_that_authorizes_promotion(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """The semantic condition, not the absence of a file.
+
+ promotion.json is still written for an aborted run - it is the record of why
+ nothing was promoted. What must hold is that nothing in it authorizes a
+ promotion from partial evidence.
+ """
+
+ cancel_event = threading.Event()
+ with pytest.raises(SystemExit):
+ _sweep(
+ tmp_path, monkeypatch, lambda _run: scored_review_row(),
+ runs=3, cancel_event=cancel_event,
+ # The candidate arm has to RUN, not merely appear in promotion
+ # metadata: _run_sweep builds cells only from args.arms, so naming it
+ # in candidate_arms alone left the candidate with no results at all -
+ # and then "insufficient_evidence" would hold because nothing ran,
+ # not because partial evidence is barred from promoting.
+ arms=["review", "candidate_review"],
+ candidate_arms=["candidate_review"],
+ after_cell=(
+ lambda run_idx, arm: cancel_event.set()
+ if run_idx == 0 and arm == "candidate_review"
+ else None
+ ),
+ )
+ rows = [
+ json.loads(line)
+ for line in (tmp_path / "out" / "results.jsonl").read_text().splitlines()
+ ]
+ assert any(r["arm"] == "candidate_review" for r in rows), (
+ "the candidate must have produced evidence, or insufficient_evidence "
+ "would hold merely because nothing ran"
+ )
+ promotion = json.loads((tmp_path / "out" / "promotion.json").read_text())
+ assert promotion["run_status"] == "aborted"
+ assert promotion["decisions"], "an aborted run still has to say what it decided"
+ for decision in promotion["decisions"]:
+ assert decision["decision"] == "insufficient_evidence"
+ assert any("partial evidence" in reason for reason in decision["reasons"])
diff --git a/eval/tests/test_workflow_bench.py b/eval/tests/test_workflow_bench.py
index d9c9c0f92..a53b2b8cd 100644
--- a/eval/tests/test_workflow_bench.py
+++ b/eval/tests/test_workflow_bench.py
@@ -14,18 +14,26 @@ import yaml
from typing import Any
from workflow_bench import runner
+from workflow_bench.evolution import CANDIDATE_ARMS
from workflow_bench.process_control import _CANCELLATION, cancellation_scope
from workflow_bench.runner import (
aggregate,
+ GraphBuildEnv,
+ arm_health,
broken_incumbent_arms,
+ unhealthy_arms,
+ unmeasured_arms,
build_parser,
infra_error_record,
+ next_graph_prefetch_target,
normalized_model_identifier,
parse_shortstat,
+ prefetch_next_graph,
render_report,
savings,
select_tasks,
systemic_outage_streak,
+ task_has_planned_paid_cells,
)
@@ -68,6 +76,15 @@ def test_aggregate_takes_medians_and_counts_resolved():
"diff_deletions": 5,
"class": "demo",
"resolved": 2,
+ # None of these are reused, so every resolution was measured this sweep.
+ "resolved_fresh": 2,
+ # Health is counted separately from resolution: all three executed and
+ # produced usable evidence, including the one that resolved nothing.
+ "fresh_attempts": 3,
+ "admissible": 3,
+ "execution_failures": 0,
+ "evidence_failures": 0,
+ "health_reasons": [],
"runs": 3,
"valid_runs": 3,
"excluded_runs": 0,
@@ -250,6 +267,9 @@ def test_shipped_scenarios_opt_out_the_cross_module_cell_and_rebuild_graph_asset
assert skipped == ["cross-module-parse-retry"]
assert all(not task.get("sandbox_copy") for task in tasks)
assert all(task["sandbox_dependencies"] for task in tasks)
+ assert all(
+ any(dep.get("source") == "gitnexus-shared/dist" for dep in task["sandbox_dependencies"]) for task in tasks
+ )
assert all(task["oracle"]["command"] and task["oracle"]["files"] for task in tasks)
assert all("./node_modules/.bin/vitest run" in task["oracle"]["command"] for task in tasks)
assert all("npx vitest" not in task["oracle"]["command"] for task in tasks)
@@ -519,6 +539,369 @@ def test_run_evolution_script_is_the_shared_ci_and_local_entrypoint():
assert printed.stderr # rewrite notice goes to stderr
+def test_planned_paid_cells_treat_missing_reuse_as_paid():
+ task = {"id": "review-pr-2718-defect"}
+ assert task_has_planned_paid_cells(
+ task,
+ arms=["ce_review", "review", "candidate_review"],
+ runs=3,
+ reusable_rows={},
+ reuse_source=None,
+ )
+ reuse_source = Path("/tmp/seed")
+ rows = {
+ (task["id"], arm, run_idx): {}
+ for run_idx in range(3)
+ for arm in ("ce_review", "review", "candidate_review")
+ }
+ assert not task_has_planned_paid_cells(
+ task,
+ arms=["ce_review", "review", "candidate_review"],
+ runs=3,
+ reusable_rows=rows,
+ reuse_source=reuse_source,
+ )
+ del rows[(task["id"], "candidate_review", 0)]
+ assert task_has_planned_paid_cells(
+ task,
+ arms=["ce_review", "review", "candidate_review"],
+ runs=3,
+ reusable_rows=rows,
+ reuse_source=reuse_source,
+ )
+
+
+def test_next_graph_prefetch_skips_ready_shas_and_fully_reused_tasks(tmp_path: Path):
+ first = {"id": "review-a"}
+ second = {"id": "review-b"}
+ third = {"id": "review-c"}
+ reuse_source = tmp_path / "seed"
+ reused_second = {
+ (second["id"], arm, 0): {} for arm in ("ce_review", "review", "candidate_review")
+ }
+ target = next_graph_prefetch_target(
+ [
+ (first, {"repo_identity": "/repo", "resolved_sha": "aaa"}),
+ (second, {"repo_identity": "/repo", "resolved_sha": "bbb"}),
+ (third, {"repo_identity": "/repo", "resolved_sha": "ccc"}),
+ ],
+ arms=["ce_review", "review", "candidate_review"],
+ runs=1,
+ reusable_rows=reused_second,
+ reuse_source=reuse_source,
+ ready_keys={("/repo", "aaa")},
+ )
+ assert target is not None
+ task, binding, key = target
+ assert task["id"] == "review-c"
+ assert key == ("/repo", "ccc")
+ assert binding["resolved_sha"] == "ccc"
+
+
+def test_prefetch_next_graph_runs_ensure_on_a_background_thread(monkeypatch):
+ started = threading.Event()
+ seen: list[tuple[str, str]] = []
+
+ def fake_ensure(**kwargs):
+ seen.append(kwargs["graph_key"])
+ started.set()
+
+ monkeypatch.setattr("workflow_bench.runner.ensure_task_graph", fake_ensure)
+ cancel = threading.Event()
+ job = prefetch_next_graph(
+ task={"id": "review-b"},
+ binding={"repo_identity": "/repo", "resolved_sha": "bbb"},
+ graph_key=("/repo", "bbb"),
+ env=GraphBuildEnv(
+ trees=Path("/tmp"),
+ task_asset_cache=None,
+ claude_bin="claude",
+ bwrap_bin="bwrap",
+ sandbox_backend="bwrap",
+ runtime_mounts=(),
+ clone_templates={},
+ clone_template_errors={},
+ graph_snapshots={},
+ graph_snapshot_errors={},
+ ),
+ cancel_event=cancel,
+ )
+ assert job.key == ("/repo", "bbb")
+ assert started.wait(timeout=2)
+ job.join()
+ assert seen == [("/repo", "bbb")]
+
+
+def test_a_reused_resolution_does_not_count_as_this_sweeps_health():
+ """resolved counts evidence; resolved_fresh counts evidence measured today.
+
+ broken_incumbent_arms reads resolved_fresh because a reused row proves last
+ generation's environment worked. Counting it would make an arm whose cells
+ were all reused look healthy in exactly the run where a broken environment
+ should have been caught.
+ """
+
+ reused = [record(resolved=True, reused=True), record(resolved=True, reused=True)]
+ agg = aggregate(reused)
+ assert agg["resolved"] == 2
+ assert agg["resolved_fresh"] == 0
+ assert broken_incumbent_arms({"t": {"review": agg}}, {"review"}) == ["review"]
+
+ mixed = aggregate([record(resolved=True, reused=True), record(resolved=True)])
+ assert mixed["resolved_fresh"] == 1
+ assert broken_incumbent_arms({"t": {"review": mixed}}, {"review"}) == []
+
+
+def test_graph_build_env_ready_keys_covers_successes_and_failures():
+ """A key that failed is attempted, not pending.
+
+ next_graph_prefetch_target skips keys already in ready_keys. If a failed
+ build were omitted, the sweep would prefetch it again every iteration and
+ pay a full clone and offline index each time for a build that cannot
+ succeed.
+ """
+
+ env = GraphBuildEnv(
+ trees=Path("/tmp"),
+ task_asset_cache=None,
+ claude_bin="claude",
+ bwrap_bin="bwrap",
+ sandbox_backend="bwrap",
+ runtime_mounts=(),
+ clone_templates={("/repo", "aaa"): (Path("/tmp/a"), "aaa")},
+ clone_template_errors={("/repo", "bbb"): OSError("clone failed")},
+ graph_snapshots={("/repo", "ccc"): object()},
+ graph_snapshot_errors={("/repo", "ddd"): OSError("index failed")},
+ )
+ assert env.ready_keys() == {
+ ("/repo", "aaa"),
+ ("/repo", "bbb"),
+ ("/repo", "ccc"),
+ ("/repo", "ddd"),
+ }
+
+
+def _cell(**overrides) -> dict[str, Any]:
+ """One results.jsonl row, healthy unless told otherwise."""
+
+ base = record(resolved=True)
+ base.update({"error_kind": None, "review_evidence_valid": True, "transcript_missing": False})
+ base.update(overrides)
+ return base
+
+
+def _arms(**by_arm) -> dict[str, dict[str, dict[str, Any]]]:
+ return {"task0": {arm: aggregate(rows) for arm, rows in by_arm.items()}}
+
+
+def test_a_reviewer_that_scores_badly_is_not_an_unhealthy_harness():
+ """Reconstructed from Actions run 33962002890's logged observations.
+
+ Every completed cell was resolved=False with error_kind=oracle-failed, at a
+ median score of 0.212 — the reviews ran, wrote artifacts and were scored.
+ That is a valid negative for the quality gate to judge. Diagnosing it as a
+ broken environment is the confusion this classification exists to end.
+ """
+
+ scored_but_wrong = [_cell(resolved=False, error_kind="oracle-failed") for _ in range(3)]
+ results = _arms(review=scored_but_wrong, ce_review=list(scored_but_wrong))
+ assert unhealthy_arms(results, {"review", "ce_review"}) == []
+ health = arm_health(results, {"review"})["review"]
+ assert health.admissible == 3 and health.fresh_attempts == 3
+ assert (health.execution_failures, health.evidence_failures) == (0, 0)
+
+
+def test_an_all_zero_score_is_still_a_valid_negative():
+ zeroed = [_cell(resolved=False, error_kind="oracle-failed", review_weighted_f1=0.0) for _ in range(3)]
+ assert unhealthy_arms(_arms(review=zeroed), {"review"}) == []
+
+
+def test_artifacts_that_were_never_written_are_an_unhealthy_harness():
+ """Reconstructed from Actions run 33912693948.
+
+ All 41 artifacts came back 0 bytes because the mount made an atomic write
+ impossible. The reviews could not produce evidence at all — the opposite of
+ the case above, and the one a health check must catch. The old caller
+ excluded review arms entirely, so it could not have.
+ """
+
+ unwritable = [_cell(resolved=False, ok=False, error_kind="review-evidence-invalid") for _ in range(3)]
+ flagged = unhealthy_arms(_arms(review=unwritable), {"review"})
+ assert [h.arm for h in flagged] == ["review"]
+ assert flagged[0].evidence_failures == 3
+ assert "review-evidence-invalid" in flagged[0].reasons
+
+
+def test_one_admissible_cell_leaves_an_arm_degraded_not_healthy():
+ """Mixed outcomes are DEGRADED. One usable measurement does not erase two failures.
+
+ Not fatal - the sweep still produced evidence - but calling it healthy is
+ how a partly-broken environment passes review.
+ """
+
+ mixed = [
+ _cell(resolved=False, error_kind="oracle-failed"),
+ _cell(resolved=False, ok=False, error_kind="session-error"),
+ _cell(resolved=False, ok=False, error_kind="infra-error"),
+ ]
+ results = _arms(review=mixed)
+ health = arm_health(results, {"review"})["review"]
+ assert health.status == "DEGRADED"
+ assert unhealthy_arms(results, {"review"}) == [], "degraded is diagnostic, not fatal"
+ assert health.execution_failures == 2, "failures must stay visible, not be erased"
+ assert health.admissible == 1
+
+
+def test_a_row_that_fails_both_ways_is_only_subtracted_once():
+ """run_arm can produce a row that is an execution AND an evidence failure.
+
+ It keeps the first error_kind — a session-error survives — and still sets
+ review_evidence_valid=False when the artifact will not parse. Counting that
+ row against admissible twice zeroed an arm that held a real measurement,
+ which arm_health reports as UNUSABLE and the measurement gate then fails on.
+ """
+
+ both = _cell(resolved=False, ok=False, error_kind="session-error", review_evidence_valid=False)
+ results = _arms(review=[both, _cell(resolved=True, error_kind="oracle-failed")])
+ health = arm_health(results, {"review"})["review"]
+ assert (health.execution_failures, health.evidence_failures) == (1, 1)
+ assert health.fresh_attempts == 2
+ assert health.admissible == 1
+ assert health.status == "DEGRADED"
+ assert unhealthy_arms(results, {"review"}) == []
+
+
+def test_reused_rows_alone_leave_current_health_unknown():
+ """Historical success cannot certify this sweep's environment."""
+
+ reused = [_cell(reused=True) for _ in range(3)]
+ results = _arms(review=reused)
+ assert unmeasured_arms(results, {"review"}) == ["review"]
+ assert unhealthy_arms(results, {"review"}) == []
+ assert arm_health(results, {"review"})["review"].measured is False
+
+
+def test_the_paid_canary_survives_a_prior_run_with_more_run_indices():
+ """The canary counts planned cells, not every key reuse selection returned.
+
+ Reuse selection accepts any non-negative prior `run`, so a results directory
+ produced with --runs 5 leaves keys this sweep never plans. Comparing against
+ those made the "arm is fully reused" test false exactly when it was true,
+ and the incumbent went a whole sweep without one measured cell.
+ """
+
+ tasks = [{"id": "task0"}, {"id": "task1"}]
+ reusable = {(task["id"], "review", run): {} for task in tasks for run in range(5)}
+
+ dropped = runner.drop_canary_reuse_key(reusable, arm="review", tasks=tasks, runs=3)
+
+ assert dropped == ("task0", "review", 0)
+ assert dropped not in reusable
+ # A second call is a no-op: the arm now has its paid cell.
+ assert runner.drop_canary_reuse_key(reusable, arm="review", tasks=tasks, runs=3) is None
+
+
+def test_an_arm_with_a_planned_paid_cell_keeps_every_reusable_row():
+ tasks = [{"id": "task0"}]
+ reusable = {("task0", "review", 0): {}}
+
+ assert runner.drop_canary_reuse_key(reusable, arm="review", tasks=tasks, runs=2) is None
+ assert len(reusable) == 1
+
+
+def test_reused_successes_do_not_mask_fresh_execution_failures():
+ rows = [_cell(reused=True), _cell(reused=True), _cell(ok=False, error_kind="session-error")]
+ flagged = unhealthy_arms(_arms(review=rows), {"review"})
+ assert [h.arm for h in flagged] == ["review"]
+ assert flagged[0].fresh_attempts == 1 and flagged[0].execution_failures == 1
+
+
+def test_a_parseable_artifact_does_not_excuse_a_failed_session():
+ """Artifact parseability must not override an execution failure."""
+
+ rows = [_cell(ok=False, error_kind="session-error", review_evidence_valid=True) for _ in range(2)]
+ flagged = unhealthy_arms(_arms(review=rows), {"review"})
+ assert [h.arm for h in flagged] == ["review"]
+ assert flagged[0].execution_failures == 2
+
+
+def test_a_single_unusable_review_is_caught_below_the_breaker_threshold():
+ """The decisive regression for the finalization guard.
+
+ A fixture of 41 empty artifacts would abort through the outage breaker -
+ review-evidence-invalid is systemic and the limit is 5 - so it proves
+ nothing about this path. One fresh unusable cell is under that threshold,
+ which leaves the finalization check as the only thing that can catch it.
+ """
+
+ streak = 0
+ for _ in range(1):
+ streak = runner.systemic_outage_streak("review-evidence-invalid", streak)
+ assert streak < runner.DEFAULT_OUTAGE_STREAK, "fixture must not reach the breaker"
+
+ results = _arms(review=[_cell(resolved=False, ok=False, error_kind="review-evidence-invalid")])
+ with pytest.raises(SystemExit) as exc:
+ runner.enforce_measurement_health(results, {"review"})
+ assert exc.value.code == 1
+
+
+def test_finalization_reports_every_arm_and_names_no_cause(capsys):
+ """Status for each arm; an empty artifact does not become an EROFS diagnosis."""
+
+ results = _arms(
+ review=[_cell(resolved=False, ok=False, error_kind="review-evidence-invalid")],
+ ce_review=[_cell(resolved=False, error_kind="oracle-failed")],
+ )
+ with pytest.raises(SystemExit):
+ runner.enforce_measurement_health(results, {"review", "ce_review"})
+ out = capsys.readouterr().out
+ assert "review: UNUSABLE" in out
+ assert "ce_review: OBSERVED_OK" in out
+ assert "cause=undetermined" in out
+ assert "EROFS" not in out and "mount" not in out
+
+
+def test_valid_negatives_do_not_abort_finalization(capsys):
+ """The 16h run's shape must survive the real guard, not just the classifier."""
+
+ scored_but_wrong = [_cell(resolved=False, error_kind="oracle-failed") for _ in range(3)]
+ health = runner.enforce_measurement_health(
+ _arms(review=scored_but_wrong, ce_review=list(scored_but_wrong)), {"review", "ce_review"}
+ )
+ assert {h.status for h in health.values()} == {"OBSERVED_OK"}
+ assert "UNUSABLE" not in capsys.readouterr().out
+
+
+def test_reused_only_arm_is_reported_unknown_by_finalization(capsys):
+ runner.enforce_measurement_health(_arms(review=[_cell(reused=True)]), {"review"})
+ assert "review: UNKNOWN" in capsys.readouterr().out
+
+
+def test_run_sweep_calls_the_health_guard_and_not_the_legacy_helper():
+ """Pins the wiring the caller correction exposed.
+
+ Reads the compiled code object's global references rather than the source
+ text: deleting the call removes the name and fails this test, which is the
+ mutation check. It does NOT prove the guard runs end to end - _run_sweep
+ needs bwrap and a sandbox, so no test here drives it.
+ """
+
+ referenced = runner._run_sweep.__code__.co_names
+ assert "enforce_measurement_health" in referenced
+ assert "broken_incumbent_arms" not in referenced
+
+
+def test_ce_review_is_classified_even_though_it_is_not_a_candidate_arm():
+ """ce_review is a comparator, absent from CANDIDATE_ARMS.
+
+ Dropping the `- {"review"}` exclusion alone would have left it unchecked.
+ """
+
+ assert "ce_review" not in set(CANDIDATE_ARMS.values())
+ health = arm_health(_arms(ce_review=[_cell()]), {"review", "ce_review"})
+ assert "ce_review" in health
+
+
def _packed_cells(tasks: int, runs: int, arms: tuple[str, ...]) -> list[tuple[str, int, str]]:
return [(f"t{t}", r, a) for t in range(tasks) for r in range(runs) for a in arms]
diff --git a/eval/tests/test_workflow_bench_sessions.py b/eval/tests/test_workflow_bench_sessions.py
index 1f66f071f..5d1692fbb 100644
--- a/eval/tests/test_workflow_bench_sessions.py
+++ b/eval/tests/test_workflow_bench_sessions.py
@@ -12,9 +12,9 @@ from types import SimpleNamespace
import pytest
-from workflow_bench import evolve, runner, runner_sessions, runtime_mounts
+from workflow_bench import evolve, runner, runner_artifacts, runner_sessions, runtime_mounts
from workflow_bench.evolution import skill_fingerprint
-from workflow_bench.process_control import ManagedProcessResult
+from workflow_bench.process_control import ManagedProcessError, ManagedProcessResult
from workflow_bench.proposer_sandbox import SandboxError
from workflow_bench.runner import snapshot_plan_docs
@@ -120,11 +120,16 @@ def skill_events(skill_input: dict, *, tool_id: str = "skill-1", is_error: bool
def fake_sandbox(root: Path) -> SimpleNamespace:
+ # private_root is NOT the clone. Conflating them puts the review artifact
+ # directory inside the workspace, which the real sandbox never does and
+ # which hides whether the workspace was left untouched.
+ private_root = root.parent / f"{root.name}-sandbox-private"
+ private_root.mkdir(exist_ok=True)
return SimpleNamespace(
backend="test-double",
claude_bin="claude",
clone=root,
- private_root=root,
+ private_root=private_root,
command_prefix=[],
command_prefix_for=lambda **_kwargs: [],
settings_json="{}",
@@ -1249,7 +1254,7 @@ def test_planning_cannot_change_source_tests_or_downstream_skill(monkeypatch, tm
@pytest.mark.parametrize(
("attack", "expected_detail"),
[
- ("workspace", "unauthorized workspace path"),
+ ("workspace", "changed the read-only workspace"),
("skill", "changed the evaluated skill fingerprint"),
],
)
@@ -1265,9 +1270,11 @@ def test_review_phase_rejects_workspace_or_skill_mutation(
expected_skill_digest = "expected-skill-fingerprint"
def adversarial_review(prompt, *args, **kwargs):
- (tmp_path / "review-output.json").write_text(
- '{"schema_version":1,"verdict":"approve","findings":[]}'
- )
+ # Write where the contract now says: the artifact directory outside the
+ # workspace, which is the only place the agent can write atomically.
+ artifact = runner.review_output_path(sandbox, runner.REVIEW_OUTPUT)
+ artifact.parent.mkdir(parents=True, exist_ok=True)
+ artifact.write_text('{"schema_version":1,"verdict":"approve","findings":[]}')
if attack == "workspace":
source.write_text("review silently changed source")
return session_record()
@@ -1298,6 +1305,97 @@ def test_review_phase_rejects_workspace_or_skill_mutation(
assert expected_detail in rec["error_detail"]
+def test_a_cancelled_clone_copy_does_not_fall_back_to_an_uncancellable_copytree(monkeypatch, tmp_path):
+ """The reflink fallback is for a filesystem, not for a teardown.
+
+ run_managed reports cancellation as a non-OK result rather than raising, so
+ the fallback treated it like an unsupported reflink and started a copytree
+ that cannot be cancelled — waiting out exactly the full copy the outage
+ breaker set the cancellation event to avoid.
+ """
+
+ source = tmp_path / "template"
+ (source / ".git").mkdir(parents=True)
+ parent = tmp_path / "clones"
+ parent.mkdir()
+ copied: list[object] = []
+
+ monkeypatch.setattr(
+ runner_artifacts,
+ "run_managed",
+ lambda *_a, **_k: ManagedProcessResult(
+ state="cancelled",
+ returncode=None,
+ stdout_tail="",
+ stderr_tail="",
+ duration_s=0.1,
+ ),
+ )
+ monkeypatch.setattr(runner_artifacts.shutil, "copytree", lambda *a, **k: copied.append(a))
+
+ with pytest.raises(ManagedProcessError):
+ runner.copy_isolated_tree(source, parent)
+ assert copied == []
+ assert list(parent.iterdir()) == [], "the partial target must be cleaned up"
+
+
+@pytest.mark.parametrize("arm", ["review", "ce_review"])
+def test_run_arm_mounts_the_review_artifact_directory_outside_the_workspace(monkeypatch, tmp_path, arm):
+ """A writable FILE inside a read-only directory is not a writable path.
+
+ The Write tool creates `.tmp..` beside the target and
+ renames it, so a read-only parent fails the temp create with EROFS and the
+ artifact stays 0 bytes. The mount target must be the directory, and it must
+ sit outside the read-only workspace.
+
+ Driven through run_arm rather than rebuilt here: an expected tuple assembled
+ in the test passes whatever run_arm actually mounts, which is the one thing
+ this needs to prove.
+ """
+
+ assert not runner.SANDBOX_REVIEW_OUTPUT.startswith(runner.SANDBOX_WORKSPACE + "/")
+ assert runner.SANDBOX_REVIEW_OUTPUT != runner.SANDBOX_WORKSPACE
+
+ verify_calls: list[dict] = []
+ sandbox = fake_sandbox(tmp_path)
+ sandbox.command_prefix_for = lambda **kwargs: verify_calls.append(kwargs) or []
+
+ def review_session(prompt, *args, **kwargs):
+ artifact = runner.review_output_path(sandbox, runner.REVIEW_OUTPUT)
+ artifact.parent.mkdir(parents=True, exist_ok=True)
+ artifact.write_text('{"schema_version":1,"verdict":"approve","findings":[]}')
+ return session_record()
+
+ monkeypatch.setattr(runner, "run_claude", review_session)
+ monkeypatch.setattr(runner, "skill_fingerprint", lambda *_a, **_k: "skill-digest")
+ monkeypatch.setattr(runner, "run_verify", lambda *a, **k: (True, "ok"))
+
+ runner.run_arm(
+ arm,
+ {"prompt": "p", "verify": "true"},
+ tmp_path,
+ bench_args(),
+ sandbox=sandbox,
+ expected_skill_digest="skill-digest",
+ )
+
+ review_output = runner.review_output_path(sandbox, runner.REVIEW_OUTPUT)
+ expected = (runner.ReadOnlyMount(source=review_output.parent, target=runner.SANDBOX_REVIEW_OUTPUT),)
+ # The EROFS bug is about the AGENT's write, so the mount that has to be the
+ # directory is the writable one on the review session — not the read-only
+ # exposure the verify command gets afterwards. Assert both: they are
+ # separate arguments to separate command prefixes.
+ writable = [call["extra_writable_mounts"] for call in verify_calls if "extra_writable_mounts" in call]
+ assert writable, "the review session must be given a writable artifact mount"
+ assert writable[-1] == expected, "mount the directory, not the file"
+ read_only = [call["extra_read_only_mounts"] for call in verify_calls if "extra_read_only_mounts" in call]
+ assert read_only, "the verify invocation must be given the artifact mount"
+ assert read_only[-1] == expected, "mount the directory, not the file"
+ assert not expected[0].target.startswith(f"{runner.SANDBOX_WORKSPACE}/")
+ # The artifact the harness later reads is the one inside that mount.
+ assert review_output.parent in review_output.parents
+
+
def _git(repo, *args, check=True):
return subprocess.run(["git", "-C", str(repo), *args], check=check, capture_output=True, text=True)
@@ -1348,3 +1446,34 @@ def test_make_worktree_clone_has_no_tags_but_keeps_all_branches(tmp_path):
current = _git(target, "rev-parse", "HEAD").stdout.strip()
assert current == other_sha
+
+
+def test_copy_isolated_tree_does_not_share_git_objects_or_refs(tmp_path):
+ repo = tmp_path / "repo"
+ repo.mkdir()
+ _git(repo, "init", "--quiet")
+ _git(repo, "checkout", "--quiet", "-b", "main")
+ sha = _git_commit(repo, "base")
+ clones = tmp_path / "clones"
+ clones.mkdir()
+ template = runner.make_worktree(repo, sha, clones)
+ (template / "marker.txt").write_text("template\n")
+
+ copy = runner.copy_isolated_tree(template, clones)
+ assert copy != template
+ assert (copy / "marker.txt").read_text() == "template\n"
+ (copy / "marker.txt").write_text("copy\n")
+ assert (template / "marker.txt").read_text() == "template\n"
+ copy_head = _git(copy, "rev-parse", "HEAD").stdout.strip()
+ template_head = _git(template, "rev-parse", "HEAD").stdout.strip()
+ assert copy_head == template_head == sha
+ # An equal initial HEAD is also what a shared ref namespace looks like, so
+ # write a ref and prove the template cannot see it. A linked worktree would
+ # pass every assertion above, including the alternates check — its `.git` is
+ # a file, so the directory inspected below simply does not exist.
+ _git(copy, "branch", "copy-only")
+ assert _git(copy, "show-ref", "--verify", "refs/heads/copy-only").returncode == 0
+ assert _git(template, "show-ref", "--verify", "refs/heads/copy-only", check=False).returncode != 0
+ assert (copy / ".git").is_dir()
+ alternates = copy / ".git" / "objects" / "info" / "alternates"
+ assert not alternates.exists()
diff --git a/eval/workflow_bench/README.md b/eval/workflow_bench/README.md
index 79bd4bfea..a0008fa7d 100644
--- a/eval/workflow_bench/README.md
+++ b/eval/workflow_bench/README.md
@@ -149,8 +149,10 @@ UNSAFE_NO_BWRAP=1 RUNS=1 ./workflow_bench/run-evolution.sh
This mode runs review sessions directly in disposable host worktrees and is
**not** a security boundary: it does not isolate the network or create a PID
namespace, and a session that can `chmod` can undo the workspace lock. The
-harness still drops write bits on the clone except `review-output.json` so
-accidental `npm install` / analyze writes cannot invalidate review evidence.
+harness drops write bits on the whole clone, with no carve-out, so accidental
+`npm install` / analyze writes cannot invalidate review evidence. The review
+artifact is not in the clone at all: it lives in a writable directory bound at
+`/review-output`, outside the workspace.
Sandbox cleanup restores owner write bits before deleting the private TMPDIR,
because a session that `copytree`s the locked clone would otherwise leave
non-empty 0555 directories that `rmtree` cannot remove. Historical review
@@ -168,6 +170,36 @@ router thresholds as an incumbent policy, not permanent truth. Candidate
changes run offline in the same throwaway clones as the incumbent; production
skills never rewrite themselves from a live task.
+On the self-hosted evolution box, `run-evolution.sh` passes
+`--max-runtime-from-instance-window` and the CLI derives its own cap from
+`/proc/uptime` at startup (24h EventBridge window minus a 90-minute upload
+reserve), in the same breath as it starts the clock that cap is measured
+against — a budget computed anywhere earlier is spent by the seconds between. A `workflow_dispatch` that lands on an
+already-running instance therefore exits in-process instead of vanishing when
+the box stops — a cancelled GitHub job skips even `if: always()`, which is
+how run 33962002890 lost 51 finished sessions. Local runs are uncapped.
+
+A review generation is 6 tasks × 3 arms × 3 runs. Serial workers=1 at ~19
+minutes per session is a 16-hour job (run 33962002890). Two harness changes
+cut that without shrinking the gate:
+
+- **Comparator reuse.** `evolve.py` forwards the seed / prior generation as
+ `--reuse-results`. Incumbent `review` and `ce_review` rows are copied into
+ the new `results.jsonl` when model, effort, task SHA, prompt digest, oracle
+ bytes, incumbent skill digest, CE plugin digest, and sandbox backend still
+ match. Candidate arms always run. A weekly generation with an unchanged
+ incumbent therefore pays 18 sessions, not 54. A promotion, model change,
+ task-corpus change, or harness `RUNTIME_DIGEST` change invalidates the
+ lock and re-runs the comparators.
+- **Sanitized clone templates.** Each unique task SHA is cloned and
+ sanitized once. Cells copy that parentless snapshot (reflink when the
+ filesystem allows) instead of `git clone --no-local` plus repack/prune/fsck
+ 54 times. Isolation is a private `.git`, not a second copy of full history.
+
+Dispatch defaults to `--workers 3` so those 18 paid cells can overlap. Size
+workers to the host: a cell that loses CPU and hits the session ceiling is
+an excluded run the gate refuses.
+
Build an overlay that mirrors only the canonical repo-local skill paths:
```text
@@ -276,9 +308,12 @@ without weakening today's deterministic promotion boundary.
The evolution workflow runs an offline containment preflight with the pinned
Claude Code 2.1.214 binary before starting a paid proposer or benchmark. The
-review canary seals the workspace read-only and exposes only the pre-created
-`review-output.json` as writable. Runtime mount placeholders are prepared in
-the disposable clone before sealing it; existing config bytes are preserved.
+review canary seals the workspace read-only and writes nothing into it: the
+artifact directory is bound at `/review-output` outside the workspace, and the
+file itself is deliberately absent until the session creates it, so its absence
+distinguishes "never written" from "written badly". Runtime mount placeholders
+are prepared in the disposable clone before sealing it; existing config bytes
+are preserved.
Any pre-existing result entry, including a symlink, is rejected. Required
canaries fail when their runtime or Bubblewrap is unavailable.
@@ -368,10 +403,10 @@ paired benchmark as any other candidate.
For ad-hoc use, run the driver on the existing re-evaluation triggers
(model/harness change or 90-day staleness). The repository workflow runs a
-deliberate weekly drift check: scheduled concurrency stays serial unless
-`GITNEXUS_EVOLUTION_WORKERS` is raised after a funded host-sized proof, and
-`--workers` is bounded to 1–8 before paid work starts. `--generations` remains
-the only loop bound.
+deliberate weekly drift check: dispatch defaults to three concurrent cells
+of one task; scheduled concurrency still requires
+`GITNEXUS_EVOLUTION_WORKERS=3` after a clean proof. `--workers` is bounded
+to 1–8 before paid work starts. `--generations` remains the only loop bound.
## Free-model setup (no paid tokens)
diff --git a/eval/workflow_bench/comparator_reuse.py b/eval/workflow_bench/comparator_reuse.py
new file mode 100644
index 000000000..5dcef1289
--- /dev/null
+++ b/eval/workflow_bench/comparator_reuse.py
@@ -0,0 +1,604 @@
+"""Reuse frozen comparator cells when the current sweep is still the same experiment.
+
+Weekly skill evolution re-runs incumbent ``review`` / ``ce_review`` (and the
+implementation incumbents) even when the model, effort, tasks, oracles,
+incumbent skill bytes, and CE plugin have not changed. Those arms are the
+baseline the gate compares a *new* candidate against — they are not the
+thing being evolved. Replaying them burns two-thirds of a generation.
+
+This module selects prior ``results.jsonl`` rows that are safe to carry
+forward. Candidate arms are never reused. A mismatch on any bound field
+falls through to a paid cell. Missing artifacts also fall through: a reused
+row that the proposer cannot read is worse than spending the tokens again.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import os
+import re
+import stat
+from collections.abc import Iterator, Mapping, Sequence
+from contextlib import contextmanager
+from dataclasses import dataclass
+from datetime import UTC, datetime, timedelta
+from pathlib import Path, PurePosixPath
+from typing import Any
+
+from .evolution import CANDIDATE_ARMS, EVIDENCE_MAX_AGE_DAYS
+from .proposer_sandbox import SandboxError
+from .runner_sessions import MAX_TRANSCRIPT_BYTES, PARENT_EVENT_STREAM_SOURCE
+from .runtime_mounts import CE_ARMS
+from .task_assets import COPY_CHUNK_BYTES, _write_all
+
+REUSABLE_COMPARATOR_ARMS = frozenset(
+ {
+ "review",
+ "ce_review",
+ "workflow",
+ "workflow_direct",
+ "ce_workflow",
+ "ce_workflow_direct",
+ "baseline",
+ "baseline_nomcp",
+ }
+)
+# Must stay aligned with runner.EXCLUDED_ERROR_KINDS plus review-invalid.
+# A reused row becomes promotion evidence; excluded kinds cannot enter that set.
+REUSE_EXCLUDED_ERROR_KINDS = frozenset(
+ {
+ "session-error",
+ "infra-error",
+ "evidence-unverified",
+ "cleanup-failure",
+ "review-evidence-invalid",
+ "cancelled",
+ }
+)
+_TRANSCRIPT_NAME = re.compile(r"[A-Za-z0-9._-]{1,200}")
+CellKey = tuple[str, str, int]
+
+
+@dataclass(frozen=True)
+class TaskReuseBinding:
+ """Per-task identity the prior row must still match."""
+
+ task_base_sha: str
+ task_prompt_digest: str
+ oracle_digest: str
+ oracle_command_digest: str
+ oracle_manifest_digest: str
+ # The cell's environment is part of its identity: a comparator measured
+ # against different task assets or different sandbox dependencies is a
+ # measurement of a different machine, not a baseline for this sweep.
+ task_asset_manifest_digest: str | None = None
+ sandbox_dependency_manifest_digest: str | None = None
+
+
+@dataclass(frozen=True)
+class ComparatorReuseExpectation:
+ """Sweep-wide lock for comparator reuse. Any drift pays for a fresh cell."""
+
+ model: str
+ effort: str
+ sandbox_backend: str
+ runtime_digest: str | None
+ now: datetime
+ max_age: timedelta
+ tasks: Mapping[str, TaskReuseBinding]
+ skill_digests: Mapping[str, str | None]
+ ce_plugin_version: str | None
+ ce_plugin_manifest_digest: str | None
+
+
+def load_result_rows(path: Path) -> list[dict[str, Any]]:
+ """Load ``results.jsonl``; skip malformed lines the same way evolve does."""
+
+ rows: list[dict[str, Any]] = []
+ for line in path.read_text().splitlines():
+ if not line.strip():
+ continue
+ try:
+ row = json.loads(line)
+ except json.JSONDecodeError:
+ continue
+ if isinstance(row, dict):
+ rows.append(row)
+ return rows
+
+
+def current_runtime_digest() -> str | None:
+ """Harness lockfile digest exported by ``run-evolution.sh``, if present."""
+
+ value = os.environ.get("RUNTIME_DIGEST", "").strip()
+ return value or None
+
+
+def row_is_reusable_comparator(row: Mapping[str, Any], expected: ComparatorReuseExpectation) -> bool:
+ """True when ``row`` is a complete, still-valid comparator measurement."""
+
+ arm = row.get("arm")
+ if not isinstance(arm, str) or arm in CANDIDATE_ARMS or arm not in REUSABLE_COMPARATOR_ARMS:
+ return False
+ if row.get("error_kind") in REUSE_EXCLUDED_ERROR_KINDS:
+ return False
+ if row.get("error_kind") not in (None, ""):
+ return False
+ if row.get("ok") is not True:
+ return False
+ if row.get("transcript_missing") is True:
+ return False
+ if row.get("candidate_overlay_digest") not in (None, ""):
+ return False
+ # Age against the ORIGINAL measurement, not the copy time: materialize_reused_row
+ # restamps recorded_at, so a chained row would otherwise refresh its own clock
+ # and never expire. Bound both directions - a future stamp is corrupt, not fresh.
+ recorded = _parse_recorded_at(row.get("reused_from_recorded_at") or row.get("recorded_at"))
+ if recorded is None:
+ return False
+ age = expected.now - recorded
+ if age > expected.max_age or age < timedelta(0):
+ return False
+ if row.get("model") != expected.model and row.get("benchmark_model") != expected.model:
+ return False
+ if row.get("effort") != expected.effort:
+ return False
+ if row.get("sandbox_backend") != expected.sandbox_backend:
+ return False
+ # Fail closed. A row with no runtime_digest was measured by a harness that
+ # did not record one, which is exactly the drift this lock exists to catch;
+ # treating the absence as agreement made every legacy row reusable forever.
+ prior_runtime = row.get("runtime_digest")
+ if not isinstance(prior_runtime, str) or not prior_runtime:
+ return False
+ if not expected.runtime_digest or prior_runtime != expected.runtime_digest:
+ return False
+
+ task_id = row.get("task")
+ binding = expected.tasks.get(task_id) if isinstance(task_id, str) else None
+ if binding is None:
+ return False
+ if row.get("task_base_sha") != binding.task_base_sha:
+ return False
+ if row.get("task_prompt_digest") != binding.task_prompt_digest:
+ return False
+ if row.get("oracle_digest") != binding.oracle_digest:
+ return False
+ if row.get("oracle_command_digest") != binding.oracle_command_digest:
+ return False
+ if row.get("oracle_manifest_digest") != binding.oracle_manifest_digest:
+ return False
+ # Fail closed on both sides, as the runtime digest does: an unbound
+ # expectation means this sweep could not determine its own environment, and
+ # a row without the field was measured before it was recorded.
+ for field, bound in (
+ ("task_asset_manifest_digest", binding.task_asset_manifest_digest),
+ ("sandbox_dependency_manifest_digest", binding.sandbox_dependency_manifest_digest),
+ ):
+ prior = row.get(field)
+ if not isinstance(prior, str) or not prior or not bound or prior != bound:
+ return False
+
+ if arm in CE_ARMS:
+ if row.get("ce_plugin_version") != expected.ce_plugin_version:
+ return False
+ if row.get("ce_plugin_manifest_digest") != expected.ce_plugin_manifest_digest:
+ return False
+ else:
+ expected_skill = expected.skill_digests.get(arm)
+ if not expected_skill or row.get("skill_digest") != expected_skill:
+ return False
+
+ if arm in {"review", "ce_review"}:
+ if row.get("review_evidence_valid") is not True:
+ return False
+ # The artifact, not just the score derived from it. materialize_reused_row
+ # copies it only when the name is present, so without this a row whose
+ # artifact copy never happened could be carried forward as a scored
+ # review that a proposer then cannot read - evidence by assertion.
+ review_artifact = row.get("review_artifact")
+ if not isinstance(review_artifact, str) or not review_artifact:
+ return False
+ if not isinstance(row.get("review_score"), dict):
+ return False
+ if row.get("review_weighted_f1") is None:
+ return False
+
+ artifacts = row.get("transcript_artifacts")
+ if not isinstance(artifacts, list) or not artifacts:
+ return False
+ try:
+ for artifact in artifacts:
+ _transcript_metadata(artifact)
+ except SandboxError:
+ return False
+ return True
+
+
+def select_reusable_comparator_rows(
+ rows: Sequence[Mapping[str, Any]],
+ *,
+ expected: ComparatorReuseExpectation,
+) -> dict[CellKey, dict[str, Any]]:
+ """Index reusable rows by ``(task, arm, run)``. Conflicting duplicates drop the key."""
+
+ chosen: dict[CellKey, dict[str, Any]] = {}
+ blocked: set[CellKey] = set()
+ for row in rows:
+ if not row_is_reusable_comparator(row, expected):
+ continue
+ task_id = row["task"]
+ arm = row["arm"]
+ run = row.get("run")
+ if not isinstance(run, int) or isinstance(run, bool) or run < 0:
+ continue
+ key = (str(task_id), str(arm), run)
+ if key in blocked:
+ continue
+ previous = chosen.get(key)
+ if previous is None:
+ chosen[key] = dict(row)
+ continue
+ if _row_identity(previous) != _row_identity(row):
+ blocked.add(key)
+ chosen.pop(key, None)
+ return chosen
+
+
+def materialize_reused_row(
+ row: Mapping[str, Any],
+ *,
+ source_dir: Path,
+ dest_dir: Path,
+) -> dict[str, Any]:
+ """Copy digest-bound artifacts into this sweep's evidence dir and stamp reuse."""
+
+ source, _ = _resolved_directory(source_dir, label="reuse source")
+ dest, _ = _resolved_directory(dest_dir, label="reuse destination")
+ if source == dest:
+ raise SandboxError("comparator reuse cannot read and write the same results directory")
+
+ materialized = dict(row)
+ materialized["reused"] = True
+ # Keep the FIRST measurement time across a chain. Overwriting it with the
+ # previous copy's stamp let a row refresh its own clock every generation and
+ # outlive the max_age bound entirely.
+ materialized["reused_from_recorded_at"] = row.get("reused_from_recorded_at") or row.get("recorded_at")
+ materialized["recorded_at"] = datetime.now(UTC).isoformat()
+
+ artifacts = row.get("transcript_artifacts")
+ if not isinstance(artifacts, list) or not artifacts:
+ raise SandboxError("reused row is missing transcript_artifacts")
+
+ # Every path below is resolved against a held descriptor, never re-walked
+ # from a name. Both roots are already symlink-free (_resolved_directory
+ # resolved them), and pinning them here means the components under them
+ # cannot be swapped out from under a check that already passed.
+ with (
+ _open_pinned_root(source_dir, label="reuse source") as source_fd,
+ _open_pinned_root(dest_dir, label="reuse destination") as dest_fd,
+ ):
+ copied_artifacts: list[dict[str, Any]] = []
+ for artifact in artifacts:
+ copied_artifacts.append(_copy_transcript_artifact(source_fd, dest_fd, artifact))
+ materialized["transcript_artifacts"] = copied_artifacts
+
+ review_name = row.get("review_artifact")
+ if isinstance(review_name, str) and review_name:
+ _copy_named_artifact(source_fd, dest_fd, review_name, label="review artifact")
+
+ task = row.get("task")
+ arm = row.get("arm")
+ run = row.get("run")
+ if isinstance(task, str) and isinstance(arm, str) and isinstance(run, int) and not isinstance(run, bool):
+ patch_name = f"{task}-{arm}-run{run}.patch"
+ if _is_regular_at(patch_name, dir_fd=source_fd):
+ _copy_named_artifact(source_fd, dest_fd, patch_name, label="patch artifact")
+ return materialized
+
+
+def default_reuse_max_age() -> timedelta:
+ return timedelta(days=EVIDENCE_MAX_AGE_DAYS)
+
+
+def _row_identity(row: Mapping[str, Any]) -> tuple[Any, ...]:
+ return (
+ row.get("skill_digest"),
+ row.get("oracle_digest"),
+ row.get("review_weighted_f1"),
+ row.get("ce_plugin_manifest_digest"),
+ row.get("recorded_at"),
+ )
+
+
+def _parse_recorded_at(value: Any) -> datetime | None:
+ if not isinstance(value, str) or not value:
+ return None
+ try:
+ parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
+ except ValueError:
+ return None
+ if parsed.tzinfo is None:
+ parsed = parsed.replace(tzinfo=UTC)
+ return parsed.astimezone(UTC)
+
+
+def _transcript_metadata(metadata: Any) -> tuple[str, str, int]:
+ if not isinstance(metadata, dict) or set(metadata) != {"path", "sha256", "bytes", "source"}:
+ raise SandboxError("transcript artifact metadata must contain only path, sha256, bytes, and source")
+ relative = metadata["path"]
+ digest = metadata["sha256"]
+ size = metadata["bytes"]
+ if metadata["source"] != PARENT_EVENT_STREAM_SOURCE:
+ raise SandboxError("transcript artifact source is not the parent event stream")
+ if not isinstance(relative, str) or not isinstance(digest, str) or not re.fullmatch(r"[0-9a-f]{64}", digest):
+ raise SandboxError("transcript artifact metadata is malformed")
+ if not isinstance(size, int) or isinstance(size, bool) or size < 0 or size > MAX_TRANSCRIPT_BYTES:
+ raise SandboxError("transcript artifact byte count is out of range")
+ relative_path = PurePosixPath(relative)
+ if (
+ relative_path.is_absolute()
+ or len(relative_path.parts) != 2
+ or relative_path.parts[0] != "transcripts"
+ or any(part in {"", ".", ".."} for part in relative_path.parts)
+ or _TRANSCRIPT_NAME.fullmatch(relative_path.parts[1]) is None
+ ):
+ raise SandboxError(f"unsafe transcript artifact path: {relative!r}")
+ return relative, digest, size
+
+
+def _resolved_directory(path: Path, *, label: str) -> tuple[Path, tuple[int, int]]:
+ """An existing, non-symlink directory, resolved through its parents.
+
+ Deliberately weaker than proposer_sandbox's same-shaped helper, which
+ refuses every symlink hop in the path. That one guards a MOUNT ROOT, where
+ a hop changes what an untrusted session is handed. This one guards a DATA
+ directory whose contents are validated individually anyway - every file
+ read goes through ``_regular_file`` (lstat, symlinks rejected) and every
+ write through ``O_NOFOLLOW`` - so a symlinked parent grants nothing those
+ guards do not already cover, while refusing one would reject ordinary
+ setups such as a symlinked artifacts directory or macOS's /var.
+
+ Separately named because they make different promises. Do not merge them
+ without first deciding which promise the reuse path should make.
+ """
+
+ resolved = path.expanduser()
+ try:
+ metadata = resolved.lstat()
+ except OSError as exc:
+ raise SandboxError(f"{label} is unavailable: {resolved}: {exc}") from exc
+ if stat.S_ISLNK(metadata.st_mode) or not stat.S_ISDIR(metadata.st_mode):
+ raise SandboxError(f"{label} must be a real directory: {resolved}")
+ return resolved.resolve(), (metadata.st_dev, metadata.st_ino)
+
+
+@contextmanager
+def _open_pinned_root(path: Path, *, label: str) -> Iterator[int]:
+ """Open a checked root and prove it is still the directory that was checked.
+
+ The symlink POLICY above is deliberate and unchanged: parent hops stay
+ allowed, so a symlinked artifacts directory or macOS's /var still works.
+ What is closed here is separate from that policy - the gap between checking
+ a name and using it. lstat names one directory and resolve() re-walks the
+ same name afterwards, so a prior sweep that renames its results root and
+ drops a symlink in its place is resolved to somewhere else entirely, and
+ O_NOFOLLOW on the open cannot see a link that resolve() already followed.
+
+ Comparing the opened descriptor's identity to the checked one costs an
+ fstat and rejects nothing that holds still: a stable directory always
+ matches itself. It matters for reuse specifically because the failure is
+ silent - rows would be copied out of the wrong directory and folded into a
+ comparator baseline as though they were this sweep's own evidence.
+ """
+
+ resolved, expected = _resolved_directory(path, label=label)
+ with _open_real_directory(resolved, label=label) as fd:
+ opened = os.fstat(fd)
+ if (opened.st_dev, opened.st_ino) != expected:
+ raise SandboxError(f"{label} was replaced between the check and the open: {resolved}")
+ yield fd
+
+
+def _copy_transcript_artifact(source_fd: int, dest_fd: int, metadata: Mapping[str, Any]) -> dict[str, Any]:
+ relative, expected_digest, expected_size = _transcript_metadata(metadata)
+ name = PurePosixPath(relative).name
+ # Both `transcripts` components are opened as descriptors, not checked as
+ # names. An lstat that passes and a pathname that is used afterwards are two
+ # different directories whenever a concurrent writer renames the first one
+ # away — which the reuse directory, written by a prior sweep, invites.
+ with (
+ _open_real_directory("transcripts", dir_fd=dest_fd, label="transcript destination", create=True) as dest_dir_fd,
+ _open_real_directory("transcripts", dir_fd=source_fd, label="transcript source") as source_dir_fd,
+ ):
+ os.fchmod(dest_dir_fd, 0o700)
+ # One descriptor for the whole transfer, and ONE read of it. Hashing the
+ # source and then reading it again to copy leaves the recorded digest
+ # describing bytes that are not the bytes written: the descriptor stops
+ # the pathname being substituted, not the inode being rewritten, and
+ # this directory belongs to a sweep that may still be writing. Digest
+ # what is copied, then judge it.
+ with _open_regular(name, dir_fd=source_dir_fd, label="transcript") as artifact_fd:
+ digest, copied_bytes = _copy_owner_only(
+ artifact_fd, name, dir_fd=dest_dir_fd, max_bytes=expected_size
+ )
+ if copied_bytes != expected_size or digest != expected_digest:
+ # The destination now holds bytes no expectation vouches for.
+ os.unlink(name, dir_fd=dest_dir_fd)
+ drift = "size" if copied_bytes != expected_size else "digest"
+ raise SandboxError(f"reused transcript {drift} drifted: {relative}")
+ return {"path": relative, "sha256": digest, "bytes": expected_size, "source": PARENT_EVENT_STREAM_SOURCE}
+
+
+def _copy_named_artifact(source_fd: int, dest_fd: int, name: str, *, label: str) -> None:
+ relative = PurePosixPath(name)
+ if relative.is_absolute() or len(relative.parts) != 1 or relative.parts[0] in {"", ".", ".."}:
+ raise SandboxError(f"unsafe {label} path: {name!r}")
+ with _open_regular(name, dir_fd=source_fd, label=label) as artifact_fd:
+ # No expectation is recorded for these, so the digest is discarded - but
+ # "no recorded size" is not "no limit". The source is a prior sweep
+ # directory that can change between sweeps, so a replaced artifact could
+ # be arbitrarily large; MAX_TRANSCRIPT_BYTES is the ceiling the capture
+ # path already enforces on evidence of this kind.
+ _, copied = _copy_owner_only(artifact_fd, name, dir_fd=dest_fd, max_bytes=MAX_TRANSCRIPT_BYTES)
+ if copied > MAX_TRANSCRIPT_BYTES:
+ os.unlink(name, dir_fd=dest_fd)
+ raise SandboxError(f"reused {label} exceeds {MAX_TRANSCRIPT_BYTES} bytes: {name}")
+
+
+def _require_openat() -> None:
+ """openat is what makes a checked directory and a used directory the same one.
+
+ Without it the only alternative is to re-walk the name after the check,
+ which is exactly the race this module is guarding. Refusing is safe: the
+ caller in runner treats a SandboxError from reuse as "run a paid cell", so
+ a platform without openat pays for the cells rather than copying through a
+ directory nobody verified. The sweep itself is Linux-only anyway (bwrap,
+ /proc/uptime); this is about the unit tests and about failing loudly.
+ """
+
+ if os.open not in os.supports_dir_fd or os.lstat not in os.supports_dir_fd:
+ raise SandboxError("comparator reuse requires POSIX openat support (os.supports_dir_fd)")
+
+
+def _is_regular_at(name: str, *, dir_fd: int) -> bool:
+ """True when `name` under the pinned directory is a regular non-symlink file."""
+
+ try:
+ metadata = os.lstat(name, dir_fd=dir_fd)
+ except OSError:
+ return False
+ return stat.S_ISREG(metadata.st_mode)
+
+
+@contextmanager
+def _open_real_directory(
+ path: Path | str,
+ *,
+ dir_fd: int | None = None,
+ label: str,
+ create: bool = False,
+) -> Iterator[int]:
+ """Open one directory that is not a symlink, and hold it for every use below.
+
+ ``O_DIRECTORY | O_NOFOLLOW`` makes the check and the open a single syscall,
+ so unlike an ``lstat`` followed by a path, there is no window in which the
+ directory can be replaced. ``_resolved_directory`` still tolerates a
+ symlinked reuse ROOT — it hands this function the already-resolved path —
+ but every component below it is pinned.
+ """
+
+ _require_openat()
+ if create:
+ try:
+ os.mkdir(path, 0o700, dir_fd=dir_fd)
+ except FileExistsError:
+ # Already there is the ordinary case — a second artifact from the
+ # same row. What it already IS still has to be proven, and the
+ # O_DIRECTORY|O_NOFOLLOW open below is what proves it, so there is
+ # nothing to do here.
+ pass
+ except OSError as exc:
+ raise SandboxError(f"{label} cannot be created: {path}: {exc}") from exc
+ try:
+ descriptor = os.open(
+ path,
+ os.O_RDONLY | getattr(os, "O_DIRECTORY", 0) | getattr(os, "O_NOFOLLOW", 0),
+ dir_fd=dir_fd,
+ )
+ except FileNotFoundError as exc:
+ # Absent is a different fact from present-but-not-a-real-directory, and
+ # the caller falls through to a paid cell on either.
+ raise SandboxError(f"{label} is missing: {path}") from exc
+ except OSError as exc:
+ raise SandboxError(f"{label} must be a real directory: {path}: {exc}") from exc
+ try:
+ # O_DIRECTORY is the check on Linux; the fstat covers a platform whose
+ # os module does not define it, where the flag degrades to 0.
+ if not stat.S_ISDIR(os.fstat(descriptor).st_mode):
+ raise SandboxError(f"{label} must be a real directory: {path}")
+ yield descriptor
+ finally:
+ os.close(descriptor)
+
+
+@contextmanager
+def _open_regular(name: str, *, dir_fd: int, label: str) -> Iterator[int]:
+ """Open a regular non-symlink file under a pinned directory, and hold it.
+
+ Checking a name and then re-opening it is a race the reuse directory is
+ exposed to: it is written by a previous sweep and read by this one, so a
+ concurrent writer can replace a validated file with a symlink in between.
+ Resolving against ``dir_fd`` removes the directory half, ``O_NOFOLLOW``
+ refuses the leaf link, and the fstat comparison proves the open descriptor
+ is the inode that was checked — the same guarantee
+ evolution._bounded_regular_bytes makes for evidence files.
+ """
+
+ _require_openat()
+ try:
+ before = os.lstat(name, dir_fd=dir_fd)
+ except OSError as exc:
+ raise SandboxError(f"{label} is missing: {name}: {exc}") from exc
+ if stat.S_ISLNK(before.st_mode) or not stat.S_ISREG(before.st_mode):
+ raise SandboxError(f"{label} must be a regular non-symlink file: {name}")
+ try:
+ descriptor = os.open(name, os.O_RDONLY | getattr(os, "O_NOFOLLOW", 0), dir_fd=dir_fd)
+ except OSError as exc:
+ raise SandboxError(f"{label} is unreadable: {name}: {exc}") from exc
+ try:
+ opened = os.fstat(descriptor)
+ if not stat.S_ISREG(opened.st_mode) or (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino):
+ raise SandboxError(f"{label} changed while opening: {name}")
+ yield descriptor
+ finally:
+ os.close(descriptor)
+
+
+def _copy_owner_only(source: int, name: str, *, dir_fd: int, max_bytes: int | None = None) -> tuple[str, int]:
+ """Copy one open file into the pinned directory; return what was written.
+
+ The digest is taken from the same buffers that are written, so it describes
+ the copy rather than a state the source was in at some earlier read.
+
+ ``max_bytes`` bounds the copy itself. The source is a prior sweep directory
+ this module already treats as concurrently writable, so a transcript
+ appended to after its metadata was recorded would otherwise be streamed to
+ EOF and only then compared against its declared size - filling the
+ destination, or never reaching EOF at all, long before the drift check could
+ reject it. Stopping one byte past the ceiling keeps that comparison
+ meaningful while bounding the work.
+ """
+
+ # O_CREAT|O_EXCL is the existence check, and unlike a stat beforehand it is
+ # atomic: a file appearing between check and open cannot slip through.
+ try:
+ descriptor = os.open(
+ name,
+ os.O_WRONLY | os.O_CREAT | os.O_EXCL | getattr(os, "O_NOFOLLOW", 0),
+ 0o600,
+ dir_fd=dir_fd,
+ )
+ except FileExistsError as exc:
+ raise SandboxError(f"reuse destination already exists: {name}") from exc
+ try:
+ os.fchmod(descriptor, 0o600)
+ os.lseek(source, 0, os.SEEK_SET)
+ digest = hashlib.sha256()
+ written = 0
+ limit = None if max_bytes is None else max_bytes + 1
+ while True:
+ want = COPY_CHUNK_BYTES if limit is None else min(COPY_CHUNK_BYTES, limit - written)
+ if want <= 0:
+ break
+ chunk = os.read(source, want)
+ if not chunk:
+ break
+ digest.update(chunk)
+ written += len(chunk)
+ _write_all(descriptor, chunk)
+ os.fsync(descriptor)
+ return digest.hexdigest(), written
+ finally:
+ os.close(descriptor)
diff --git a/eval/workflow_bench/evolve.py b/eval/workflow_bench/evolve.py
index 71a2299cb..3c17325be 100644
--- a/eval/workflow_bench/evolve.py
+++ b/eval/workflow_bench/evolve.py
@@ -48,6 +48,7 @@ import yaml
from . import runner
from . import runner_sessions
+from .comparator_reuse import current_runtime_digest
from .model_gateway import (
ANTHROPIC_API_KEY_ENV,
attach_openai_gateway,
@@ -698,8 +699,17 @@ def run_proposer(
bwrap_bin: Path,
sandbox_backend: str = "bwrap",
progress_label: str | None = None,
+ started_monotonic: float | None = None,
) -> dict[str, Any]:
- """Run one proposer in confinement and copy only validated outputs out."""
+ """Run one proposer in confinement and copy only validated outputs out.
+
+ ``started_monotonic`` is the sweep clock, not a precomputed budget. The
+ per-session ``--timeout`` is sized for a whole generation, so a proposer
+ started with only the sweep minimum left would otherwise run far past the
+ instance window; the clock is passed rather than the leftover because the
+ clone, the sanitize pass and the sandbox setup below all happen before the
+ session starts, and a number sampled by the caller is already stale by then.
+ """
with tempfile.TemporaryDirectory(prefix="wfevolve-") as tmp:
clone = runner.make_worktree(REPO_ROOT, "HEAD", Path(tmp))
@@ -732,11 +742,45 @@ def run_proposer(
host_text = getattr(sandbox, "host_text", lambda value: value)
environment_builder = getattr(sandbox, "environment", build_sandbox_environment)
backend = getattr(sandbox, "backend", "bwrap")
+ # Sampled here, after the setup above: this is the last
+ # moment before the session starts, so it is the only reading
+ # the session's own timeout can honestly be clamped to.
+ remaining_seconds = (
+ None
+ if started_monotonic is None
+ else remaining_runtime_seconds(
+ max_runtime_seconds=args.max_runtime_seconds,
+ started_monotonic=started_monotonic,
+ )
+ )
+ # An exhausted cap must stop the run, not buy one more second.
+ # remaining_runtime_seconds floors at 0, and max(1, ...) turned
+ # that 0 into a one-second paid session: the admission check
+ # happens before cloning, sanitizing and sandbox setup, so those
+ # unbounded steps can spend the rest of the window and leave
+ # nothing for the upload reserve this cap exists to protect.
+ if remaining_seconds is not None and remaining_seconds < 1:
+ # The caller stops the run on a not-ok record, which is the
+ # right outcome: an exhausted cap should end the generation,
+ # not start a session it cannot afford to finish.
+ return {
+ "ok": False,
+ "error_kind": "runtime-cap-exhausted",
+ "error_detail": (
+ "the wall-clock cap elapsed during proposer setup "
+ "(clone, sanitize, sandbox), before the session started"
+ ),
+ "duration_s": 0.0,
+ "num_turns": 0,
+ "cost_usd": None,
+ }
record = runner.run_claude(
host_text(prompt),
clone,
claude_bin=sandbox.claude_bin,
- timeout=args.timeout,
+ timeout=(
+ args.timeout if remaining_seconds is None else min(args.timeout, remaining_seconds)
+ ),
model=args.proposer_model,
effort=args.effort,
env=model_session_environment(
@@ -823,6 +867,119 @@ def _timeout_arm_key(arm: str) -> str:
return CANDIDATE_ARMS.get(arm, arm)
+EVENTBRIDGE_INSTANCE_WINDOW_SECONDS = 86_400
+EVENTBRIDGE_STOP_RESERVE_SECONDS = 5_400
+MIN_INSTANCE_SWEEP_SECONDS = 600
+
+
+def instance_window_budget_seconds(
+ uptime_seconds: float,
+ *,
+ window_seconds: int = EVENTBRIDGE_INSTANCE_WINDOW_SECONDS,
+ reserve_seconds: int = EVENTBRIDGE_STOP_RESERVE_SECONDS,
+ min_seconds: int = MIN_INSTANCE_SWEEP_SECONDS,
+) -> int:
+ """Seconds a sweep may run before an EventBridge 24h instance stop.
+
+ The dedicated evolution box is started ~15 minutes before the Saturday
+ cron and stopped 24h later. A ``workflow_dispatch`` that lands on an
+ already-running box inherits the leftover uptime, not a fresh day.
+ Run 33962002890 dispatched Friday 10:57 UTC and was still on its last
+ review cell when the Saturday 03:00 stop cancelled the runner — 51
+ finished sessions never uploaded because a cancelled job skips even
+ ``if: always()``. Capping the in-process sweep so it *fails* (instead
+ of vanishing) leaves the reserve for the upload step.
+ """
+
+ if window_seconds < 1 or reserve_seconds < 0 or min_seconds < 1:
+ raise ValueError("instance window and minimum must be positive; reserve must be non-negative")
+ if not math.isfinite(uptime_seconds) or uptime_seconds < 0:
+ raise ValueError("uptime must be a finite non-negative number")
+ leftover = int(window_seconds - uptime_seconds - reserve_seconds)
+ if leftover < min_seconds:
+ raise ValueError(
+ f"instance window has only {leftover}s left after a {reserve_seconds}s "
+ f"upload reserve (uptime {uptime_seconds:.0f}s of {window_seconds}s); "
+ f"need at least {min_seconds}s"
+ )
+ return leftover
+
+
+def _instance_uptime_or_none() -> float | None:
+ """The uptime read main() takes before it knows whether it needs it.
+
+ Deferring the read until after argument parsing would put the parse back
+ inside the interval the cap is supposed to cover, so it happens first and
+ an unreadable /proc/uptime is only an error if the flag turns out to be set.
+ """
+
+ try:
+ return read_instance_uptime_seconds()
+ except ValueError:
+ return None
+
+
+def read_instance_uptime_seconds(uptime_path: Path = Path("/proc/uptime")) -> float:
+ """Host uptime, the clock the EventBridge stop is scheduled against."""
+
+ try:
+ return float(uptime_path.read_text().split()[0])
+ except (OSError, IndexError, ValueError) as exc:
+ raise ValueError(f"cannot read instance uptime from {uptime_path}: {exc}") from exc
+
+
+def instance_window_budget_from_uptime(
+ uptime_seconds: float,
+ *,
+ window_seconds: int | None = None,
+ reserve_seconds: int | None = None,
+) -> int:
+ """Apply the EventBridge window env overrides to an already-read uptime.
+
+ Separate from the read so ``main`` can take the uptime in the same breath
+ as its own clock: the budget and the clock it is measured against have to
+ describe one instant, or the interval between them is spent by nobody and
+ charged to the sweep.
+ """
+
+ window = (
+ window_seconds
+ if window_seconds is not None
+ else int(os.environ.get("EVENTBRIDGE_INSTANCE_WINDOW_SECONDS", str(EVENTBRIDGE_INSTANCE_WINDOW_SECONDS)))
+ )
+ reserve = (
+ reserve_seconds
+ if reserve_seconds is not None
+ else int(os.environ.get("EVENTBRIDGE_STOP_RESERVE_SECONDS", str(EVENTBRIDGE_STOP_RESERVE_SECONDS)))
+ )
+ return instance_window_budget_seconds(uptime_seconds, window_seconds=window, reserve_seconds=reserve)
+
+
+
+
+def remaining_runtime_seconds(*, max_runtime_seconds: int | None, started_monotonic: float) -> int | None:
+ """Seconds left in an optional wall-clock cap, or None when uncapped."""
+
+ if max_runtime_seconds is None:
+ return None
+ if max_runtime_seconds < 1:
+ raise ValueError("max runtime must be positive")
+ leftover = max_runtime_seconds - (time.monotonic() - started_monotonic)
+ return max(0, int(leftover))
+
+
+def capped_timeout_seconds(requested: int, remaining: int | None) -> int:
+ """Clamp one managed-process timeout to the leftover instance window."""
+
+ if requested < 1:
+ raise ValueError("requested timeout must be positive")
+ if remaining is None:
+ return requested
+ if remaining < 1:
+ raise ValueError("no time remains in the instance window")
+ return min(requested, remaining)
+
+
def generation_timeout_seconds(
*,
task_count: int,
@@ -877,6 +1034,7 @@ def runner_argv(
task_bindings: list[dict[str, Any]],
target_base_digests: dict[str, str],
proposer_model: str | None = None,
+ reuse_results: Path | None = None,
) -> list[str]:
incumbent_arms = resolve_incumbent_arms(overlay_dir, args.arms)
paired_arms = executed_benchmark_arms(incumbent_arms)
@@ -927,6 +1085,8 @@ def runner_argv(
argv += ["--ce-plugin-dir", str(args.ce_plugin_dir), "--ce-plugin-version", args.ce_plugin_version]
if args.unsafe_no_bwrap:
argv.append("--unsafe-no-bwrap")
+ if reuse_results is not None:
+ argv += ["--reuse-results", str(reuse_results)]
return argv
@@ -944,6 +1104,12 @@ def runner_environment(args: argparse.Namespace) -> dict[str, str]:
# actually show progress rather than a burst at the end.
"PYTHONUNBUFFERED": "1",
}
+ # process_control replaces the child environment wholesale, so a digest the
+ # workflow exported reaches the runner only if it is forwarded here. Without
+ # this the runner stamps no runtime_digest and the reuse lock never engages.
+ runtime_digest = current_runtime_digest()
+ if runtime_digest:
+ env["RUNTIME_DIGEST"] = runtime_digest
if args.auth_token:
env[ANTHROPIC_API_KEY_ENV] = args.auth_token
return env
@@ -1199,6 +1365,13 @@ def _require_finite_metric(value: Any, name: str, *, nullable: bool = False, max
raise ValueError(f"promotion has invalid {name}")
+def _positive_int(value: str) -> int:
+ parsed = int(value)
+ if parsed < 1:
+ raise argparse.ArgumentTypeError(f"{value} is not a positive integer")
+ return parsed
+
+
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--tasks", required=True, type=Path)
@@ -1237,7 +1410,8 @@ def build_parser() -> argparse.ArgumentParser:
"--seed-results",
type=Path,
default=None,
- help="prior wfbench results dir used as generation-0 proposer evidence",
+ help="prior wfbench results dir used as generation-0 proposer evidence "
+ "and as --reuse-results for unchanged incumbent/CE cells",
)
parser.add_argument(
"--initial-overlay",
@@ -1271,6 +1445,19 @@ def build_parser() -> argparse.ArgumentParser:
default=runner_sessions.SESSION_TIMEOUT_SECONDS,
help="per session, seconds",
)
+ parser.add_argument(
+ "--max-runtime-seconds",
+ type=_positive_int,
+ default=None,
+ help="wall-clock cap for the whole evolve process (CI derives this from "
+ "instance uptime so the sweep exits before EventBridge stops the box)",
+ )
+ parser.add_argument(
+ "--max-runtime-from-instance-window",
+ action="store_true",
+ help="derive --max-runtime-seconds from /proc/uptime at startup, so the "
+ "budget and the clock it is measured against describe one instant",
+ )
parser.add_argument("--base-url", default=None)
parser.add_argument(
"--anthropic-api-key",
@@ -1308,8 +1495,26 @@ def build_parser() -> argparse.ArgumentParser:
def main() -> int:
+ # These two lines are the cap, and they are adjacent on purpose: the clock
+ # the sweep is measured against, and the uptime the budget is derived from.
+ # run-evolution.sh used to compute the budget in its own `uv run python -c`
+ # and pass a number, so the script's remaining work and this interpreter's
+ # startup were spent by nobody and charged to the sweep — out of the upload
+ # reserve the cap exists to protect. Nothing can be spent between them now.
+ started_monotonic = time.monotonic()
+ instance_uptime = _instance_uptime_or_none()
parser = build_parser()
args = parser.parse_args()
+ if args.max_runtime_from_instance_window:
+ if args.max_runtime_seconds is not None:
+ parser.error("--max-runtime-from-instance-window and --max-runtime-seconds are mutually exclusive")
+ if instance_uptime is None:
+ parser.error("--max-runtime-from-instance-window needs a readable /proc/uptime")
+ try:
+ args.max_runtime_seconds = instance_window_budget_from_uptime(instance_uptime)
+ except ValueError as exc:
+ parser.error(str(exc))
+ print(f"capping the sweep to {args.max_runtime_seconds}s so the instance-window reserve can upload evidence")
if args.generations < 1:
parser.error("--generations must be positive")
if args.runs < 1 or args.timeout < 1:
@@ -1379,6 +1584,7 @@ def main() -> int:
try:
return _run_generations(
args,
+ started_monotonic=started_monotonic,
selected_task_rows=selected_task_rows,
skipped_expensive=skipped_expensive,
selected_tasks=selected_tasks,
@@ -1394,6 +1600,7 @@ def main() -> int:
def _run_generations(
args: argparse.Namespace,
*,
+ started_monotonic: float,
selected_task_rows: list[dict[str, Any]],
skipped_expensive: list[str],
selected_tasks: list[dict[str, Any]],
@@ -1465,6 +1672,19 @@ def _run_generations(
incumbent_arms=requested_arms,
prior_proposal=staged_prior_included,
)
+ # Check the window before the paid session, not after it. A
+ # generation that cannot fit its sweep should not buy a proposal
+ # first and discover the deadline on the way out.
+ before_proposer = remaining_runtime_seconds(
+ max_runtime_seconds=args.max_runtime_seconds,
+ started_monotonic=started_monotonic,
+ )
+ if before_proposer is not None and before_proposer < MIN_INSTANCE_SWEEP_SECONDS:
+ print(
+ f"[gen {generation}] stopping with {before_proposer}s left before the "
+ f"instance window ends; not starting a proposer session"
+ )
+ return 1
print(f"[gen {generation}] proposing…")
record = run_proposer(
prompt,
@@ -1475,6 +1695,11 @@ def _run_generations(
bwrap_bin=bwrap_bin,
sandbox_backend=sandbox_backend,
progress_label=f"gen {generation} proposer",
+ # The clock, not the reading taken above: run_proposer clones,
+ # sanitizes and builds a sandbox before the session starts, so
+ # before_proposer is stale by then. It still decides whether to
+ # start at all — it just cannot decide how long to allow.
+ started_monotonic=started_monotonic,
)
# Redact any API token echoed into the session record (e.g. an
# error_detail stderr_tail) before it enters the uploaded artifact.
@@ -1513,6 +1738,16 @@ def _run_generations(
print(f"[gen {generation}] promotion targets contain uncommitted or drifted bytes")
return 1
print(f"[gen {generation}] benchmarking candidate…")
+ leftover = remaining_runtime_seconds(
+ max_runtime_seconds=args.max_runtime_seconds,
+ started_monotonic=started_monotonic,
+ )
+ if leftover is not None and leftover < MIN_INSTANCE_SWEEP_SECONDS:
+ print(
+ f"[gen {generation}] stopping with {leftover}s left before the "
+ f"instance window ends; partial evidence is in {out_root}/"
+ )
+ return 1
benchmark_argv = runner_argv(
args,
bench_dir,
@@ -1520,20 +1755,30 @@ def _run_generations(
task_bindings=selected_tasks,
target_base_digests=target_base_digests,
proposer_model=generation_proposer_model,
+ reuse_results=evidence_dir,
)
benchmark_command = (
benchmark_argv
if sandbox_backend == "host-unsafe"
else pid_namespace_command(benchmark_argv, bwrap_bin=bwrap_bin)
)
- bench = run_managed(
- benchmark_command,
- timeout=generation_timeout_seconds(
+ sweep_timeout = capped_timeout_seconds(
+ generation_timeout_seconds(
task_count=len(selected_task_rows),
runs=args.runs,
session_timeout=args.timeout,
incumbent_arms=incumbent_arms,
),
+ leftover,
+ )
+ if leftover is not None:
+ print(
+ f"[gen {generation}] sweep timeout {sweep_timeout}s "
+ f"(instance window leftover {leftover}s)"
+ )
+ bench = run_managed(
+ benchmark_command,
+ timeout=sweep_timeout,
env=runner_environment(args),
require_pid_namespace=sandbox_backend == "bwrap",
# The sweep is the multi-hour phase; without this its per-run
@@ -1545,7 +1790,13 @@ def _run_generations(
# The sweep runs with GITNEXUS_BENCH_ANTHROPIC_API_KEY in its environment,
# so its detail/stderr tail is a token-bearing sink like any other.
detail = redacted_failure(args, str(bench.detail or bench.stderr_tail[-1000:]))
- print(f"[gen {generation}] benchmark run failed ({bench.state}, exit {bench.returncode}): {detail}")
+ if leftover is not None and bench.state == "timeout":
+ print(
+ f"[gen {generation}] benchmark hit the instance-window budget "
+ f"({sweep_timeout}s); partial evidence is in {bench_dir}: {detail}"
+ )
+ else:
+ print(f"[gen {generation}] benchmark run failed ({bench.state}, exit {bench.returncode}): {detail}")
return 1
promotion = json.loads((bench_dir / "promotion.json").read_text())
for line in summarize_gate(promotion):
diff --git a/eval/workflow_bench/proposer_sandbox.py b/eval/workflow_bench/proposer_sandbox.py
index a2dfaa54a..2b0347da1 100644
--- a/eval/workflow_bench/proposer_sandbox.py
+++ b/eval/workflow_bench/proposer_sandbox.py
@@ -22,6 +22,15 @@ from .process_control import ManagedProcessResult, run_managed
MAX_EVIDENCE_FILE_BYTES = 256 * 1024
MAX_BUNDLE_BYTES = 2 * 1024 * 1024
SANDBOX_WORKSPACE = "/workspace"
+# The review artifact lives OUTSIDE the workspace, in its own writable
+# directory. A writable FILE inside a read-only directory is not writable to
+# anything that writes atomically: the Write tool creates
+# `.tmp..` beside the target and renames it, so a read-only
+# parent fails the temp create with EROFS and the artifact is never written.
+# Binding a writable directory outside /workspace lets the rename land while
+# the workspace itself stays entirely read-only.
+SANDBOX_REVIEW_OUTPUT = "/review-output"
+REVIEW_OUTPUT_DIRNAME = "review-output"
SANDBOX_HOME = "/home/agent"
SANDBOX_TMP = "/tmp"
SANDBOX_CLAUDE = "/opt/claude/claude"
@@ -118,36 +127,41 @@ REVIEW_RUNTIME_DIRECTORIES = (
)
-def prepare_review_workspace(sandbox: SandboxSession, artifact_name: str) -> Path:
- """Prepare disposable mount targets; never truncate a pre-existing entry."""
+def review_output_path(sandbox: SandboxSession, artifact_name: str) -> Path:
+ """Host path of the review artifact: a private directory, not the clone.
+
+ One source of truth for the location, so the mount, the parse and the
+ artifact copy cannot drift apart.
+ """
- clone = _real_directory(sandbox.clone, label="review clone")
if PurePosixPath(artifact_name).name != artifact_name or "\\" in artifact_name or artifact_name in ("", ".", ".."):
raise SandboxError("review artifact must be a root filename")
- output = clone / artifact_name
- # No agent runs while this private clone is being prepared. On POSIX the
- # directory descriptor additionally binds the exclusive create to its owner.
- directory_fd = None
- try:
- if os.name != "nt":
- directory_fd = os.open(clone, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW)
- fd = os.open(
- artifact_name if directory_fd is not None else output,
- os.O_WRONLY | os.O_CREAT | os.O_EXCL | getattr(os, "O_NOFOLLOW", 0),
- 0o600,
- dir_fd=directory_fd,
- )
- try:
- if not stat.S_ISREG(os.fstat(fd).st_mode):
- raise SandboxError("review artifact must be a regular file")
- finally:
- os.close(fd)
- except FileExistsError as exc:
- raise SandboxError("review artifact already exists") from exc
- finally:
- if directory_fd is not None:
- os.close(directory_fd)
+ return Path(sandbox.private_root) / REVIEW_OUTPUT_DIRNAME / artifact_name
+
+def prepare_review_workspace(sandbox: SandboxSession, artifact_name: str) -> Path:
+ """Prepare disposable mount targets; never truncate a pre-existing entry.
+
+ Creates the artifact's own directory and returns the path the agent is
+ expected to write. The file itself is deliberately NOT pre-created: the
+ agent writes it atomically (temp file beside the target, then rename), so
+ the directory is what has to be writable, and an existing empty file would
+ only be something for the write to trip over. Absence is meaningful — it is
+ how ``parse_review_output`` tells "never written" from "written badly".
+ """
+
+ output = review_output_path(sandbox, artifact_name)
+ # No agent runs while this private root is being prepared, and the
+ # exclusive create is what proves the directory is ours rather than
+ # something a previous cell left behind.
+ try:
+ output.parent.mkdir(mode=0o700, parents=False, exist_ok=False)
+ except FileExistsError as exc:
+ raise SandboxError("review artifact directory already exists") from exc
+ except OSError as exc:
+ raise SandboxError(f"review artifact directory is unavailable: {output.parent}") from exc
+
+ clone = _real_directory(sandbox.clone, label="review clone")
if sandbox.backend != "bwrap":
return output
created: list[str] = []
@@ -214,6 +228,10 @@ class SandboxSession:
ReadOnlyMount(self.clone, SANDBOX_WORKSPACE),
ReadOnlyMount(self.home, SANDBOX_HOME),
ReadOnlyMount(self.temp, SANDBOX_TMP),
+ # The review artifact directory is a real mount on bwrap, so the
+ # host-unsafe backend has to translate it too. Without this the
+ # review prompt names a path that exists on neither backend.
+ ReadOnlyMount(Path(self.private_root) / REVIEW_OUTPUT_DIRNAME, SANDBOX_REVIEW_OUTPUT),
]
for mount in sorted(mappings, key=lambda item: len(item.target), reverse=True):
target = mount.target.rstrip("/")
@@ -234,9 +252,16 @@ class SandboxSession:
SANDBOX_WORKSPACE,
SANDBOX_HOME,
SANDBOX_TMP,
+ SANDBOX_REVIEW_OUTPUT,
]
ordered = sorted(set(targets), key=len, reverse=True)
- pattern = re.compile("|".join(re.escape(target) for target in ordered))
+ # Only translate at a path boundary. "/review-output" occurs twice in
+ # "/review-output/review-output.json" - once as the directory and once
+ # inside the filename - and rewriting the second turned the artifact path
+ # into nonsense. A target must be followed by "/", whitespace, a quote or
+ # end of string to be a path rather than a prefix of a longer name.
+ boundary = r"""(?=[/\s"']|$)"""
+ pattern = re.compile("(?:" + "|".join(re.escape(target) for target in ordered) + ")" + boundary)
return pattern.sub(lambda match: self.host_path(match.group(0)), value)
@property
@@ -549,12 +574,17 @@ def build_claude_settings(*, sandbox_enabled: bool = True) -> str:
"allowLocalBinding": False,
},
"filesystem": {
- "allowWrite": [SANDBOX_WORKSPACE, SANDBOX_TMP, SANDBOX_HOME],
+ # SANDBOX_REVIEW_OUTPUT is the review artifact directory. The
+ # bwrap bind alone is not enough: this policy is a second,
+ # independent gate the CLI applies to its own tools, and a path
+ # missing here is unwritable however the mount is shaped.
+ "allowWrite": [SANDBOX_WORKSPACE, SANDBOX_TMP, SANDBOX_HOME, SANDBOX_REVIEW_OUTPUT],
"denyRead": ["/"],
"allowRead": [
SANDBOX_WORKSPACE,
SANDBOX_TMP,
SANDBOX_HOME,
+ SANDBOX_REVIEW_OUTPUT,
"/usr",
"/bin",
"/lib",
diff --git a/eval/workflow_bench/review_scoring.py b/eval/workflow_bench/review_scoring.py
index b8ad0151b..eaf328e37 100644
--- a/eval/workflow_bench/review_scoring.py
+++ b/eval/workflow_bench/review_scoring.py
@@ -12,6 +12,9 @@ from typing import Any, Mapping, Sequence
from .oracle_assets import TaskOracleSnapshot
REVIEW_OUTPUT = "review-output.json"
+# Task verify/oracle commands read the artifact location from here rather
+# than hardcoding a path, so one command works under bwrap and host-unsafe.
+REVIEW_OUTPUT_ENV_VAR = "GITNEXUS_BENCH_REVIEW_OUTPUT"
REVIEW_SCHEMA_VERSION = 1
MAX_REVIEW_BYTES = 256 * 1024
MAX_FINDINGS = 100
@@ -112,13 +115,33 @@ def _parse_review_finding(raw: Any, index: int) -> ReviewFinding:
def parse_review_output(path: Path) -> tuple[str, tuple[ReviewFinding, ...]]:
- metadata = path.lstat()
+ # Distinguish these. Folding empty, malformed and encoding failures into one
+ # message is how a sandbox that left the artifact at 0 bytes read for 15
+ # runs as an encoding fault: json.loads("") raises, and every such cell
+ # reported "not valid UTF-8 JSON". A path the agent never created was not in
+ # that fold — the lstat below sat outside the try and raised
+ # FileNotFoundError — but it reached the caller as a bare OSError rather
+ # than saying what was wrong, which is why it is named here too.
+ try:
+ metadata = path.lstat()
+ except FileNotFoundError as exc:
+ raise ValueError("review output was never written") from exc
+ except OSError as exc:
+ raise ValueError(f"review output is unreadable: {exc.strerror}") from exc
if path.is_symlink() or not path.is_file() or metadata.st_size > MAX_REVIEW_BYTES:
raise ValueError("review output must be a bounded regular non-symlink file")
+ if metadata.st_size == 0:
+ raise ValueError("review output is empty")
try:
- raw = json.loads(path.read_text())
- except (OSError, UnicodeError, json.JSONDecodeError) as exc:
- raise ValueError("review output is not valid UTF-8 JSON") from exc
+ text = path.read_text(encoding="utf-8")
+ except OSError as exc:
+ raise ValueError(f"review output is unreadable: {exc.strerror}") from exc
+ except UnicodeError as exc:
+ raise ValueError("review output is not valid UTF-8") from exc
+ try:
+ raw = json.loads(text)
+ except json.JSONDecodeError as exc:
+ raise ValueError(f"review output is not valid JSON: {exc.msg} at line {exc.lineno}") from exc
if not isinstance(raw, Mapping) or set(raw) != {"schema_version", "verdict", "findings"}:
raise ValueError("review output requires exactly schema_version, verdict, and findings")
if raw["schema_version"] != REVIEW_SCHEMA_VERSION:
diff --git a/eval/workflow_bench/run-evolution.sh b/eval/workflow_bench/run-evolution.sh
index c9d2a77a1..066df111a 100755
--- a/eval/workflow_bench/run-evolution.sh
+++ b/eval/workflow_bench/run-evolution.sh
@@ -15,6 +15,8 @@
# MODEL PROPOSER_MODEL EFFORT GENERATIONS RUNS WORKERS PROVIDER
# EVOLUTION_PROFILE CE_PLUGIN_DIR CE_PLUGIN_VERSION
# INCLUDE_EXPENSIVE SEED_RESULTS CLAUDE_BIN OUT_ROOT
+# CI (passes --max-runtime-from-instance-window; the CLI reads /proc/uptime)
+# EVENTBRIDGE_INSTANCE_WINDOW_SECONDS EVENTBRIDGE_STOP_RESERVE_SECONDS
# UNSAFE_NO_BWRAP=1 (local review diagnostics only)
# GITNEXUS_BENCH_ANTHROPIC_API_KEY (legacy GITNEXUS_BENCH_AUTH_TOKEN)
# GITNEXUS_BENCH_OPENAI_API_KEY
@@ -193,6 +195,18 @@ if ((${#passthrough[@]})); then
cmd+=("${passthrough[@]}")
fi
+# A cancelled GitHub job skips even `if: always()`, so evidence dies with the
+# runner. The evolution box is EventBridge-stopped 24h after boot; a Friday
+# dispatch inherits leftover uptime. Cap the sweep so it fails in-process and
+# the upload step still runs (run 33962002890). The CLI reads /proc/uptime
+# itself, in the same breath as it starts the clock the cap is measured
+# against; computing a number here — in a separate interpreter, before the
+# provenance work and the exec below — charged the sweep for every second
+# this script spent afterwards.
+if [[ -n "${CI:-}" && -r /proc/uptime ]]; then
+ cmd+=(--max-runtime-from-instance-window)
+fi
+
if ((dry_run)); then
printf '%q ' "${cmd[@]}"
printf '\n'
@@ -220,5 +234,9 @@ SOURCE_SHA="${source_sha}" RUNTIME_DIGEST="${runtime_digest}" SANDBOX_BACKEND="$
}, null, 2) + "\n")' "${out_root}/runtime-provenance.json"
export PYTHONUNBUFFERED=1
+# The runner stamps this on every results.jsonl row and refuses to reuse a
+# comparator cell when a prior row's digest disagrees. Keep it on the evolve
+# process, not only in the provenance JSON sidecar.
+export RUNTIME_DIGEST="${runtime_digest}"
cd "${eval_dir}"
exec "${cmd[@]}"
diff --git a/eval/workflow_bench/runner.py b/eval/workflow_bench/runner.py
index 9f9dd9ef3..a2223a20a 100644
--- a/eval/workflow_bench/runner.py
+++ b/eval/workflow_bench/runner.py
@@ -57,6 +57,17 @@ from typing import Any
import yaml
+from .comparator_reuse import (
+ REUSE_EXCLUDED_ERROR_KINDS,
+ CellKey,
+ ComparatorReuseExpectation,
+ TaskReuseBinding,
+ current_runtime_digest,
+ default_reuse_max_age,
+ load_result_rows,
+ materialize_reused_row,
+ select_reusable_comparator_rows,
+)
from .evolution import (
CANDIDATE_ARMS,
EVIDENCE_MAX_AGE_DAYS,
@@ -88,12 +99,19 @@ from .oracle_assets import (
)
from .process_control import cancellation_scope, ManagedProcessError
from .promotion_apply import committed_destination_base_digests
-from .review_scoring import REVIEW_OUTPUT, expected_findings, parse_review_output, score_review
+from .review_scoring import (
+ REVIEW_OUTPUT,
+ REVIEW_OUTPUT_ENV_VAR,
+ expected_findings,
+ parse_review_output,
+ score_review,
+)
from .proposer_sandbox import (
SANDBOX_GITNEXUS as SANDBOX_GITNEXUS,
SANDBOX_GITNEXUS_REGISTRY,
SANDBOX_GITNEXUS_SHARED as SANDBOX_GITNEXUS_SHARED,
SANDBOX_NODE as SANDBOX_NODE,
+ SANDBOX_REVIEW_OUTPUT,
SANDBOX_WORKSPACE,
ReadOnlyMount,
SandboxError,
@@ -104,6 +122,7 @@ from .proposer_sandbox import (
prepare_sandbox,
prepare_review_workspace,
redact_text,
+ review_output_path,
require_claude_sandbox_helpers,
sandbox_workspace_write_boundary,
)
@@ -121,6 +140,7 @@ from .runner_artifacts import (
enforce_phase_workspace,
enforce_work_evidence,
implementation_diff_digest,
+ copy_isolated_tree,
make_worktree,
new_plan_doc,
parse_shortstat as parse_shortstat,
@@ -222,8 +242,9 @@ CE_WORK_DIRECT_PROMPT = (
# Review cell: setup applies a historical PR diff, then the model sees a
# read-only checkout. Both arms emit the same strict artifact so quality can be
# scored deterministically against labels that remain hidden until it exits.
-REVIEW_OUTPUT_CONTRACT = """
-Write /workspace/review-output.json as UTF-8 JSON with exactly this shape:
+# Concatenated, not an f-string: the JSON shape below keeps its braces doubled
+# because the finished prompt is .format()-ed with the task text.
+REVIEW_OUTPUT_CONTRACT = f"\nWrite {SANDBOX_REVIEW_OUTPUT}/{REVIEW_OUTPUT} " + """as UTF-8 JSON with exactly this shape:
{{"schema_version":1,"verdict":"approve|comment|request_changes","findings":[{{
"id":"unique stable id","severity":"critical|high|medium|low",
"path":"repository-relative changed file","line":1,"end_line":1,
@@ -563,9 +584,13 @@ def run_arm(
read_only_workspace=True,
read_only_paths=_evaluated_skill_roots(worktree, arm),
extra_writable_mounts=(
+ # The DIRECTORY, outside the workspace. Binding the file
+ # itself left the agent nowhere to put the temp file it
+ # renames into place, so every review artifact came back
+ # empty with EROFS in the transcript.
ReadOnlyMount(
- source=review_output,
- target=f"{SANDBOX_WORKSPACE}/{REVIEW_OUTPUT}",
+ source=review_output.parent,
+ target=SANDBOX_REVIEW_OUTPUT,
),
),
),
@@ -573,7 +598,8 @@ def run_arm(
with sandbox_workspace_write_boundary(
sandbox,
read_only_workspace=True,
- writable=(review_output,),
+ # Nothing in the workspace is writable now — the artifact left it.
+ writable=(),
):
review_session = run_claude(
host_text(review_prompt.format(task=task["prompt"])),
@@ -584,11 +610,9 @@ def run_arm(
sessions.append(review_session)
if review_session["ok"] and phase_before is not None:
try:
- enforce_phase_workspace(
- worktree,
- phase_before,
- allowed_artifact=worktree / REVIEW_OUTPUT,
- )
+ # The artifact is no longer in the workspace, so the review
+ # phase may now change nothing there at all.
+ enforce_phase_workspace(worktree, phase_before, allowed_artifact=None)
require_skill_fingerprint(
worktree,
arm,
@@ -643,6 +667,18 @@ def run_arm(
record = sum_sessions(sessions)
record["arm"] = arm
record["plan_produced"] = arm not in ("workflow", "ce_workflow") or plan_doc is not None
+ # The verify command runs in its OWN sandbox invocation, which knows nothing
+ # about the review session's writable mount. Expose the artifact read-only and
+ # name it through the environment, the same shape _run_hidden_oracle uses, so
+ # one command works on both backends instead of hardcoding either path.
+ verify_env = environment_builder()
+ verify_mounts: tuple[ReadOnlyMount, ...] = ()
+ if arm in ("review", "ce_review"):
+ review_artifact = review_output_path(sandbox, REVIEW_OUTPUT)
+ verify_env[REVIEW_OUTPUT_ENV_VAR] = host_text(f"{SANDBOX_REVIEW_OUTPUT}/{REVIEW_OUTPUT}")
+ verify_mounts = (
+ ReadOnlyMount(source=review_artifact.parent, target=SANDBOX_REVIEW_OUTPUT),
+ )
authored_tests_passed, authored_test_output = _verification_outcome(
run_verify(
task["verify"],
@@ -651,21 +687,25 @@ def run_arm(
command_prefix=sandbox.command_prefix_for(
read_only_workspace=True,
unshare_network=True,
+ extra_read_only_mounts=verify_mounts,
),
- env=environment_builder(),
+ env=verify_env,
require_pid_namespace=getattr(sandbox, "require_pid_namespace", True),
)
)
review_score: dict[str, Any] | None = None
if arm in ("review", "ce_review"):
try:
- verdict, findings = parse_review_output(worktree / REVIEW_OUTPUT)
+ verdict, findings = parse_review_output(review_output_path(sandbox, REVIEW_OUTPUT))
labels = expected_findings(oracle_snapshot) if oracle_snapshot is not None else ()
review_score = score_review(verdict, findings, labels)
except (OSError, ValueError) as exc:
record["ok"] = False
record["error_kind"] = record["error_kind"] or "review-evidence-invalid"
- record["error_detail"] = str(exc)
+ # Keep the FIRST detail, as error_kind already does. A phase-
+ # boundary violation is why the artifact is unparseable; reporting
+ # the parse failure over it buries the cause under the symptom.
+ record["error_detail"] = record.get("error_detail") or str(exc)
record["review_score"] = review_score
record["review_evidence_valid"] = review_score is not None
if review_score is not None:
@@ -721,9 +761,53 @@ CHURN_FIELDS = ("diff_files", "diff_insertions", "diff_deletions")
# Rows where the session (or the harness) died carry no measured evidence and
# must not skew efficiency medians or resolve denominators. verify-failed and
# skill-not-invoked rows DO count: those sessions ran and spent real tokens.
-EXCLUDED_ERROR_KINDS = frozenset(
- {"session-error", "infra-error", "evidence-unverified", "cleanup-failure", "review-evidence-invalid", "cancelled"}
-)
+# One definition, in comparator_reuse: reuse eligibility and aggregate
+# exclusion must never drift apart. The dependency only runs this way -
+# comparator_reuse importing back from runner is a circular import.
+EXCLUDED_ERROR_KINDS = REUSE_EXCLUDED_ERROR_KINDS
+
+
+
+# Health classification. These answer "did the harness work", which is a
+# different question from "did the agent get the right answer" - a review can be
+# wrong about a hard corpus while every process, mount and capture behaved.
+#
+# EXECUTION: the process or its tooling did not complete. Nothing was measured.
+# EVIDENCE: it completed, but what it produced cannot be trusted or scored.
+# Everything else - including resolved=False and a zero score - is a VALID
+# NEGATIVE: an admissible measurement that the quality gate then judges.
+EXECUTION_FAILURE_KINDS = frozenset({"session-error", "infra-error", "cleanup-failure", "cancelled"})
+EVIDENCE_FAILURE_KINDS = frozenset({"review-evidence-invalid", "evidence-unverified", "skill-not-invoked"})
+
+
+def execution_failed(record: Mapping[str, Any]) -> bool:
+ """The process or its tooling did not complete."""
+
+ return record.get("error_kind") in EXECUTION_FAILURE_KINDS
+
+
+def evidence_failed(record: Mapping[str, Any]) -> bool:
+ """It completed, but what it produced cannot be trusted or scored."""
+
+ return (
+ record.get("error_kind") in EVIDENCE_FAILURE_KINDS
+ or record.get("review_evidence_valid") is False
+ or record.get("transcript_missing") is True
+ )
+
+
+def task_prompt_digest(task: Mapping[str, Any]) -> str:
+ """The prompt digest, computed once for the row and the reuse expectation.
+
+ row_is_reusable_comparator compares the value a prior row stored against the
+ value this sweep derives, as exact strings. Two inline copies of this hash
+ had already drifted - one picked up a str() cast the other lacked - and a
+ further divergence (normalising whitespace on one side, say) would silently
+ stop rows matching, or match rows that should not.
+ """
+
+ return hashlib.sha256(str(task["prompt"]).encode()).hexdigest()
+
# A sustained upstream outage shows up as a run of session/infra/cleanup
# failures. (cleanup-failure overwrites the primary error_kind, so a
@@ -806,7 +890,13 @@ def sweep_task_cells(
raise ValueError("workers must be positive")
for wave_start in range(0, len(cells), workers):
if cancel_event.is_set():
- return outage_streak, True
+ # False: this flag means the OUTAGE breaker tripped, and the
+ # caller turns it into exit 1 with "Sweep aborted". Cancellation
+ # stops the sweep too, but it is the operator's Ctrl-C, not a
+ # systemic failure - reporting True relabelled every interrupted
+ # run an outage and returned 1 where the contract says 130. The
+ # caller tests cancel_event itself for the stop decision.
+ return outage_streak, False
wave = list(cells[wave_start : wave_start + workers])
for run_idx, arm in wave:
on_start(run_idx, arm)
@@ -833,7 +923,7 @@ def sweep_task_cells(
# an earlier row trips the breaker; only later waves are skipped.
on_record(run_idx, arm, record)
if cancel_event.is_set():
- return outage_streak, True
+ return outage_streak, False
for record in records:
kind = (
"review-evidence-invalid"
@@ -847,14 +937,15 @@ def sweep_task_cells(
"failures — aborting the remaining sweep; report and promotion are written "
"from partial evidence and the run exits non-zero."
)
+ # Signal in-flight background work too. The breaker exists to
+ # SHORTEN a doomed run; without this a graph prefetch keeps
+ # building and the unconditional join blocks the abort for the
+ # length of a full clone and offline index.
+ cancel_event.set()
return outage_streak, True
return outage_streak, False
-# A sustained upstream outage shows up as a run of session/infra/cleanup
-# failures. (cleanup-failure overwrites the primary error_kind, so a
-# session-error whose worktree cleanup also failed still counts.) A task's own
-# resolved=False is real signal, not an outage, so it never trips the breaker.
# How far ahead of the in-order fold pointer cells may be submitted, as a
# multiple of the worker count. This is the wall-clock/wasted-cell trade, and it
# is a real one - measured against the review corpus at workers=3, with failures
@@ -1082,6 +1173,8 @@ class TaskCellContext:
candidate_overlay: Path | None
overlay_digest: str | None
sandbox_backend: str = "bwrap"
+ clone_template: Path | None = None
+ sanitized_head: str | None = None
def run_cell(ctx: TaskCellContext, run_idx: int, arm: str) -> dict[str, Any]:
@@ -1106,8 +1199,14 @@ def run_cell(ctx: TaskCellContext, run_idx: int, arm: str) -> dict[str, Any]:
raise RuntimeError("sanitized graph snapshot is unavailable")
if ctx.asset_snapshot is None:
raise RuntimeError("task asset snapshot is unavailable")
- worktree = make_worktree(ctx.repo, ctx.task_sha, ctx.trees_dir)
- sanitized_head = sanitize_clone_for_hidden_oracles(worktree)
+ if ctx.clone_template is not None:
+ if not ctx.sanitized_head:
+ raise RuntimeError("clone template is missing its sanitized HEAD")
+ worktree = copy_isolated_tree(ctx.clone_template, ctx.trees_dir)
+ sanitized_head = ctx.sanitized_head
+ else:
+ worktree = make_worktree(ctx.repo, ctx.task_sha, ctx.trees_dir)
+ sanitized_head = sanitize_clone_for_hidden_oracles(worktree)
ctx.graph_snapshot.materialize(worktree, sanitized_head=sanitized_head)
dependency_mounts = stage_task_assets(
task,
@@ -1212,7 +1311,7 @@ def run_cell(ctx: TaskCellContext, run_idx: int, arm: str) -> dict[str, Any]:
oracle_snapshot=ctx.oracle_snapshot,
)
if execution_arm in ("review", "ce_review"):
- review_source = worktree / REVIEW_OUTPUT
+ review_source = review_output_path(sandbox, REVIEW_OUTPUT)
if review_source.is_file() and not review_source.is_symlink():
review_artifact = ctx.out_dir / f"{task['id']}-{arm}-run{run_idx}.review.json"
review_artifact.write_bytes(_bounded_regular_bytes(review_source, limit=256 * 1024))
@@ -1252,9 +1351,10 @@ def run_cell(ctx: TaskCellContext, run_idx: int, arm: str) -> dict[str, Any]:
"task_base_sha": ctx.task_sha,
"sanitized_task_sha": sanitized_head,
"variant_head_sha": orig_sha,
- "task_prompt_digest": hashlib.sha256(task["prompt"].encode()).hexdigest(),
+ "task_prompt_digest": task_prompt_digest(task),
"skill_digest": expected_skill_digest,
"candidate_overlay_digest": (ctx.overlay_digest if arm in CANDIDATE_ARMS else None),
+ "runtime_digest": current_runtime_digest(),
"recorded_at": datetime.now(UTC).isoformat(),
}
)
@@ -1441,7 +1541,30 @@ def aggregate(records: list[dict[str, Any]]) -> dict[str, Any]:
out["cost_usd"] = (
None if (not valid or any(cost is None for cost in valid_costs)) else statistics.median(valid_costs)
)
+ fresh = [r for r in records if not r.get("reused")]
+ out["fresh_attempts"] = len(fresh)
+ out["execution_failures"] = sum(1 for r in fresh if execution_failed(r))
+ out["evidence_failures"] = sum(1 for r in fresh if evidence_failed(r))
+ # Admissible means the harness delivered a trustworthy measurement. It says
+ # nothing about whether the answer was right, which is the whole point.
+ #
+ # Count the rows that failed NEITHER way rather than subtracting both
+ # counters: run_arm keeps a pre-existing session error and still marks the
+ # review evidence invalid, so one row can land in both. Subtracting it twice
+ # drove an arm holding real measurements to admissible=0, which arm_health
+ # reads as UNUSABLE and enforce_measurement_health then fails the sweep on.
+ out["admissible"] = sum(1 for r in fresh if not execution_failed(r) and not evidence_failed(r))
+ out["health_reasons"] = sorted(
+ {
+ str(r.get("error_kind"))
+ for r in fresh
+ if r.get("error_kind") in EXECUTION_FAILURE_KINDS or r.get("error_kind") in EVIDENCE_FAILURE_KINDS
+ }
+ )
out["resolved"] = sum(1 for r in records if r["resolved"])
+ # Reused rows are last generation's measurement. The health canary below has
+ # to ask whether THIS environment worked, so it needs the freshly-run count.
+ out["resolved_fresh"] = sum(1 for r in records if r["resolved"] and not r.get("reused"))
out["runs"] = len(records)
out["valid_runs"] = len(valid)
out["excluded_runs"] = len(records) - len(valid)
@@ -1499,11 +1622,135 @@ def savings(baseline: dict[str, Any], workflow: dict[str, Any]) -> dict[str, Any
return out
+@dataclass(frozen=True)
+class ArmHealth:
+ """What the harness observed for one arm this sweep, before any judgement."""
+
+ arm: str
+ fresh_attempts: int
+ admissible: int
+ execution_failures: int
+ evidence_failures: int
+ reasons: tuple[str, ...]
+
+ @property
+ def measured(self) -> bool:
+ """False when only reused rows exist - current health is UNKNOWN, not good."""
+
+ return self.fresh_attempts > 0
+
+ @property
+ def status(self) -> str:
+ """UNKNOWN / OBSERVED_OK / DEGRADED / UNUSABLE.
+
+ DEGRADED is the distinction that matters: an arm with both admissible
+ measurements and observed failures produced usable evidence but did not
+ run reliably. Reporting that as healthy is how a partly-broken sweep
+ looks fine. It is diagnostic here - only UNUSABLE is fatal - so this
+ patch changes what is reported, not what is eligible.
+ """
+
+ if not self.measured:
+ return "UNKNOWN"
+ failures = self.execution_failures + self.evidence_failures
+ if self.admissible == 0 and failures > 0:
+ return "UNUSABLE"
+ if failures > 0:
+ return "DEGRADED"
+ return "OBSERVED_OK"
+
+ @property
+ def unhealthy(self) -> bool:
+ """Every fresh attempt failed to execute or to produce usable evidence.
+
+ Deliberately not "resolved zero tasks". A reviewer can be wrong about
+ every task in a hard corpus with the harness working perfectly; that is
+ a valid negative and belongs to the quality gate, not here.
+ """
+
+ return self.status == "UNUSABLE"
+
+
+def arm_health(results: dict[str, dict[str, dict[str, Any]]], arms: set[str]) -> dict[str, ArmHealth]:
+ """Fold per-task aggregates into one health observation per arm."""
+
+ health: dict[str, ArmHealth] = {}
+ for arm in sorted(arms):
+ rows = [task_arms[arm] for task_arms in results.values() if arm in task_arms]
+ if not rows:
+ continue
+ reasons: set[str] = set()
+ for row in rows:
+ reasons.update(row.get("health_reasons") or ())
+ health[arm] = ArmHealth(
+ arm=arm,
+ fresh_attempts=sum(int(r.get("fresh_attempts", 0)) for r in rows),
+ admissible=sum(int(r.get("admissible", 0)) for r in rows),
+ execution_failures=sum(int(r.get("execution_failures", 0)) for r in rows),
+ evidence_failures=sum(int(r.get("evidence_failures", 0)) for r in rows),
+ reasons=tuple(sorted(reasons)),
+ )
+ return health
+
+
+def unhealthy_arms(results: dict[str, dict[str, dict[str, Any]]], arms: set[str]) -> list[ArmHealth]:
+ """Arms whose every fresh attempt failed to execute or to produce evidence."""
+
+ return [h for h in arm_health(results, arms).values() if h.unhealthy]
+
+
+def unmeasured_arms(results: dict[str, dict[str, dict[str, Any]]], arms: set[str]) -> list[str]:
+ """Arms with no fresh attempt at all - reported as unknown, never as healthy."""
+
+ return [h.arm for h in arm_health(results, arms).values() if not h.measured]
+
+
+def enforce_measurement_health(
+ results: dict[str, dict[str, dict[str, Any]]], arms: set[str]
+) -> dict[str, ArmHealth]:
+ """Report every arm's measurement status; abort only on UNUSABLE.
+
+ Runs after report.md and promotion.json are written, so a failing sweep
+ still leaves its evidence behind. Reports cause as undetermined: an empty
+ artifact establishes that evidence is unusable, not why - naming a mount
+ failure here would be a guess the recorded rows do not support.
+ """
+
+ health = arm_health(results, arms)
+ for arm in sorted(health):
+ observed = health[arm]
+ reasons = f" reason={','.join(observed.reasons)}" if observed.reasons else ""
+ print(
+ f"[measurement-health] {arm}: {observed.status} "
+ f"fresh_attempts={observed.fresh_attempts} admissible={observed.admissible} "
+ f"execution_failures={observed.execution_failures} "
+ f"evidence_failures={observed.evidence_failures}{reasons}"
+ )
+ unusable = [h for h in health.values() if h.unhealthy]
+ if unusable:
+ detail = "; ".join(f"{h.arm} ({h.fresh_attempts} fresh attempt(s))" for h in unusable)
+ print(
+ f"[measurement-health] {detail} produced no usable measurement this sweep. "
+ "cause=undetermined — see error_detail in results.jsonl. Exiting non-zero rather "
+ "than reporting a quiet no-promotion."
+ )
+ raise SystemExit(1)
+ return health
+
+
def broken_incumbent_arms(
results: dict[str, dict[str, dict[str, Any]]],
incumbent_arms: set[str],
) -> list[str]:
- """Incumbent arms that resolved nothing across every task they ran.
+ """LEGACY, NON-AUTHORITATIVE. Superseded by ``enforce_measurement_health``.
+
+ Kept only so its historical behaviour stays documented and testable while
+ the replacement settles; it has no production caller. Do not wire it into a
+ health decision - it infers a broken environment from a resolution count,
+ which a reviewer facing a hard corpus falsifies. Remove once the
+ measurement-health path has run in CI.
+
+ Incumbent arms that resolved nothing across every task they ran.
An incumbent arm is the currently-shipped, presumably-working skill: if it
resolves NOTHING across every task it ran, that reads as an environment or
@@ -1521,9 +1768,19 @@ def broken_incumbent_arms(
some-runs-resolved-zero case since here nothing completed at all.
aggregate() never marks an excluded/unverifiable row resolved=True, so
resolved == 0 alone already covers both cases.
+
+ A reused row proves last generation's environment worked, not this one's, so
+ the count consulted here is ``resolved_fresh``. Without that, an arm whose
+ cells were all reused always looks healthy and the canary can never fire —
+ which is exactly when a broken environment would go unnoticed. The sweep
+ keeps one paid cell per incumbent arm so this count is never vacuous.
"""
present = incumbent_arms & {arm for arms in results.values() for arm in arms}
- return sorted(arm for arm in present if all(arms[arm]["resolved"] == 0 for arms in results.values() if arm in arms))
+ return sorted(
+ arm
+ for arm in present
+ if all(arms[arm].get("resolved_fresh", arms[arm]["resolved"]) == 0 for arms in results.values() if arm in arms)
+ )
def _cost_cell(value: Any) -> str:
@@ -1755,6 +2012,15 @@ def build_parser() -> argparse.ArgumentParser:
parser.add_argument("--task-bindings-json", default=None, help=argparse.SUPPRESS)
parser.add_argument("--promotion-target-bases-json", default=None, help=argparse.SUPPRESS)
parser.add_argument("--unsafe-no-bwrap", action="store_true", help=argparse.SUPPRESS)
+ parser.add_argument(
+ "--reuse-results",
+ type=Path,
+ default=None,
+ help="prior wfbench results dir whose incumbent/CE rows may be reused "
+ "when model, effort, tasks, oracles, skill bytes, and CE plugin still "
+ "match. Candidate arms always run. Used by evolve.py so a weekly "
+ "generation does not re-pay for an unchanged comparator.",
+ )
return parser
@@ -1882,6 +2148,244 @@ def main() -> None:
gateway.__exit__(None, None, None)
+def _comparator_reuse_expectation(
+ *,
+ args: argparse.Namespace,
+ tasks: Sequence[Any],
+ task_bindings: Sequence[Mapping[str, Any]],
+ oracle_snapshots: Sequence[Any],
+ asset_snapshots: Mapping[str, Any],
+ sandbox_backend: str,
+ ce_plugin_snapshot: CePluginSnapshot | None,
+) -> ComparatorReuseExpectation:
+ """Bind this sweep's immutable identity for comparator-row reuse."""
+
+ skill_digests: dict[str, str | None] = {}
+ for arm in args.arms:
+ execution = CANDIDATE_ARMS.get(arm, arm)
+ if execution in EVALUATED_ARM_SKILLS:
+ skill_digests[arm] = skill_fingerprint(HARNESS_ROOT, execution)
+ else:
+ skill_digests[arm] = None
+ task_locks: dict[str, TaskReuseBinding] = {}
+ for task, binding, oracle in zip(tasks, task_bindings, oracle_snapshots, strict=True):
+ task_locks[str(task["id"])] = TaskReuseBinding(
+ task_base_sha=str(binding["resolved_sha"]),
+ task_prompt_digest=task_prompt_digest(task),
+ oracle_digest=oracle.digest,
+ oracle_command_digest=oracle.command_digest,
+ oracle_manifest_digest=oracle.manifest_digest,
+ task_asset_manifest_digest=getattr(
+ asset_snapshots.get(str(task["id"])), "manifest_digest", None
+ ),
+ sandbox_dependency_manifest_digest=getattr(
+ asset_snapshots.get(str(task["id"])), "dependency_manifest_digest", None
+ ),
+ )
+ return ComparatorReuseExpectation(
+ model=args.model,
+ effort=args.effort,
+ sandbox_backend=sandbox_backend,
+ runtime_digest=current_runtime_digest(),
+ now=datetime.now(UTC),
+ max_age=default_reuse_max_age(),
+ tasks=task_locks,
+ skill_digests=skill_digests,
+ ce_plugin_version=ce_plugin_snapshot.version if ce_plugin_snapshot is not None else None,
+ ce_plugin_manifest_digest=(
+ ce_plugin_snapshot.manifest_digest if ce_plugin_snapshot is not None else None
+ ),
+ )
+
+
+def task_has_planned_paid_cells(
+ task: Mapping[str, Any],
+ *,
+ arms: Sequence[str],
+ runs: int,
+ reusable_rows: Mapping[tuple[str, str, int], object],
+ reuse_source: Path | None,
+) -> bool:
+ """True when at least one planned cell is not a reusable comparator row."""
+
+ task_id = str(task["id"])
+ for run_idx in range(runs):
+ for arm in arms:
+ if reuse_source is None or (task_id, arm, run_idx) not in reusable_rows:
+ return True
+ return False
+
+
+def drop_canary_reuse_key(
+ reusable_rows: dict[CellKey, dict[str, Any]],
+ *,
+ arm: str,
+ tasks: Sequence[Mapping[str, Any]],
+ runs: int,
+) -> CellKey | None:
+ """Drop one reusable cell so an incumbent arm still measures THIS sweep.
+
+ An arm reused end to end measures nothing about today's environment, and
+ arm_health would then be reading last week's health.
+
+ Counted against the cells this sweep PLANS, not every key reuse selection
+ returned: selection accepts any non-negative prior run index, so a results
+ directory produced with more runs than this invocation leaves extra keys.
+ Comparing against those made the check false exactly when it mattered, and
+ the canary silently stopped firing while every planned cell stayed reused.
+
+ Returns the dropped key, or None when the arm already has a paid cell.
+ """
+
+ planned_keys = [(str(task["id"]), arm, run_idx) for task in tasks for run_idx in range(runs)]
+ arm_keys = sorted(key for key in planned_keys if key in reusable_rows)
+ if not arm_keys or len(arm_keys) != len(planned_keys):
+ return None
+ dropped = arm_keys[0]
+ del reusable_rows[dropped]
+ return dropped
+
+
+def next_graph_prefetch_target(
+ remaining: Sequence[tuple[Mapping[str, Any], Mapping[str, Any]]],
+ *,
+ arms: Sequence[str],
+ runs: int,
+ reusable_rows: Mapping[tuple[str, str, int], object],
+ reuse_source: Path | None,
+ ready_keys: set[tuple[str, str]],
+) -> tuple[Mapping[str, Any], Mapping[str, Any], tuple[str, str]] | None:
+ """Next later task that still needs a clone template and sanitized graph."""
+
+ for task, binding in remaining:
+ if not task_has_planned_paid_cells(
+ task,
+ arms=arms,
+ runs=runs,
+ reusable_rows=reusable_rows,
+ reuse_source=reuse_source,
+ ):
+ continue
+ key = (str(binding["repo_identity"]), str(binding["resolved_sha"]))
+ if key in ready_keys:
+ continue
+ return task, binding, key
+ return None
+
+
+@dataclass(frozen=True)
+class GraphBuildEnv:
+ """Per-sweep state every graph build shares, and the caches it fills.
+
+ The four dicts are the sweep's memo of what has already been built, keyed by
+ (repo, sha). They are mutable by design and are written by both the sweep
+ thread and the prefetch thread, which is safe only because a build is
+ started for a key exactly once and joined before that key is read.
+ """
+
+ trees: Path
+ task_asset_cache: TaskAssetCache
+ claude_bin: Path | str
+ bwrap_bin: Path | str
+ sandbox_backend: str
+ runtime_mounts: Sequence[ReadOnlyMount]
+ clone_templates: dict[tuple[str, str], tuple[Path, str]]
+ clone_template_errors: dict[tuple[str, str], BaseException]
+ graph_snapshots: dict[tuple[str, str], SanitizedGraphSnapshot]
+ graph_snapshot_errors: dict[tuple[str, str], BaseException]
+
+ def ready_keys(self) -> set[tuple[str, str]]:
+ """Keys whose build has already been attempted, successfully or not."""
+
+ return (
+ set(self.clone_templates)
+ | set(self.clone_template_errors)
+ | set(self.graph_snapshots)
+ | set(self.graph_snapshot_errors)
+ )
+
+
+def ensure_task_graph(
+ *,
+ task: Mapping[str, Any],
+ repo: Path,
+ task_sha: str,
+ graph_key: tuple[str, str],
+ env: GraphBuildEnv,
+) -> None:
+ """Build one SHA's sanitized clone template and graph. Idempotent per key."""
+
+ if graph_key in env.graph_snapshots or graph_key in env.graph_snapshot_errors:
+ return
+ try:
+ validate_no_prebuilt_graph_assets(task)
+ if graph_key not in env.clone_templates and graph_key not in env.clone_template_errors:
+ template = make_worktree(repo, task_sha, env.trees)
+ template_head = sanitize_clone_for_hidden_oracles(template)
+ env.clone_templates[graph_key] = (template, template_head)
+ clone_template: Path | None = None
+ template_head: str | None = None
+ if graph_key in env.clone_templates:
+ clone_template, template_head = env.clone_templates[graph_key]
+ if graph_key in env.clone_template_errors:
+ env.graph_snapshot_errors[graph_key] = env.clone_template_errors[graph_key]
+ return
+ env.graph_snapshots[graph_key] = prepare_sanitized_graph(
+ task,
+ repo=repo,
+ resolved_sha=task_sha,
+ parent=env.trees,
+ cache=env.task_asset_cache,
+ claude_bin=env.claude_bin,
+ bwrap_bin=env.bwrap_bin,
+ sandbox_backend=env.sandbox_backend,
+ runtime_mounts=env.runtime_mounts,
+ clone_template=clone_template,
+ sanitized_head=template_head,
+ )
+ except (ManagedProcessError, OSError, SandboxError, RuntimeError, ValueError) as exc:
+ env.graph_snapshot_errors[graph_key] = exc
+ env.clone_template_errors.setdefault(graph_key, exc)
+
+
+@dataclass
+class GraphPrefetch:
+ """In-flight clone+graph build for a later task SHA."""
+
+ key: tuple[str, str]
+ thread: threading.Thread
+
+ def join(self) -> None:
+ self.thread.join()
+
+
+def prefetch_next_graph(
+ *,
+ task: Mapping[str, Any],
+ binding: Mapping[str, Any],
+ graph_key: tuple[str, str],
+ env: GraphBuildEnv,
+ cancel_event: threading.Event,
+) -> GraphPrefetch:
+ """Start clone+graph prep for the next unpaid SHA during paid sessions."""
+
+ repo = Path(binding["repo_identity"])
+ task_sha = str(binding["resolved_sha"])
+
+ def run() -> None:
+ if cancel_event.is_set():
+ return
+ print(f"[prefetch_next_graph] clone+graph for {task_sha}")
+ ensure_task_graph(task=task, repo=repo, task_sha=task_sha, graph_key=graph_key, env=env)
+
+ # copy_context, as the worker pool already does at _run_wave: a plain Thread
+ # does not inherit ContextVars, so without this every run_managed inside the
+ # graph build resolves _CANCELLATION to None and ignores the shared cancel.
+ thread = threading.Thread(target=copy_context().run, args=(run,), name="prefetch_next_graph", daemon=False)
+ thread.start()
+ return GraphPrefetch(key=graph_key, thread=thread)
+
+
def _run_sweep(
args: argparse.Namespace,
*,
@@ -1938,107 +2442,213 @@ def _run_sweep(
except (OSError, SandboxError, ValueError) as exc:
parser.error(str(exc))
raise AssertionError("ArgumentParser.error() returned unexpectedly")
- graph_snapshots: dict[tuple[str, str], SanitizedGraphSnapshot] = {}
- graph_snapshot_errors: dict[tuple[str, str], BaseException] = {}
- for task, task_binding, oracle_snapshot in zip(
- tasks,
- task_bindings,
- oracle_snapshots,
- strict=True,
- ):
- if outage_tripped or cancel_event.is_set():
- break
- repo = Path(task_binding["repo_identity"])
- task_sha = task_binding["resolved_sha"]
- asset_snapshot: TaskAssetSnapshot | None = None
- asset_snapshot_error: BaseException | None = None
- graph_key = (str(repo), task_sha)
- graph_snapshot: SanitizedGraphSnapshot | None = graph_snapshots.get(graph_key)
- graph_snapshot_error: BaseException | None = graph_snapshot_errors.get(graph_key)
+ # Asset snapshots are built here, up front, for two reasons. Comparator
+ # reuse has to compare this sweep's task-asset and dependency digests
+ # against the prior row's, and those digests do not exist until the
+ # snapshot does. Building them all before any cell or prefetch thread
+ # starts also keeps TaskAssetCache single-threaded, which is what its own
+ # "plain dict, read-then-write race" comment asks for.
+ asset_snapshots: dict[str, TaskAssetSnapshot] = {}
+ asset_snapshot_errors: dict[str, BaseException] = {}
+ for _task, _binding in zip(tasks, task_bindings, strict=True):
try:
- validate_no_prebuilt_graph_assets(task)
- if graph_snapshot is None and graph_snapshot_error is None:
- graph_snapshot = prepare_sanitized_graph(
- task,
- repo=repo,
- resolved_sha=task_sha,
- parent=Path(trees),
- cache=task_asset_cache,
- claude_bin=args.claude_bin,
- bwrap_bin=bwrap_bin,
- sandbox_backend=sandbox_backend,
- runtime_mounts=runtime_mounts,
- )
- graph_snapshots[graph_key] = graph_snapshot
- except (ManagedProcessError, OSError, SandboxError, RuntimeError, ValueError) as exc:
- graph_snapshot_error = exc
- graph_snapshot_errors[graph_key] = exc
- try:
- # Prepared here, once, rather than lazily inside the first cell:
- # TaskAssetCache is a plain dict, so a lazy build would be a
- # read-then-write race the moment cells stop running serially.
- asset_snapshot = task_asset_cache.prepare(
- task,
- repo=repo,
- resolved_sha=task_sha,
- expected_dependency_binding=task_binding,
+ asset_snapshots[str(_task["id"])] = task_asset_cache.prepare(
+ _task,
+ repo=Path(_binding["repo_identity"]),
+ resolved_sha=_binding["resolved_sha"],
+ expected_dependency_binding=_binding,
)
except (OSError, SandboxError, ValueError) as exc:
- asset_snapshot_error = exc
- cell_context = TaskCellContext(
- task=task,
- oracle_snapshot=oracle_snapshot,
- repo=repo,
- task_sha=task_sha,
- graph_snapshot=graph_snapshot,
- graph_snapshot_error=graph_snapshot_error,
- asset_snapshot=asset_snapshot,
- asset_snapshot_error=asset_snapshot_error,
- args=args,
- out_dir=out_dir,
- ce_plugin_snapshot=ce_plugin_snapshot,
- trees_dir=Path(trees),
- bwrap_bin=bwrap_bin,
- sandbox_backend=sandbox_backend,
- runtime_mounts=runtime_mounts,
- candidate_overlay=candidate_overlay,
- overlay_digest=overlay_digest,
- )
- per_arm: dict[str, list[dict[str, Any]]] = {a: [] for a in args.arms}
- cells = [(run_idx, arm) for run_idx in range(args.runs) for arm in args.arms]
+ asset_snapshot_errors[str(_task["id"])] = exc
- def announce(run_idx: int, arm: str) -> None:
- nonlocal started_cells
- started_cells += 1
- print(
- f"[{task['id']}][{arm}][run {run_idx}] starting "
- f"({started_cells}/{total_cells}, {(time.monotonic() - sweep_started) / 60:.0f}m elapsed)"
+ reuse_source = args.reuse_results.expanduser().resolve() if args.reuse_results is not None else None
+ reusable_rows: dict[tuple[str, str, int], dict[str, Any]] = {}
+ if reuse_source is not None:
+ if reuse_source == out_dir.resolve():
+ parser.error("--reuse-results cannot be this sweep's --out directory")
+ raise AssertionError("ArgumentParser.error() returned unexpectedly")
+ results_file = reuse_source / "results.jsonl"
+ if results_file.is_symlink() or not results_file.is_file():
+ parser.error("--reuse-results must contain a regular results.jsonl")
+ raise AssertionError("ArgumentParser.error() returned unexpectedly")
+ reusable_rows = select_reusable_comparator_rows(
+ load_result_rows(results_file),
+ expected=_comparator_reuse_expectation(
+ args=args,
+ tasks=tasks,
+ task_bindings=task_bindings,
+ oracle_snapshots=oracle_snapshots,
+ asset_snapshots=asset_snapshots,
+ sandbox_backend=sandbox_backend,
+ ce_plugin_snapshot=ce_plugin_snapshot,
+ ),
+ )
+ # Keep one paid cell per incumbent arm. A generation that reuses an
+ # arm end to end measures nothing about today's environment, and
+ # arm_health would then be reading last week's health.
+ # One cell per arm is the cheapest thing that keeps the canary real.
+ incumbent_arms = [arm for arm in args.arms if arm not in candidate_arms]
+ for arm in incumbent_arms:
+ dropped = drop_canary_reuse_key(reusable_rows, arm=arm, tasks=tasks, runs=args.runs)
+ if dropped is not None:
+ print(
+ f"reuse-results: keeping one paid {arm} cell "
+ f"({dropped[0]} run {dropped[2]}) so incumbent health is measured this sweep"
+ )
+ print(
+ f"reuse-results {reuse_source}: {len(reusable_rows)} comparator "
+ f"cell(s) match this sweep; candidate arms always run"
+ )
+ graph_env = GraphBuildEnv(
+ trees=Path(trees),
+ task_asset_cache=task_asset_cache,
+ claude_bin=args.claude_bin,
+ bwrap_bin=bwrap_bin,
+ sandbox_backend=sandbox_backend,
+ runtime_mounts=runtime_mounts,
+ clone_templates={},
+ clone_template_errors={},
+ graph_snapshots={},
+ graph_snapshot_errors={},
+ )
+ graph_prefetch: GraphPrefetch | None = None
+ sweep_rows = list(zip(tasks, task_bindings, oracle_snapshots, strict=True))
+
+ def _join_graph_prefetch() -> None:
+ nonlocal graph_prefetch
+ if graph_prefetch is not None:
+ graph_prefetch.join()
+ graph_prefetch = None
+
+ try:
+ for index, (task, task_binding, oracle_snapshot) in enumerate(sweep_rows):
+ if outage_tripped or cancel_event.is_set():
+ break
+ repo = Path(task_binding["repo_identity"])
+ task_sha = task_binding["resolved_sha"]
+ per_arm: dict[str, list[dict[str, Any]]] = {a: [] for a in args.arms}
+ planned = [(run_idx, arm) for run_idx in range(args.runs) for arm in args.arms]
+ reused_records: list[tuple[int, str, dict[str, Any]]] = []
+ paid_cells: list[tuple[int, str]] = []
+ for run_idx, arm in planned:
+ prior = reusable_rows.get((task["id"], arm, run_idx))
+ if prior is None or reuse_source is None:
+ paid_cells.append((run_idx, arm))
+ continue
+ try:
+ reused_records.append(
+ (
+ run_idx,
+ arm,
+ materialize_reused_row(prior, source_dir=reuse_source, dest_dir=out_dir),
+ )
+ )
+ except (OSError, SandboxError, ValueError) as exc:
+ print(
+ f"[{task['id']}][{arm}][run {run_idx}] comparator reuse "
+ f"failed ({exc}); running a paid cell"
+ )
+ paid_cells.append((run_idx, arm))
+
+ asset_snapshot = asset_snapshots.get(str(task["id"]))
+ asset_snapshot_error: BaseException | None = asset_snapshot_errors.get(str(task["id"]))
+ graph_key = (str(repo), task_sha)
+ if graph_prefetch is not None and graph_prefetch.key == graph_key:
+ _join_graph_prefetch()
+ if paid_cells:
+ ensure_task_graph(
+ task=task, repo=repo, task_sha=task_sha, graph_key=graph_key, env=graph_env
+ )
+ graph_snapshot = graph_env.graph_snapshots.get(graph_key)
+ graph_snapshot_error = graph_env.graph_snapshot_errors.get(graph_key)
+ clone_template, template_head = graph_env.clone_templates.get(graph_key, (None, None))
+ cell_context = TaskCellContext(
+ task=task,
+ oracle_snapshot=oracle_snapshot,
+ repo=repo,
+ task_sha=task_sha,
+ graph_snapshot=graph_snapshot,
+ graph_snapshot_error=graph_snapshot_error,
+ asset_snapshot=asset_snapshot,
+ asset_snapshot_error=asset_snapshot_error,
+ args=args,
+ out_dir=out_dir,
+ ce_plugin_snapshot=ce_plugin_snapshot,
+ trees_dir=Path(trees),
+ bwrap_bin=bwrap_bin,
+ sandbox_backend=sandbox_backend,
+ runtime_mounts=runtime_mounts,
+ candidate_overlay=candidate_overlay,
+ overlay_digest=overlay_digest,
+ clone_template=clone_template,
+ sanitized_head=template_head,
)
- def keep(run_idx: int, arm: str, record: dict[str, Any]) -> None:
- per_arm[arm].append(record)
- with results_path.open("a") as fh:
- # Redact any API token a session-error stderr_tail echoed
- # into error_detail before it enters the uploaded
- # results.jsonl artifact (transcripts are redacted; this
- # sink was not).
- fh.write(redact_text(json.dumps(record), credential_secrets(args)) + "\n")
- print(cell_progress_line(task["id"], arm, run_idx, record))
- failure = cell_failure_detail_line(task["id"], arm, run_idx, record, credential_secrets(args))
- if failure:
- print(failure)
+ def announce(run_idx: int, arm: str) -> None:
+ nonlocal started_cells
+ started_cells += 1
+ print(
+ f"[{task['id']}][{arm}][run {run_idx}] starting "
+ f"({started_cells}/{total_cells}, {(time.monotonic() - sweep_started) / 60:.0f}m elapsed)"
+ )
- outage_streak, outage_tripped = sweep_task_cells(
- cells,
- workers=args.workers,
- run=partial(run_cell, cell_context),
- on_start=announce,
- on_record=keep,
- outage_streak=outage_streak,
- outage_limit=args.outage_streak,
- cancel_event=cancel_event,
- )
- results[task["id"]] = {a: aggregate(rs) for a, rs in per_arm.items() if rs}
+ def keep(run_idx: int, arm: str, record: dict[str, Any]) -> None:
+ per_arm[arm].append(record)
+ with results_path.open("a") as fh:
+ # Redact any API token a session-error stderr_tail echoed
+ # into error_detail before it enters the uploaded
+ # results.jsonl artifact (transcripts are redacted; this
+ # sink was not).
+ fh.write(redact_text(json.dumps(record), credential_secrets(args)) + "\n")
+ print(cell_progress_line(task["id"], arm, run_idx, record))
+ failure = cell_failure_detail_line(task["id"], arm, run_idx, record, credential_secrets(args))
+ if failure:
+ print(failure)
+
+ for run_idx, arm, record in reused_records:
+ started_cells += 1
+ print(
+ f"[{task['id']}][{arm}][run {run_idx}] reused comparator "
+ f"({started_cells}/{total_cells}, {(time.monotonic() - sweep_started) / 60:.0f}m elapsed)"
+ )
+ keep(run_idx, arm, record)
+ if paid_cells and graph_prefetch is None and not cancel_event.is_set():
+ target = next_graph_prefetch_target(
+ [(later_task, later_binding) for later_task, later_binding, _ in sweep_rows[index + 1 :]],
+ arms=args.arms,
+ runs=args.runs,
+ reusable_rows=reusable_rows,
+ reuse_source=reuse_source,
+ ready_keys=graph_env.ready_keys(),
+ )
+ if target is not None:
+ later_task, later_binding, later_key = target
+ graph_prefetch = prefetch_next_graph(
+ task=later_task,
+ binding=later_binding,
+ graph_key=later_key,
+ env=graph_env,
+ cancel_event=cancel_event,
+ )
+ # A reused success is evidence the pipeline can produce a good row,
+ # so it resets the consecutive-failure count the same way a paid
+ # success does. Leaving reused rows out let a streak carry across
+ # them and trip on stale history.
+ for _run_idx, _arm, _record in reused_records:
+ outage_streak = systemic_outage_streak(_record.get("error_kind"), outage_streak)
+ outage_streak, outage_tripped = sweep_task_cells(
+ paid_cells,
+ workers=args.workers,
+ run=partial(run_cell, cell_context),
+ on_start=announce,
+ on_record=keep,
+ outage_streak=outage_streak,
+ outage_limit=args.outage_streak,
+ cancel_event=cancel_event,
+ )
+ results[task["id"]] = {a: aggregate(rs) for a, rs in per_arm.items() if rs}
+ finally:
+ _join_graph_prefetch()
selection_report = [
"## Run provenance",
@@ -2056,10 +2666,12 @@ def _run_sweep(
selection_report.append(
f"Compound Engineering plugin: `{ce_plugin_snapshot.version}` (`{ce_plugin_snapshot.manifest_digest}`)"
)
- if cancel_event.is_set():
- selection_report.append("Sweep cancelled: partial evidence; promotion is disabled.")
- elif outage_tripped:
+ # outage first: the breaker now sets cancel_event to stop in-flight background
+ # work, so testing cancellation first would relabel every outage a cancellation.
+ if outage_tripped:
selection_report.append("Sweep aborted: partial evidence; promotion is disabled.")
+ elif cancel_event.is_set():
+ selection_report.append("Sweep cancelled: partial evidence; promotion is disabled.")
report = render_report(results) + "\n\n" + "\n".join(selection_report) + "\n"
(out_dir / "report.md").write_text(report)
if candidate_arms:
@@ -2092,26 +2704,22 @@ def _run_sweep(
}
(out_dir / "promotion.json").write_text(json.dumps(promotion, indent=2) + "\n")
print(f"\n{report}\n\nWritten to {out_dir}/")
- # A reviewer may legitimately match none of a difficult hidden corpus;
- # unlike an implementation arm, zero exact resolutions is quality signal,
- # not proof that the harness failed.
- broken_incumbents = broken_incumbent_arms(results, set(CANDIDATE_ARMS.values()) - {"review"})
- if broken_incumbents:
- # Fail loudly rather than let a broken environment read as a quiet
- # "no promotion, incumbent stands."
- print(
- f"[harness-health] incumbent arm(s) {', '.join(broken_incumbents)} resolved zero "
- "tasks across every valid run — this looks like an environment/harness failure, "
- "not a normal candidate miss. See the errors column in report.md and error_detail "
- "in results.jsonl. Exiting non-zero rather than reporting a quiet no-promotion."
- )
- raise SystemExit(1)
- if cancel_event.is_set():
- raise SystemExit(130)
+ # Health is judged on whether fresh attempts EXECUTED and produced usable
+ # evidence - never on how many tasks they resolved. Review arms used to be
+ # excluded here because "resolved zero" is quality signal for a reviewer
+ # facing a hard corpus; with the inference corrected they are included
+ # again, which is what lets an all-artifacts-empty run be caught at all.
+ # ce_review is named explicitly: it is a comparator, not a candidate, so it
+ # is absent from CANDIDATE_ARMS and would otherwise go unclassified.
+ enforce_measurement_health(results, set(CANDIDATE_ARMS.values()) | {"review", "ce_review"})
if outage_tripped:
# Non-zero exit so a driver (evolve.py) treats the partial benchmark as a
# failed run and halts instead of proposing from outage-truncated evidence.
+ # Checked before cancel_event because the breaker sets it (see above), and
+ # an outage must keep exit 1 rather than becoming the 130 of a Ctrl-C.
raise SystemExit(1)
+ if cancel_event.is_set():
+ raise SystemExit(130)
if __name__ == "__main__":
diff --git a/eval/workflow_bench/runner_artifacts.py b/eval/workflow_bench/runner_artifacts.py
index 7f05e14b1..3780b9e46 100644
--- a/eval/workflow_bench/runner_artifacts.py
+++ b/eval/workflow_bench/runner_artifacts.py
@@ -217,11 +217,25 @@ def enforce_phase_workspace(
worktree: Path,
before: dict[str, str],
*,
- allowed_artifact: Path,
+ allowed_artifact: Path | None,
) -> None:
- """Require a phase to change only its one explicit workspace artifact."""
+ """Require a phase to change only its one explicit workspace artifact.
+
+ ``allowed_artifact=None`` is the stricter contract: the phase must leave
+ the workspace byte-identical. That is what a review phase whose artifact
+ lives outside the workspace has to satisfy — there is nothing in there it
+ is entitled to touch.
+ """
root = worktree.expanduser().absolute()
+ if allowed_artifact is None:
+ after = workspace_snapshot(root)
+ changed = sorted(
+ path for path in before.keys() | after.keys() if before.get(path) != after.get(path)
+ )
+ if changed:
+ raise ValueError(f"phase changed the read-only workspace: {', '.join(changed[:5])}")
+ return
artifact = allowed_artifact.expanduser().absolute()
try:
relative = PurePosixPath(artifact.relative_to(root).as_posix())
@@ -325,6 +339,65 @@ def new_plan_doc(worktree: Path, before: dict[Path, str]) -> Path:
return changed[0]
+def _assert_self_contained_git_objects(clone: Path) -> None:
+ """Refuse clones that share pack/object bytes with another repository."""
+
+ alternates = clone / ".git" / "objects" / "info" / "alternates"
+ if alternates.exists():
+ raise RuntimeError(f"clone unexpectedly has an external object alternate: {alternates}")
+ objects = clone / ".git" / "objects"
+ if not objects.is_dir():
+ raise RuntimeError(f"clone is missing a git object store: {clone}")
+ for obj in objects.rglob("*"):
+ if obj.is_file() and obj.stat().st_nlink > 1:
+ raise RuntimeError(f"clone object is hardlinked to host storage: {obj}")
+
+
+def copy_isolated_tree(source: Path, parent: Path) -> Path:
+ """Copy a sanitized clone without sharing git objects or a ref namespace.
+
+ ``git clone --no-local`` of GitNexus plus ``sanitize_clone_for_hidden_oracles``
+ (repack/prune/fsck) is minutes per cell. After sanitization the snapshot is
+ one parentless commit; copying that tree is the isolation boundary the
+ contamination bug actually required (a private ``.git``), not a second
+ fetch of full history. Prefer ``cp --reflink=auto`` so XFS/btrfs pay COW;
+ fall back to a full copy on filesystems that cannot reflink.
+ """
+
+ try:
+ source_meta = source.expanduser().lstat()
+ except OSError as exc:
+ raise RuntimeError(f"clone template is unavailable: {source}: {exc}") from exc
+ if stat.S_ISLNK(source_meta.st_mode) or not stat.S_ISDIR(source_meta.st_mode):
+ raise RuntimeError(f"clone template must be a real directory: {source}")
+ source = source.expanduser().resolve()
+ target = Path(tempfile.mkdtemp(prefix="wfbench-", dir=parent))
+ target.rmdir()
+ try:
+ copied = run_managed(
+ ["cp", "-a", "--reflink=auto", str(source), str(target)],
+ timeout=600,
+ )
+ if not copied.ok:
+ # The fallback is for a filesystem that cannot reflink, which shows
+ # up as a normal nonzero exit. A cancellation or timeout is reported
+ # the same way (run_managed returns it rather than raising), and
+ # copytree cannot be cancelled — so falling back there makes the
+ # outage breaker wait out the full copy it set the event to avoid.
+ if copied.state != "exited":
+ raise ManagedProcessError(["cp", "-a", "--reflink=auto", str(source), str(target)], copied)
+ shutil.copytree(source, target, symlinks=True, copy_function=shutil.copy2)
+ _assert_self_contained_git_objects(target)
+ return target
+ except BaseException as primary:
+ if target.exists():
+ try:
+ shutil.rmtree(target)
+ except OSError as cleanup:
+ primary.add_note(f"clone copy cleanup also failed: {type(cleanup).__name__}: {cleanup}")
+ raise
+
+
def make_worktree(repo: Path, ref: str, parent: Path) -> Path:
"""Create a self-contained clone per benchmark arm."""
@@ -344,12 +417,7 @@ def make_worktree(repo: Path, ref: str, parent: Path) -> Path:
],
timeout=600,
)
- alternates = target / ".git" / "objects" / "info" / "alternates"
- if alternates.exists():
- raise RuntimeError(f"clone unexpectedly has an external object alternate: {alternates}")
- for obj in (target / ".git" / "objects").rglob("*"):
- if obj.is_file() and obj.stat().st_nlink > 1:
- raise RuntimeError(f"clone object is hardlinked to host storage: {obj}")
+ _assert_self_contained_git_objects(target)
for candidate in (ref, f"origin/{ref}"):
proc = run_managed(
["git", "-C", str(target), "checkout", "--detach", "--quiet", candidate],
diff --git a/eval/workflow_bench/runner_sessions.py b/eval/workflow_bench/runner_sessions.py
index 9daef00b4..053d5e48d 100644
--- a/eval/workflow_bench/runner_sessions.py
+++ b/eval/workflow_bench/runner_sessions.py
@@ -57,6 +57,24 @@ MAX_PROGRESS_PENDING = 256
MAX_PROGRESS_TOOL_ID_CHARS = 256
MAX_TOOL_PREVIEW_CHARS = 800
_SAFE_TOOL_NAME = re.compile(r"[A-Za-z0-9._:-]{1,64}")
+_GHA_WORKFLOW_COMMAND = re.compile(r"(^|[\n\r])::")
+_GHA_HASH_COMMAND = re.compile(r"##\[")
+_GHA_COMPILER_ANNOTATION = re.compile(r"\((\d+),(\d+)\):\s+error\b", re.IGNORECASE)
+
+
+def neutralize_ci_log_text(text: str) -> str:
+ """Stop GitHub Actions from promoting tool output into check annotations.
+
+ Run 33962002890 logged in-sandbox ``tsc`` failures as
+ ``file.ts(line,col): error TS2307``, which Actions parsed as workflow
+ annotations on ``.github``. The same parser treats ``::error::`` and
+ ``##[error]`` as commands. Progress previews are evidence, not CI
+ signaling, so rewrite those forms before they hit the job log.
+ """
+
+ text = _GHA_WORKFLOW_COMMAND.sub(r"\1[:]", text)
+ text = _GHA_HASH_COMMAND.sub("# [", text)
+ return _GHA_COMPILER_ANNOTATION.sub(r"(\1,\2): compiler-error", text)
def _safe_tool_name(value: Any) -> str:
@@ -177,7 +195,9 @@ class SessionProgress:
def _say(self, message: str) -> None:
# Queue only: the stdout drain thread calls observe() and must not
# block on a full log pipe (process_control.stdout_observer contract).
- self._pending_messages.append(f"[{self.label} {self._elapsed()}] {message}")
+ self._pending_messages.append(
+ neutralize_ci_log_text(f"[{self.label} {self._elapsed()}] {message}")
+ )
self._last_spoke = time.monotonic()
def _emit_pending(self) -> None:
diff --git a/eval/workflow_bench/sanitized_graph.py b/eval/workflow_bench/sanitized_graph.py
index 3aea9e8c1..fc3d766a2 100644
--- a/eval/workflow_bench/sanitized_graph.py
+++ b/eval/workflow_bench/sanitized_graph.py
@@ -24,7 +24,7 @@ from .proposer_sandbox import (
build_sandbox_environment,
prepare_sandbox,
)
-from .runner_artifacts import make_worktree, remove_clone
+from .runner_artifacts import copy_isolated_tree, make_worktree, remove_clone
from .task_assets import TaskAssetCache, TaskAssetSnapshot, _is_harness_sandbox_copy
GRAPH_ASSET_PATHS = (
@@ -360,14 +360,29 @@ def prepare_sanitized_graph(
bwrap_bin: Path | str,
runtime_mounts: Sequence[ReadOnlyMount],
sandbox_backend: str = "bwrap",
+ clone_template: Path | None = None,
+ sanitized_head: str | None = None,
) -> SanitizedGraphSnapshot:
- """Sanitize, index offline once, scrub, and freeze graph assets for all arms."""
+ """Sanitize, index offline once, scrub, and freeze graph assets for all arms.
+
+ When ``clone_template`` is an already-sanitized snapshot, this copies it
+ (the copy is scrubbed and indexed) so the template stays a clean cell
+ seed. Callers that already paid for ``make_worktree`` + sanitization
+ should pass that template rather than cloning GitNexus again.
+ """
validate_no_prebuilt_graph_assets(task)
- seed = make_worktree(repo, resolved_sha, parent)
+ if clone_template is not None:
+ if not isinstance(sanitized_head, str) or not sanitized_head:
+ raise SandboxError("clone template requires the sanitized HEAD")
+ seed = copy_isolated_tree(clone_template, parent)
+ else:
+ seed = make_worktree(repo, resolved_sha, parent)
+ sanitized_head = None
primary: BaseException | None = None
try:
- sanitized_head = sanitize_clone_for_hidden_oracles(seed)
+ if sanitized_head is None:
+ sanitized_head = sanitize_clone_for_hidden_oracles(seed)
_scrub_source_references(seed)
_neutralize_target_index_inputs(seed)
with prepare_sandbox(
diff --git a/eval/workflow_bench/tasks.review.scenarios.yaml b/eval/workflow_bench/tasks.review.scenarios.yaml
index 8ff323ed1..4d68ae871 100644
--- a/eval/workflow_bench/tasks.review.scenarios.yaml
+++ b/eval/workflow_bench/tasks.review.scenarios.yaml
@@ -9,14 +9,17 @@ tasks:
sandbox_copy: [eval/workflow_bench/review_cases/pr-2718.patch]
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2718.patch && rm -rf eval/workflow_bench
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2718. Report only actionable defects introduced by the local diff.
- verify: test -s review-output.json
+ verify: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2718-defect.labels.json, target: review-labels.json }]
sandbox_dependencies: &deps
- { source: node_modules, target: node_modules }
- { source: gitnexus/node_modules, target: gitnexus/node_modules }
- { source: gitnexus-shared/node_modules, target: gitnexus-shared/node_modules }
+ # Host-built types/JS. Historical clones have no dist/, so `tsc` in the
+ # read-only workspace otherwise reports TS2307/TS6379 (run 33962002890).
+ - { source: gitnexus-shared/dist, target: gitnexus-shared/dist }
- <<: *review_case
id: review-pr-2794-defect
@@ -25,7 +28,7 @@ tasks:
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2794.patch && rm -rf eval/workflow_bench
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2794. Report only actionable defects introduced by the local diff.
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2794-defect.labels.json, target: review-labels.json }]
sandbox_dependencies: *deps
@@ -36,7 +39,7 @@ tasks:
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2108.patch && rm -rf eval/workflow_bench
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2108. Report only actionable defects introduced by the local diff.
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2108-defect.labels.json, target: review-labels.json }]
sandbox_dependencies: *deps
@@ -47,7 +50,7 @@ tasks:
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2258.patch && rm -rf eval/workflow_bench
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2258. Report only actionable defects introduced by the local diff.
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2258-defect.labels.json, target: review-labels.json }]
sandbox_dependencies: *deps
@@ -59,7 +62,7 @@ tasks:
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2258b.patch && rm -rf eval/workflow_bench
prompt: Review this historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2258. Report only actionable defects introduced by the local diff.
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2258-clean.labels.json, target: review-labels.json }]
sandbox_dependencies: *deps
@@ -71,6 +74,6 @@ tasks:
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2773.patch && rm -rf eval/workflow_bench
prompt: Review this historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2773. Report only actionable defects introduced by the local diff.
oracle:
- command: test -s review-output.json
+ command: test -s "$GITNEXUS_BENCH_REVIEW_OUTPUT"
files: [{ source: review-pr-2773-clean.labels.json, target: review-labels.json }]
sandbox_dependencies: *deps
diff --git a/eval/workflow_bench/tasks.scenarios.yaml b/eval/workflow_bench/tasks.scenarios.yaml
index af4755f88..59448ae4f 100644
--- a/eval/workflow_bench/tasks.scenarios.yaml
+++ b/eval/workflow_bench/tasks.scenarios.yaml
@@ -44,6 +44,11 @@ tasks:
target: gitnexus/node_modules
- source: gitnexus-shared/node_modules
target: gitnexus-shared/node_modules
+ # Host-built types/JS. The clone has no dist/, and node_modules/gitnexus-shared
+ # is a relative symlink into that unbuilt tree — without this mount, in-sandbox
+ # `tsc --noEmit` / vitest fail with TS2307 / TS6379 (run 33962002890).
+ - source: gitnexus-shared/dist
+ target: gitnexus-shared/dist
prompt: >
Add -j as a short alias for --json on the gitnexus status command
(gitnexus/src/cli/index.ts), and cover the alias with a unit test in
diff --git a/gitnexus/test/unit/skill-evolution-workflow.test.ts b/gitnexus/test/unit/skill-evolution-workflow.test.ts
index b10313e42..6916c7028 100644
--- a/gitnexus/test/unit/skill-evolution-workflow.test.ts
+++ b/gitnexus/test/unit/skill-evolution-workflow.test.ts
@@ -270,12 +270,13 @@ describe('gitnexus skill-evolution workflow contract', () => {
});
it('passes the cell concurrency through to the benchmark', () => {
- // The lane is serial unless told otherwise: concurrency only pays off when
- // the runner has the vCPUs for it, and a cell starved of CPU drifts toward
- // its session timeout, which the gate counts as an excluded run.
+ // Dispatch defaults to 3. Scheduled runs still fall back to serial unless
+ // GITNEXUS_EVOLUTION_WORKERS is set — a cell starved of CPU that hits the
+ // session ceiling is an excluded run the gate refuses.
expect(evolveJob?.env?.WORKERS).toBe(
"${{ inputs.workers || vars.GITNEXUS_EVOLUTION_WORKERS || '1' }}",
);
+ expect(workflow).toMatch(/workers:\n(?:[^\n]*\n){0,4} default: '3'/);
});
it('seeds from the newest usable completed main run, including failed runs', () => {
@@ -572,6 +573,27 @@ exit 1`);
// uploads. The job must finish inside that window even when the schedule
// fires late (the 2026-08-01 run was queued 65 minutes after the cron).
expect(jobBudget as number).toBeLessThanOrEqual(21 * 60);
+ // A Friday dispatch inherits leftover uptime. The shared entrypoint — not
+ // the workflow YAML — must cap the sweep so it fails in-process and the
+ // always() upload still runs (run 33962002890).
+ const script = readFileSync(
+ path.join(REPO_ROOT, 'eval/workflow_bench/run-evolution.sh'),
+ 'utf8',
+ );
+ // The flag, not a precomputed number: the CLI reads /proc/uptime in the
+ // same breath as it starts the clock the cap is measured against, so
+ // nothing between the two can be charged to the sweep.
+ expect(script).toContain('--max-runtime-from-instance-window');
+ expect(script).not.toContain('instance_window_budget_from_proc');
+ expect(script).toContain('export RUNTIME_DIGEST');
+ // The workflow's half of that contract is calling the entrypoint, not
+ // naming the flag. The YAML never mentions --max-runtime-from-instance-window
+ // at all — only the prose at gitnexus-skill-evolution.yml:179-184 describing
+ // the cap — so an assertion on flag text there would test a comment, and a
+ // loop step that had stopped invoking the script would still pass it.
+ expect(stepRun('Run the propose → benchmark → gate loop')).toContain(
+ './workflow_bench/run-evolution.sh --apply',
+ );
});
it('uploads benchmark evidence unconditionally, on a path it addresses itself', () => {
From 4757c0d3cbd65e8ec90ec13f771e50a29f425d2e Mon Sep 17 00:00:00 2001
From: "dependabot[bot]" <49699333+dependabot[bot]@users.noreply.github.com>
Date: Tue, 8 Sep 2026 10:58:41 +0100
Subject: [PATCH 06/53] chore(deps)(deps-dev): bump @types/node in /gitnexus
(#3215)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 26.4.0 to 26.4.1.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)
---
updated-dependencies:
- dependency-name: "@types/node"
dependency-version: 26.4.1
dependency-type: direct:development
update-type: version-update:semver-patch
...
Signed-off-by: dependabot[bot]
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar
---
gitnexus/package-lock.json | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/gitnexus/package-lock.json b/gitnexus/package-lock.json
index 6e521f664..caf8a1df0 100644
--- a/gitnexus/package-lock.json
+++ b/gitnexus/package-lock.json
@@ -1888,9 +1888,9 @@
"license": "MIT"
},
"node_modules/@types/node": {
- "version": "26.4.0",
- "resolved": "https://registry.npmjs.org/@types/node/-/node-26.4.0.tgz",
- "integrity": "sha512-faiGnoIrLH/V8cibOMEAZ8pMw6oXqSukl29ra4mN8GdaB2ZewzeaLj+INpV5N+Z1eKWzY+IzaIZH2EIR6YZRNQ==",
+ "version": "26.4.1",
+ "resolved": "https://registry.npmjs.org/@types/node/-/node-26.4.1.tgz",
+ "integrity": "sha512-k97ENvZWtvA6yqz5/FS6a7duDgOPEeOQOc2iKS/nY6mX6qJUKtLnWzQS+Xj6tXweyj6ZcTAK2Qecetnvi9nCLA==",
"devOptional": true,
"license": "MIT",
"dependencies": {
From 96132bd13a386388d09c1515c4f13ad0fd5e1cff Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Gerg=C5=91=20Magyar?=
Date: Tue, 8 Sep 2026 13:48:30 +0100
Subject: [PATCH 07/53] perf(scope-resolution): stop re-scanning the ParsedFile
store once per language (#3211)
* perf(scope-resolution): stop re-scanning the ParsedFile store once per language
Scope resolution calls `loadParsedFilesForPaths` once per language, and every
call walks every shard in the store. The skip decision needs the envelope's
path listing, and that listing is only trustworthy after the payload digest
has been checked -- so a pass that wants 50 Python files still opens and
SHA-256s all 413 shards / 301MB of a TypeScript-dominated store to prove it
can skip them. A pass wanting a SINGLE file costs 335ms. The cost scales with
language count, not with the files that language has, so a polyglot repo pays
it worst.
`tryLoadV8Cache` now returns the listing it already parsed for that skip
decision, and the store memoizes it per run. Later passes skip on the
memoized listing without reopening the file.
Measured on a 2234-file, 3-language repo, min-of-5:
before python 411ms typescript 2872ms javascript 507ms = 3834ms
after python 415ms typescript 2805ms javascript 248ms = 3484ms
-350ms here, roughly -250ms per additional language elsewhere. The first pass
is unchanged by construction -- it is what populates the memo. End-to-end the
graph is byte-identical: 51,288 nodes / 163,094 edges / 2106 clusters /
759 flows on a true incremental run.
Keyed on size+mtime as well as name. Shard names are content-addressed, so a
name collision across different content should be impossible, but that
invariant lives in the parse-cache keying rather than here and one stat per
shard is a few ms against the hundreds this saves. The memo holds one store
directory at a time, so a new repo in a long-lived MCP process drops the
previous set instead of accumulating.
The failure mode a listing memo introduces is a FALSE SKIP: a pass concludes a
shard holds nothing it wants and those files silently never reach the graph --
an exit-0 wrong answer, not a crash. The new test walks four passes with
disjoint wants over one store, plus a shard written after the memo is warm;
it fails when the skip is forced.
Also records the full scopeResolution breakdown in bench/. The headline is
that `emit` is 7161ms of the 14.7s phase and ~21% of the edit loop, spread
across a fan of passes with no hot inner loop -- so the win there is not
running them for unchanged files, which is a design rather than a patch.
Co-Authored-By: Claude Opus 5 (1M context)
* Address PR review feedback (#3211)
Strengthen the shard-listing memo test so a later miss asserts fs.open and
v8.deserialize never run for the skipped shard. Key-set checks alone still
passed if the memo never skipped.
Co-authored-by: Cursor
* chore(autofix): apply prettier + eslint fixes via /autofix command
---------
Co-authored-by: Gergo Magyar
Co-authored-by: Claude Opus 5 (1M context)
Co-authored-by: Cursor
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
---
gitnexus/bench/analyze-phase-breakdown.md | 86 ++++++++++++++++-----
gitnexus/src/storage/parsedfile-store.ts | 59 +++++++++++++-
gitnexus/src/storage/v8-sidecar.ts | 45 ++++++++---
gitnexus/test/unit/parsedfile-store.test.ts | 58 ++++++++++++++
4 files changed, 216 insertions(+), 32 deletions(-)
diff --git a/gitnexus/bench/analyze-phase-breakdown.md b/gitnexus/bench/analyze-phase-breakdown.md
index 6706c1905..e04d48628 100644
--- a/gitnexus/bench/analyze-phase-breakdown.md
+++ b/gitnexus/bench/analyze-phase-breakdown.md
@@ -169,31 +169,73 @@ Two things block it today, and neither is small:
would then leave the live index with fresh File content and stale symbols,
instead of untouched.
-## scopeResolution is memory-traffic bound, not algorithmic
+## scopeResolution — 14.7s, and it is two different problems
-`--cpu-prof` of a one-file-edit run, top main-thread self time:
+`PROF_SCOPE_RESOLUTION=1` already exists and reports the internal split. Marks
+injected around the phase supply the rest. On a one-file edit:
+
+| step | ms | share |
+| ---------------------------- | -------: | ------: |
+| ParsedFile store rehydration | 4426 | 30% |
+| **`emit`** | **7161** | **49%** |
+| `resolve` | 966 | 7% |
+| `finalize` | 552 | 4% |
+| `extract` | 369 | 3% |
+| per-language teardown, misc | ~750 | 5% |
+
+`extract` is small because the parse cache works: 2121/2121 pre-extracted hits
+on the TypeScript pass. Nothing here re-parses.
+
+### Rehydration: three full store scans, one per language
+
+The corpus resolves three languages — python (54 files), typescript (2121),
+javascript (59) — and `loadParsedFilesForPaths` walks **every** shard on each
+pass. The store is 413 shards / 301 MB. A pass that wants a single file costs
+335 ms, because the skip decision needs the envelope's path listing and that
+listing is only trustworthy after the payload digest is checked.
+
+So the fixed cost is paid per language, and **it scales with language count**,
+not with how many files that language has. Measured, min-of-5, three passes:
```
-4294ms (garbage collector)
-1552ms v8.deserialize
- 756ms crypto update
- 728ms runScopeResolution
- 620ms scope-resolution/pipeline/reconcile-ownership
- 380ms v8-sidecar walk
- 323ms internString
+before python 411ms typescript 2872ms javascript 507ms = 3834ms
+after python 415ms typescript 2805ms javascript 248ms = 3484ms
```
-then a tail of passes at 200–750ms (`emitReceiverBoundCalls`,
-`emitCallableValueFlow`, `buildGraphNodeLookup`, `resolveReferenceSites`).
+Memoizing each shard's authenticated listing for the run removes the repeat
+scans (−350 ms here; roughly −250 ms per additional language elsewhere). The
+first pass is unchanged by construction — it is what populates the memo.
-No dominant hot function, nothing quadratic. The cost is rehydrating every
-file's `ParsedFile` from the durable `.v8` shards into main-thread memory and
-re-running every pass over them.
+What is left is a floor. The digest is **not** the cost: SHA-256 over all
+301 MB takes 134 ms (2.25 GB/s, hardware-accelerated), so swapping it for
+CRC-32 would buy ~90 ms and cost an envelope-format bump. The remainder is
+`v8.deserialize` (~1240 ms) plus the cross-shard string intern walk
+(~1085 ms), and the intern walk is not optional — dropping it regresses
+retained heap ~59%, which is #2649's constraint.
-**So the win is not caching resolution output** — that stores more of exactly
-what is already the memory problem, and #2649 (large-repo OOM) is the standing
-constraint. The win is skipping rehydration and re-resolution for files whose
-inputs provably did not change, as a streaming/bounded design.
+### `emit` is the real target
+
+7.2s, 49% of the phase and ~21% of the whole edit loop. It is a fan of passes,
+each walking every reference site in the repo:
+
+```
+1828ms emitCallableValueFlow 480ms emitFreeCallFallback
+1590ms emitReceiverBoundCalls 426ms emitReferencesViaLookup
+ 681ms resolveDefGraphId 313ms emitUniqueNamePropertyAccesses
+ 529ms lookupCore 280ms emitReturnShapeMemberAccesses
+ 472ms tryEmitEdge 222ms emitPropertyDispatchCalls
+ 449ms getScope
+```
+
+No hot inner loop, nothing quadratic, no single pass worth more than 12% of the
+phase. Micro-optimizing any of them is not the win.
+
+The win is not running them. A one-file edit re-emits all 163,094 edges to
+write a 3,980-node subgraph. Every pass above runs over all 2121 TypeScript
+files because the pipeline rebuilds the full in-memory graph every run and only
+the DB _write_ is incremental. Making `emit` incremental means knowing which
+files' edges can change when one file's registry contribution changes — which
+is a design, not a patch, and #2649 rules out "just cache the emitted edges".
---
@@ -232,8 +274,12 @@ and 0.8s of post-FTS event-loop drain between them.
FTS is closed as an optimization target except for the overlap above, which is a
scheduling change to `run-analyze.ts`'s open/close discipline rather than
-anything about FTS. `scopeResolution` is the single largest step, memory-traffic
-bound, and still untouched.
+anything about FTS.
+
+`scopeResolution` is now broken down. Its rehydration half has a floor and one
+repeat-scan win that is taken. Its other half — `emit`, 7.2s — is the largest
+remaining target in the whole edit loop, and the only way at it is incremental
+resolution.
Every optimization in this document that looked compelling from the armchair
died under measurement. Measure first, and check the banner.
diff --git a/gitnexus/src/storage/parsedfile-store.ts b/gitnexus/src/storage/parsedfile-store.ts
index dce10b03d..68ed2e06e 100644
--- a/gitnexus/src/storage/parsedfile-store.ts
+++ b/gitnexus/src/storage/parsedfile-store.ts
@@ -169,6 +169,7 @@ export const getParsedFileStoreDir = (storagePath: string): string =>
/** Remove any prior run's shards so a fresh parse starts clean. Idempotent. */
export const clearParsedFileStore = async (storagePath: string): Promise => {
await fs.rm(getParsedFileStoreDir(storagePath), { recursive: true, force: true });
+ forgetShardListings();
};
const isV8ShardName = (name: string): boolean => name.endsWith('.v8') && !name.includes('.v8.');
@@ -245,13 +246,52 @@ const listV8Shards = async (dir: string): Promise => {
}
};
+/**
+ * Per-run memo of each shard's authenticated path listing, so the SECOND and
+ * later passes over the store can decide "this shard holds nothing I want"
+ * without reopening it.
+ *
+ * Scope resolution calls `loadParsedFilesForPaths` once per language, and each
+ * call walks every shard in the store. The skip decision needs the envelope's
+ * path listing, and that listing is only trustworthy once the payload digest
+ * has been checked — so today a pass that wants 50 Python files still reads and
+ * SHA-256s all ~300MB of a TypeScript-dominated store to prove it can skip it.
+ * Measured on a 2234-file repo: a pass wanting a single file costs 335ms, and
+ * the three language passes together spend 4426ms in here. The cost scales with
+ * LANGUAGE COUNT, so a polyglot repo pays it worst.
+ *
+ * Keyed by size+mtime as well as name. Shard names are content-addressed
+ * (parse-chunk hash + worker id), so a name collision across different content
+ * should be impossible — but that invariant lives in the parse-cache keying,
+ * not here, and one `stat` per shard is a few ms against the hundreds this
+ * saves. The memo holds one store directory at a time: a different `dir` (a new
+ * repo in a long-lived MCP process, or a wiped store) drops the previous set
+ * rather than accumulating.
+ */
+let shardListingMemo: { dir: string; byIdentity: Map } | undefined;
+
+const shardIdentity = (file: string, size: number, mtimeMs: number): string =>
+ `${file}\0${size}\0${mtimeMs}`;
+
+const memoFor = (dir: string): Map => {
+ if (shardListingMemo?.dir !== dir) shardListingMemo = { dir, byIdentity: new Map() };
+ return shardListingMemo.byIdentity;
+};
+
+/** Drop the memo when the store it describes is removed. */
+const forgetShardListings = (): void => {
+ shardListingMemo = undefined;
+};
+
export const loadParsedFilesForPaths = async (
storagePath: string,
wantPaths: ReadonlySet,
): Promise