GitNexus/eval/workflow_bench/runner.py
Gergő Magyar 8b5057f325
feat(skills): GitNexus Engineering Tool Kits (#2566)
* feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill

Adds .claude/skills/ce-plan: a planning-only skill that builds
implementation-ready plans from GitNexus graph navigation (query/context/
impact/trace), bounded statement-level PDG slices (pdg_query, impact
mode:pdg, explain), and targeted source verification, with a context
ledger to prevent repeated reads and a machine-readable implementation
context pack (stable contract for a future ce-implement). Whitelisted in
.gitignore and registered in AGENTS.md and CLAUDE.md outside the
auto-managed gitnexus block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions)

Tool contract: impact mode:'pdg' shape now includes the schema-required
direction param; CDG branch sense documented as the result 'label' field
(reason is cypher/raw-edge only); explain caveats corrected to its real
false-negative classes (cross-function TAINT_PATH is modeled).

Consistency: PDG slice homed in working memory (ledger keeps one-liners);
depth knob defined and category-overrides-baseline ordering stated;
call_depth (consumed by nothing) and content-hash bookkeeping dropped;
Never section folded into Hard rules; Phase 3 deduplicated to a pointer;
allowed-repeat escalations defined; budget/discard accounting clarified;
verification-commands gathering added to Phase 4; open_questions added to
the context pack.

From scenario runs: plans now pin the verified-at HEAD commit and index
freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed],
quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and
support an out:<path> destination override; output path defined as the
Phase 1 target repo root.

Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata
bumps; future ce-implement qualified as future.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): rename ce-plan → gitnexus-plan; add cross-CLI (Codex) entrypoints

Renames the skill dir, frontmatter, output filename convention, plan H1
(GitNexus Engineering Plan), the future executor handle
(gitnexus-implement), the .gitignore whitelist entry, and all
AGENTS.md/CLAUDE.md references. Follows the pr-swarm-review cross-CLI
pattern: SKILL.md is the canonical CLI-neutral spec, AGENTS.md § Engineering
planning is the Codex/any-agent entrypoint, and the README documents the
optional user-level ~/.codex/prompts/gitnexus-plan.md slash command plus an
invocation matrix. Skill prose de-branded from Claude Code (agent-neutral
verification layer).

Also fixes two post-review README contradictions: the anti-reread claim now
names the ledger's allowed escalations, and 'read-only by contract' is now
'planning-only' (the skill writes exactly one repo file — the plan); the
scope-creep rule and template §12 now agree on where deferred follow-ups
land. Drops the stale plugin-collision limitation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): document Codex user-level install path for gitnexus-plan

Codex discovers SKILL.md skills from ~/.agents/skills (same path the other
gitnexus-* skills install to); README now documents the cp install plus the
optional ~/.codex/prompts slash-command file, with the prompt body preferring
the repo copy and falling back to the user-level install.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan freshness gate + active PDG-layer refresh

Freshness is now a Phase 1 gate, not advisory: under the default
freshness:strict, a stale index is refreshed once per planning session via
node .gitnexus/run.cjs analyze --index-only (appending --pdg when the task
will reach the PDG phase), then the context resource is re-read. A missing
PDG layer likewise triggers the one permitted --index-only --pdg refresh
and re-probe instead of a passive recommendation. freshness:accept (or a
failed/impractical refresh) preserves the old behavior: plan on the stale
graph, source-weighted, labelled in the plan header. --index-only is the
load-bearing flag choice — it suppresses all file generation, so the
planning-only contract holds (only the .gitnexus store changes). Ledger
gains an index_refresh record; plan header states fresh / refreshed /
refresh-skipped-with-reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan runner build check before freshness refresh

When the target repo builds the analyzer from its own source (bin → dist/
mapping, as gitnexus/ does), the Phase 1 freshness gate now verifies dist/
is current before running the analyze refresh — rebuilding via the
package's build script when any analyzer source file is newer than the
built entrypoint — and prefers that freshly built CLI. Otherwise a stale
dist re-indexes with outdated extraction logic and the 'fresh' index lies.
Rebuilds are recorded in the ledger's index_refresh; the PDG-phase refresh
inherits the same check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): add gitnexus-work executor and gitnexus-lfg pipeline

gitnexus-work executes a gitnexus-plan as verified atomic commits: consumes
the §11 implementation_context pack, drift-checks the plan's evidence pin
against HEAD, re-verifies assumptions before relying on them, runs impact
before every symbol edit and detect_changes before every commit (repo
mandates), builds tests from the plan's scenarios, and routes structural
drift back to gitnexus-plan Deepen mode instead of coding around it.

gitnexus-lfg is a thin orchestrator: gitnexus-plan → blocking user gate
(deepen / proceed / stop, deepen loops allowed) → gitnexus-work → review
via the existing gitnexus-pr-review skill (open PR, else branch diff vs
default). One bounded fix cycle for review findings; never pushes or opens
a PR on its own.

gitnexus-plan gains a Deepen mode (re-run freshness gate, escalate to
depth:deep, re-verify graph/inferred/assumed claims toward verified,
rewrite the same file); its 'future gitnexus-implement' placeholder is
retired in favor of gitnexus-work. Registered via .gitignore whitelists,
AGENTS.md 1.10.0 (section renamed to Engineering planning & execution),
CLAUDE.md 1.5.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply cross-skill review findings to the gitnexus skill family

Two P1s: gitnexus-plan Deepen mode now re-anchors before re-pinning
(diffs the old evidence pin over every [verified]-claim file and re-reads
or downgrades before the header moves — moving the pin without this
laundered stale claims as verified); the index-refresh budget is stated
once in Phase 1 (one --index-only refresh plus at most one Phase 3 --pdg
upgrade per session, Deepen = its own session) with ledger and pdg-slice
deferring to it.

Contract fixes: gitnexus-work's drift check now covers every file the
pack cites (not just files_to_modify) and parses the full pack incl.
primary/related symbols and acceptance_criteria (walked in Phase 4
alongside §13); a pre-completed check skips §7 steps already landed and
Deepen gains a reconcile-execution-state step, closing the mid-execution
route-back loop; pack assumptions must name what to check and how.

lfg: Lane 4 passes the merge-base to detect_changes compare (two-dot
diff misattributes upstream commits when default advanced), branch-diff
is the stated normal case, oversized review findings route to the plan
gate instead of overflowing direct mode, the one-fix-cycle cap is
explicit on re-run, and headless runs end at the plan gate with the plan
as deliverable. work: blank mode narrowed to *gitnexus-plan*.md with a
re-execution guard, direct-mode discipline spelled out, branch
meaningfulness defined against the plan slug, and the plan document is
committed as the branch's docs commit (review diff includes it).
Planning-only contract now names the dist/ rebuild as the second
permitted state change; Phase 5.1 names the four claim tags; stale
AGENTS.md anchors fixed.

Known latent issue left untouched: gitnexus/gitnexus-pr-review pairs a
three-dot example with a two-dot detect_changes compare — that skill is
also shipped by the plugin, so fixing it here would drift the copies;
lfg compensates by passing the merge-base.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ship the engineering skill family with the gitnexus package

npm i -g gitnexus users now get gitnexus-plan / gitnexus-work / gitnexus-lfg:
the three skills are added to gitnexus/skills/ in directory form (SKILL.md +
references/), which installSkillsTo already enumerates dynamically and copies
recursively to every editor target (~/.agents/skills for Codex, Cursor,
OpenCode, Qoder, ...) on gitnexus setup — uninstall enumerates the same root,
so removal stays clean. The Claude Code plugin channel
(gitnexus-claude-plugin/skills/) carries the same copies plus the standard
per-skill mcp.json.

Global-install support in the skill text: gitnexus-plan Phase 1 now resolves
the analyzer runner explicitly — node .gitnexus/run.cjs analyze when the
project has a runner, else gitnexus analyze (installed CLI), else
npx gitnexus analyze — and all analyze mentions route through it, satisfying
the skills-steering policy (#1939/#1945) which sweeps the plugin copies.

New drift guard test/unit/shipped-skills-sync.test.ts asserts the npm and
plugin copies stay byte-identical to the canonical .claude/skills/ family
(plugin = canonical + mcp.json), same discipline as run.cjs ↔
resolve-invocation.ts. skills-steering + shipped-skills-sync: 11/11 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench — measure the skill workflow's token savings

Benchmarks gitnexus-plan → gitnexus-work against a baseline agent
(--disallowedTools Skill) on identical tasks, in fresh detached worktrees,
using real headless Claude Code sessions; every number comes from the CLI's
--output-format json usage report (field names validated against a live
2.1.207 session). Reports per-arm medians (input/cache/output tokens, cost,
wall time, turns), a savings row, and resolve status from a per-task verify
command — savings on failed tasks are flagged, not celebrated. Per-task
setup hook prepares fresh worktrees (deps); --permission-mode
bypassPermissions (default) lets sessions run unattended in the throwaway
trees.

Free-model support: --base-url/--auth-token/--model route headless sessions
through any Anthropic-compatible endpoint; free-model.litellm.yaml is a
ready litellm-proxy template for OpenRouter :free variants or local Ollama,
so benchmarking burns no paid tokens (README documents rate limits and the
small-model skill-following caveat).

Harness validated end-to-end with a stub CLI (worktree lifecycle, both
arms, plan→work chaining, verify, aggregation, report) and 4 pytest units
for the pure aggregation/savings/report helpers. AGENTS.md 1.11.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record first workflow_bench calibration run

Trivial-task calibration (add -V alias): both arms resolved; workflow arm
~4.3x baseline cost — the documented overhead-dominated regime, recorded so
the regime boundary is empirical rather than asserted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench scenario matrix — arm variants, task classes, churn

Ground-base measurement across scenarios: tasks.scenarios.yaml spans four
labeled classes (trivial → investigation-bug → investigation-feature →
cross-module) with deterministic verifies (prescribed test files). New arms:
workflow_direct (gitnexus-work direct mode — the middle option that locates
the routing boundary lfg's gate and work's triage encode) and baseline_nomcp
(no skills AND no graph tools — separates workflow-discipline value from
GitNexus-tool value; off by default). Records now carry task class and diff
churn (files/+ins/−del vs the starting commit) as an over-engineering proxy;
the report renders a class column and per-arm savings rows vs baseline.
5 pytest units + stub-CLI e2e of the full three-arm matrix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record workflow_bench ground base; fix churn measurement bias

Ground base (3 classes x 3 arms, n=1/cell): every arm resolved every task —
pass/fail quality saturates at this difficulty, making the comparison pure
cost. Full plan→work never amortized its ~$9-11 fixed cost on tasks a
baseline finishes in ≤35 turns (−211% to −333% cost); workflow_direct sits
near baseline (−15% to −55%, once faster wall) with more test coverage.
Routing implication recorded: direct mode/plain agent below this scale,
full workflow for cross-module / multi-session / plan-as-deliverable work.
The cross-module cell and multi-run variance are the next measurements.

Churn fix: git add --intent-to-add -A before diffing (arms that never
commit no longer undercount new files) and :(exclude)docs/plans (the
committed plan doc no longer inflates workflow churn); this run's churn
numbers predate the fix and are omitted from the recorded table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(skills): cost-optimize the workflow from measured ground base

Every optimization targets a measured fixed-cost component
(eval/workflow_bench ground base: workflow arm −211% to −333% vs baseline,
all tasks resolved):

- Plan form is category-priced: compact form (core sections w/ § anchors
  preserved, ≤80 lines excl. pack, mini-pack subset of the context pack)
  for narrow/default categories; the full 13 sections only for deep work
  (refactor/security/performance/concurrency/architecture). A compact plan
  outgrowing its cap reclassifies to full rather than overflowing.
- Freshness gate is category-priced: compact categories default to accept
  (source-weighted, refresh only when a graph claim becomes load-bearing);
  strict stays the default for full-plan categories — the rebuild+re-index
  was the largest single fixed cost.
- Turn economy: per-category tool-call budgets (~10 to ~45; architecture
  uncapped); budget exhaustion routes open questions to §12 instead of
  more digging.
- gitnexus-work fast path: HEAD == evidence pin → skip all citation
  re-reading (the pin's entire point); mini-pack fields tolerated.
- lfg Lane 1 boundary triage: tasks below the measured ~35-turn boundary
  get offered gitnexus-work direct mode before the plan lane is spent.

Copies re-synced (npm skills/, plugin, ~/.agents); steering + sync guards
green. Re-measurement of the workflow arm follows to verify the numbers
actually improve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record optimization re-measurement — inv-bug workflow cell −20% cost

Same task, same conditions, post-830a0459 skills: $14.56→$11.70 (−20%),
83→72 turns, cache_read −24%; verified in-transcript that the compact form,
turn budget, and skipped rebuild/re-index all fired. Wall +15% from a work-
session test-debugging tail (n=1 variance). Regime unchanged (~3.5x baseline
on this class) — routing rule stands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): per-arm clone isolation — worktree ref-namespace leak contaminated an arm

The cross-module workflow_direct cell reported an impossible 28-turn solve
with churn byte-identical to the workflow arm: git worktree add shares the
repo's ref namespace, so the workflow arm's slug branch (created by
gitnexus-work Phase 2) survived worktree removal and the direct arm found
and adopted the completed work. Arms now get isolated git clone --shared
copies (object store via alternates, refs clone-local — agent branches and
stashes die with the clone; origin/<ref> fallback for non-default refs).
Leaked branch deleted; baseline arm verified clean (0 branch references in
its transcript); cell marked invalidated pending re-run.

Records the valid cross-module cells: workflow $18.32 vs baseline $18.03
(premium −1.6%, vs −211%..−333% on smaller classes) — fixed costs amortize
at this scale, with a less destructive diff and a plan artifact as bonus;
resolve rate still tied. Churn fingerprinting is what caught the
contamination — noted in the README as an integrity check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): complete cross-module cell — direct mode wins 47% cost / 56% wall

Clean clone-isolated re-run: workflow_direct resolved the hardest class at
$9.53/52 turns/15m vs $18.03/98/34m baseline and $18.32/107/37m full
workflow. The measured story across all four classes: the execution
discipline (gitnexus-work) is the consistent sweet spot and delivers real
token savings on hard tasks; the planning pass buys its artifact, not
same-session savings. Resolve rate tied everywhere (n=1/cell caveat).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): add trajectory-gated skill evolution (#2431)

- Pair prompt candidates with incumbent workflow arms
- Gate promotions on pinned-model quality and efficiency
- Expire router evidence and document its lifecycle

* fix(eval): allow pr-review skill candidates

* feat(skills): rename and generalize GitNexus review

* feat(eval): external-comparator and review arms for workflow_bench

- ce_workflow / ce_workflow_direct: compound-engineering ce-plan/ce-work
  arms prompted with the same structure as the gitnexus arms
- review / ce_review: gitnexus-review vs ce-code-review on an identical
  diff applied by the task's setup
- plan handoff is snapshot-based: committed example plans in docs/plans/
  tie on clone mtimes and broke the name-glob pick (executed a stale plan)
- verify output tail is recorded per run and the final working-tree patch
  is kept, so failed rows are diagnosable after the clone is destroyed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills,eval): address #2431 review — data-safe rename migration, fail-closed bench evidence

- setup: never delete a legacy renamed skill dir — the installer cannot
  prove ownership (users customize or hand-write skills under these
  names); warn with the path instead, and the test now asserts survival
- workflow_bench: fail closed when a session's --output-format json
  report is empty, malformed, or missing usage fields — an exit-0 shell
  with no parseable usage no longer counts as measured evidence
  (5 parametrized regression tests)
- workflow_bench: document the trust model prominently (task setup/verify
  are shell-executed, sessions run bypassPermissions with the parent env,
  candidate overlays are prompt injection surface) in README + docstring
- free-model.litellm.yaml: master_key from LITELLM_MASTER_KEY env instead
  of a static token; loopback-binding warning
- ci: run the eval workflow_bench pytest suite on ubuntu (pytest+pyyaml
  only — no full eval stack)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): demand observed foreground verification in headless work-arm prompts

In a headless -p session there is no later turn: a work arm backgrounded
its slow test run, scheduled wakeups that can never fire, and reported
done while two of its tests failed. All four work-arm prompts (both
skill families, symmetric) now require verification output to be
observed inside the session.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ask plan depth up front instead of offering deepen afterwards

gitnexus-plan Phase 0 now asks one blocking question in interactive
sessions — quick / standard / deep, mapped onto the existing depth/form/
freshness knobs — when the invocation carries no explicit depth signal.
Explicit knobs and headless runs skip the question (category posture
unchanged, so benchmarks and automation behave as before).

gitnexus-lfg's plan gate slims to proceed/stop: depth was already the
user's up-front choice, so deepening is no longer offered by default —
an explicit deepen request at the gate and executor route-backs still
run Deepen mode, which remains the mechanism for strengthening an
existing plan document.

All shipped copies resynced (npm skills/, Claude plugin); AGENTS.md
1.13.0 and CLAUDE.md 1.7.0 pointers updated, including the analyzer's
regenerated index-stats block at this branch's head.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): taint pass, expert lenses, and post-work index refresh

gitnexus-review gains a PDG-backed taint-and-dependence pass (explain +
pdg_query, --pdg folded into the stale refresh on trust-boundary diffs) and
an Expert lenses section: domain reviewers derived from the graph's
clusters plus four cross-cutting lenses (architectural fit, language
conformance per the repo's own contract, Definition of Done, simplicity),
dispatched once after the evidence-gathering steps and scaled to the diff.
gitnexus-work Phase 4 now refreshes the knowledge graph after the DoD walk
via the resolved-runner ladder with analyze --index-only, so the lfg review
lane and later sessions query the finished work without dirtying the tree.
lfg's threshold-governance paragraph moves to its README; eval citations
are tagged as measured in the GitNexus repo. All shipped copies re-synced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(cli): remove legacy gitnexus-pr-review on uninstall; cover the rename migration

uninstall's removal set now includes LEGACY_SKILL_DIR_NAMES derived from
RENAMED_SKILL_DIRS, so a pre-rename install is cleaned up instead of
orphaned. The rename warning gains behavioral coverage (fires with a legacy
dir present, silent without), and shipped-skills-sync asserts legacy names
stay absent from every shipped tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): metric provenance, error-kind rows, skill-invocation verification, gate noise floor

The promotion gate defaults to cost_usd (the only metric that includes
subagent spend); token metrics carry an explicit main-loop-only warning in
the report and promotion.json. Rows are classified by error_kind
(session-error / verify-failed / infra-error), excluded from efficiency
medians, and the gate requires equal valid-run counts. Each session's
transcript is scanned for the expected Skill invocation and fails closed on
a verified miss; a one-run resolution edge no longer promotes (noise
floor). Per-run timeouts and setup failures record an infra-error row
instead of aborting the sweep. Overlays touching skills no candidate arm
exercises are rejected up front.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fix skill routing paths, version headers, and skill rosters

Routing tables point at the tracked direct skill paths (matching the
post-#2434 generator output), AGENTS.md/CLAUDE.md headers match their
latest changelog rows, the 1.12.0 row describes what the migration actually
does, package/cursor READMEs list the full shipped skill roster, and the
swarm READMEs describe /gitnexus-review's expert lenses instead of calling
it single-agent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: drift-guard workflow for skill copies; pin eval pip deps; track docs/plans

ci.yml ignores '**.md', so an md-only skill edit would merge without the
shipped-skills-sync test running — skill-sync.yml triggers exactly on the
guarded trees. The eval job's pip install is version-pinned, and
docs/plans/ is unignored so gitnexus-plan output can be committed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep the runner-invocation literal in gitnexus-review; add concurrency block to skill-sync

skills-steering requires skills with a stale-index hint to carry the exact
'node .gitnexus/run.cjs analyze' form — restore it with the fallback ladder
as a parenthetical instead of replacing it. skill-sync.yml gains the
top-level concurrency block the workflow-convention check enforces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): token-economy guidance for expert lenses

Merge lenses that ground in the same material into one reviewer, and use
cheaper model/effort tiers for mechanical lenses where the harness offers
them, reserving the strongest engine for adversarial judgment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(eval): isolate transcript home on Windows

Ensure workflow_bench transcript tests set USERPROFILE alongside HOME so Path.home() resolves to the temporary test home on Windows.

* docs(skills): fold PR #2522 execution learnings into review/work/plan

Eight incident-backed hardenings from running the full skill cycle
(review -> plan -> work, 28-finding fix series) on PR #2522:

gitnexus-review:
- Expert lenses execute the code under review on candidate failing shapes
  (empirical probe outranks source reading — every HIGH the language
  lenses found came from a probe, not a read).
- Step 7 re-runs the exact CI check for refreshed baselines/fingerprints
  (a stale committed artifact is invisible in the diff; caught a red
  benchmarks arm).
- Step 8 treats version/invalidation constants as review surface
  (INCREMENTAL_SCHEMA_VERSION class recurred verbatim from #2494).

gitnexus-work:
- Step 4 proves regression tests discriminate against the pre-fix tree.
- Step 5 rebuilds executed build output before every verification run
  (parse workers load dist/; a correct fix 'failed' until rebuilt).
- Step 6 makes stage -> detect_changes -> commit one unbroken sequence.

gitnexus-plan:
- Phase 0 seeded-evidence mode: plan FROM a completed review's verified
  findings instead of re-running the graph ladder.
- Template §7: fingerprint/golden-guarded output rebaselines once, at the
  series tip.

All distribution copies resynced; shipped-skills-sync + skills-steering
24/24 locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): close the skill-evolution loop with an automated proposer driver

workflow_bench.evolve adds the three arrows the README described as manual:
a proposer session that turns loser trajectories (results.jsonl rows,
transcripts, patches, the learning queue) into ONE bounded candidate
overlay, a driver that iterates propose -> paired benchmark -> deterministic
gate up to --generations, and an --apply step that copies a promoted
overlay onto the canonical skills and shipped mirrors as a working-tree
diff. The trust boundary is unchanged: overlays re-validate through
candidate_overlay_files before any benchmark or apply consumes them, and
committing, CI, and the PR merge stay human.

learnings.jsonl is gitignored: it is machine-local evidence, like the
session transcripts it complements.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): route live-task friction into the evolution learning queue

Each family skill gains a short 'Skill feedback' section: on friction with
the skill's own instructions, append one JSON line to
eval/workflow_bench/learnings.jsonl (GitNexus repo only) — never self-edit
the skill from a live task. The proposer in workflow_bench.evolve consumes
the queue as hints; a learning reaches a shipped skill only by beating the
incumbent on the paired benchmark. All shipped mirrors re-copied byte-
identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(tests): run the evolve helper tests in the eval pytest job

test_evolve.py needs only pytest+pyyaml, same as the harness tests the job
already runs — without this line the new module had no CI coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): comment-triggered GitNexus review agent for PRs

'@gitnexus review' from a maintainer (OWNER/MEMBER/COLLABORATOR; the action
re-validates write access) runs the repo's gitnexus-review skill headlessly
against the PR and posts the review as a sticky comment — remote triggering
with no local setup. Read-only by construction: contents: read token,
Write/Edit and web tools disallowed, Bash allowlisted to git reads and the
gitnexus CLI; analyze parses PR code with tree-sitter, never executes it.
Requires the ANTHROPIC_API_KEY repository secret; activates once the file
is on the default branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): dispatch lane + existing OAuth secret for the review agent

Align with claude.yml: same action pin and the CLAUDE_CODE_OAUTH_TOKEN
secret the repo already carries — no new secret to configure. Add a
workflow_dispatch lane (PR number input) so the agent can be triggered from
the Actions UI and tested before the issue_comment trigger reaches the
default branch. Allowlist gh pr view/diff and gh api, which the review
skill uses to pin PR SHAs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): close a fork-PR RCE vector in the review agent's tool allowlist

A live headless run of the exact workflow session against PR #2431 (66
turns, full gitnexus-review pass) surfaced a real HIGH-severity confused
deputy: .gitnexus/ is gitignored, not blocked — a fork PR can commit its
own .gitnexus/run.cjs, issue_comment checks out PR-head content, and the
skill's runner ladder tries 'node .gitnexus/run.cjs analyze' first. That
would execute fork-controlled JS inside a job holding
CLAUDE_CODE_OAUTH_TOKEN and a write-scoped GITHUB_TOKEN — the opposite of
the 'PR code is read, never executed' claim in the workflow's own header.

Fix: drop the run.cjs allowlist entry so analyze always resolves through
npx gitnexus (npm registry, not the checked-out tree); the skill's
documented fallback mode covers the resulting graceful degradation. Also
drop 'gh api' (not read-only — accepts -X POST/PATCH/DELETE) and downgrade
pull-requests: write to read (comment posting only needs issues: write;
the prompt already forbids formal review submission).

Same session flagged a latent evolve.py bug: select_evidence's cost sort
used dict.get's missing-key default, which doesn't cover an explicit JSON
null in a foreign --seed-results row and crashes proposer setup with
TypeError. Guarded with 'or 0.0' and added a regression test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden PR review and evolution trust boundaries

* ci: follow workflow concurrency convention

* fix(eval): make terminating error paths explicit

* fix: unblock hardened review runtime checks

* test: make containment canaries deterministic

* test: expose Claude canary tool failures

* fix: adapt clean shell environment for Claude

* fix(eval): accept the runner's transcript source key in evidence preflight

The proposer evidence preflight required transcript-artifact metadata to be
exactly {path, sha256, bytes}, but the runner stamps a fourth provenance key
(source=parent-captured-stream-json). Any --seed-results or generation>=2 run
therefore aborted with SandboxError before proposing or promoting. Pin the
producer literal as PARENT_EVENT_STREAM_SOURCE and validate it in the metadata
check, and round-trip real producer output through sum_sessions into the
preflight so the schema can't drift again.

* fix(eval): treat an unmeasured session cost as unavailable, not $0

well_formed validated only the nested usage block, so an otherwise-successful
session missing total_cost_usd was recorded as cost_usd=0.0 — and cost_usd is
the default promotion metric (lower wins), so a cost-less session scored as
free and could win promotion it never earned. Extract cost via measured_cost()
(None on absent/garbage, a measured 0.0 preserved), propagate None through
sum_sessions/aggregate/savings/report, and have the gate refuse to rank on a
metric that was not measured on every run in both arms.

* fix(eval): warn when ranking on the main-loop-only num_turns metric

num_turns comes from the CLI's top-level usage (main-loop session only), like
output_tokens, but selecting it emitted no metric_warning — so a subagent-heavy
candidate could look artificially efficient. Add num_turns to
MAIN_LOOP_ONLY_METRICS and broaden the warning to cover turns.

* fix(eval): fail closed when an overlay adds a file with no committed base

An overlay adding a new .md under gitnexus-{plan,work} passes the structural
overlay checks but has no committed base for committed_destination_base_digests
to bind against, so it raised an uncaught ValueError that crashed the evolve
driver (and runner --candidate-overlay) mid-run. Catch it at both call sites:
evolve reports NOT PROMOTED and exits, runner routes it through parser.error.

* feat(eval): circuit-break the runner sweep on a systemic outage

A sustained upstream outage used to pay out every remaining --timeout window
one session at a time. Track consecutive session/infra/cleanup failures via a
pure systemic_outage_streak helper; after --outage-streak (default 5) in a row,
stop the sweep, still write report.md/promotion.json from partial evidence, and
exit non-zero so evolve.py halts instead of proposing from truncated evidence.
A task's own resolved=False never trips the breaker.

* fix(cli): report a dirty working tree as stale in gitnexus status

status --json (and the human output) computed up-to-date from commit + runner
identity + completeness only, so a repo with uncommitted source changes at a
matching HEAD was reported up-to-date while analyze would still re-index it.
A graph-backed agent gating on that JSON could skip re-analysis on a stale
graph. Extract analyze's dirty-tree check into a shared isWorkingTreeDirty()
in storage/git and fold it into the status freshness decision.

* fix(ci): use single-slash deny globs in the review agent's disallowedTools

github.workspace already expands to an absolute path, so Read(/${{ github.workspace }}/**)
and Read(//proc/**),(//sys/**),(//dev/**) produced double-slash patterns that a
normalizing matcher may not match — silently no-opping the deny layer. Not
exploitable (the allowlist is the primary control and never grants those
paths), but the globs should be well-formed. Update the pinned test strings.

* ci: install gitnexus-shared with npm ci from the committed lockfile

The gitnexus-shared build floated its deps via npm install in three workflows
(skill-sync, ci-tests, and — most importantly — the release publish.yml) while
every other install step uses npm ci. The lockfile is committed and in sync, so
switch all three to npm ci for reproducible, locked installs.

* test(cli): make the shipped-skills drift guard reject symlinks

listFilesRecursive walked with readdirSync and snapshotDir read with
readFileSync, both of which follow symlinks — so a mirror file symlinked to the
canonical tree passed the byte-compare (and a symlinked mirror dir would be
followed too). Reject a symlinked root via lstat and any symlinked entry via
Dirent.isSymbolicLink, with negative tests (skipped on Windows).

* test(eval): guard the candidate-skill vs mirror-root coverage invariant

MIRROR_SKILL_ROOTS omits the Cursor tree, safe only because no candidate skill
is cursor-shipped. Pin that invariant: every CANDIDATE_SKILLS entry must exist
under canonical + every mirror root and must not ship to Cursor, so adding a
cursor-shipped skill to the candidate set (the PR #2488 asymmetric-sync class)
fails loudly instead of syncing three of four trees.

* docs(ci): describe the review agent's staged post-merge rollout

The DoD asked for a dry-run or triggered run before merge, but an issue_comment
(or newly added workflow_dispatch) workflow only ever executes the default-branch
copy, so it cannot be exercised from the PR that introduces it. Reword the DoD
and the activation checklist to a staged rollout: merge registered-but-disabled,
validate same-repo and fork execution post-merge, then enable the variable.

* fix: pin plugin skill mcp.json to the release version via #2445 tooling

The ten plugin skill mcp.json launched `npx -y gitnexus@latest mcp` on every
skill connect — non-reproducible and a supply-chain surface, and (unlike the
persisted setup config) never pinned. Extend sync-plugin-manifests.mjs with an
mcp surface kind that stamps the gitnexus@<version> launch arg, pin all ten to
1.6.9 now, and keep them byte-identical so the drift guard stays green. The
release lifecycle + publish.yml --check now re-stamp them like the four manifest
surfaces; only READMEs stay on @latest as docs.

* test(eval): prove the proposer's built-in file tools are confined

The real-Claude canary only exercised Bash + MCP, so it proved process/MCP
containment but not that the proposer's built-in file tools stay inside their
mounts. Add a canary over the exact PROPOSER_ALLOWED_TOOLS surface and the same
read-only /evidence mount as run_proposer (allowlist extracted to a shared
constant so it can't drift): Read reaches /evidence, a Write into the read-only
evidence mount is denied, and a Write lands in the output tree.

* fix(eval): apply the candidate overlay after task setup for fair arms

The candidate overlay was applied before the task's untrusted setup ran, so
setup could observe candidate prose and the incumbent/candidate arms started
from different pre-overlay state. Reorder within the sandbox: capture the base
(pre-overlay) skill digest, run setup against the base skills, verify setup did
not tamper them, then apply the overlay and capture the post-overlay digest the
model must preserve. apply_candidate_overlay stages path-specific overlay files,
so setup's uncommitted changes stay out of the baseline and churn is unchanged.

Graph freshness for the review arm is handled by the status dirty-tree fix plus
the review skill's stale-triggered re-index, not by reordering the cached
per-task-sha graph materialization (which is mechanically blocked).

* test(eval): end-to-end containment proof of the autonomous proposer

Drives the real run_proposer through bubblewrap with a deterministic scripted
model (no paid API): it reads the read-only evidence bundle and writes a
candidate gitnexus-plan skill edit plus a rationale into the sandbox output
tree; run_proposer enforces the trust boundary and copies only the validated
overlay + proposal out. This exercises the autonomous-proposal stage of the
self-evolution loop end-to-end in the eval/containment CI job (the gate and
apply stages are covered by test_workflow_bench_evolution and
test_promotion_apply). Env-gated on GITNEXUS_REQUIRE_CLAUDE_CANARY, so it runs
only where the pinned Claude binary and user namespaces are available.

* fix(eval): let the proposer author its overlay via Bash

Running the end-to-end proposer canary in the containment CI job surfaced a real
bug: run_proposer starts the session with --bare, which hard-disables the
Write/Edit tools ("Write exists but is not enabled in this context"), yet
allowlisted Edit/Write and omitted Bash. The proposer therefore had no working
way to write its candidate overlay — the self-evolution loop could never produce
a candidate. The sandbox settings already pre-authorize Bash
(autoAllowBashIfSandboxed) and confine writes to workspace/tmp/home, so switch
PROPOSER_ALLOWED_TOOLS to Read/Grep/Glob/Bash and tell the proposer to author
files with Bash. The end-to-end test now drives the real run_proposer through
bubblewrap and asserts a validated overlay + proposal are produced (this also
replaces the earlier file-tool canary, whose Write/Edit premise was moot).

* test(eval): author the proposer overlay with newline-free Bash content

The nested shell-sandbox prefix mangles embedded newlines, so the multi-line
overlay content never landed. Use single-line content for the deterministic
proposer canary.

* test(eval): drop the unverifiable end-to-end proposer canary

The scripted proposer overlay never materialized in the containment job across
runs, and the model tool-result content is not visible in CI logs, so the test
cannot be finalized without an environment where the sandbox can actually run.
Keep the verified production fix (Bash-authoring in run_proposer); the proposer
sandbox/containment stays covered by the existing Bash+MCP and process-tree
canaries.

* test(cli): drop run-analyze.ts from the windowsHide spawn-family list

U7 moved run-analyze.ts's only child_process call (the git status --porcelain
dirty check) into storage/git.ts (already covered by this test, with
windowsHide). run-analyze.ts no longer imports a spawn-family function, so the
windowsHide-regression test's 'must have >=1 spawn call' invariant failed for
it. Remove it from SRC_FILES.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Zander Raycraft <zanderjraycraft@gmail.com>
Co-authored-by: Azizur Rahman <azizur100389@gmail.com>
2026-07-19 15:07:24 +01:00

1331 lines
60 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Benchmark the gitnexus-plan/work workflow against a baseline agent.
Usage:
uv run --locked --extra dev python -m workflow_bench.runner \
--tasks workflow_bench/tasks.scenarios.yaml --runs 3 \
--model claude-sonnet-4-20250514
Each task runs in a fresh detached git worktree of the target repo, once per
arm per run:
* ``workflow`` — two headless Claude Code sessions: gitnexus-plan, then
gitnexus-work on the produced plan.
* ``candidate_workflow`` / ``candidate_workflow_direct`` — the matching
workflow arm with a prompt-only candidate overlay committed in its clone.
* ``baseline`` — one headless session with the same task text and the Skill
tool disallowed (so it cannot borrow the workflow), everything else equal.
Token usage, cost, duration, and turn counts come from the CLI's own
``--output-format json`` report — nothing is estimated. Caveat: the report's
top-level ``usage`` counts ONLY the main-loop session; ``total_cost_usd`` is
the only reported number that includes subagent spend. A task's model-visible
``verify`` command is retained as an authored-test quality signal; ``resolved``
also requires its harness-owned hidden behavioral oracle. Token savings on
unresolved runs are reported but flagged, because saving tokens by failing is
not a saving.
Trust model: task files and candidate prompts are executable input. Every
setup, verifier, and model session runs inside a preflighted Linux Bubblewrap
boundary with an allowlisted environment, isolated home, PID namespace,
self-contained clone, and task-declared read-only dependencies. Unsupported
or unavailable containment fails before model invocation (README § Trust
model).
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
import re
import secrets
import stat
import statistics
import tempfile
import time
from collections.abc import Mapping
from dataclasses import replace
from datetime import UTC, datetime, timedelta
from pathlib import Path
from typing import Any
import yaml
from .evolution import (
CANDIDATE_ARMS,
EVIDENCE_MAX_AGE_DAYS,
EVALUATED_ARM_SKILLS,
MAIN_LOOP_ONLY_METRICS,
MAIN_LOOP_ONLY_WARNING,
PROMOTION_METRICS,
apply_candidate_overlay,
candidate_overlay_digest,
evaluate_candidate,
required_candidate_arms,
skill_fingerprint,
)
from .oracle_assets import (
ORACLE_ENV_VAR,
TaskOracleSnapshot,
capture_task_oracles,
sanitize_clone_for_hidden_oracles,
staged_task_oracle,
)
from .process_control import ManagedProcessError
from .promotion_apply import committed_destination_base_digests
from .proposer_sandbox import (
SANDBOX_GITNEXUS as SANDBOX_GITNEXUS,
SANDBOX_GITNEXUS_REGISTRY,
SANDBOX_GITNEXUS_SHARED as SANDBOX_GITNEXUS_SHARED,
SANDBOX_WORKSPACE,
ReadOnlyMount,
SandboxError,
SandboxSession,
build_sandbox_environment,
preflight_bubblewrap,
prepare_sandbox,
require_claude_sandbox_helpers,
)
from .runner_artifacts import (
IMPLEMENTATION_ARMS,
MAX_PATCH_BYTES as MAX_PATCH_BYTES,
MAX_WORKSPACE_SNAPSHOT_ENTRIES as MAX_WORKSPACE_SNAPSHOT_ENTRIES,
MAX_WORKSPACE_SNAPSHOT_FILE_BYTES as MAX_WORKSPACE_SNAPSHOT_FILE_BYTES,
MAX_WORKSPACE_SNAPSHOT_PATH_BYTES as MAX_WORKSPACE_SNAPSHOT_PATH_BYTES,
_bounded_regular_bytes,
_prepare_untracked_for_diff,
_sandbox_git,
capture_patch,
diff_churn,
enforce_phase_workspace,
enforce_work_evidence,
implementation_diff_digest,
make_worktree,
new_plan_doc,
parse_shortstat as parse_shortstat,
remove_clone,
require_skill_fingerprint,
run_verify,
snapshot_plan_docs,
VerificationResult,
workspace_snapshot,
)
from .runner_sessions import (
BUILTIN_AGENT_TOOLS as BUILTIN_AGENT_TOOLS,
GITNEXUS_MUTATING_TOOLS as GITNEXUS_MUTATING_TOOLS,
GITNEXUS_READ_ONLY_TOOLS as GITNEXUS_READ_ONLY_TOOLS,
MAX_TRANSCRIPT_BYTES as MAX_TRANSCRIPT_BYTES,
SANDBOX_GITNEXUS_ENTRYPOINT as SANDBOX_GITNEXUS_ENTRYPOINT,
USAGE_FIELDS,
allowed_agent_tools,
run_claude,
sandbox_mcp_config,
sum_sessions,
)
from .runner_tasks import (
normalized_model_identifier,
resolve_task_bindings,
select_tasks,
selected_task_bindings as selected_task_bindings,
)
from .sanitized_graph import (
SanitizedGraphSnapshot,
prepare_sanitized_graph,
validate_no_prebuilt_graph_assets,
)
from .runtime_mounts import (
CE_ARMS,
HARNESS_ROOT as HARNESS_ROOT,
PINNED_GITNEXUS_VERSION as PINNED_GITNEXUS_VERSION,
ce_plugin_dir_for_arm,
ce_plugin_mounts_for_arm,
staged_ce_plugin_snapshot,
trusted_gitnexus_runtime_mounts,
validate_ce_plugin_inputs,
)
from .task_assets import TaskAssetCache, TaskAssetSnapshot, stage_task_assets
PLAN_PROMPT = (
"Use the gitnexus-plan skill for: {task}\n"
"Headless run: make reasonable choices without asking; the plan document "
"is the deliverable."
)
# Appended to every work-arm prompt. In a headless `claude -p` session there
# is no later turn: backgrounded test runs and scheduled wakeups never come
# back, so a session that "waits" for verification ends unverified (observed:
# a work arm backgrounded its slow tests, scheduled three wakeups that never
# fired, and reported done while two tests failed).
HEADLESS_VERIFY = (
" Verification must be observed inside this session: run the typecheck "
"and test commands in the foreground to completion and report their "
"actual output — never background them or wait on scheduled wakeups."
)
WORK_PROMPT = (
"Use the gitnexus-work skill to execute the plan at {plan}.\n"
"Headless run: proceed without asking; report Definition of Done status "
"at the end." + HEADLESS_VERIFY
)
WORK_DIRECT_PROMPT = (
"Use the gitnexus-work skill for: {task}\n"
"Headless run: proceed without asking. The user explicitly declines a "
"separate planning pass — execute in direct mode with the skill's "
"execution discipline." + HEADLESS_VERIFY
)
BASELINE_PROMPT = (
"{task}\n\n"
"Implement the change in this repository and verify it by running the "
"relevant tests. Work autonomously without asking questions."
)
# External-comparator arms: the compound-engineering plugin's plan/work family,
# prompted with the same structure as the gitnexus arms so only the skill
# family differs. The plugin ships user-level, so clones need no repo files.
CE_PLAN_PROMPT = (
"Use the ce-plan skill (compound-engineering plugin) for: {task}\n"
"Headless run: make reasonable choices without asking; the plan document "
"is the deliverable."
)
CE_WORK_PROMPT = (
"Use the ce-work skill (compound-engineering plugin) to execute the plan "
"at {plan}.\n"
"Headless run: proceed without asking; report completion status at the "
"end." + HEADLESS_VERIFY
)
CE_WORK_DIRECT_PROMPT = (
"Use the ce-work skill (compound-engineering plugin) for: {task}\n"
"Headless run: proceed without asking. The user explicitly declines a "
"separate planning pass — execute directly with the skill's execution "
"discipline." + HEADLESS_VERIFY
)
# Review cell: the task's `setup` applies the diff under review as local
# changes; both arms review the same working tree and write to the same file
# so `verify` can gate on a produced review.
REVIEW_PROMPT = (
"Use the gitnexus-review skill to review the local uncommitted changes "
"in this repository. {task}\n"
"Headless run: proceed without asking; do not post to GitHub or anywhere "
"external; write the complete review to review-output.md in the "
"repository root."
)
CE_REVIEW_PROMPT = (
"Use the ce-code-review skill (compound-engineering plugin) to review "
"the local uncommitted changes in this repository. {task}\n"
"Headless run: proceed without asking; do not post to GitHub or anywhere "
"external; write the complete review to review-output.md in the "
"repository root."
)
# Skill each arm's session(s) must actually invoke; a session that never ran
# its skill is a silent no-op arm, not a data point (checked via transcript).
ARM_EXPECTED_SKILLS: dict[str, tuple[str, ...]] = {
"workflow": ("gitnexus-plan", "gitnexus-work"),
"ce_workflow": ("ce-plan", "ce-work"),
"workflow_direct": ("gitnexus-work",),
"ce_workflow_direct": ("ce-work",),
"review": ("gitnexus-review",),
"ce_review": ("ce-code-review",),
}
def _require_implementation_fingerprint(
session: dict[str, Any],
worktree: Path,
arm: str,
expected: str | None,
) -> None:
"""Bind a just-finished implementation session to its original skill bytes."""
try:
require_skill_fingerprint(
worktree,
arm,
expected,
phase="implementation",
)
except ValueError as exc:
if session.get("error_kind") is None:
session["ok"] = False
session["error_kind"] = "implementation-evidence-invalid"
session["error_detail"] = str(exc)
else:
session.setdefault("evidence_diagnostics", []).append(str(exc))
def _verification_outcome(result: VerificationResult | tuple[bool, str]) -> tuple[bool, str]:
if isinstance(result, VerificationResult):
if result.process.state != "exited":
# Hidden-oracle output can contain mounted test bytes. Preserve
# terminal-state evidence without letting candidate-controlled
# stdout/stderr enter results.jsonl through the exception string.
safe_process = replace(
result.process,
stdout_tail="",
stderr_tail="",
detail=result.process.detail or "verifier infrastructure failed",
)
raise ManagedProcessError(result.command, safe_process)
return result.passed, result.output
return result
def _run_hidden_oracle(
snapshot: TaskOracleSnapshot,
worktree: Path,
args: argparse.Namespace,
sandbox: SandboxSession,
) -> tuple[bool, str]:
"""Stage a captured oracle after the model exits, execute it, then erase it."""
if worktree.expanduser().absolute() != sandbox.clone.expanduser().absolute():
raise SandboxError("hidden oracle sandbox does not bind the credited worktree")
mount_name = f".wfbench-oracle-{secrets.token_hex(16)}"
mount_point = worktree / mount_name
mount_point.mkdir(mode=0o700)
primary: BaseException | None = None
try:
with staged_task_oracle(sandbox.private_root, snapshot) as stage_root:
oracle_env = build_sandbox_environment()
# A private RO bind at a random workspace sibling preserves each
# oracle's ../gitnexus import as the candidate implementation. The
# empty mountpoint exists only post-model and is removed before the
# credited patch is captured.
oracle_mount = f"{SANDBOX_WORKSPACE}/{mount_name}"
oracle_env[ORACLE_ENV_VAR] = oracle_mount
passed, _output = _verification_outcome(
run_verify(
snapshot.command,
sandbox.clone,
args.timeout,
command_prefix=sandbox.command_prefix_for(
read_only_workspace=True,
unshare_network=True,
extra_read_only_mounts=(ReadOnlyMount(source=stage_root, target=oracle_mount),),
),
env=oracle_env,
require_pid_namespace=True,
)
)
# Candidate code executes in this process. Never persist its stdout
# or stderr: it can read the mounted hidden test bytes and print them.
return passed, "hidden oracle passed" if passed else "hidden oracle failed"
except BaseException as exc:
primary = exc
raise
finally:
try:
metadata = mount_point.lstat()
if stat.S_ISLNK(metadata.st_mode) or not stat.S_ISDIR(metadata.st_mode):
raise SandboxError("hidden oracle mountpoint changed type during verification")
mount_point.rmdir()
except (OSError, SandboxError) as cleanup:
if primary is None:
raise
primary.add_note(f"hidden oracle mountpoint cleanup also failed: {cleanup}")
def _evaluated_skill_roots(worktree: Path, arm: str) -> tuple[Path, ...]:
"""Repo-local prompt roots that must remain immutable during a session."""
return tuple(worktree / ".claude" / "skills" / name for name in EVALUATED_ARM_SKILLS.get(arm, ()))
def isolated_gitnexus_registry_mount(worktree: Path, parent: Path) -> ReadOnlyMount:
"""Create a one-clone registry that cannot route MCP to any host repo."""
metadata_path = worktree / ".gitnexus" / "gitnexus.json"
if not metadata_path.exists():
metadata_path = worktree / ".gitnexus" / "meta.json"
mode = metadata_path.lstat().st_mode
if stat.S_ISLNK(mode) or not stat.S_ISREG(mode):
raise SandboxError(f"benchmark index metadata must be regular and non-symlink: {metadata_path}")
raw = _bounded_regular_bytes(metadata_path, limit=2 * 1024 * 1024)
try:
metadata = json.loads(raw)
except json.JSONDecodeError as exc:
raise SandboxError(f"benchmark index metadata is malformed: {metadata_path}") from exc
if not isinstance(metadata, dict):
raise SandboxError(f"benchmark index metadata must be an object: {metadata_path}")
indexed_at = metadata.get("indexedAt")
last_commit = metadata.get("lastCommit")
if not isinstance(indexed_at, str) or not indexed_at or not isinstance(last_commit, str) or not last_commit:
raise SandboxError("benchmark index metadata is missing indexedAt or lastCommit")
parent = parent.expanduser().absolute()
registry = Path(tempfile.mkdtemp(prefix="wfbench-registry-", dir=parent))
registry.chmod(0o700)
entry: dict[str, Any] = {
"name": "benchmark-target",
"path": SANDBOX_WORKSPACE,
"storagePath": f"{SANDBOX_WORKSPACE}/.gitnexus",
"indexedAt": indexed_at,
"lastCommit": last_commit,
}
for field in ("remoteUrl", "stats", "branch"):
if field in metadata:
entry[field] = metadata[field]
registry_file = registry / "registry.json"
descriptor = os.open(
registry_file,
os.O_WRONLY | os.O_CREAT | os.O_EXCL | getattr(os, "O_NOFOLLOW", 0),
0o600,
)
try:
os.fchmod(descriptor, 0o600)
payload = (json.dumps([entry], sort_keys=True, separators=(",", ":")) + "\n").encode()
view = memoryview(payload)
while view:
written = os.write(descriptor, view)
view = view[written:]
finally:
os.close(descriptor)
return ReadOnlyMount(source=registry, target=SANDBOX_GITNEXUS_REGISTRY)
def run_arm(
arm: str,
task: dict[str, Any],
worktree: Path,
args: argparse.Namespace,
*,
sandbox: SandboxSession,
transcript_output_dir: Path | None = None,
transcript_output_prefix: str | None = None,
expected_skill_digest: str | None = None,
enforce_phase_boundary: bool = False,
ce_plugin_dir: str | None = None,
oracle_snapshot: TaskOracleSnapshot | None = None,
) -> dict[str, Any]:
sessions: list[dict[str, Any]] = []
env = build_sandbox_environment(
auth_token=args.auth_token,
base_url=args.base_url,
)
common = {
"claude_bin": sandbox.claude_bin,
"timeout": args.timeout,
"model": args.model,
"env": env,
"permission_mode": "dontAsk",
"command_prefix": sandbox.command_prefix_for(
read_only_paths=_evaluated_skill_roots(worktree, arm),
),
"require_pid_namespace": True,
"bare": True,
"settings_json": sandbox.settings_json,
"strict_mcp_config": True,
"mcp_config_json": sandbox_mcp_config(),
"transcript_projects": sandbox.transcript_projects,
"transcript_cwd": Path(SANDBOX_WORKSPACE),
"transcript_wait_seconds": 5,
"transcript_output_dir": transcript_output_dir,
"transcript_output_prefix": transcript_output_prefix,
"transcript_secrets": tuple(secret for secret in (args.auth_token,) if secret),
}
if ce_plugin_dir is not None:
common["plugin_dirs"] = (ce_plugin_dir,)
expected_skills = ARM_EXPECTED_SKILLS.get(arm, ())
plan_doc: Path | None = None
if arm in ("workflow", "ce_workflow"):
plan_prompt = PLAN_PROMPT if arm == "workflow" else CE_PLAN_PROMPT
work_prompt = WORK_PROMPT if arm == "workflow" else CE_WORK_PROMPT
pre = snapshot_plan_docs(worktree)
phase_before = workspace_snapshot(worktree) if enforce_phase_boundary else None
plan_session = run_claude(
plan_prompt.format(task=task["prompt"]),
worktree,
expected_skill=expected_skills[0],
**{**common, "allowed_tools": allowed_agent_tools(implementation=False)},
)
sessions.append(plan_session)
if plan_session["ok"]:
try:
plan_doc = new_plan_doc(worktree, pre)
if phase_before is not None:
enforce_phase_workspace(
worktree,
phase_before,
allowed_artifact=plan_doc,
)
require_skill_fingerprint(
worktree,
arm,
expected_skill_digest,
phase="planning",
)
except ValueError as exc:
plan_session["ok"] = False
plan_session["error_kind"] = "plan-evidence-invalid"
plan_session["error_detail"] = str(exc)
else:
work_session = run_claude(
work_prompt.format(plan=plan_doc.relative_to(worktree)),
worktree,
expected_skill=expected_skills[1],
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
)
_require_implementation_fingerprint(
work_session,
worktree,
arm,
expected_skill_digest,
)
sessions.append(work_session)
elif arm == "ce_workflow_direct":
work_session = run_claude(
CE_WORK_DIRECT_PROMPT.format(task=task["prompt"]),
worktree,
expected_skill=expected_skills[0],
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
)
_require_implementation_fingerprint(
work_session,
worktree,
arm,
expected_skill_digest,
)
sessions.append(work_session)
elif arm in ("review", "ce_review"):
review_prompt = REVIEW_PROMPT if arm == "review" else CE_REVIEW_PROMPT
phase_before = workspace_snapshot(worktree) if enforce_phase_boundary else None
review_session = run_claude(
review_prompt.format(task=task["prompt"]),
worktree,
expected_skill=expected_skills[0],
**{**common, "allowed_tools": allowed_agent_tools(implementation=False)},
)
sessions.append(review_session)
if review_session["ok"] and phase_before is not None:
try:
enforce_phase_workspace(
worktree,
phase_before,
allowed_artifact=worktree / "review-output.md",
)
require_skill_fingerprint(
worktree,
arm,
expected_skill_digest,
phase="review",
)
except ValueError as exc:
review_session["ok"] = False
review_session["error_kind"] = "review-evidence-invalid"
review_session["error_detail"] = str(exc)
elif arm == "workflow_direct":
work_session = run_claude(
WORK_DIRECT_PROMPT.format(task=task["prompt"]),
worktree,
expected_skill=expected_skills[0],
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
)
_require_implementation_fingerprint(
work_session,
worktree,
arm,
expected_skill_digest,
)
sessions.append(work_session)
elif arm == "baseline_nomcp":
# Isolates the workflow-discipline question from the GitNexus-tools
# question: no skills AND no graph tools.
sessions.append(
run_claude(
BASELINE_PROMPT.format(task=task["prompt"]),
worktree,
disallowed_tools=["Skill", "mcp__gitnexus"],
**{
**common,
"mcp_config_json": '{"mcpServers":{}}',
"allowed_tools": allowed_agent_tools(
implementation=True,
include_mcp=False,
),
},
)
)
else:
sessions.append(
run_claude(
BASELINE_PROMPT.format(task=task["prompt"]),
worktree,
disallowed_tools=["Skill"],
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
)
)
record = sum_sessions(sessions)
record["arm"] = arm
record["plan_produced"] = arm not in ("workflow", "ce_workflow") or plan_doc is not None
authored_tests_passed, authored_test_output = _verification_outcome(
run_verify(
task["verify"],
worktree,
args.timeout,
command_prefix=sandbox.command_prefix_for(
read_only_workspace=True,
unshare_network=True,
),
env=build_sandbox_environment(),
require_pid_namespace=True,
)
)
if oracle_snapshot is None:
oracle_passed, oracle_output = False, "hidden oracle snapshot unavailable"
else:
oracle_passed, oracle_output = _run_hidden_oracle(
oracle_snapshot,
worktree,
args,
sandbox,
)
record["authored_tests_passed"] = authored_tests_passed
record["authored_test_output"] = authored_test_output
record["oracle_passed"] = oracle_passed
record["oracle_output"] = oracle_output
record["resolved"] = record["ok"] and authored_tests_passed and oracle_passed
# Compatibility alias for existing report consumers. The authored tests are
# now an explicit signal and can never self-certify resolution.
record["verify_output"] = authored_test_output
if oracle_snapshot is not None:
record.update(
{
"oracle_digest": oracle_snapshot.digest,
"oracle_command_digest": oracle_snapshot.command_digest,
"oracle_manifest_digest": oracle_snapshot.manifest_digest,
}
)
if record["error_kind"] is None and not authored_tests_passed:
# The sessions completed — the produced change just failed the task's
# verify command. Kept distinct from session-error so aggregates can
# exclude infrastructure deaths without hiding real failures.
record["error_kind"] = "verify-failed"
elif record["error_kind"] is None and not oracle_passed:
record["error_kind"] = "oracle-failed" if oracle_snapshot is not None else "oracle-unavailable"
return record
# ─── Pure aggregation/report helpers (unit-tested) ──────────────────────────
CHURN_FIELDS = ("diff_files", "diff_insertions", "diff_deletions")
# Rows where the session (or the harness) died carry no measured evidence and
# must not skew efficiency medians or resolve denominators. verify-failed and
# skill-not-invoked rows DO count: those sessions ran and spent real tokens.
EXCLUDED_ERROR_KINDS = frozenset({"session-error", "infra-error", "evidence-unverified", "cleanup-failure"})
# A sustained upstream outage shows up as a run of session/infra/cleanup
# failures. (cleanup-failure overwrites the primary error_kind, so a
# session-error whose worktree cleanup also failed still counts.) A task's own
# resolved=False is real signal, not an outage, so it never trips the breaker.
SYSTEMIC_ERROR_KINDS = frozenset({"session-error", "infra-error", "cleanup-failure"})
DEFAULT_OUTAGE_STREAK = 5
def systemic_outage_streak(error_kind: str | None, prior_streak: int) -> int:
"""Consecutive systemic-failure count: +1 on a systemic kind, else reset to 0."""
return prior_streak + 1 if error_kind in SYSTEMIC_ERROR_KINDS else 0
def infra_error_record(exc: BaseException) -> dict[str, Any]:
"""Row for a run the harness itself killed (timeout, setup failure)."""
if isinstance(exc, ManagedProcessError):
process = exc.result
detail = f"{process.state}: {process.detail or process.stderr_tail[-1500:]}"
else:
detail = f"{type(exc).__name__}: {exc}"
record: dict[str, Any] = dict.fromkeys(USAGE_FIELDS, 0)
record.update(
{
"ok": False,
"resolved": False,
"error_kind": "infra-error",
"error_detail": detail[:2000],
"session_ids": [],
"cost_usd": 0.0,
"duration_s": 0.0,
"num_turns": 0,
"plan_produced": False,
"authored_tests_passed": False,
"authored_test_output": "",
"oracle_passed": False,
"oracle_output": "",
"verify_output": "",
"skill_invoked": None,
"transcript_missing": False,
}
)
return record
def aggregate(records: list[dict[str, Any]]) -> dict[str, Any]:
"""Median metrics + resolve rate across repeated runs of one task+arm.
Session/infra-error rows are excluded from the medians (they measured
nothing); ``valid_runs``/``excluded_runs`` make the exclusion visible.
"""
valid = [r for r in records if r.get("error_kind") not in EXCLUDED_ERROR_KINDS]
metrics = (*USAGE_FIELDS, "duration_s", "num_turns", *CHURN_FIELDS)
out: dict[str, Any] = {m: statistics.median(r.get(m, 0) for r in (valid or [{}])) for m in metrics}
# cost_usd can be None (unmeasured) on an otherwise-valid run; a single
# unmeasured run makes the whole median unavailable so the gate won't rank
# a candidate on a cost that was never actually captured.
valid_costs = [r.get("cost_usd") for r in valid]
out["cost_usd"] = None if (not valid or any(cost is None for cost in valid_costs)) else statistics.median(valid_costs)
out["resolved"] = sum(1 for r in records if r["resolved"])
out["runs"] = len(records)
out["valid_runs"] = len(valid)
out["excluded_runs"] = len(records) - len(valid)
out["transcripts_missing"] = sum(1 for r in records if r.get("transcript_missing"))
out["class"] = records[0].get("class", "")
return out
def savings(baseline: dict[str, Any], workflow: dict[str, Any]) -> dict[str, Any]:
"""Percent saved by the workflow arm per metric (positive = cheaper)."""
out: dict[str, Any] = {}
for metric in (*USAGE_FIELDS, "cost_usd", "duration_s"):
base = baseline.get(metric)
arm = workflow.get(metric)
if base is None or arm is None:
out[metric] = None
else:
out[metric] = round(100 * (base - arm) / base, 1) if base else 0.0
return out
def _na(value: Any) -> Any:
"""Render an unmeasured metric as ``n/a`` instead of a misleading number."""
return "n/a" if value is None else value
def _cost_cell(value: Any) -> str:
return "n/a" if value is None else f"{value:.4f}"
def render_report(results: dict[str, dict[str, dict[str, Any]]]) -> str:
"""results: {task_id: {arm: aggregate}} → markdown report."""
lines = [
"# gitnexus workflow benchmark",
"",
"Medians across runs; savings rows = (baseline arm) / baseline per arm.",
"A negative saving means that arm spent more than baseline. churn =",
"files/+insertions/deletions vs the worktree's starting commit.",
"",
"**WARNING:** token columns count only each arm's main-loop session —",
"subagent spend is invisible to them and flatters subagent-heavy arms.",
"cost $ is the only column that includes subagent spend; to rank token",
"efficiency, sum usage from the session transcripts instead",
"(dedup events sharing one message.id).",
"",
"| task | class | arm | resolved | input | cache_create | cache_read | output | cost $ | wall s | turns | churn |",
"| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |",
]
for task_id, arms in results.items():
for arm, agg in arms.items():
excluded = agg.get("excluded_runs", 0)
resolved_cell = f"{agg['resolved']}/{agg.get('valid_runs', agg['runs'])}"
if excluded:
resolved_cell += f" ({excluded} excluded)"
lines.append(
f"| {task_id} | {agg['class']} | {arm} | {resolved_cell} "
f"| {agg['input_tokens']:.0f} | {agg['cache_creation_input_tokens']:.0f} "
f"| {agg['cache_read_input_tokens']:.0f} | {agg['output_tokens']:.0f} "
f"| {_cost_cell(agg['cost_usd'])} | {agg['duration_s']:.0f} | {agg['num_turns']:.0f} "
f"| {agg['diff_files']:.0f}/+{agg['diff_insertions']:.0f}/{agg['diff_deletions']:.0f} |"
)
for arm in arms:
if arm != "baseline" and "baseline" in arms:
s = savings(arms["baseline"], arms[arm])
lines.append(
f"| {task_id} | {arms[arm]['class']} | **{arm} savings %** | — "
f"| {s['input_tokens']} | {s['cache_creation_input_tokens']} "
f"| {s['cache_read_input_tokens']} | {s['output_tokens']} "
f"| {_na(s['cost_usd'])} | {s['duration_s']} | — | — |"
)
lines.append("")
all_aggs = [agg for arms in results.values() for agg in arms.values()]
excluded_total = sum(agg.get("excluded_runs", 0) for agg in all_aggs)
if excluded_total:
lines.append(
f"{excluded_total} run(s) hit session/infra errors or had unverifiable "
"evidence and were excluded "
"from medians and resolve denominators — see error_kind in results.jsonl."
)
missing_total = sum(agg.get("transcripts_missing", 0) for agg in all_aggs)
if missing_total:
lines.append(
f"{missing_total} run(s) had no locatable session transcript or it was "
"unreadable, so they were excluded from promotion evidence "
"(skill_invoked=null in results.jsonl)."
)
lines.append(
"Session ids for every run are in results.jsonl — open the matching "
"transcript to see where each arm spent its tokens."
)
return "\n".join(lines)
# ─── Main ────────────────────────────────────────────────────────────────────
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--tasks", required=True, type=Path)
parser.add_argument("--runs", type=int, default=1)
parser.add_argument(
"--outage-streak",
type=int,
default=DEFAULT_OUTAGE_STREAK,
help="abort the sweep after this many consecutive session/infra/cleanup "
"failures (0 disables the circuit breaker)",
)
parser.add_argument(
"--arms",
nargs="+",
default=["workflow", "workflow_direct", "baseline"],
choices=[
"workflow",
"candidate_workflow",
"workflow_direct",
"candidate_workflow_direct",
"ce_workflow",
"ce_workflow_direct",
"review",
"ce_review",
"baseline",
"baseline_nomcp",
],
)
parser.add_argument("--claude-bin", default="claude")
parser.add_argument(
"--ce-plugin-dir",
type=Path,
default=None,
help="operator-supplied Compound Engineering plugin directory; required for ce_* arms",
)
parser.add_argument(
"--ce-plugin-version",
default=None,
help="exact Compound Engineering plugin version; required for ce_* arms",
)
parser.add_argument("--timeout", type=int, default=3600, help="per session, seconds")
parser.add_argument("--out", type=Path, default=None)
parser.add_argument(
"--model",
required=True,
help="named, versioned model passed to every `claude --model` invocation",
)
parser.add_argument(
"--proposer-model",
default=None,
help="model that generated the candidate overlay (recorded for provenance)",
)
parser.add_argument(
"--base-url",
default=None,
help="ANTHROPIC_BASE_URL override — point at an Anthropic-compatible "
"proxy (see free-model.litellm.yaml) to run on a free model",
)
parser.add_argument(
"--auth-token",
default=os.environ.get("GITNEXUS_BENCH_AUTH_TOKEN"),
help="ANTHROPIC_API_KEY for the --base-url endpoint (prefer GITNEXUS_BENCH_AUTH_TOKEN env)",
)
parser.add_argument(
"--include-expensive",
action="store_true",
help="include scenarios marked expensive: true (excluded by default)",
)
parser.add_argument(
"--candidate-overlay",
type=Path,
default=None,
help="directory mirroring .claude/skills/gitnexus-{plan,work}; applied only to candidate_* arms",
)
parser.add_argument(
"--promotion-metric",
choices=PROMOTION_METRICS,
default="cost_usd",
help="efficiency metric used by the deterministic candidate gate; "
"cost_usd (default) is the only CLI-reported number that includes "
"subagent spend — token metrics count only the main loop",
)
parser.add_argument("--promotion-min-runs", type=int, default=3)
parser.add_argument("--promotion-min-improvement", type=float, default=5.0)
parser.add_argument("--promotion-max-task-regression", type=float, default=20.0)
parser.add_argument("--task-bindings-json", default=None, help=argparse.SUPPRESS)
parser.add_argument("--promotion-target-bases-json", default=None, help=argparse.SUPPRESS)
return parser
def main() -> None:
parser = build_parser()
args = parser.parse_args()
try:
args.model = normalized_model_identifier(args.model)
args.proposer_model = (
normalized_model_identifier(args.proposer_model, flag="--proposer-model")
if args.proposer_model is not None
else None
)
task_document = yaml.safe_load(args.tasks.read_text())
if not isinstance(task_document, Mapping) or not isinstance(task_document.get("tasks"), list):
raise ValueError("task file must contain a tasks list")
tasks, skipped_expensive = select_tasks(
task_document["tasks"],
include_expensive=args.include_expensive,
)
oracle_snapshots = capture_task_oracles(tasks)
expected_task_bindings = json.loads(args.task_bindings_json) if args.task_bindings_json else None
if expected_task_bindings is not None and not isinstance(expected_task_bindings, list):
raise ValueError("--task-bindings-json must contain a list")
supplied_promotion_target_bases = (
json.loads(args.promotion_target_bases_json) if args.promotion_target_bases_json else {}
)
if not isinstance(supplied_promotion_target_bases, dict) or not all(
isinstance(path, str) and isinstance(digest, str)
for path, digest in supplied_promotion_target_bases.items()
):
raise ValueError("--promotion-target-bases-json must contain a string mapping")
ce_plugin_config = validate_ce_plugin_inputs(
args.arms,
args.ce_plugin_dir,
args.ce_plugin_version,
)
except (OSError, SandboxError, ValueError, yaml.YAMLError) as exc:
parser.error(str(exc))
raise AssertionError("ArgumentParser.error() returned unexpectedly")
candidate_arms = [arm for arm in args.arms if arm in CANDIDATE_ARMS]
if candidate_arms and args.candidate_overlay is None:
parser.error("candidate_* arms require --candidate-overlay")
if args.candidate_overlay is not None and not candidate_arms:
parser.error("--candidate-overlay requires at least one candidate_* arm")
for candidate_arm in candidate_arms:
incumbent_arm = CANDIDATE_ARMS[candidate_arm]
if incumbent_arm not in args.arms:
parser.error(f"{candidate_arm} must be paired with {incumbent_arm}")
if args.runs < 1 or args.promotion_min_runs < 1:
parser.error("--runs and --promotion-min-runs must be positive")
candidate_overlay = args.candidate_overlay.expanduser().absolute() if args.candidate_overlay is not None else None
overlay_digest = candidate_overlay_digest(candidate_overlay) if candidate_overlay is not None else None
if candidate_overlay is not None:
required_candidates = required_candidate_arms(candidate_overlay)
required_arms = [arm for candidate in required_candidates for arm in (CANDIDATE_ARMS[candidate], candidate)]
if args.arms != required_arms:
parser.error("candidate overlay requires exactly these paired arms: " + " ".join(required_arms))
try:
promotion_target_bases = committed_destination_base_digests(candidate_overlay)
except ValueError as exc:
# Overlay adds a promotion target with no committed base — a clean
# CLI error, not a traceback.
parser.error(str(exc))
raise AssertionError("ArgumentParser.error() returned unexpectedly")
if supplied_promotion_target_bases and supplied_promotion_target_bases != promotion_target_bases:
parser.error("--promotion-target-bases-json does not match the committed incumbent")
else:
if supplied_promotion_target_bases:
parser.error("--promotion-target-bases-json requires --candidate-overlay")
promotion_target_bases = {}
try:
bwrap_bin = preflight_bubblewrap()
require_claude_sandbox_helpers()
runtime_mounts = trusted_gitnexus_runtime_mounts()
except SandboxError as exc:
parser.error(str(exc))
raise AssertionError("ArgumentParser.error() returned unexpectedly")
out_dir = args.out or Path("results") / time.strftime("wfbench-%Y%m%d-%H%M%S")
out_dir.mkdir(parents=True, exist_ok=True)
results_path = out_dir / "results.jsonl"
selected_ids = [task["id"] for task in tasks]
print(
f"selected {len(selected_ids)} task(s): {', '.join(selected_ids)}; "
f"skipped {len(skipped_expensive)} expensive task(s): "
f"{', '.join(skipped_expensive) if skipped_expensive else 'none'}"
)
results: dict[str, dict[str, dict[str, Any]]] = {}
outage_streak = 0
outage_tripped = False
with (
tempfile.TemporaryDirectory(prefix="wfbench-trees-") as trees,
TaskAssetCache(Path(trees) / ".task-assets") as task_asset_cache,
staged_ce_plugin_snapshot(
ce_plugin_config,
destination_parent=Path(trees),
) as ce_plugin_snapshot,
):
try:
task_bindings = resolve_task_bindings(
tasks,
expected_task_bindings,
oracle_snapshots=oracle_snapshots,
task_asset_cache=task_asset_cache,
)
except (OSError, SandboxError, ValueError) as exc:
parser.error(str(exc))
raise AssertionError("ArgumentParser.error() returned unexpectedly")
oracle_mask = Path(trees) / ".oracle-mask"
oracle_mask.mkdir(mode=0o500)
oracle_mask.chmod(0o500)
graph_snapshots: dict[tuple[str, str], SanitizedGraphSnapshot] = {}
graph_snapshot_errors: dict[tuple[str, str], BaseException] = {}
for task, task_binding, oracle_snapshot in zip(
tasks,
task_bindings,
oracle_snapshots,
strict=True,
):
if outage_tripped:
break
repo = Path(task_binding["repo_identity"])
task_sha = task_binding["resolved_sha"]
asset_snapshot: TaskAssetSnapshot | None = None
asset_snapshot_error: BaseException | None = None
graph_key = (str(repo), task_sha)
graph_snapshot: SanitizedGraphSnapshot | None = graph_snapshots.get(graph_key)
graph_snapshot_error: BaseException | None = graph_snapshot_errors.get(graph_key)
try:
validate_no_prebuilt_graph_assets(task)
if graph_snapshot is None and graph_snapshot_error is None:
graph_snapshot = prepare_sanitized_graph(
task,
repo=repo,
resolved_sha=task_sha,
parent=Path(trees),
cache=task_asset_cache,
claude_bin=args.claude_bin,
bwrap_bin=bwrap_bin,
runtime_mounts=runtime_mounts,
)
graph_snapshots[graph_key] = graph_snapshot
except (ManagedProcessError, OSError, SandboxError, RuntimeError, ValueError) as exc:
graph_snapshot_error = exc
graph_snapshot_errors[graph_key] = exc
per_arm: dict[str, list[dict[str, Any]]] = {a: [] for a in args.arms}
for run_idx in range(args.runs):
if outage_tripped:
break
for arm in args.arms:
if outage_tripped:
break
worktree: Path | None = None
record: dict[str, Any] | None = None
cleanup_error: OSError | None = None
try:
if asset_snapshot_error is not None:
raise RuntimeError(f"task asset snapshot preparation failed: {asset_snapshot_error}")
if graph_snapshot_error is not None:
raise RuntimeError(f"sanitized graph snapshot preparation failed: {graph_snapshot_error}")
if graph_snapshot is None:
raise RuntimeError("sanitized graph snapshot is unavailable")
if asset_snapshot is None:
try:
asset_snapshot = task_asset_cache.prepare(
task,
repo=repo,
resolved_sha=task_sha,
expected_dependency_binding=task_binding,
)
except (OSError, SandboxError, ValueError) as exc:
asset_snapshot_error = exc
raise
worktree = make_worktree(repo, task_sha, Path(trees))
sanitized_head = sanitize_clone_for_hidden_oracles(worktree)
graph_snapshot.materialize(worktree, sanitized_head=sanitized_head)
dependency_mounts = stage_task_assets(
task,
repo=repo,
clone=worktree,
snapshot=asset_snapshot,
)
registry_mount = isolated_gitnexus_registry_mount(worktree, Path(trees))
hidden_harness = worktree / "eval" / "workflow_bench"
oracle_visibility_mounts: list[ReadOnlyMount] = []
if hidden_harness.exists() or hidden_harness.is_symlink():
hidden_metadata = hidden_harness.lstat()
if stat.S_ISLNK(hidden_metadata.st_mode) or not stat.S_ISDIR(hidden_metadata.st_mode):
raise SandboxError(
"benchmark harness path must be a real directory before it can be hidden"
)
oracle_visibility_mounts.append(
ReadOnlyMount(
source=oracle_mask,
target=f"{SANDBOX_WORKSPACE}/eval/workflow_bench",
)
)
execution_arm = CANDIDATE_ARMS.get(arm, arm)
ce_mounts = ce_plugin_mounts_for_arm(execution_arm, ce_plugin_snapshot)
with prepare_sandbox(
clone=worktree,
claude_bin=args.claude_bin,
bwrap_bin=bwrap_bin,
read_only_mounts=[
*dependency_mounts,
*runtime_mounts,
registry_mount,
*ce_mounts,
*oracle_visibility_mounts,
],
preflight=False,
) as sandbox:
# Capture the BASE (pre-overlay) skill digest — identical
# for the incumbent and candidate arms — then run the
# task's untrusted setup against those base skills. The
# candidate overlay is applied only afterwards, so setup
# can never observe candidate prose and both arms share
# byte-identical pre-overlay state.
base_skill_digest = skill_fingerprint(worktree, execution_arm)
if task.get("setup"):
setup_command = ["/bin/sh", "-lc", str(task["setup"])]
setup = sandbox.run(
setup_command,
timeout=600,
env=build_sandbox_environment(),
)
if not setup.ok:
raise ManagedProcessError(setup_command, setup)
# Tamper-evidence: setup must not have rewritten the base
# skills, verified before any candidate overlay lands.
require_skill_fingerprint(
worktree,
execution_arm,
base_skill_digest,
phase="task setup",
)
if arm in CANDIDATE_ARMS:
assert candidate_overlay is not None
applied_digest = apply_candidate_overlay(
candidate_overlay,
worktree,
sandbox=sandbox,
)
if applied_digest != overlay_digest:
raise RuntimeError("candidate overlay changed during the benchmark run")
# The digest the model must preserve during its run is the
# post-overlay skill surface (candidate skills for
# candidate arms; unchanged base skills otherwise).
expected_skill_digest = skill_fingerprint(worktree, execution_arm)
orig_sha = _sandbox_git(sandbox, ["rev-parse", "HEAD"]).strip()
if not re.fullmatch(r"[0-9a-fA-F]{40,64}", orig_sha):
raise RuntimeError("sandboxed candidate setup did not produce an immutable commit")
before_work_digest = (
implementation_diff_digest(sandbox, orig_sha)
if execution_arm in IMPLEMENTATION_ARMS
else ""
)
record = run_arm(
execution_arm,
task,
worktree,
args,
sandbox=sandbox,
transcript_output_dir=out_dir,
transcript_output_prefix=f"{task['id']}-{arm}-run{run_idx}",
expected_skill_digest=expected_skill_digest,
enforce_phase_boundary=True,
ce_plugin_dir=ce_plugin_dir_for_arm(execution_arm, ce_plugin_snapshot),
oracle_snapshot=oracle_snapshot,
)
_prepare_untracked_for_diff(sandbox)
after_work_digest = (
implementation_diff_digest(
sandbox,
orig_sha,
prepare_untracked=False,
)
if execution_arm in IMPLEMENTATION_ARMS
else ""
)
record.update(
diff_churn(
sandbox,
orig_sha,
prepare_untracked=False,
)
)
enforce_work_evidence(
record,
arm=execution_arm,
before_digest=before_work_digest,
after_digest=after_work_digest,
)
patch_bytes = capture_patch(sandbox, worktree, orig_sha)
record["arm"] = arm
record.update(
{
"model": args.model,
"benchmark_model": args.model,
"proposer_model": args.proposer_model,
"task_ref": task.get("ref", "HEAD"),
"task_base_sha": task_sha,
"sanitized_task_sha": sanitized_head,
"variant_head_sha": orig_sha,
"task_prompt_digest": hashlib.sha256(task["prompt"].encode()).hexdigest(),
"skill_digest": expected_skill_digest,
"candidate_overlay_digest": (overlay_digest if arm in CANDIDATE_ARMS else None),
"recorded_at": datetime.now(UTC).isoformat(),
}
)
# Final working-tree patch — the clone is destroyed, so
# this is the only artifact for diagnosing verify fails.
patch_path = out_dir / f"{task['id']}-{arm}-run{run_idx}.patch"
patch_path.write_bytes(patch_bytes)
except (
ManagedProcessError,
SandboxError,
OSError,
RuntimeError,
ValueError,
) as exc:
# One hung session or failed setup must not abort the
# sweep — record the run as infra-error and move on so
# report.md/promotion.json still get written.
record = infra_error_record(exc)
record["arm"] = arm
print(f"[{task['id']}][{arm}][run {run_idx}] infra-error: {exc}")
finally:
if worktree is not None and worktree.exists():
try:
remove_clone(worktree)
except OSError as exc:
cleanup_error = exc
assert record is not None
if cleanup_error is not None:
primary_kind = record.get("error_kind")
primary_detail = record.get("error_detail")
record["resolved"] = False
record["ok"] = False
record["error_kind"] = "cleanup-failure"
record["error_detail"] = (
f"primary={primary_kind}: {primary_detail}; cleanup: "
f"{type(cleanup_error).__name__}: {cleanup_error}"
)[:2000]
record.update(
{
"task": task["id"],
"class": task.get("class", ""),
"run": run_idx,
"task_asset_snapshot_digest": (
asset_snapshot.digest if asset_snapshot is not None else None
),
"task_asset_manifest_digest": (
asset_snapshot.manifest_digest if asset_snapshot is not None else None
),
"sandbox_dependency_content_digest": (
asset_snapshot.dependency_content_digest if asset_snapshot is not None else None
),
"sandbox_dependency_manifest_digest": (
asset_snapshot.dependency_manifest_digest if asset_snapshot is not None else None
),
"sanitized_graph_snapshot_digest": (
graph_snapshot.digest if graph_snapshot is not None else None
),
"sanitized_graph_manifest_digest": (
graph_snapshot.manifest_digest if graph_snapshot is not None else None
),
"oracle_digest": oracle_snapshot.digest,
"oracle_command_digest": oracle_snapshot.command_digest,
"oracle_manifest_digest": oracle_snapshot.manifest_digest,
"ce_plugin_version": (
ce_plugin_snapshot.version
if arm in CE_ARMS and ce_plugin_snapshot is not None
else None
),
"ce_plugin_manifest_digest": (
ce_plugin_snapshot.manifest_digest
if arm in CE_ARMS and ce_plugin_snapshot is not None
else None
),
}
)
per_arm[arm].append(record)
with results_path.open("a") as fh:
fh.write(json.dumps(record) + "\n")
print(
f"[{task['id']}][{arm}][run {run_idx}] resolved={record['resolved']} "
f"in={record['input_tokens']} out={record['output_tokens']} "
f"cost=${_na(record['cost_usd'])}"
)
outage_streak = systemic_outage_streak(record.get("error_kind"), outage_streak)
if args.outage_streak and outage_streak >= args.outage_streak:
outage_tripped = True
print(
f"[systemic-outage] {outage_streak} consecutive session/infra/cleanup "
"failures — aborting the remaining sweep; report and promotion are written "
"from partial evidence and the run exits non-zero."
)
break
results[task["id"]] = {a: aggregate(rs) for a, rs in per_arm.items() if rs}
selection_report = [
"## Run provenance",
"",
f"Benchmark model: `{args.model}`",
f"Proposer model: `{args.proposer_model}`",
f"Selected tasks ({len(selected_ids)}): {', '.join(selected_ids)}",
(
f"Skipped expensive tasks ({len(skipped_expensive)}): "
+ (", ".join(skipped_expensive) if skipped_expensive else "none")
),
]
if ce_plugin_snapshot is not None:
selection_report.append(
f"Compound Engineering plugin: `{ce_plugin_snapshot.version}` (`{ce_plugin_snapshot.manifest_digest}`)"
)
report = render_report(results) + "\n\n" + "\n".join(selection_report) + "\n"
(out_dir / "report.md").write_text(report)
if candidate_arms:
promotion_generated_at = datetime.now(UTC)
promotion = {
# Schema 3 is the first promotion evidence that requires hidden,
# byte-bound behavioral oracles. Older self-authored-only rows are
# intentionally ineligible for application.
"schema_version": 3,
"generated_at": promotion_generated_at.isoformat(),
"evidence_expires_at": (promotion_generated_at + timedelta(days=EVIDENCE_MAX_AGE_DAYS)).isoformat(),
"benchmark_model": args.model,
"proposer_model": args.proposer_model,
"candidate_origin": ("model-proposer" if args.proposer_model is not None else "manual-initial-overlay"),
"candidate_overlay": str(candidate_overlay),
"candidate_overlay_digest": overlay_digest,
"target_base_digests": promotion_target_bases,
"required_candidate_arms": candidate_arms,
"selected_tasks": task_bindings,
"ce_plugin": ce_plugin_snapshot.provenance if ce_plugin_snapshot is not None else None,
"policy": {
"metric": args.promotion_metric,
"metric_warning": (MAIN_LOOP_ONLY_WARNING if args.promotion_metric in MAIN_LOOP_ONLY_METRICS else None),
"min_runs": args.promotion_min_runs,
"min_improvement_pct": args.promotion_min_improvement,
"max_task_regression_pct": args.promotion_max_task_regression,
"quality_rule": "no per-task resolution-rate regression",
"max_age_days": EVIDENCE_MAX_AGE_DAYS,
},
"decisions": [
evaluate_candidate(
results,
incumbent_arm=CANDIDATE_ARMS[candidate_arm],
candidate_arm=candidate_arm,
model=args.model,
metric=args.promotion_metric,
min_runs=args.promotion_min_runs,
min_improvement_pct=args.promotion_min_improvement,
max_task_regression_pct=args.promotion_max_task_regression,
)
for candidate_arm in candidate_arms
],
}
(out_dir / "promotion.json").write_text(json.dumps(promotion, indent=2) + "\n")
print(f"\n{report}\n\nWritten to {out_dir}/")
if outage_tripped:
# Non-zero exit so a driver (evolve.py) treats the partial benchmark as a
# failed run and halts instead of proposing from outage-truncated evidence.
raise SystemExit(1)
if __name__ == "__main__":
main()