mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-09-20 00:11:37 +00:00
* feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill Adds .claude/skills/ce-plan: a planning-only skill that builds implementation-ready plans from GitNexus graph navigation (query/context/ impact/trace), bounded statement-level PDG slices (pdg_query, impact mode:pdg, explain), and targeted source verification, with a context ledger to prevent repeated reads and a machine-readable implementation context pack (stable contract for a future ce-implement). Whitelisted in .gitignore and registered in AGENTS.md and CLAUDE.md outside the auto-managed gitnexus block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions) Tool contract: impact mode:'pdg' shape now includes the schema-required direction param; CDG branch sense documented as the result 'label' field (reason is cypher/raw-edge only); explain caveats corrected to its real false-negative classes (cross-function TAINT_PATH is modeled). Consistency: PDG slice homed in working memory (ledger keeps one-liners); depth knob defined and category-overrides-baseline ordering stated; call_depth (consumed by nothing) and content-hash bookkeeping dropped; Never section folded into Hard rules; Phase 3 deduplicated to a pointer; allowed-repeat escalations defined; budget/discard accounting clarified; verification-commands gathering added to Phase 4; open_questions added to the context pack. From scenario runs: plans now pin the verified-at HEAD commit and index freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed], quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and support an out:<path> destination override; output path defined as the Phase 1 target repo root. Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata bumps; future ce-implement qualified as future. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): rename ce-plan → gitnexus-plan; add cross-CLI (Codex) entrypoints Renames the skill dir, frontmatter, output filename convention, plan H1 (GitNexus Engineering Plan), the future executor handle (gitnexus-implement), the .gitignore whitelist entry, and all AGENTS.md/CLAUDE.md references. Follows the pr-swarm-review cross-CLI pattern: SKILL.md is the canonical CLI-neutral spec, AGENTS.md § Engineering planning is the Codex/any-agent entrypoint, and the README documents the optional user-level ~/.codex/prompts/gitnexus-plan.md slash command plus an invocation matrix. Skill prose de-branded from Claude Code (agent-neutral verification layer). Also fixes two post-review README contradictions: the anti-reread claim now names the ledger's allowed escalations, and 'read-only by contract' is now 'planning-only' (the skill writes exactly one repo file — the plan); the scope-creep rule and template §12 now agree on where deferred follow-ups land. Drops the stale plugin-collision limitation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(skills): document Codex user-level install path for gitnexus-plan Codex discovers SKILL.md skills from ~/.agents/skills (same path the other gitnexus-* skills install to); README now documents the cp install plus the optional ~/.codex/prompts slash-command file, with the prompt body preferring the repo copy and falling back to the user-level install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): gitnexus-plan freshness gate + active PDG-layer refresh Freshness is now a Phase 1 gate, not advisory: under the default freshness:strict, a stale index is refreshed once per planning session via node .gitnexus/run.cjs analyze --index-only (appending --pdg when the task will reach the PDG phase), then the context resource is re-read. A missing PDG layer likewise triggers the one permitted --index-only --pdg refresh and re-probe instead of a passive recommendation. freshness:accept (or a failed/impractical refresh) preserves the old behavior: plan on the stale graph, source-weighted, labelled in the plan header. --index-only is the load-bearing flag choice — it suppresses all file generation, so the planning-only contract holds (only the .gitnexus store changes). Ledger gains an index_refresh record; plan header states fresh / refreshed / refresh-skipped-with-reason. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): gitnexus-plan runner build check before freshness refresh When the target repo builds the analyzer from its own source (bin → dist/ mapping, as gitnexus/ does), the Phase 1 freshness gate now verifies dist/ is current before running the analyze refresh — rebuilding via the package's build script when any analyzer source file is newer than the built entrypoint — and prefers that freshly built CLI. Otherwise a stale dist re-indexes with outdated extraction logic and the 'fresh' index lies. Rebuilds are recorded in the ledger's index_refresh; the PDG-phase refresh inherits the same check. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): add gitnexus-work executor and gitnexus-lfg pipeline gitnexus-work executes a gitnexus-plan as verified atomic commits: consumes the §11 implementation_context pack, drift-checks the plan's evidence pin against HEAD, re-verifies assumptions before relying on them, runs impact before every symbol edit and detect_changes before every commit (repo mandates), builds tests from the plan's scenarios, and routes structural drift back to gitnexus-plan Deepen mode instead of coding around it. gitnexus-lfg is a thin orchestrator: gitnexus-plan → blocking user gate (deepen / proceed / stop, deepen loops allowed) → gitnexus-work → review via the existing gitnexus-pr-review skill (open PR, else branch diff vs default). One bounded fix cycle for review findings; never pushes or opens a PR on its own. gitnexus-plan gains a Deepen mode (re-run freshness gate, escalate to depth:deep, re-verify graph/inferred/assumed claims toward verified, rewrite the same file); its 'future gitnexus-implement' placeholder is retired in favor of gitnexus-work. Registered via .gitignore whitelists, AGENTS.md 1.10.0 (section renamed to Engineering planning & execution), CLAUDE.md 1.5.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): apply cross-skill review findings to the gitnexus skill family Two P1s: gitnexus-plan Deepen mode now re-anchors before re-pinning (diffs the old evidence pin over every [verified]-claim file and re-reads or downgrades before the header moves — moving the pin without this laundered stale claims as verified); the index-refresh budget is stated once in Phase 1 (one --index-only refresh plus at most one Phase 3 --pdg upgrade per session, Deepen = its own session) with ledger and pdg-slice deferring to it. Contract fixes: gitnexus-work's drift check now covers every file the pack cites (not just files_to_modify) and parses the full pack incl. primary/related symbols and acceptance_criteria (walked in Phase 4 alongside §13); a pre-completed check skips §7 steps already landed and Deepen gains a reconcile-execution-state step, closing the mid-execution route-back loop; pack assumptions must name what to check and how. lfg: Lane 4 passes the merge-base to detect_changes compare (two-dot diff misattributes upstream commits when default advanced), branch-diff is the stated normal case, oversized review findings route to the plan gate instead of overflowing direct mode, the one-fix-cycle cap is explicit on re-run, and headless runs end at the plan gate with the plan as deliverable. work: blank mode narrowed to *gitnexus-plan*.md with a re-execution guard, direct-mode discipline spelled out, branch meaningfulness defined against the plan slug, and the plan document is committed as the branch's docs commit (review diff includes it). Planning-only contract now names the dist/ rebuild as the second permitted state change; Phase 5.1 names the four claim tags; stale AGENTS.md anchors fixed. Known latent issue left untouched: gitnexus/gitnexus-pr-review pairs a three-dot example with a two-dot detect_changes compare — that skill is also shipped by the plugin, so fixing it here would drift the copies; lfg compensates by passing the merge-base. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): ship the engineering skill family with the gitnexus package npm i -g gitnexus users now get gitnexus-plan / gitnexus-work / gitnexus-lfg: the three skills are added to gitnexus/skills/ in directory form (SKILL.md + references/), which installSkillsTo already enumerates dynamically and copies recursively to every editor target (~/.agents/skills for Codex, Cursor, OpenCode, Qoder, ...) on gitnexus setup — uninstall enumerates the same root, so removal stays clean. The Claude Code plugin channel (gitnexus-claude-plugin/skills/) carries the same copies plus the standard per-skill mcp.json. Global-install support in the skill text: gitnexus-plan Phase 1 now resolves the analyzer runner explicitly — node .gitnexus/run.cjs analyze when the project has a runner, else gitnexus analyze (installed CLI), else npx gitnexus analyze — and all analyze mentions route through it, satisfying the skills-steering policy (#1939/#1945) which sweeps the plugin copies. New drift guard test/unit/shipped-skills-sync.test.ts asserts the npm and plugin copies stay byte-identical to the canonical .claude/skills/ family (plugin = canonical + mcp.json), same discipline as run.cjs ↔ resolve-invocation.ts. skills-steering + shipped-skills-sync: 11/11 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(eval): workflow_bench — measure the skill workflow's token savings Benchmarks gitnexus-plan → gitnexus-work against a baseline agent (--disallowedTools Skill) on identical tasks, in fresh detached worktrees, using real headless Claude Code sessions; every number comes from the CLI's --output-format json usage report (field names validated against a live 2.1.207 session). Reports per-arm medians (input/cache/output tokens, cost, wall time, turns), a savings row, and resolve status from a per-task verify command — savings on failed tasks are flagged, not celebrated. Per-task setup hook prepares fresh worktrees (deps); --permission-mode bypassPermissions (default) lets sessions run unattended in the throwaway trees. Free-model support: --base-url/--auth-token/--model route headless sessions through any Anthropic-compatible endpoint; free-model.litellm.yaml is a ready litellm-proxy template for OpenRouter :free variants or local Ollama, so benchmarking burns no paid tokens (README documents rate limits and the small-model skill-following caveat). Harness validated end-to-end with a stub CLI (worktree lifecycle, both arms, plan→work chaining, verify, aggregation, report) and 4 pytest units for the pure aggregation/savings/report helpers. AGENTS.md 1.11.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(eval): record first workflow_bench calibration run Trivial-task calibration (add -V alias): both arms resolved; workflow arm ~4.3x baseline cost — the documented overhead-dominated regime, recorded so the regime boundary is empirical rather than asserted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(eval): workflow_bench scenario matrix — arm variants, task classes, churn Ground-base measurement across scenarios: tasks.scenarios.yaml spans four labeled classes (trivial → investigation-bug → investigation-feature → cross-module) with deterministic verifies (prescribed test files). New arms: workflow_direct (gitnexus-work direct mode — the middle option that locates the routing boundary lfg's gate and work's triage encode) and baseline_nomcp (no skills AND no graph tools — separates workflow-discipline value from GitNexus-tool value; off by default). Records now carry task class and diff churn (files/+ins/−del vs the starting commit) as an over-engineering proxy; the report renders a class column and per-arm savings rows vs baseline. 5 pytest units + stub-CLI e2e of the full three-arm matrix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(eval): record workflow_bench ground base; fix churn measurement bias Ground base (3 classes x 3 arms, n=1/cell): every arm resolved every task — pass/fail quality saturates at this difficulty, making the comparison pure cost. Full plan→work never amortized its ~$9-11 fixed cost on tasks a baseline finishes in ≤35 turns (−211% to −333% cost); workflow_direct sits near baseline (−15% to −55%, once faster wall) with more test coverage. Routing implication recorded: direct mode/plain agent below this scale, full workflow for cross-module / multi-session / plan-as-deliverable work. The cross-module cell and multi-run variance are the next measurements. Churn fix: git add --intent-to-add -A before diffing (arms that never commit no longer undercount new files) and :(exclude)docs/plans (the committed plan doc no longer inflates workflow churn); this run's churn numbers predate the fix and are omitted from the recorded table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * perf(skills): cost-optimize the workflow from measured ground base Every optimization targets a measured fixed-cost component (eval/workflow_bench ground base: workflow arm −211% to −333% vs baseline, all tasks resolved): - Plan form is category-priced: compact form (core sections w/ § anchors preserved, ≤80 lines excl. pack, mini-pack subset of the context pack) for narrow/default categories; the full 13 sections only for deep work (refactor/security/performance/concurrency/architecture). A compact plan outgrowing its cap reclassifies to full rather than overflowing. - Freshness gate is category-priced: compact categories default to accept (source-weighted, refresh only when a graph claim becomes load-bearing); strict stays the default for full-plan categories — the rebuild+re-index was the largest single fixed cost. - Turn economy: per-category tool-call budgets (~10 to ~45; architecture uncapped); budget exhaustion routes open questions to §12 instead of more digging. - gitnexus-work fast path: HEAD == evidence pin → skip all citation re-reading (the pin's entire point); mini-pack fields tolerated. - lfg Lane 1 boundary triage: tasks below the measured ~35-turn boundary get offered gitnexus-work direct mode before the plan lane is spent. Copies re-synced (npm skills/, plugin, ~/.agents); steering + sync guards green. Re-measurement of the workflow arm follows to verify the numbers actually improve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(eval): record optimization re-measurement — inv-bug workflow cell −20% cost Same task, same conditions, post-830a0459 skills: $14.56→$11.70 (−20%), 83→72 turns, cache_read −24%; verified in-transcript that the compact form, turn budget, and skipped rebuild/re-index all fired. Wall +15% from a work- session test-debugging tail (n=1 variance). Regime unchanged (~3.5x baseline on this class) — routing rule stands. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(eval): per-arm clone isolation — worktree ref-namespace leak contaminated an arm The cross-module workflow_direct cell reported an impossible 28-turn solve with churn byte-identical to the workflow arm: git worktree add shares the repo's ref namespace, so the workflow arm's slug branch (created by gitnexus-work Phase 2) survived worktree removal and the direct arm found and adopted the completed work. Arms now get isolated git clone --shared copies (object store via alternates, refs clone-local — agent branches and stashes die with the clone; origin/<ref> fallback for non-default refs). Leaked branch deleted; baseline arm verified clean (0 branch references in its transcript); cell marked invalidated pending re-run. Records the valid cross-module cells: workflow $18.32 vs baseline $18.03 (premium −1.6%, vs −211%..−333% on smaller classes) — fixed costs amortize at this scale, with a less destructive diff and a plan artifact as bonus; resolve rate still tied. Churn fingerprinting is what caught the contamination — noted in the README as an integrity check. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(eval): complete cross-module cell — direct mode wins 47% cost / 56% wall Clean clone-isolated re-run: workflow_direct resolved the hardest class at $9.53/52 turns/15m vs $18.03/98/34m baseline and $18.32/107/37m full workflow. The measured story across all four classes: the execution discipline (gitnexus-work) is the consistent sweet spot and delivers real token savings on hard tasks; the planning pass buys its artifact, not same-session savings. Resolve rate tied everywhere (n=1/cell caveat). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(eval): add trajectory-gated skill evolution (#2431) - Pair prompt candidates with incumbent workflow arms - Gate promotions on pinned-model quality and efficiency - Expire router evidence and document its lifecycle * fix(eval): allow pr-review skill candidates * feat(skills): rename and generalize GitNexus review * feat(eval): external-comparator and review arms for workflow_bench - ce_workflow / ce_workflow_direct: compound-engineering ce-plan/ce-work arms prompted with the same structure as the gitnexus arms - review / ce_review: gitnexus-review vs ce-code-review on an identical diff applied by the task's setup - plan handoff is snapshot-based: committed example plans in docs/plans/ tie on clone mtimes and broke the name-glob pick (executed a stale plan) - verify output tail is recorded per run and the final working-tree patch is kept, so failed rows are diagnosable after the clone is destroyed Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills,eval): address #2431 review — data-safe rename migration, fail-closed bench evidence - setup: never delete a legacy renamed skill dir — the installer cannot prove ownership (users customize or hand-write skills under these names); warn with the path instead, and the test now asserts survival - workflow_bench: fail closed when a session's --output-format json report is empty, malformed, or missing usage fields — an exit-0 shell with no parseable usage no longer counts as measured evidence (5 parametrized regression tests) - workflow_bench: document the trust model prominently (task setup/verify are shell-executed, sessions run bypassPermissions with the parent env, candidate overlays are prompt injection surface) in README + docstring - free-model.litellm.yaml: master_key from LITELLM_MASTER_KEY env instead of a static token; loopback-binding warning - ci: run the eval workflow_bench pytest suite on ubuntu (pytest+pyyaml only — no full eval stack) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(eval): demand observed foreground verification in headless work-arm prompts In a headless -p session there is no later turn: a work arm backgrounded its slow test run, scheduled wakeups that can never fire, and reported done while two of its tests failed. All four work-arm prompts (both skill families, symmetric) now require verification output to be observed inside the session. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): ask plan depth up front instead of offering deepen afterwards gitnexus-plan Phase 0 now asks one blocking question in interactive sessions — quick / standard / deep, mapped onto the existing depth/form/ freshness knobs — when the invocation carries no explicit depth signal. Explicit knobs and headless runs skip the question (category posture unchanged, so benchmarks and automation behave as before). gitnexus-lfg's plan gate slims to proceed/stop: depth was already the user's up-front choice, so deepening is no longer offered by default — an explicit deepen request at the gate and executor route-backs still run Deepen mode, which remains the mechanism for strengthening an existing plan document. All shipped copies resynced (npm skills/, Claude plugin); AGENTS.md 1.13.0 and CLAUDE.md 1.7.0 pointers updated, including the analyzer's regenerated index-stats block at this branch's head. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): taint pass, expert lenses, and post-work index refresh gitnexus-review gains a PDG-backed taint-and-dependence pass (explain + pdg_query, --pdg folded into the stale refresh on trust-boundary diffs) and an Expert lenses section: domain reviewers derived from the graph's clusters plus four cross-cutting lenses (architectural fit, language conformance per the repo's own contract, Definition of Done, simplicity), dispatched once after the evidence-gathering steps and scaled to the diff. gitnexus-work Phase 4 now refreshes the knowledge graph after the DoD walk via the resolved-runner ladder with analyze --index-only, so the lfg review lane and later sessions query the finished work without dirtying the tree. lfg's threshold-governance paragraph moves to its README; eval citations are tagged as measured in the GitNexus repo. All shipped copies re-synced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(cli): remove legacy gitnexus-pr-review on uninstall; cover the rename migration uninstall's removal set now includes LEGACY_SKILL_DIR_NAMES derived from RENAMED_SKILL_DIRS, so a pre-rename install is cleaned up instead of orphaned. The rename warning gains behavioral coverage (fires with a legacy dir present, silent without), and shipped-skills-sync asserts legacy names stay absent from every shipped tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(eval): metric provenance, error-kind rows, skill-invocation verification, gate noise floor The promotion gate defaults to cost_usd (the only metric that includes subagent spend); token metrics carry an explicit main-loop-only warning in the report and promotion.json. Rows are classified by error_kind (session-error / verify-failed / infra-error), excluded from efficiency medians, and the gate requires equal valid-run counts. Each session's transcript is scanned for the expected Skill invocation and fails closed on a verified miss; a one-run resolution edge no longer promotes (noise floor). Per-run timeouts and setup failures record an infra-error row instead of aborting the sweep. Overlays touching skills no candidate arm exercises are rejected up front. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: fix skill routing paths, version headers, and skill rosters Routing tables point at the tracked direct skill paths (matching the post-#2434 generator output), AGENTS.md/CLAUDE.md headers match their latest changelog rows, the 1.12.0 row describes what the migration actually does, package/cursor READMEs list the full shipped skill roster, and the swarm READMEs describe /gitnexus-review's expert lenses instead of calling it single-agent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: drift-guard workflow for skill copies; pin eval pip deps; track docs/plans ci.yml ignores '**.md', so an md-only skill edit would merge without the shipped-skills-sync test running — skill-sync.yml triggers exactly on the guarded trees. The eval job's pip install is version-pinned, and docs/plans/ is unignored so gitnexus-plan output can be committed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep the runner-invocation literal in gitnexus-review; add concurrency block to skill-sync skills-steering requires skills with a stale-index hint to carry the exact 'node .gitnexus/run.cjs analyze' form — restore it with the fallback ladder as a parenthetical instead of replacing it. skill-sync.yml gains the top-level concurrency block the workflow-convention check enforces. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(skills): token-economy guidance for expert lenses Merge lenses that ground in the same material into one reviewer, and use cheaper model/effort tiers for mechanical lenses where the harness offers them, reserving the strongest engine for adversarial judgment. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(eval): isolate transcript home on Windows Ensure workflow_bench transcript tests set USERPROFILE alongside HOME so Path.home() resolves to the temporary test home on Windows. * docs(skills): fold PR #2522 execution learnings into review/work/plan Eight incident-backed hardenings from running the full skill cycle (review -> plan -> work, 28-finding fix series) on PR #2522: gitnexus-review: - Expert lenses execute the code under review on candidate failing shapes (empirical probe outranks source reading — every HIGH the language lenses found came from a probe, not a read). - Step 7 re-runs the exact CI check for refreshed baselines/fingerprints (a stale committed artifact is invisible in the diff; caught a red benchmarks arm). - Step 8 treats version/invalidation constants as review surface (INCREMENTAL_SCHEMA_VERSION class recurred verbatim from #2494). gitnexus-work: - Step 4 proves regression tests discriminate against the pre-fix tree. - Step 5 rebuilds executed build output before every verification run (parse workers load dist/; a correct fix 'failed' until rebuilt). - Step 6 makes stage -> detect_changes -> commit one unbroken sequence. gitnexus-plan: - Phase 0 seeded-evidence mode: plan FROM a completed review's verified findings instead of re-running the graph ladder. - Template §7: fingerprint/golden-guarded output rebaselines once, at the series tip. All distribution copies resynced; shipped-skills-sync + skills-steering 24/24 locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(eval): close the skill-evolution loop with an automated proposer driver workflow_bench.evolve adds the three arrows the README described as manual: a proposer session that turns loser trajectories (results.jsonl rows, transcripts, patches, the learning queue) into ONE bounded candidate overlay, a driver that iterates propose -> paired benchmark -> deterministic gate up to --generations, and an --apply step that copies a promoted overlay onto the canonical skills and shipped mirrors as a working-tree diff. The trust boundary is unchanged: overlays re-validate through candidate_overlay_files before any benchmark or apply consumes them, and committing, CI, and the PR merge stay human. learnings.jsonl is gitignored: it is machine-local evidence, like the session transcripts it complements. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(skills): route live-task friction into the evolution learning queue Each family skill gains a short 'Skill feedback' section: on friction with the skill's own instructions, append one JSON line to eval/workflow_bench/learnings.jsonl (GitNexus repo only) — never self-edit the skill from a live task. The proposer in workflow_bench.evolve consumes the queue as hints; a learning reaches a shipped skill only by beating the incumbent on the paired benchmark. All shipped mirrors re-copied byte- identical. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(tests): run the evolve helper tests in the eval pytest job test_evolve.py needs only pytest+pyyaml, same as the harness tests the job already runs — without this line the new module had no CI coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): comment-triggered GitNexus review agent for PRs '@gitnexus review' from a maintainer (OWNER/MEMBER/COLLABORATOR; the action re-validates write access) runs the repo's gitnexus-review skill headlessly against the PR and posts the review as a sticky comment — remote triggering with no local setup. Read-only by construction: contents: read token, Write/Edit and web tools disallowed, Bash allowlisted to git reads and the gitnexus CLI; analyze parses PR code with tree-sitter, never executes it. Requires the ANTHROPIC_API_KEY repository secret; activates once the file is on the default branch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): dispatch lane + existing OAuth secret for the review agent Align with claude.yml: same action pin and the CLAUDE_CODE_OAUTH_TOKEN secret the repo already carries — no new secret to configure. Add a workflow_dispatch lane (PR number input) so the agent can be triggered from the Actions UI and tested before the issue_comment trigger reaches the default branch. Allowlist gh pr view/diff and gh api, which the review skill uses to pin PR SHAs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): close a fork-PR RCE vector in the review agent's tool allowlist A live headless run of the exact workflow session against PR #2431 (66 turns, full gitnexus-review pass) surfaced a real HIGH-severity confused deputy: .gitnexus/ is gitignored, not blocked — a fork PR can commit its own .gitnexus/run.cjs, issue_comment checks out PR-head content, and the skill's runner ladder tries 'node .gitnexus/run.cjs analyze' first. That would execute fork-controlled JS inside a job holding CLAUDE_CODE_OAUTH_TOKEN and a write-scoped GITHUB_TOKEN — the opposite of the 'PR code is read, never executed' claim in the workflow's own header. Fix: drop the run.cjs allowlist entry so analyze always resolves through npx gitnexus (npm registry, not the checked-out tree); the skill's documented fallback mode covers the resulting graceful degradation. Also drop 'gh api' (not read-only — accepts -X POST/PATCH/DELETE) and downgrade pull-requests: write to read (comment posting only needs issues: write; the prompt already forbids formal review submission). Same session flagged a latent evolve.py bug: select_evidence's cost sort used dict.get's missing-key default, which doesn't cover an explicit JSON null in a foreign --seed-results row and crashes proposer setup with TypeError. Guarded with 'or 0.0' and added a regression test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden PR review and evolution trust boundaries * ci: follow workflow concurrency convention * fix(eval): make terminating error paths explicit * fix: unblock hardened review runtime checks * test: make containment canaries deterministic * test: expose Claude canary tool failures * fix: adapt clean shell environment for Claude * fix(eval): accept the runner's transcript source key in evidence preflight The proposer evidence preflight required transcript-artifact metadata to be exactly {path, sha256, bytes}, but the runner stamps a fourth provenance key (source=parent-captured-stream-json). Any --seed-results or generation>=2 run therefore aborted with SandboxError before proposing or promoting. Pin the producer literal as PARENT_EVENT_STREAM_SOURCE and validate it in the metadata check, and round-trip real producer output through sum_sessions into the preflight so the schema can't drift again. * fix(eval): treat an unmeasured session cost as unavailable, not $0 well_formed validated only the nested usage block, so an otherwise-successful session missing total_cost_usd was recorded as cost_usd=0.0 — and cost_usd is the default promotion metric (lower wins), so a cost-less session scored as free and could win promotion it never earned. Extract cost via measured_cost() (None on absent/garbage, a measured 0.0 preserved), propagate None through sum_sessions/aggregate/savings/report, and have the gate refuse to rank on a metric that was not measured on every run in both arms. * fix(eval): warn when ranking on the main-loop-only num_turns metric num_turns comes from the CLI's top-level usage (main-loop session only), like output_tokens, but selecting it emitted no metric_warning — so a subagent-heavy candidate could look artificially efficient. Add num_turns to MAIN_LOOP_ONLY_METRICS and broaden the warning to cover turns. * fix(eval): fail closed when an overlay adds a file with no committed base An overlay adding a new .md under gitnexus-{plan,work} passes the structural overlay checks but has no committed base for committed_destination_base_digests to bind against, so it raised an uncaught ValueError that crashed the evolve driver (and runner --candidate-overlay) mid-run. Catch it at both call sites: evolve reports NOT PROMOTED and exits, runner routes it through parser.error. * feat(eval): circuit-break the runner sweep on a systemic outage A sustained upstream outage used to pay out every remaining --timeout window one session at a time. Track consecutive session/infra/cleanup failures via a pure systemic_outage_streak helper; after --outage-streak (default 5) in a row, stop the sweep, still write report.md/promotion.json from partial evidence, and exit non-zero so evolve.py halts instead of proposing from truncated evidence. A task's own resolved=False never trips the breaker. * fix(cli): report a dirty working tree as stale in gitnexus status status --json (and the human output) computed up-to-date from commit + runner identity + completeness only, so a repo with uncommitted source changes at a matching HEAD was reported up-to-date while analyze would still re-index it. A graph-backed agent gating on that JSON could skip re-analysis on a stale graph. Extract analyze's dirty-tree check into a shared isWorkingTreeDirty() in storage/git and fold it into the status freshness decision. * fix(ci): use single-slash deny globs in the review agent's disallowedTools github.workspace already expands to an absolute path, so Read(/${{ github.workspace }}/**) and Read(//proc/**),(//sys/**),(//dev/**) produced double-slash patterns that a normalizing matcher may not match — silently no-opping the deny layer. Not exploitable (the allowlist is the primary control and never grants those paths), but the globs should be well-formed. Update the pinned test strings. * ci: install gitnexus-shared with npm ci from the committed lockfile The gitnexus-shared build floated its deps via npm install in three workflows (skill-sync, ci-tests, and — most importantly — the release publish.yml) while every other install step uses npm ci. The lockfile is committed and in sync, so switch all three to npm ci for reproducible, locked installs. * test(cli): make the shipped-skills drift guard reject symlinks listFilesRecursive walked with readdirSync and snapshotDir read with readFileSync, both of which follow symlinks — so a mirror file symlinked to the canonical tree passed the byte-compare (and a symlinked mirror dir would be followed too). Reject a symlinked root via lstat and any symlinked entry via Dirent.isSymbolicLink, with negative tests (skipped on Windows). * test(eval): guard the candidate-skill vs mirror-root coverage invariant MIRROR_SKILL_ROOTS omits the Cursor tree, safe only because no candidate skill is cursor-shipped. Pin that invariant: every CANDIDATE_SKILLS entry must exist under canonical + every mirror root and must not ship to Cursor, so adding a cursor-shipped skill to the candidate set (the PR #2488 asymmetric-sync class) fails loudly instead of syncing three of four trees. * docs(ci): describe the review agent's staged post-merge rollout The DoD asked for a dry-run or triggered run before merge, but an issue_comment (or newly added workflow_dispatch) workflow only ever executes the default-branch copy, so it cannot be exercised from the PR that introduces it. Reword the DoD and the activation checklist to a staged rollout: merge registered-but-disabled, validate same-repo and fork execution post-merge, then enable the variable. * fix: pin plugin skill mcp.json to the release version via #2445 tooling The ten plugin skill mcp.json launched `npx -y gitnexus@latest mcp` on every skill connect — non-reproducible and a supply-chain surface, and (unlike the persisted setup config) never pinned. Extend sync-plugin-manifests.mjs with an mcp surface kind that stamps the gitnexus@<version> launch arg, pin all ten to 1.6.9 now, and keep them byte-identical so the drift guard stays green. The release lifecycle + publish.yml --check now re-stamp them like the four manifest surfaces; only READMEs stay on @latest as docs. * test(eval): prove the proposer's built-in file tools are confined The real-Claude canary only exercised Bash + MCP, so it proved process/MCP containment but not that the proposer's built-in file tools stay inside their mounts. Add a canary over the exact PROPOSER_ALLOWED_TOOLS surface and the same read-only /evidence mount as run_proposer (allowlist extracted to a shared constant so it can't drift): Read reaches /evidence, a Write into the read-only evidence mount is denied, and a Write lands in the output tree. * fix(eval): apply the candidate overlay after task setup for fair arms The candidate overlay was applied before the task's untrusted setup ran, so setup could observe candidate prose and the incumbent/candidate arms started from different pre-overlay state. Reorder within the sandbox: capture the base (pre-overlay) skill digest, run setup against the base skills, verify setup did not tamper them, then apply the overlay and capture the post-overlay digest the model must preserve. apply_candidate_overlay stages path-specific overlay files, so setup's uncommitted changes stay out of the baseline and churn is unchanged. Graph freshness for the review arm is handled by the status dirty-tree fix plus the review skill's stale-triggered re-index, not by reordering the cached per-task-sha graph materialization (which is mechanically blocked). * test(eval): end-to-end containment proof of the autonomous proposer Drives the real run_proposer through bubblewrap with a deterministic scripted model (no paid API): it reads the read-only evidence bundle and writes a candidate gitnexus-plan skill edit plus a rationale into the sandbox output tree; run_proposer enforces the trust boundary and copies only the validated overlay + proposal out. This exercises the autonomous-proposal stage of the self-evolution loop end-to-end in the eval/containment CI job (the gate and apply stages are covered by test_workflow_bench_evolution and test_promotion_apply). Env-gated on GITNEXUS_REQUIRE_CLAUDE_CANARY, so it runs only where the pinned Claude binary and user namespaces are available. * fix(eval): let the proposer author its overlay via Bash Running the end-to-end proposer canary in the containment CI job surfaced a real bug: run_proposer starts the session with --bare, which hard-disables the Write/Edit tools ("Write exists but is not enabled in this context"), yet allowlisted Edit/Write and omitted Bash. The proposer therefore had no working way to write its candidate overlay — the self-evolution loop could never produce a candidate. The sandbox settings already pre-authorize Bash (autoAllowBashIfSandboxed) and confine writes to workspace/tmp/home, so switch PROPOSER_ALLOWED_TOOLS to Read/Grep/Glob/Bash and tell the proposer to author files with Bash. The end-to-end test now drives the real run_proposer through bubblewrap and asserts a validated overlay + proposal are produced (this also replaces the earlier file-tool canary, whose Write/Edit premise was moot). * test(eval): author the proposer overlay with newline-free Bash content The nested shell-sandbox prefix mangles embedded newlines, so the multi-line overlay content never landed. Use single-line content for the deterministic proposer canary. * test(eval): drop the unverifiable end-to-end proposer canary The scripted proposer overlay never materialized in the containment job across runs, and the model tool-result content is not visible in CI logs, so the test cannot be finalized without an environment where the sandbox can actually run. Keep the verified production fix (Bash-authoring in run_proposer); the proposer sandbox/containment stays covered by the existing Bash+MCP and process-tree canaries. * test(cli): drop run-analyze.ts from the windowsHide spawn-family list U7 moved run-analyze.ts's only child_process call (the git status --porcelain dirty check) into storage/git.ts (already covered by this test, with windowsHide). run-analyze.ts no longer imports a spawn-family function, so the windowsHide-regression test's 'must have >=1 spawn call' invariant failed for it. Remove it from SRC_FILES. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Zander Raycraft <zanderjraycraft@gmail.com> Co-authored-by: Azizur Rahman <azizur100389@gmail.com>
1331 lines
60 KiB
Python
1331 lines
60 KiB
Python
"""Benchmark the gitnexus-plan/work workflow against a baseline agent.
|
||
|
||
Usage:
|
||
uv run --locked --extra dev python -m workflow_bench.runner \
|
||
--tasks workflow_bench/tasks.scenarios.yaml --runs 3 \
|
||
--model claude-sonnet-4-20250514
|
||
|
||
Each task runs in a fresh detached git worktree of the target repo, once per
|
||
arm per run:
|
||
|
||
* ``workflow`` — two headless Claude Code sessions: gitnexus-plan, then
|
||
gitnexus-work on the produced plan.
|
||
* ``candidate_workflow`` / ``candidate_workflow_direct`` — the matching
|
||
workflow arm with a prompt-only candidate overlay committed in its clone.
|
||
* ``baseline`` — one headless session with the same task text and the Skill
|
||
tool disallowed (so it cannot borrow the workflow), everything else equal.
|
||
|
||
Token usage, cost, duration, and turn counts come from the CLI's own
|
||
``--output-format json`` report — nothing is estimated. Caveat: the report's
|
||
top-level ``usage`` counts ONLY the main-loop session; ``total_cost_usd`` is
|
||
the only reported number that includes subagent spend. A task's model-visible
|
||
``verify`` command is retained as an authored-test quality signal; ``resolved``
|
||
also requires its harness-owned hidden behavioral oracle. Token savings on
|
||
unresolved runs are reported but flagged, because saving tokens by failing is
|
||
not a saving.
|
||
|
||
Trust model: task files and candidate prompts are executable input. Every
|
||
setup, verifier, and model session runs inside a preflighted Linux Bubblewrap
|
||
boundary with an allowlisted environment, isolated home, PID namespace,
|
||
self-contained clone, and task-declared read-only dependencies. Unsupported
|
||
or unavailable containment fails before model invocation (README § Trust
|
||
model).
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import argparse
|
||
import hashlib
|
||
import json
|
||
import os
|
||
import re
|
||
import secrets
|
||
import stat
|
||
import statistics
|
||
import tempfile
|
||
import time
|
||
from collections.abc import Mapping
|
||
from dataclasses import replace
|
||
from datetime import UTC, datetime, timedelta
|
||
from pathlib import Path
|
||
from typing import Any
|
||
|
||
import yaml
|
||
|
||
from .evolution import (
|
||
CANDIDATE_ARMS,
|
||
EVIDENCE_MAX_AGE_DAYS,
|
||
EVALUATED_ARM_SKILLS,
|
||
MAIN_LOOP_ONLY_METRICS,
|
||
MAIN_LOOP_ONLY_WARNING,
|
||
PROMOTION_METRICS,
|
||
apply_candidate_overlay,
|
||
candidate_overlay_digest,
|
||
evaluate_candidate,
|
||
required_candidate_arms,
|
||
skill_fingerprint,
|
||
)
|
||
from .oracle_assets import (
|
||
ORACLE_ENV_VAR,
|
||
TaskOracleSnapshot,
|
||
capture_task_oracles,
|
||
sanitize_clone_for_hidden_oracles,
|
||
staged_task_oracle,
|
||
)
|
||
from .process_control import ManagedProcessError
|
||
from .promotion_apply import committed_destination_base_digests
|
||
from .proposer_sandbox import (
|
||
SANDBOX_GITNEXUS as SANDBOX_GITNEXUS,
|
||
SANDBOX_GITNEXUS_REGISTRY,
|
||
SANDBOX_GITNEXUS_SHARED as SANDBOX_GITNEXUS_SHARED,
|
||
SANDBOX_WORKSPACE,
|
||
ReadOnlyMount,
|
||
SandboxError,
|
||
SandboxSession,
|
||
build_sandbox_environment,
|
||
preflight_bubblewrap,
|
||
prepare_sandbox,
|
||
require_claude_sandbox_helpers,
|
||
)
|
||
from .runner_artifacts import (
|
||
IMPLEMENTATION_ARMS,
|
||
MAX_PATCH_BYTES as MAX_PATCH_BYTES,
|
||
MAX_WORKSPACE_SNAPSHOT_ENTRIES as MAX_WORKSPACE_SNAPSHOT_ENTRIES,
|
||
MAX_WORKSPACE_SNAPSHOT_FILE_BYTES as MAX_WORKSPACE_SNAPSHOT_FILE_BYTES,
|
||
MAX_WORKSPACE_SNAPSHOT_PATH_BYTES as MAX_WORKSPACE_SNAPSHOT_PATH_BYTES,
|
||
_bounded_regular_bytes,
|
||
_prepare_untracked_for_diff,
|
||
_sandbox_git,
|
||
capture_patch,
|
||
diff_churn,
|
||
enforce_phase_workspace,
|
||
enforce_work_evidence,
|
||
implementation_diff_digest,
|
||
make_worktree,
|
||
new_plan_doc,
|
||
parse_shortstat as parse_shortstat,
|
||
remove_clone,
|
||
require_skill_fingerprint,
|
||
run_verify,
|
||
snapshot_plan_docs,
|
||
VerificationResult,
|
||
workspace_snapshot,
|
||
)
|
||
from .runner_sessions import (
|
||
BUILTIN_AGENT_TOOLS as BUILTIN_AGENT_TOOLS,
|
||
GITNEXUS_MUTATING_TOOLS as GITNEXUS_MUTATING_TOOLS,
|
||
GITNEXUS_READ_ONLY_TOOLS as GITNEXUS_READ_ONLY_TOOLS,
|
||
MAX_TRANSCRIPT_BYTES as MAX_TRANSCRIPT_BYTES,
|
||
SANDBOX_GITNEXUS_ENTRYPOINT as SANDBOX_GITNEXUS_ENTRYPOINT,
|
||
USAGE_FIELDS,
|
||
allowed_agent_tools,
|
||
run_claude,
|
||
sandbox_mcp_config,
|
||
sum_sessions,
|
||
)
|
||
from .runner_tasks import (
|
||
normalized_model_identifier,
|
||
resolve_task_bindings,
|
||
select_tasks,
|
||
selected_task_bindings as selected_task_bindings,
|
||
)
|
||
from .sanitized_graph import (
|
||
SanitizedGraphSnapshot,
|
||
prepare_sanitized_graph,
|
||
validate_no_prebuilt_graph_assets,
|
||
)
|
||
from .runtime_mounts import (
|
||
CE_ARMS,
|
||
HARNESS_ROOT as HARNESS_ROOT,
|
||
PINNED_GITNEXUS_VERSION as PINNED_GITNEXUS_VERSION,
|
||
ce_plugin_dir_for_arm,
|
||
ce_plugin_mounts_for_arm,
|
||
staged_ce_plugin_snapshot,
|
||
trusted_gitnexus_runtime_mounts,
|
||
validate_ce_plugin_inputs,
|
||
)
|
||
from .task_assets import TaskAssetCache, TaskAssetSnapshot, stage_task_assets
|
||
|
||
PLAN_PROMPT = (
|
||
"Use the gitnexus-plan skill for: {task}\n"
|
||
"Headless run: make reasonable choices without asking; the plan document "
|
||
"is the deliverable."
|
||
)
|
||
# Appended to every work-arm prompt. In a headless `claude -p` session there
|
||
# is no later turn: backgrounded test runs and scheduled wakeups never come
|
||
# back, so a session that "waits" for verification ends unverified (observed:
|
||
# a work arm backgrounded its slow tests, scheduled three wakeups that never
|
||
# fired, and reported done while two tests failed).
|
||
HEADLESS_VERIFY = (
|
||
" Verification must be observed inside this session: run the typecheck "
|
||
"and test commands in the foreground to completion and report their "
|
||
"actual output — never background them or wait on scheduled wakeups."
|
||
)
|
||
WORK_PROMPT = (
|
||
"Use the gitnexus-work skill to execute the plan at {plan}.\n"
|
||
"Headless run: proceed without asking; report Definition of Done status "
|
||
"at the end." + HEADLESS_VERIFY
|
||
)
|
||
WORK_DIRECT_PROMPT = (
|
||
"Use the gitnexus-work skill for: {task}\n"
|
||
"Headless run: proceed without asking. The user explicitly declines a "
|
||
"separate planning pass — execute in direct mode with the skill's "
|
||
"execution discipline." + HEADLESS_VERIFY
|
||
)
|
||
BASELINE_PROMPT = (
|
||
"{task}\n\n"
|
||
"Implement the change in this repository and verify it by running the "
|
||
"relevant tests. Work autonomously without asking questions."
|
||
)
|
||
# External-comparator arms: the compound-engineering plugin's plan/work family,
|
||
# prompted with the same structure as the gitnexus arms so only the skill
|
||
# family differs. The plugin ships user-level, so clones need no repo files.
|
||
CE_PLAN_PROMPT = (
|
||
"Use the ce-plan skill (compound-engineering plugin) for: {task}\n"
|
||
"Headless run: make reasonable choices without asking; the plan document "
|
||
"is the deliverable."
|
||
)
|
||
CE_WORK_PROMPT = (
|
||
"Use the ce-work skill (compound-engineering plugin) to execute the plan "
|
||
"at {plan}.\n"
|
||
"Headless run: proceed without asking; report completion status at the "
|
||
"end." + HEADLESS_VERIFY
|
||
)
|
||
CE_WORK_DIRECT_PROMPT = (
|
||
"Use the ce-work skill (compound-engineering plugin) for: {task}\n"
|
||
"Headless run: proceed without asking. The user explicitly declines a "
|
||
"separate planning pass — execute directly with the skill's execution "
|
||
"discipline." + HEADLESS_VERIFY
|
||
)
|
||
# Review cell: the task's `setup` applies the diff under review as local
|
||
# changes; both arms review the same working tree and write to the same file
|
||
# so `verify` can gate on a produced review.
|
||
REVIEW_PROMPT = (
|
||
"Use the gitnexus-review skill to review the local uncommitted changes "
|
||
"in this repository. {task}\n"
|
||
"Headless run: proceed without asking; do not post to GitHub or anywhere "
|
||
"external; write the complete review to review-output.md in the "
|
||
"repository root."
|
||
)
|
||
CE_REVIEW_PROMPT = (
|
||
"Use the ce-code-review skill (compound-engineering plugin) to review "
|
||
"the local uncommitted changes in this repository. {task}\n"
|
||
"Headless run: proceed without asking; do not post to GitHub or anywhere "
|
||
"external; write the complete review to review-output.md in the "
|
||
"repository root."
|
||
)
|
||
|
||
|
||
# Skill each arm's session(s) must actually invoke; a session that never ran
|
||
# its skill is a silent no-op arm, not a data point (checked via transcript).
|
||
ARM_EXPECTED_SKILLS: dict[str, tuple[str, ...]] = {
|
||
"workflow": ("gitnexus-plan", "gitnexus-work"),
|
||
"ce_workflow": ("ce-plan", "ce-work"),
|
||
"workflow_direct": ("gitnexus-work",),
|
||
"ce_workflow_direct": ("ce-work",),
|
||
"review": ("gitnexus-review",),
|
||
"ce_review": ("ce-code-review",),
|
||
}
|
||
|
||
|
||
def _require_implementation_fingerprint(
|
||
session: dict[str, Any],
|
||
worktree: Path,
|
||
arm: str,
|
||
expected: str | None,
|
||
) -> None:
|
||
"""Bind a just-finished implementation session to its original skill bytes."""
|
||
|
||
try:
|
||
require_skill_fingerprint(
|
||
worktree,
|
||
arm,
|
||
expected,
|
||
phase="implementation",
|
||
)
|
||
except ValueError as exc:
|
||
if session.get("error_kind") is None:
|
||
session["ok"] = False
|
||
session["error_kind"] = "implementation-evidence-invalid"
|
||
session["error_detail"] = str(exc)
|
||
else:
|
||
session.setdefault("evidence_diagnostics", []).append(str(exc))
|
||
|
||
|
||
def _verification_outcome(result: VerificationResult | tuple[bool, str]) -> tuple[bool, str]:
|
||
if isinstance(result, VerificationResult):
|
||
if result.process.state != "exited":
|
||
# Hidden-oracle output can contain mounted test bytes. Preserve
|
||
# terminal-state evidence without letting candidate-controlled
|
||
# stdout/stderr enter results.jsonl through the exception string.
|
||
safe_process = replace(
|
||
result.process,
|
||
stdout_tail="",
|
||
stderr_tail="",
|
||
detail=result.process.detail or "verifier infrastructure failed",
|
||
)
|
||
raise ManagedProcessError(result.command, safe_process)
|
||
return result.passed, result.output
|
||
return result
|
||
|
||
|
||
def _run_hidden_oracle(
|
||
snapshot: TaskOracleSnapshot,
|
||
worktree: Path,
|
||
args: argparse.Namespace,
|
||
sandbox: SandboxSession,
|
||
) -> tuple[bool, str]:
|
||
"""Stage a captured oracle after the model exits, execute it, then erase it."""
|
||
|
||
if worktree.expanduser().absolute() != sandbox.clone.expanduser().absolute():
|
||
raise SandboxError("hidden oracle sandbox does not bind the credited worktree")
|
||
mount_name = f".wfbench-oracle-{secrets.token_hex(16)}"
|
||
mount_point = worktree / mount_name
|
||
mount_point.mkdir(mode=0o700)
|
||
primary: BaseException | None = None
|
||
try:
|
||
with staged_task_oracle(sandbox.private_root, snapshot) as stage_root:
|
||
oracle_env = build_sandbox_environment()
|
||
# A private RO bind at a random workspace sibling preserves each
|
||
# oracle's ../gitnexus import as the candidate implementation. The
|
||
# empty mountpoint exists only post-model and is removed before the
|
||
# credited patch is captured.
|
||
oracle_mount = f"{SANDBOX_WORKSPACE}/{mount_name}"
|
||
oracle_env[ORACLE_ENV_VAR] = oracle_mount
|
||
passed, _output = _verification_outcome(
|
||
run_verify(
|
||
snapshot.command,
|
||
sandbox.clone,
|
||
args.timeout,
|
||
command_prefix=sandbox.command_prefix_for(
|
||
read_only_workspace=True,
|
||
unshare_network=True,
|
||
extra_read_only_mounts=(ReadOnlyMount(source=stage_root, target=oracle_mount),),
|
||
),
|
||
env=oracle_env,
|
||
require_pid_namespace=True,
|
||
)
|
||
)
|
||
# Candidate code executes in this process. Never persist its stdout
|
||
# or stderr: it can read the mounted hidden test bytes and print them.
|
||
return passed, "hidden oracle passed" if passed else "hidden oracle failed"
|
||
except BaseException as exc:
|
||
primary = exc
|
||
raise
|
||
finally:
|
||
try:
|
||
metadata = mount_point.lstat()
|
||
if stat.S_ISLNK(metadata.st_mode) or not stat.S_ISDIR(metadata.st_mode):
|
||
raise SandboxError("hidden oracle mountpoint changed type during verification")
|
||
mount_point.rmdir()
|
||
except (OSError, SandboxError) as cleanup:
|
||
if primary is None:
|
||
raise
|
||
primary.add_note(f"hidden oracle mountpoint cleanup also failed: {cleanup}")
|
||
|
||
|
||
def _evaluated_skill_roots(worktree: Path, arm: str) -> tuple[Path, ...]:
|
||
"""Repo-local prompt roots that must remain immutable during a session."""
|
||
|
||
return tuple(worktree / ".claude" / "skills" / name for name in EVALUATED_ARM_SKILLS.get(arm, ()))
|
||
|
||
|
||
def isolated_gitnexus_registry_mount(worktree: Path, parent: Path) -> ReadOnlyMount:
|
||
"""Create a one-clone registry that cannot route MCP to any host repo."""
|
||
|
||
metadata_path = worktree / ".gitnexus" / "gitnexus.json"
|
||
if not metadata_path.exists():
|
||
metadata_path = worktree / ".gitnexus" / "meta.json"
|
||
mode = metadata_path.lstat().st_mode
|
||
if stat.S_ISLNK(mode) or not stat.S_ISREG(mode):
|
||
raise SandboxError(f"benchmark index metadata must be regular and non-symlink: {metadata_path}")
|
||
raw = _bounded_regular_bytes(metadata_path, limit=2 * 1024 * 1024)
|
||
try:
|
||
metadata = json.loads(raw)
|
||
except json.JSONDecodeError as exc:
|
||
raise SandboxError(f"benchmark index metadata is malformed: {metadata_path}") from exc
|
||
if not isinstance(metadata, dict):
|
||
raise SandboxError(f"benchmark index metadata must be an object: {metadata_path}")
|
||
indexed_at = metadata.get("indexedAt")
|
||
last_commit = metadata.get("lastCommit")
|
||
if not isinstance(indexed_at, str) or not indexed_at or not isinstance(last_commit, str) or not last_commit:
|
||
raise SandboxError("benchmark index metadata is missing indexedAt or lastCommit")
|
||
|
||
parent = parent.expanduser().absolute()
|
||
registry = Path(tempfile.mkdtemp(prefix="wfbench-registry-", dir=parent))
|
||
registry.chmod(0o700)
|
||
entry: dict[str, Any] = {
|
||
"name": "benchmark-target",
|
||
"path": SANDBOX_WORKSPACE,
|
||
"storagePath": f"{SANDBOX_WORKSPACE}/.gitnexus",
|
||
"indexedAt": indexed_at,
|
||
"lastCommit": last_commit,
|
||
}
|
||
for field in ("remoteUrl", "stats", "branch"):
|
||
if field in metadata:
|
||
entry[field] = metadata[field]
|
||
registry_file = registry / "registry.json"
|
||
descriptor = os.open(
|
||
registry_file,
|
||
os.O_WRONLY | os.O_CREAT | os.O_EXCL | getattr(os, "O_NOFOLLOW", 0),
|
||
0o600,
|
||
)
|
||
try:
|
||
os.fchmod(descriptor, 0o600)
|
||
payload = (json.dumps([entry], sort_keys=True, separators=(",", ":")) + "\n").encode()
|
||
view = memoryview(payload)
|
||
while view:
|
||
written = os.write(descriptor, view)
|
||
view = view[written:]
|
||
finally:
|
||
os.close(descriptor)
|
||
return ReadOnlyMount(source=registry, target=SANDBOX_GITNEXUS_REGISTRY)
|
||
|
||
|
||
def run_arm(
|
||
arm: str,
|
||
task: dict[str, Any],
|
||
worktree: Path,
|
||
args: argparse.Namespace,
|
||
*,
|
||
sandbox: SandboxSession,
|
||
transcript_output_dir: Path | None = None,
|
||
transcript_output_prefix: str | None = None,
|
||
expected_skill_digest: str | None = None,
|
||
enforce_phase_boundary: bool = False,
|
||
ce_plugin_dir: str | None = None,
|
||
oracle_snapshot: TaskOracleSnapshot | None = None,
|
||
) -> dict[str, Any]:
|
||
sessions: list[dict[str, Any]] = []
|
||
env = build_sandbox_environment(
|
||
auth_token=args.auth_token,
|
||
base_url=args.base_url,
|
||
)
|
||
common = {
|
||
"claude_bin": sandbox.claude_bin,
|
||
"timeout": args.timeout,
|
||
"model": args.model,
|
||
"env": env,
|
||
"permission_mode": "dontAsk",
|
||
"command_prefix": sandbox.command_prefix_for(
|
||
read_only_paths=_evaluated_skill_roots(worktree, arm),
|
||
),
|
||
"require_pid_namespace": True,
|
||
"bare": True,
|
||
"settings_json": sandbox.settings_json,
|
||
"strict_mcp_config": True,
|
||
"mcp_config_json": sandbox_mcp_config(),
|
||
"transcript_projects": sandbox.transcript_projects,
|
||
"transcript_cwd": Path(SANDBOX_WORKSPACE),
|
||
"transcript_wait_seconds": 5,
|
||
"transcript_output_dir": transcript_output_dir,
|
||
"transcript_output_prefix": transcript_output_prefix,
|
||
"transcript_secrets": tuple(secret for secret in (args.auth_token,) if secret),
|
||
}
|
||
if ce_plugin_dir is not None:
|
||
common["plugin_dirs"] = (ce_plugin_dir,)
|
||
expected_skills = ARM_EXPECTED_SKILLS.get(arm, ())
|
||
plan_doc: Path | None = None
|
||
if arm in ("workflow", "ce_workflow"):
|
||
plan_prompt = PLAN_PROMPT if arm == "workflow" else CE_PLAN_PROMPT
|
||
work_prompt = WORK_PROMPT if arm == "workflow" else CE_WORK_PROMPT
|
||
pre = snapshot_plan_docs(worktree)
|
||
phase_before = workspace_snapshot(worktree) if enforce_phase_boundary else None
|
||
plan_session = run_claude(
|
||
plan_prompt.format(task=task["prompt"]),
|
||
worktree,
|
||
expected_skill=expected_skills[0],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=False)},
|
||
)
|
||
sessions.append(plan_session)
|
||
if plan_session["ok"]:
|
||
try:
|
||
plan_doc = new_plan_doc(worktree, pre)
|
||
if phase_before is not None:
|
||
enforce_phase_workspace(
|
||
worktree,
|
||
phase_before,
|
||
allowed_artifact=plan_doc,
|
||
)
|
||
require_skill_fingerprint(
|
||
worktree,
|
||
arm,
|
||
expected_skill_digest,
|
||
phase="planning",
|
||
)
|
||
except ValueError as exc:
|
||
plan_session["ok"] = False
|
||
plan_session["error_kind"] = "plan-evidence-invalid"
|
||
plan_session["error_detail"] = str(exc)
|
||
else:
|
||
work_session = run_claude(
|
||
work_prompt.format(plan=plan_doc.relative_to(worktree)),
|
||
worktree,
|
||
expected_skill=expected_skills[1],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
|
||
)
|
||
_require_implementation_fingerprint(
|
||
work_session,
|
||
worktree,
|
||
arm,
|
||
expected_skill_digest,
|
||
)
|
||
sessions.append(work_session)
|
||
elif arm == "ce_workflow_direct":
|
||
work_session = run_claude(
|
||
CE_WORK_DIRECT_PROMPT.format(task=task["prompt"]),
|
||
worktree,
|
||
expected_skill=expected_skills[0],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
|
||
)
|
||
_require_implementation_fingerprint(
|
||
work_session,
|
||
worktree,
|
||
arm,
|
||
expected_skill_digest,
|
||
)
|
||
sessions.append(work_session)
|
||
elif arm in ("review", "ce_review"):
|
||
review_prompt = REVIEW_PROMPT if arm == "review" else CE_REVIEW_PROMPT
|
||
phase_before = workspace_snapshot(worktree) if enforce_phase_boundary else None
|
||
review_session = run_claude(
|
||
review_prompt.format(task=task["prompt"]),
|
||
worktree,
|
||
expected_skill=expected_skills[0],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=False)},
|
||
)
|
||
sessions.append(review_session)
|
||
if review_session["ok"] and phase_before is not None:
|
||
try:
|
||
enforce_phase_workspace(
|
||
worktree,
|
||
phase_before,
|
||
allowed_artifact=worktree / "review-output.md",
|
||
)
|
||
require_skill_fingerprint(
|
||
worktree,
|
||
arm,
|
||
expected_skill_digest,
|
||
phase="review",
|
||
)
|
||
except ValueError as exc:
|
||
review_session["ok"] = False
|
||
review_session["error_kind"] = "review-evidence-invalid"
|
||
review_session["error_detail"] = str(exc)
|
||
elif arm == "workflow_direct":
|
||
work_session = run_claude(
|
||
WORK_DIRECT_PROMPT.format(task=task["prompt"]),
|
||
worktree,
|
||
expected_skill=expected_skills[0],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
|
||
)
|
||
_require_implementation_fingerprint(
|
||
work_session,
|
||
worktree,
|
||
arm,
|
||
expected_skill_digest,
|
||
)
|
||
sessions.append(work_session)
|
||
elif arm == "baseline_nomcp":
|
||
# Isolates the workflow-discipline question from the GitNexus-tools
|
||
# question: no skills AND no graph tools.
|
||
sessions.append(
|
||
run_claude(
|
||
BASELINE_PROMPT.format(task=task["prompt"]),
|
||
worktree,
|
||
disallowed_tools=["Skill", "mcp__gitnexus"],
|
||
**{
|
||
**common,
|
||
"mcp_config_json": '{"mcpServers":{}}',
|
||
"allowed_tools": allowed_agent_tools(
|
||
implementation=True,
|
||
include_mcp=False,
|
||
),
|
||
},
|
||
)
|
||
)
|
||
else:
|
||
sessions.append(
|
||
run_claude(
|
||
BASELINE_PROMPT.format(task=task["prompt"]),
|
||
worktree,
|
||
disallowed_tools=["Skill"],
|
||
**{**common, "allowed_tools": allowed_agent_tools(implementation=True)},
|
||
)
|
||
)
|
||
record = sum_sessions(sessions)
|
||
record["arm"] = arm
|
||
record["plan_produced"] = arm not in ("workflow", "ce_workflow") or plan_doc is not None
|
||
authored_tests_passed, authored_test_output = _verification_outcome(
|
||
run_verify(
|
||
task["verify"],
|
||
worktree,
|
||
args.timeout,
|
||
command_prefix=sandbox.command_prefix_for(
|
||
read_only_workspace=True,
|
||
unshare_network=True,
|
||
),
|
||
env=build_sandbox_environment(),
|
||
require_pid_namespace=True,
|
||
)
|
||
)
|
||
if oracle_snapshot is None:
|
||
oracle_passed, oracle_output = False, "hidden oracle snapshot unavailable"
|
||
else:
|
||
oracle_passed, oracle_output = _run_hidden_oracle(
|
||
oracle_snapshot,
|
||
worktree,
|
||
args,
|
||
sandbox,
|
||
)
|
||
record["authored_tests_passed"] = authored_tests_passed
|
||
record["authored_test_output"] = authored_test_output
|
||
record["oracle_passed"] = oracle_passed
|
||
record["oracle_output"] = oracle_output
|
||
record["resolved"] = record["ok"] and authored_tests_passed and oracle_passed
|
||
# Compatibility alias for existing report consumers. The authored tests are
|
||
# now an explicit signal and can never self-certify resolution.
|
||
record["verify_output"] = authored_test_output
|
||
if oracle_snapshot is not None:
|
||
record.update(
|
||
{
|
||
"oracle_digest": oracle_snapshot.digest,
|
||
"oracle_command_digest": oracle_snapshot.command_digest,
|
||
"oracle_manifest_digest": oracle_snapshot.manifest_digest,
|
||
}
|
||
)
|
||
if record["error_kind"] is None and not authored_tests_passed:
|
||
# The sessions completed — the produced change just failed the task's
|
||
# verify command. Kept distinct from session-error so aggregates can
|
||
# exclude infrastructure deaths without hiding real failures.
|
||
record["error_kind"] = "verify-failed"
|
||
elif record["error_kind"] is None and not oracle_passed:
|
||
record["error_kind"] = "oracle-failed" if oracle_snapshot is not None else "oracle-unavailable"
|
||
return record
|
||
|
||
|
||
# ─── Pure aggregation/report helpers (unit-tested) ──────────────────────────
|
||
|
||
|
||
CHURN_FIELDS = ("diff_files", "diff_insertions", "diff_deletions")
|
||
|
||
# Rows where the session (or the harness) died carry no measured evidence and
|
||
# must not skew efficiency medians or resolve denominators. verify-failed and
|
||
# skill-not-invoked rows DO count: those sessions ran and spent real tokens.
|
||
EXCLUDED_ERROR_KINDS = frozenset({"session-error", "infra-error", "evidence-unverified", "cleanup-failure"})
|
||
|
||
# A sustained upstream outage shows up as a run of session/infra/cleanup
|
||
# failures. (cleanup-failure overwrites the primary error_kind, so a
|
||
# session-error whose worktree cleanup also failed still counts.) A task's own
|
||
# resolved=False is real signal, not an outage, so it never trips the breaker.
|
||
SYSTEMIC_ERROR_KINDS = frozenset({"session-error", "infra-error", "cleanup-failure"})
|
||
DEFAULT_OUTAGE_STREAK = 5
|
||
|
||
|
||
def systemic_outage_streak(error_kind: str | None, prior_streak: int) -> int:
|
||
"""Consecutive systemic-failure count: +1 on a systemic kind, else reset to 0."""
|
||
return prior_streak + 1 if error_kind in SYSTEMIC_ERROR_KINDS else 0
|
||
|
||
|
||
def infra_error_record(exc: BaseException) -> dict[str, Any]:
|
||
"""Row for a run the harness itself killed (timeout, setup failure)."""
|
||
if isinstance(exc, ManagedProcessError):
|
||
process = exc.result
|
||
detail = f"{process.state}: {process.detail or process.stderr_tail[-1500:]}"
|
||
else:
|
||
detail = f"{type(exc).__name__}: {exc}"
|
||
record: dict[str, Any] = dict.fromkeys(USAGE_FIELDS, 0)
|
||
record.update(
|
||
{
|
||
"ok": False,
|
||
"resolved": False,
|
||
"error_kind": "infra-error",
|
||
"error_detail": detail[:2000],
|
||
"session_ids": [],
|
||
"cost_usd": 0.0,
|
||
"duration_s": 0.0,
|
||
"num_turns": 0,
|
||
"plan_produced": False,
|
||
"authored_tests_passed": False,
|
||
"authored_test_output": "",
|
||
"oracle_passed": False,
|
||
"oracle_output": "",
|
||
"verify_output": "",
|
||
"skill_invoked": None,
|
||
"transcript_missing": False,
|
||
}
|
||
)
|
||
return record
|
||
|
||
|
||
def aggregate(records: list[dict[str, Any]]) -> dict[str, Any]:
|
||
"""Median metrics + resolve rate across repeated runs of one task+arm.
|
||
|
||
Session/infra-error rows are excluded from the medians (they measured
|
||
nothing); ``valid_runs``/``excluded_runs`` make the exclusion visible.
|
||
"""
|
||
valid = [r for r in records if r.get("error_kind") not in EXCLUDED_ERROR_KINDS]
|
||
metrics = (*USAGE_FIELDS, "duration_s", "num_turns", *CHURN_FIELDS)
|
||
out: dict[str, Any] = {m: statistics.median(r.get(m, 0) for r in (valid or [{}])) for m in metrics}
|
||
# cost_usd can be None (unmeasured) on an otherwise-valid run; a single
|
||
# unmeasured run makes the whole median unavailable so the gate won't rank
|
||
# a candidate on a cost that was never actually captured.
|
||
valid_costs = [r.get("cost_usd") for r in valid]
|
||
out["cost_usd"] = None if (not valid or any(cost is None for cost in valid_costs)) else statistics.median(valid_costs)
|
||
out["resolved"] = sum(1 for r in records if r["resolved"])
|
||
out["runs"] = len(records)
|
||
out["valid_runs"] = len(valid)
|
||
out["excluded_runs"] = len(records) - len(valid)
|
||
out["transcripts_missing"] = sum(1 for r in records if r.get("transcript_missing"))
|
||
out["class"] = records[0].get("class", "")
|
||
return out
|
||
|
||
|
||
def savings(baseline: dict[str, Any], workflow: dict[str, Any]) -> dict[str, Any]:
|
||
"""Percent saved by the workflow arm per metric (positive = cheaper)."""
|
||
out: dict[str, Any] = {}
|
||
for metric in (*USAGE_FIELDS, "cost_usd", "duration_s"):
|
||
base = baseline.get(metric)
|
||
arm = workflow.get(metric)
|
||
if base is None or arm is None:
|
||
out[metric] = None
|
||
else:
|
||
out[metric] = round(100 * (base - arm) / base, 1) if base else 0.0
|
||
return out
|
||
|
||
|
||
def _na(value: Any) -> Any:
|
||
"""Render an unmeasured metric as ``n/a`` instead of a misleading number."""
|
||
return "n/a" if value is None else value
|
||
|
||
|
||
def _cost_cell(value: Any) -> str:
|
||
return "n/a" if value is None else f"{value:.4f}"
|
||
|
||
|
||
def render_report(results: dict[str, dict[str, dict[str, Any]]]) -> str:
|
||
"""results: {task_id: {arm: aggregate}} → markdown report."""
|
||
lines = [
|
||
"# gitnexus workflow benchmark",
|
||
"",
|
||
"Medians across runs; savings rows = (baseline − arm) / baseline per arm.",
|
||
"A negative saving means that arm spent more than baseline. churn =",
|
||
"files/+insertions/−deletions vs the worktree's starting commit.",
|
||
"",
|
||
"**WARNING:** token columns count only each arm's main-loop session —",
|
||
"subagent spend is invisible to them and flatters subagent-heavy arms.",
|
||
"cost $ is the only column that includes subagent spend; to rank token",
|
||
"efficiency, sum usage from the session transcripts instead",
|
||
"(dedup events sharing one message.id).",
|
||
"",
|
||
"| task | class | arm | resolved | input | cache_create | cache_read | output | cost $ | wall s | turns | churn |",
|
||
"| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |",
|
||
]
|
||
for task_id, arms in results.items():
|
||
for arm, agg in arms.items():
|
||
excluded = agg.get("excluded_runs", 0)
|
||
resolved_cell = f"{agg['resolved']}/{agg.get('valid_runs', agg['runs'])}"
|
||
if excluded:
|
||
resolved_cell += f" ({excluded} excluded)"
|
||
lines.append(
|
||
f"| {task_id} | {agg['class']} | {arm} | {resolved_cell} "
|
||
f"| {agg['input_tokens']:.0f} | {agg['cache_creation_input_tokens']:.0f} "
|
||
f"| {agg['cache_read_input_tokens']:.0f} | {agg['output_tokens']:.0f} "
|
||
f"| {_cost_cell(agg['cost_usd'])} | {agg['duration_s']:.0f} | {agg['num_turns']:.0f} "
|
||
f"| {agg['diff_files']:.0f}/+{agg['diff_insertions']:.0f}/−{agg['diff_deletions']:.0f} |"
|
||
)
|
||
for arm in arms:
|
||
if arm != "baseline" and "baseline" in arms:
|
||
s = savings(arms["baseline"], arms[arm])
|
||
lines.append(
|
||
f"| {task_id} | {arms[arm]['class']} | **{arm} savings %** | — "
|
||
f"| {s['input_tokens']} | {s['cache_creation_input_tokens']} "
|
||
f"| {s['cache_read_input_tokens']} | {s['output_tokens']} "
|
||
f"| {_na(s['cost_usd'])} | {s['duration_s']} | — | — |"
|
||
)
|
||
lines.append("")
|
||
all_aggs = [agg for arms in results.values() for agg in arms.values()]
|
||
excluded_total = sum(agg.get("excluded_runs", 0) for agg in all_aggs)
|
||
if excluded_total:
|
||
lines.append(
|
||
f"{excluded_total} run(s) hit session/infra errors or had unverifiable "
|
||
"evidence and were excluded "
|
||
"from medians and resolve denominators — see error_kind in results.jsonl."
|
||
)
|
||
missing_total = sum(agg.get("transcripts_missing", 0) for agg in all_aggs)
|
||
if missing_total:
|
||
lines.append(
|
||
f"{missing_total} run(s) had no locatable session transcript or it was "
|
||
"unreadable, so they were excluded from promotion evidence "
|
||
"(skill_invoked=null in results.jsonl)."
|
||
)
|
||
lines.append(
|
||
"Session ids for every run are in results.jsonl — open the matching "
|
||
"transcript to see where each arm spent its tokens."
|
||
)
|
||
return "\n".join(lines)
|
||
|
||
|
||
# ─── Main ────────────────────────────────────────────────────────────────────
|
||
|
||
|
||
def build_parser() -> argparse.ArgumentParser:
|
||
parser = argparse.ArgumentParser(description=__doc__)
|
||
parser.add_argument("--tasks", required=True, type=Path)
|
||
parser.add_argument("--runs", type=int, default=1)
|
||
parser.add_argument(
|
||
"--outage-streak",
|
||
type=int,
|
||
default=DEFAULT_OUTAGE_STREAK,
|
||
help="abort the sweep after this many consecutive session/infra/cleanup "
|
||
"failures (0 disables the circuit breaker)",
|
||
)
|
||
parser.add_argument(
|
||
"--arms",
|
||
nargs="+",
|
||
default=["workflow", "workflow_direct", "baseline"],
|
||
choices=[
|
||
"workflow",
|
||
"candidate_workflow",
|
||
"workflow_direct",
|
||
"candidate_workflow_direct",
|
||
"ce_workflow",
|
||
"ce_workflow_direct",
|
||
"review",
|
||
"ce_review",
|
||
"baseline",
|
||
"baseline_nomcp",
|
||
],
|
||
)
|
||
parser.add_argument("--claude-bin", default="claude")
|
||
parser.add_argument(
|
||
"--ce-plugin-dir",
|
||
type=Path,
|
||
default=None,
|
||
help="operator-supplied Compound Engineering plugin directory; required for ce_* arms",
|
||
)
|
||
parser.add_argument(
|
||
"--ce-plugin-version",
|
||
default=None,
|
||
help="exact Compound Engineering plugin version; required for ce_* arms",
|
||
)
|
||
parser.add_argument("--timeout", type=int, default=3600, help="per session, seconds")
|
||
parser.add_argument("--out", type=Path, default=None)
|
||
parser.add_argument(
|
||
"--model",
|
||
required=True,
|
||
help="named, versioned model passed to every `claude --model` invocation",
|
||
)
|
||
parser.add_argument(
|
||
"--proposer-model",
|
||
default=None,
|
||
help="model that generated the candidate overlay (recorded for provenance)",
|
||
)
|
||
parser.add_argument(
|
||
"--base-url",
|
||
default=None,
|
||
help="ANTHROPIC_BASE_URL override — point at an Anthropic-compatible "
|
||
"proxy (see free-model.litellm.yaml) to run on a free model",
|
||
)
|
||
parser.add_argument(
|
||
"--auth-token",
|
||
default=os.environ.get("GITNEXUS_BENCH_AUTH_TOKEN"),
|
||
help="ANTHROPIC_API_KEY for the --base-url endpoint (prefer GITNEXUS_BENCH_AUTH_TOKEN env)",
|
||
)
|
||
parser.add_argument(
|
||
"--include-expensive",
|
||
action="store_true",
|
||
help="include scenarios marked expensive: true (excluded by default)",
|
||
)
|
||
parser.add_argument(
|
||
"--candidate-overlay",
|
||
type=Path,
|
||
default=None,
|
||
help="directory mirroring .claude/skills/gitnexus-{plan,work}; applied only to candidate_* arms",
|
||
)
|
||
parser.add_argument(
|
||
"--promotion-metric",
|
||
choices=PROMOTION_METRICS,
|
||
default="cost_usd",
|
||
help="efficiency metric used by the deterministic candidate gate; "
|
||
"cost_usd (default) is the only CLI-reported number that includes "
|
||
"subagent spend — token metrics count only the main loop",
|
||
)
|
||
parser.add_argument("--promotion-min-runs", type=int, default=3)
|
||
parser.add_argument("--promotion-min-improvement", type=float, default=5.0)
|
||
parser.add_argument("--promotion-max-task-regression", type=float, default=20.0)
|
||
parser.add_argument("--task-bindings-json", default=None, help=argparse.SUPPRESS)
|
||
parser.add_argument("--promotion-target-bases-json", default=None, help=argparse.SUPPRESS)
|
||
return parser
|
||
|
||
|
||
def main() -> None:
|
||
parser = build_parser()
|
||
args = parser.parse_args()
|
||
try:
|
||
args.model = normalized_model_identifier(args.model)
|
||
args.proposer_model = (
|
||
normalized_model_identifier(args.proposer_model, flag="--proposer-model")
|
||
if args.proposer_model is not None
|
||
else None
|
||
)
|
||
task_document = yaml.safe_load(args.tasks.read_text())
|
||
if not isinstance(task_document, Mapping) or not isinstance(task_document.get("tasks"), list):
|
||
raise ValueError("task file must contain a tasks list")
|
||
tasks, skipped_expensive = select_tasks(
|
||
task_document["tasks"],
|
||
include_expensive=args.include_expensive,
|
||
)
|
||
oracle_snapshots = capture_task_oracles(tasks)
|
||
expected_task_bindings = json.loads(args.task_bindings_json) if args.task_bindings_json else None
|
||
if expected_task_bindings is not None and not isinstance(expected_task_bindings, list):
|
||
raise ValueError("--task-bindings-json must contain a list")
|
||
supplied_promotion_target_bases = (
|
||
json.loads(args.promotion_target_bases_json) if args.promotion_target_bases_json else {}
|
||
)
|
||
if not isinstance(supplied_promotion_target_bases, dict) or not all(
|
||
isinstance(path, str) and isinstance(digest, str)
|
||
for path, digest in supplied_promotion_target_bases.items()
|
||
):
|
||
raise ValueError("--promotion-target-bases-json must contain a string mapping")
|
||
ce_plugin_config = validate_ce_plugin_inputs(
|
||
args.arms,
|
||
args.ce_plugin_dir,
|
||
args.ce_plugin_version,
|
||
)
|
||
except (OSError, SandboxError, ValueError, yaml.YAMLError) as exc:
|
||
parser.error(str(exc))
|
||
raise AssertionError("ArgumentParser.error() returned unexpectedly")
|
||
|
||
candidate_arms = [arm for arm in args.arms if arm in CANDIDATE_ARMS]
|
||
if candidate_arms and args.candidate_overlay is None:
|
||
parser.error("candidate_* arms require --candidate-overlay")
|
||
if args.candidate_overlay is not None and not candidate_arms:
|
||
parser.error("--candidate-overlay requires at least one candidate_* arm")
|
||
for candidate_arm in candidate_arms:
|
||
incumbent_arm = CANDIDATE_ARMS[candidate_arm]
|
||
if incumbent_arm not in args.arms:
|
||
parser.error(f"{candidate_arm} must be paired with {incumbent_arm}")
|
||
if args.runs < 1 or args.promotion_min_runs < 1:
|
||
parser.error("--runs and --promotion-min-runs must be positive")
|
||
|
||
candidate_overlay = args.candidate_overlay.expanduser().absolute() if args.candidate_overlay is not None else None
|
||
overlay_digest = candidate_overlay_digest(candidate_overlay) if candidate_overlay is not None else None
|
||
if candidate_overlay is not None:
|
||
required_candidates = required_candidate_arms(candidate_overlay)
|
||
required_arms = [arm for candidate in required_candidates for arm in (CANDIDATE_ARMS[candidate], candidate)]
|
||
if args.arms != required_arms:
|
||
parser.error("candidate overlay requires exactly these paired arms: " + " ".join(required_arms))
|
||
try:
|
||
promotion_target_bases = committed_destination_base_digests(candidate_overlay)
|
||
except ValueError as exc:
|
||
# Overlay adds a promotion target with no committed base — a clean
|
||
# CLI error, not a traceback.
|
||
parser.error(str(exc))
|
||
raise AssertionError("ArgumentParser.error() returned unexpectedly")
|
||
if supplied_promotion_target_bases and supplied_promotion_target_bases != promotion_target_bases:
|
||
parser.error("--promotion-target-bases-json does not match the committed incumbent")
|
||
else:
|
||
if supplied_promotion_target_bases:
|
||
parser.error("--promotion-target-bases-json requires --candidate-overlay")
|
||
promotion_target_bases = {}
|
||
|
||
try:
|
||
bwrap_bin = preflight_bubblewrap()
|
||
require_claude_sandbox_helpers()
|
||
runtime_mounts = trusted_gitnexus_runtime_mounts()
|
||
except SandboxError as exc:
|
||
parser.error(str(exc))
|
||
raise AssertionError("ArgumentParser.error() returned unexpectedly")
|
||
out_dir = args.out or Path("results") / time.strftime("wfbench-%Y%m%d-%H%M%S")
|
||
out_dir.mkdir(parents=True, exist_ok=True)
|
||
results_path = out_dir / "results.jsonl"
|
||
selected_ids = [task["id"] for task in tasks]
|
||
print(
|
||
f"selected {len(selected_ids)} task(s): {', '.join(selected_ids)}; "
|
||
f"skipped {len(skipped_expensive)} expensive task(s): "
|
||
f"{', '.join(skipped_expensive) if skipped_expensive else 'none'}"
|
||
)
|
||
|
||
results: dict[str, dict[str, dict[str, Any]]] = {}
|
||
outage_streak = 0
|
||
outage_tripped = False
|
||
with (
|
||
tempfile.TemporaryDirectory(prefix="wfbench-trees-") as trees,
|
||
TaskAssetCache(Path(trees) / ".task-assets") as task_asset_cache,
|
||
staged_ce_plugin_snapshot(
|
||
ce_plugin_config,
|
||
destination_parent=Path(trees),
|
||
) as ce_plugin_snapshot,
|
||
):
|
||
try:
|
||
task_bindings = resolve_task_bindings(
|
||
tasks,
|
||
expected_task_bindings,
|
||
oracle_snapshots=oracle_snapshots,
|
||
task_asset_cache=task_asset_cache,
|
||
)
|
||
except (OSError, SandboxError, ValueError) as exc:
|
||
parser.error(str(exc))
|
||
raise AssertionError("ArgumentParser.error() returned unexpectedly")
|
||
oracle_mask = Path(trees) / ".oracle-mask"
|
||
oracle_mask.mkdir(mode=0o500)
|
||
oracle_mask.chmod(0o500)
|
||
graph_snapshots: dict[tuple[str, str], SanitizedGraphSnapshot] = {}
|
||
graph_snapshot_errors: dict[tuple[str, str], BaseException] = {}
|
||
for task, task_binding, oracle_snapshot in zip(
|
||
tasks,
|
||
task_bindings,
|
||
oracle_snapshots,
|
||
strict=True,
|
||
):
|
||
if outage_tripped:
|
||
break
|
||
repo = Path(task_binding["repo_identity"])
|
||
task_sha = task_binding["resolved_sha"]
|
||
asset_snapshot: TaskAssetSnapshot | None = None
|
||
asset_snapshot_error: BaseException | None = None
|
||
graph_key = (str(repo), task_sha)
|
||
graph_snapshot: SanitizedGraphSnapshot | None = graph_snapshots.get(graph_key)
|
||
graph_snapshot_error: BaseException | None = graph_snapshot_errors.get(graph_key)
|
||
try:
|
||
validate_no_prebuilt_graph_assets(task)
|
||
if graph_snapshot is None and graph_snapshot_error is None:
|
||
graph_snapshot = prepare_sanitized_graph(
|
||
task,
|
||
repo=repo,
|
||
resolved_sha=task_sha,
|
||
parent=Path(trees),
|
||
cache=task_asset_cache,
|
||
claude_bin=args.claude_bin,
|
||
bwrap_bin=bwrap_bin,
|
||
runtime_mounts=runtime_mounts,
|
||
)
|
||
graph_snapshots[graph_key] = graph_snapshot
|
||
except (ManagedProcessError, OSError, SandboxError, RuntimeError, ValueError) as exc:
|
||
graph_snapshot_error = exc
|
||
graph_snapshot_errors[graph_key] = exc
|
||
per_arm: dict[str, list[dict[str, Any]]] = {a: [] for a in args.arms}
|
||
for run_idx in range(args.runs):
|
||
if outage_tripped:
|
||
break
|
||
for arm in args.arms:
|
||
if outage_tripped:
|
||
break
|
||
worktree: Path | None = None
|
||
record: dict[str, Any] | None = None
|
||
cleanup_error: OSError | None = None
|
||
try:
|
||
if asset_snapshot_error is not None:
|
||
raise RuntimeError(f"task asset snapshot preparation failed: {asset_snapshot_error}")
|
||
if graph_snapshot_error is not None:
|
||
raise RuntimeError(f"sanitized graph snapshot preparation failed: {graph_snapshot_error}")
|
||
if graph_snapshot is None:
|
||
raise RuntimeError("sanitized graph snapshot is unavailable")
|
||
if asset_snapshot is None:
|
||
try:
|
||
asset_snapshot = task_asset_cache.prepare(
|
||
task,
|
||
repo=repo,
|
||
resolved_sha=task_sha,
|
||
expected_dependency_binding=task_binding,
|
||
)
|
||
except (OSError, SandboxError, ValueError) as exc:
|
||
asset_snapshot_error = exc
|
||
raise
|
||
worktree = make_worktree(repo, task_sha, Path(trees))
|
||
sanitized_head = sanitize_clone_for_hidden_oracles(worktree)
|
||
graph_snapshot.materialize(worktree, sanitized_head=sanitized_head)
|
||
dependency_mounts = stage_task_assets(
|
||
task,
|
||
repo=repo,
|
||
clone=worktree,
|
||
snapshot=asset_snapshot,
|
||
)
|
||
registry_mount = isolated_gitnexus_registry_mount(worktree, Path(trees))
|
||
hidden_harness = worktree / "eval" / "workflow_bench"
|
||
oracle_visibility_mounts: list[ReadOnlyMount] = []
|
||
if hidden_harness.exists() or hidden_harness.is_symlink():
|
||
hidden_metadata = hidden_harness.lstat()
|
||
if stat.S_ISLNK(hidden_metadata.st_mode) or not stat.S_ISDIR(hidden_metadata.st_mode):
|
||
raise SandboxError(
|
||
"benchmark harness path must be a real directory before it can be hidden"
|
||
)
|
||
oracle_visibility_mounts.append(
|
||
ReadOnlyMount(
|
||
source=oracle_mask,
|
||
target=f"{SANDBOX_WORKSPACE}/eval/workflow_bench",
|
||
)
|
||
)
|
||
execution_arm = CANDIDATE_ARMS.get(arm, arm)
|
||
ce_mounts = ce_plugin_mounts_for_arm(execution_arm, ce_plugin_snapshot)
|
||
with prepare_sandbox(
|
||
clone=worktree,
|
||
claude_bin=args.claude_bin,
|
||
bwrap_bin=bwrap_bin,
|
||
read_only_mounts=[
|
||
*dependency_mounts,
|
||
*runtime_mounts,
|
||
registry_mount,
|
||
*ce_mounts,
|
||
*oracle_visibility_mounts,
|
||
],
|
||
preflight=False,
|
||
) as sandbox:
|
||
# Capture the BASE (pre-overlay) skill digest — identical
|
||
# for the incumbent and candidate arms — then run the
|
||
# task's untrusted setup against those base skills. The
|
||
# candidate overlay is applied only afterwards, so setup
|
||
# can never observe candidate prose and both arms share
|
||
# byte-identical pre-overlay state.
|
||
base_skill_digest = skill_fingerprint(worktree, execution_arm)
|
||
if task.get("setup"):
|
||
setup_command = ["/bin/sh", "-lc", str(task["setup"])]
|
||
setup = sandbox.run(
|
||
setup_command,
|
||
timeout=600,
|
||
env=build_sandbox_environment(),
|
||
)
|
||
if not setup.ok:
|
||
raise ManagedProcessError(setup_command, setup)
|
||
# Tamper-evidence: setup must not have rewritten the base
|
||
# skills, verified before any candidate overlay lands.
|
||
require_skill_fingerprint(
|
||
worktree,
|
||
execution_arm,
|
||
base_skill_digest,
|
||
phase="task setup",
|
||
)
|
||
if arm in CANDIDATE_ARMS:
|
||
assert candidate_overlay is not None
|
||
applied_digest = apply_candidate_overlay(
|
||
candidate_overlay,
|
||
worktree,
|
||
sandbox=sandbox,
|
||
)
|
||
if applied_digest != overlay_digest:
|
||
raise RuntimeError("candidate overlay changed during the benchmark run")
|
||
# The digest the model must preserve during its run is the
|
||
# post-overlay skill surface (candidate skills for
|
||
# candidate arms; unchanged base skills otherwise).
|
||
expected_skill_digest = skill_fingerprint(worktree, execution_arm)
|
||
orig_sha = _sandbox_git(sandbox, ["rev-parse", "HEAD"]).strip()
|
||
if not re.fullmatch(r"[0-9a-fA-F]{40,64}", orig_sha):
|
||
raise RuntimeError("sandboxed candidate setup did not produce an immutable commit")
|
||
before_work_digest = (
|
||
implementation_diff_digest(sandbox, orig_sha)
|
||
if execution_arm in IMPLEMENTATION_ARMS
|
||
else ""
|
||
)
|
||
record = run_arm(
|
||
execution_arm,
|
||
task,
|
||
worktree,
|
||
args,
|
||
sandbox=sandbox,
|
||
transcript_output_dir=out_dir,
|
||
transcript_output_prefix=f"{task['id']}-{arm}-run{run_idx}",
|
||
expected_skill_digest=expected_skill_digest,
|
||
enforce_phase_boundary=True,
|
||
ce_plugin_dir=ce_plugin_dir_for_arm(execution_arm, ce_plugin_snapshot),
|
||
oracle_snapshot=oracle_snapshot,
|
||
)
|
||
_prepare_untracked_for_diff(sandbox)
|
||
after_work_digest = (
|
||
implementation_diff_digest(
|
||
sandbox,
|
||
orig_sha,
|
||
prepare_untracked=False,
|
||
)
|
||
if execution_arm in IMPLEMENTATION_ARMS
|
||
else ""
|
||
)
|
||
record.update(
|
||
diff_churn(
|
||
sandbox,
|
||
orig_sha,
|
||
prepare_untracked=False,
|
||
)
|
||
)
|
||
enforce_work_evidence(
|
||
record,
|
||
arm=execution_arm,
|
||
before_digest=before_work_digest,
|
||
after_digest=after_work_digest,
|
||
)
|
||
patch_bytes = capture_patch(sandbox, worktree, orig_sha)
|
||
record["arm"] = arm
|
||
record.update(
|
||
{
|
||
"model": args.model,
|
||
"benchmark_model": args.model,
|
||
"proposer_model": args.proposer_model,
|
||
"task_ref": task.get("ref", "HEAD"),
|
||
"task_base_sha": task_sha,
|
||
"sanitized_task_sha": sanitized_head,
|
||
"variant_head_sha": orig_sha,
|
||
"task_prompt_digest": hashlib.sha256(task["prompt"].encode()).hexdigest(),
|
||
"skill_digest": expected_skill_digest,
|
||
"candidate_overlay_digest": (overlay_digest if arm in CANDIDATE_ARMS else None),
|
||
"recorded_at": datetime.now(UTC).isoformat(),
|
||
}
|
||
)
|
||
# Final working-tree patch — the clone is destroyed, so
|
||
# this is the only artifact for diagnosing verify fails.
|
||
patch_path = out_dir / f"{task['id']}-{arm}-run{run_idx}.patch"
|
||
patch_path.write_bytes(patch_bytes)
|
||
except (
|
||
ManagedProcessError,
|
||
SandboxError,
|
||
OSError,
|
||
RuntimeError,
|
||
ValueError,
|
||
) as exc:
|
||
# One hung session or failed setup must not abort the
|
||
# sweep — record the run as infra-error and move on so
|
||
# report.md/promotion.json still get written.
|
||
record = infra_error_record(exc)
|
||
record["arm"] = arm
|
||
print(f"[{task['id']}][{arm}][run {run_idx}] infra-error: {exc}")
|
||
finally:
|
||
if worktree is not None and worktree.exists():
|
||
try:
|
||
remove_clone(worktree)
|
||
except OSError as exc:
|
||
cleanup_error = exc
|
||
assert record is not None
|
||
if cleanup_error is not None:
|
||
primary_kind = record.get("error_kind")
|
||
primary_detail = record.get("error_detail")
|
||
record["resolved"] = False
|
||
record["ok"] = False
|
||
record["error_kind"] = "cleanup-failure"
|
||
record["error_detail"] = (
|
||
f"primary={primary_kind}: {primary_detail}; cleanup: "
|
||
f"{type(cleanup_error).__name__}: {cleanup_error}"
|
||
)[:2000]
|
||
record.update(
|
||
{
|
||
"task": task["id"],
|
||
"class": task.get("class", ""),
|
||
"run": run_idx,
|
||
"task_asset_snapshot_digest": (
|
||
asset_snapshot.digest if asset_snapshot is not None else None
|
||
),
|
||
"task_asset_manifest_digest": (
|
||
asset_snapshot.manifest_digest if asset_snapshot is not None else None
|
||
),
|
||
"sandbox_dependency_content_digest": (
|
||
asset_snapshot.dependency_content_digest if asset_snapshot is not None else None
|
||
),
|
||
"sandbox_dependency_manifest_digest": (
|
||
asset_snapshot.dependency_manifest_digest if asset_snapshot is not None else None
|
||
),
|
||
"sanitized_graph_snapshot_digest": (
|
||
graph_snapshot.digest if graph_snapshot is not None else None
|
||
),
|
||
"sanitized_graph_manifest_digest": (
|
||
graph_snapshot.manifest_digest if graph_snapshot is not None else None
|
||
),
|
||
"oracle_digest": oracle_snapshot.digest,
|
||
"oracle_command_digest": oracle_snapshot.command_digest,
|
||
"oracle_manifest_digest": oracle_snapshot.manifest_digest,
|
||
"ce_plugin_version": (
|
||
ce_plugin_snapshot.version
|
||
if arm in CE_ARMS and ce_plugin_snapshot is not None
|
||
else None
|
||
),
|
||
"ce_plugin_manifest_digest": (
|
||
ce_plugin_snapshot.manifest_digest
|
||
if arm in CE_ARMS and ce_plugin_snapshot is not None
|
||
else None
|
||
),
|
||
}
|
||
)
|
||
per_arm[arm].append(record)
|
||
with results_path.open("a") as fh:
|
||
fh.write(json.dumps(record) + "\n")
|
||
print(
|
||
f"[{task['id']}][{arm}][run {run_idx}] resolved={record['resolved']} "
|
||
f"in={record['input_tokens']} out={record['output_tokens']} "
|
||
f"cost=${_na(record['cost_usd'])}"
|
||
)
|
||
outage_streak = systemic_outage_streak(record.get("error_kind"), outage_streak)
|
||
if args.outage_streak and outage_streak >= args.outage_streak:
|
||
outage_tripped = True
|
||
print(
|
||
f"[systemic-outage] {outage_streak} consecutive session/infra/cleanup "
|
||
"failures — aborting the remaining sweep; report and promotion are written "
|
||
"from partial evidence and the run exits non-zero."
|
||
)
|
||
break
|
||
results[task["id"]] = {a: aggregate(rs) for a, rs in per_arm.items() if rs}
|
||
|
||
selection_report = [
|
||
"## Run provenance",
|
||
"",
|
||
f"Benchmark model: `{args.model}`",
|
||
f"Proposer model: `{args.proposer_model}`",
|
||
f"Selected tasks ({len(selected_ids)}): {', '.join(selected_ids)}",
|
||
(
|
||
f"Skipped expensive tasks ({len(skipped_expensive)}): "
|
||
+ (", ".join(skipped_expensive) if skipped_expensive else "none")
|
||
),
|
||
]
|
||
if ce_plugin_snapshot is not None:
|
||
selection_report.append(
|
||
f"Compound Engineering plugin: `{ce_plugin_snapshot.version}` (`{ce_plugin_snapshot.manifest_digest}`)"
|
||
)
|
||
report = render_report(results) + "\n\n" + "\n".join(selection_report) + "\n"
|
||
(out_dir / "report.md").write_text(report)
|
||
if candidate_arms:
|
||
promotion_generated_at = datetime.now(UTC)
|
||
promotion = {
|
||
# Schema 3 is the first promotion evidence that requires hidden,
|
||
# byte-bound behavioral oracles. Older self-authored-only rows are
|
||
# intentionally ineligible for application.
|
||
"schema_version": 3,
|
||
"generated_at": promotion_generated_at.isoformat(),
|
||
"evidence_expires_at": (promotion_generated_at + timedelta(days=EVIDENCE_MAX_AGE_DAYS)).isoformat(),
|
||
"benchmark_model": args.model,
|
||
"proposer_model": args.proposer_model,
|
||
"candidate_origin": ("model-proposer" if args.proposer_model is not None else "manual-initial-overlay"),
|
||
"candidate_overlay": str(candidate_overlay),
|
||
"candidate_overlay_digest": overlay_digest,
|
||
"target_base_digests": promotion_target_bases,
|
||
"required_candidate_arms": candidate_arms,
|
||
"selected_tasks": task_bindings,
|
||
"ce_plugin": ce_plugin_snapshot.provenance if ce_plugin_snapshot is not None else None,
|
||
"policy": {
|
||
"metric": args.promotion_metric,
|
||
"metric_warning": (MAIN_LOOP_ONLY_WARNING if args.promotion_metric in MAIN_LOOP_ONLY_METRICS else None),
|
||
"min_runs": args.promotion_min_runs,
|
||
"min_improvement_pct": args.promotion_min_improvement,
|
||
"max_task_regression_pct": args.promotion_max_task_regression,
|
||
"quality_rule": "no per-task resolution-rate regression",
|
||
"max_age_days": EVIDENCE_MAX_AGE_DAYS,
|
||
},
|
||
"decisions": [
|
||
evaluate_candidate(
|
||
results,
|
||
incumbent_arm=CANDIDATE_ARMS[candidate_arm],
|
||
candidate_arm=candidate_arm,
|
||
model=args.model,
|
||
metric=args.promotion_metric,
|
||
min_runs=args.promotion_min_runs,
|
||
min_improvement_pct=args.promotion_min_improvement,
|
||
max_task_regression_pct=args.promotion_max_task_regression,
|
||
)
|
||
for candidate_arm in candidate_arms
|
||
],
|
||
}
|
||
(out_dir / "promotion.json").write_text(json.dumps(promotion, indent=2) + "\n")
|
||
print(f"\n{report}\n\nWritten to {out_dir}/")
|
||
if outage_tripped:
|
||
# Non-zero exit so a driver (evolve.py) treats the partial benchmark as a
|
||
# failed run and halts instead of proposing from outage-truncated evidence.
|
||
raise SystemExit(1)
|
||
|
||
|
||
if __name__ == "__main__":
|
||
main()
|