GitNexus/eval/tests/test_promotion_apply.py
Gergő Magyar 8b5057f325
feat(skills): GitNexus Engineering Tool Kits (#2566)
* feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill

Adds .claude/skills/ce-plan: a planning-only skill that builds
implementation-ready plans from GitNexus graph navigation (query/context/
impact/trace), bounded statement-level PDG slices (pdg_query, impact
mode:pdg, explain), and targeted source verification, with a context
ledger to prevent repeated reads and a machine-readable implementation
context pack (stable contract for a future ce-implement). Whitelisted in
.gitignore and registered in AGENTS.md and CLAUDE.md outside the
auto-managed gitnexus block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions)

Tool contract: impact mode:'pdg' shape now includes the schema-required
direction param; CDG branch sense documented as the result 'label' field
(reason is cypher/raw-edge only); explain caveats corrected to its real
false-negative classes (cross-function TAINT_PATH is modeled).

Consistency: PDG slice homed in working memory (ledger keeps one-liners);
depth knob defined and category-overrides-baseline ordering stated;
call_depth (consumed by nothing) and content-hash bookkeeping dropped;
Never section folded into Hard rules; Phase 3 deduplicated to a pointer;
allowed-repeat escalations defined; budget/discard accounting clarified;
verification-commands gathering added to Phase 4; open_questions added to
the context pack.

From scenario runs: plans now pin the verified-at HEAD commit and index
freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed],
quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and
support an out:<path> destination override; output path defined as the
Phase 1 target repo root.

Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata
bumps; future ce-implement qualified as future.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): rename ce-plan → gitnexus-plan; add cross-CLI (Codex) entrypoints

Renames the skill dir, frontmatter, output filename convention, plan H1
(GitNexus Engineering Plan), the future executor handle
(gitnexus-implement), the .gitignore whitelist entry, and all
AGENTS.md/CLAUDE.md references. Follows the pr-swarm-review cross-CLI
pattern: SKILL.md is the canonical CLI-neutral spec, AGENTS.md § Engineering
planning is the Codex/any-agent entrypoint, and the README documents the
optional user-level ~/.codex/prompts/gitnexus-plan.md slash command plus an
invocation matrix. Skill prose de-branded from Claude Code (agent-neutral
verification layer).

Also fixes two post-review README contradictions: the anti-reread claim now
names the ledger's allowed escalations, and 'read-only by contract' is now
'planning-only' (the skill writes exactly one repo file — the plan); the
scope-creep rule and template §12 now agree on where deferred follow-ups
land. Drops the stale plugin-collision limitation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): document Codex user-level install path for gitnexus-plan

Codex discovers SKILL.md skills from ~/.agents/skills (same path the other
gitnexus-* skills install to); README now documents the cp install plus the
optional ~/.codex/prompts slash-command file, with the prompt body preferring
the repo copy and falling back to the user-level install.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan freshness gate + active PDG-layer refresh

Freshness is now a Phase 1 gate, not advisory: under the default
freshness:strict, a stale index is refreshed once per planning session via
node .gitnexus/run.cjs analyze --index-only (appending --pdg when the task
will reach the PDG phase), then the context resource is re-read. A missing
PDG layer likewise triggers the one permitted --index-only --pdg refresh
and re-probe instead of a passive recommendation. freshness:accept (or a
failed/impractical refresh) preserves the old behavior: plan on the stale
graph, source-weighted, labelled in the plan header. --index-only is the
load-bearing flag choice — it suppresses all file generation, so the
planning-only contract holds (only the .gitnexus store changes). Ledger
gains an index_refresh record; plan header states fresh / refreshed /
refresh-skipped-with-reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan runner build check before freshness refresh

When the target repo builds the analyzer from its own source (bin → dist/
mapping, as gitnexus/ does), the Phase 1 freshness gate now verifies dist/
is current before running the analyze refresh — rebuilding via the
package's build script when any analyzer source file is newer than the
built entrypoint — and prefers that freshly built CLI. Otherwise a stale
dist re-indexes with outdated extraction logic and the 'fresh' index lies.
Rebuilds are recorded in the ledger's index_refresh; the PDG-phase refresh
inherits the same check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): add gitnexus-work executor and gitnexus-lfg pipeline

gitnexus-work executes a gitnexus-plan as verified atomic commits: consumes
the §11 implementation_context pack, drift-checks the plan's evidence pin
against HEAD, re-verifies assumptions before relying on them, runs impact
before every symbol edit and detect_changes before every commit (repo
mandates), builds tests from the plan's scenarios, and routes structural
drift back to gitnexus-plan Deepen mode instead of coding around it.

gitnexus-lfg is a thin orchestrator: gitnexus-plan → blocking user gate
(deepen / proceed / stop, deepen loops allowed) → gitnexus-work → review
via the existing gitnexus-pr-review skill (open PR, else branch diff vs
default). One bounded fix cycle for review findings; never pushes or opens
a PR on its own.

gitnexus-plan gains a Deepen mode (re-run freshness gate, escalate to
depth:deep, re-verify graph/inferred/assumed claims toward verified,
rewrite the same file); its 'future gitnexus-implement' placeholder is
retired in favor of gitnexus-work. Registered via .gitignore whitelists,
AGENTS.md 1.10.0 (section renamed to Engineering planning & execution),
CLAUDE.md 1.5.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply cross-skill review findings to the gitnexus skill family

Two P1s: gitnexus-plan Deepen mode now re-anchors before re-pinning
(diffs the old evidence pin over every [verified]-claim file and re-reads
or downgrades before the header moves — moving the pin without this
laundered stale claims as verified); the index-refresh budget is stated
once in Phase 1 (one --index-only refresh plus at most one Phase 3 --pdg
upgrade per session, Deepen = its own session) with ledger and pdg-slice
deferring to it.

Contract fixes: gitnexus-work's drift check now covers every file the
pack cites (not just files_to_modify) and parses the full pack incl.
primary/related symbols and acceptance_criteria (walked in Phase 4
alongside §13); a pre-completed check skips §7 steps already landed and
Deepen gains a reconcile-execution-state step, closing the mid-execution
route-back loop; pack assumptions must name what to check and how.

lfg: Lane 4 passes the merge-base to detect_changes compare (two-dot
diff misattributes upstream commits when default advanced), branch-diff
is the stated normal case, oversized review findings route to the plan
gate instead of overflowing direct mode, the one-fix-cycle cap is
explicit on re-run, and headless runs end at the plan gate with the plan
as deliverable. work: blank mode narrowed to *gitnexus-plan*.md with a
re-execution guard, direct-mode discipline spelled out, branch
meaningfulness defined against the plan slug, and the plan document is
committed as the branch's docs commit (review diff includes it).
Planning-only contract now names the dist/ rebuild as the second
permitted state change; Phase 5.1 names the four claim tags; stale
AGENTS.md anchors fixed.

Known latent issue left untouched: gitnexus/gitnexus-pr-review pairs a
three-dot example with a two-dot detect_changes compare — that skill is
also shipped by the plugin, so fixing it here would drift the copies;
lfg compensates by passing the merge-base.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ship the engineering skill family with the gitnexus package

npm i -g gitnexus users now get gitnexus-plan / gitnexus-work / gitnexus-lfg:
the three skills are added to gitnexus/skills/ in directory form (SKILL.md +
references/), which installSkillsTo already enumerates dynamically and copies
recursively to every editor target (~/.agents/skills for Codex, Cursor,
OpenCode, Qoder, ...) on gitnexus setup — uninstall enumerates the same root,
so removal stays clean. The Claude Code plugin channel
(gitnexus-claude-plugin/skills/) carries the same copies plus the standard
per-skill mcp.json.

Global-install support in the skill text: gitnexus-plan Phase 1 now resolves
the analyzer runner explicitly — node .gitnexus/run.cjs analyze when the
project has a runner, else gitnexus analyze (installed CLI), else
npx gitnexus analyze — and all analyze mentions route through it, satisfying
the skills-steering policy (#1939/#1945) which sweeps the plugin copies.

New drift guard test/unit/shipped-skills-sync.test.ts asserts the npm and
plugin copies stay byte-identical to the canonical .claude/skills/ family
(plugin = canonical + mcp.json), same discipline as run.cjs ↔
resolve-invocation.ts. skills-steering + shipped-skills-sync: 11/11 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench — measure the skill workflow's token savings

Benchmarks gitnexus-plan → gitnexus-work against a baseline agent
(--disallowedTools Skill) on identical tasks, in fresh detached worktrees,
using real headless Claude Code sessions; every number comes from the CLI's
--output-format json usage report (field names validated against a live
2.1.207 session). Reports per-arm medians (input/cache/output tokens, cost,
wall time, turns), a savings row, and resolve status from a per-task verify
command — savings on failed tasks are flagged, not celebrated. Per-task
setup hook prepares fresh worktrees (deps); --permission-mode
bypassPermissions (default) lets sessions run unattended in the throwaway
trees.

Free-model support: --base-url/--auth-token/--model route headless sessions
through any Anthropic-compatible endpoint; free-model.litellm.yaml is a
ready litellm-proxy template for OpenRouter :free variants or local Ollama,
so benchmarking burns no paid tokens (README documents rate limits and the
small-model skill-following caveat).

Harness validated end-to-end with a stub CLI (worktree lifecycle, both
arms, plan→work chaining, verify, aggregation, report) and 4 pytest units
for the pure aggregation/savings/report helpers. AGENTS.md 1.11.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record first workflow_bench calibration run

Trivial-task calibration (add -V alias): both arms resolved; workflow arm
~4.3x baseline cost — the documented overhead-dominated regime, recorded so
the regime boundary is empirical rather than asserted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench scenario matrix — arm variants, task classes, churn

Ground-base measurement across scenarios: tasks.scenarios.yaml spans four
labeled classes (trivial → investigation-bug → investigation-feature →
cross-module) with deterministic verifies (prescribed test files). New arms:
workflow_direct (gitnexus-work direct mode — the middle option that locates
the routing boundary lfg's gate and work's triage encode) and baseline_nomcp
(no skills AND no graph tools — separates workflow-discipline value from
GitNexus-tool value; off by default). Records now carry task class and diff
churn (files/+ins/−del vs the starting commit) as an over-engineering proxy;
the report renders a class column and per-arm savings rows vs baseline.
5 pytest units + stub-CLI e2e of the full three-arm matrix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record workflow_bench ground base; fix churn measurement bias

Ground base (3 classes x 3 arms, n=1/cell): every arm resolved every task —
pass/fail quality saturates at this difficulty, making the comparison pure
cost. Full plan→work never amortized its ~$9-11 fixed cost on tasks a
baseline finishes in ≤35 turns (−211% to −333% cost); workflow_direct sits
near baseline (−15% to −55%, once faster wall) with more test coverage.
Routing implication recorded: direct mode/plain agent below this scale,
full workflow for cross-module / multi-session / plan-as-deliverable work.
The cross-module cell and multi-run variance are the next measurements.

Churn fix: git add --intent-to-add -A before diffing (arms that never
commit no longer undercount new files) and :(exclude)docs/plans (the
committed plan doc no longer inflates workflow churn); this run's churn
numbers predate the fix and are omitted from the recorded table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(skills): cost-optimize the workflow from measured ground base

Every optimization targets a measured fixed-cost component
(eval/workflow_bench ground base: workflow arm −211% to −333% vs baseline,
all tasks resolved):

- Plan form is category-priced: compact form (core sections w/ § anchors
  preserved, ≤80 lines excl. pack, mini-pack subset of the context pack)
  for narrow/default categories; the full 13 sections only for deep work
  (refactor/security/performance/concurrency/architecture). A compact plan
  outgrowing its cap reclassifies to full rather than overflowing.
- Freshness gate is category-priced: compact categories default to accept
  (source-weighted, refresh only when a graph claim becomes load-bearing);
  strict stays the default for full-plan categories — the rebuild+re-index
  was the largest single fixed cost.
- Turn economy: per-category tool-call budgets (~10 to ~45; architecture
  uncapped); budget exhaustion routes open questions to §12 instead of
  more digging.
- gitnexus-work fast path: HEAD == evidence pin → skip all citation
  re-reading (the pin's entire point); mini-pack fields tolerated.
- lfg Lane 1 boundary triage: tasks below the measured ~35-turn boundary
  get offered gitnexus-work direct mode before the plan lane is spent.

Copies re-synced (npm skills/, plugin, ~/.agents); steering + sync guards
green. Re-measurement of the workflow arm follows to verify the numbers
actually improve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record optimization re-measurement — inv-bug workflow cell −20% cost

Same task, same conditions, post-830a0459 skills: $14.56→$11.70 (−20%),
83→72 turns, cache_read −24%; verified in-transcript that the compact form,
turn budget, and skipped rebuild/re-index all fired. Wall +15% from a work-
session test-debugging tail (n=1 variance). Regime unchanged (~3.5x baseline
on this class) — routing rule stands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): per-arm clone isolation — worktree ref-namespace leak contaminated an arm

The cross-module workflow_direct cell reported an impossible 28-turn solve
with churn byte-identical to the workflow arm: git worktree add shares the
repo's ref namespace, so the workflow arm's slug branch (created by
gitnexus-work Phase 2) survived worktree removal and the direct arm found
and adopted the completed work. Arms now get isolated git clone --shared
copies (object store via alternates, refs clone-local — agent branches and
stashes die with the clone; origin/<ref> fallback for non-default refs).
Leaked branch deleted; baseline arm verified clean (0 branch references in
its transcript); cell marked invalidated pending re-run.

Records the valid cross-module cells: workflow $18.32 vs baseline $18.03
(premium −1.6%, vs −211%..−333% on smaller classes) — fixed costs amortize
at this scale, with a less destructive diff and a plan artifact as bonus;
resolve rate still tied. Churn fingerprinting is what caught the
contamination — noted in the README as an integrity check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): complete cross-module cell — direct mode wins 47% cost / 56% wall

Clean clone-isolated re-run: workflow_direct resolved the hardest class at
$9.53/52 turns/15m vs $18.03/98/34m baseline and $18.32/107/37m full
workflow. The measured story across all four classes: the execution
discipline (gitnexus-work) is the consistent sweet spot and delivers real
token savings on hard tasks; the planning pass buys its artifact, not
same-session savings. Resolve rate tied everywhere (n=1/cell caveat).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): add trajectory-gated skill evolution (#2431)

- Pair prompt candidates with incumbent workflow arms
- Gate promotions on pinned-model quality and efficiency
- Expire router evidence and document its lifecycle

* fix(eval): allow pr-review skill candidates

* feat(skills): rename and generalize GitNexus review

* feat(eval): external-comparator and review arms for workflow_bench

- ce_workflow / ce_workflow_direct: compound-engineering ce-plan/ce-work
  arms prompted with the same structure as the gitnexus arms
- review / ce_review: gitnexus-review vs ce-code-review on an identical
  diff applied by the task's setup
- plan handoff is snapshot-based: committed example plans in docs/plans/
  tie on clone mtimes and broke the name-glob pick (executed a stale plan)
- verify output tail is recorded per run and the final working-tree patch
  is kept, so failed rows are diagnosable after the clone is destroyed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills,eval): address #2431 review — data-safe rename migration, fail-closed bench evidence

- setup: never delete a legacy renamed skill dir — the installer cannot
  prove ownership (users customize or hand-write skills under these
  names); warn with the path instead, and the test now asserts survival
- workflow_bench: fail closed when a session's --output-format json
  report is empty, malformed, or missing usage fields — an exit-0 shell
  with no parseable usage no longer counts as measured evidence
  (5 parametrized regression tests)
- workflow_bench: document the trust model prominently (task setup/verify
  are shell-executed, sessions run bypassPermissions with the parent env,
  candidate overlays are prompt injection surface) in README + docstring
- free-model.litellm.yaml: master_key from LITELLM_MASTER_KEY env instead
  of a static token; loopback-binding warning
- ci: run the eval workflow_bench pytest suite on ubuntu (pytest+pyyaml
  only — no full eval stack)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): demand observed foreground verification in headless work-arm prompts

In a headless -p session there is no later turn: a work arm backgrounded
its slow test run, scheduled wakeups that can never fire, and reported
done while two of its tests failed. All four work-arm prompts (both
skill families, symmetric) now require verification output to be
observed inside the session.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ask plan depth up front instead of offering deepen afterwards

gitnexus-plan Phase 0 now asks one blocking question in interactive
sessions — quick / standard / deep, mapped onto the existing depth/form/
freshness knobs — when the invocation carries no explicit depth signal.
Explicit knobs and headless runs skip the question (category posture
unchanged, so benchmarks and automation behave as before).

gitnexus-lfg's plan gate slims to proceed/stop: depth was already the
user's up-front choice, so deepening is no longer offered by default —
an explicit deepen request at the gate and executor route-backs still
run Deepen mode, which remains the mechanism for strengthening an
existing plan document.

All shipped copies resynced (npm skills/, Claude plugin); AGENTS.md
1.13.0 and CLAUDE.md 1.7.0 pointers updated, including the analyzer's
regenerated index-stats block at this branch's head.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): taint pass, expert lenses, and post-work index refresh

gitnexus-review gains a PDG-backed taint-and-dependence pass (explain +
pdg_query, --pdg folded into the stale refresh on trust-boundary diffs) and
an Expert lenses section: domain reviewers derived from the graph's
clusters plus four cross-cutting lenses (architectural fit, language
conformance per the repo's own contract, Definition of Done, simplicity),
dispatched once after the evidence-gathering steps and scaled to the diff.
gitnexus-work Phase 4 now refreshes the knowledge graph after the DoD walk
via the resolved-runner ladder with analyze --index-only, so the lfg review
lane and later sessions query the finished work without dirtying the tree.
lfg's threshold-governance paragraph moves to its README; eval citations
are tagged as measured in the GitNexus repo. All shipped copies re-synced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(cli): remove legacy gitnexus-pr-review on uninstall; cover the rename migration

uninstall's removal set now includes LEGACY_SKILL_DIR_NAMES derived from
RENAMED_SKILL_DIRS, so a pre-rename install is cleaned up instead of
orphaned. The rename warning gains behavioral coverage (fires with a legacy
dir present, silent without), and shipped-skills-sync asserts legacy names
stay absent from every shipped tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): metric provenance, error-kind rows, skill-invocation verification, gate noise floor

The promotion gate defaults to cost_usd (the only metric that includes
subagent spend); token metrics carry an explicit main-loop-only warning in
the report and promotion.json. Rows are classified by error_kind
(session-error / verify-failed / infra-error), excluded from efficiency
medians, and the gate requires equal valid-run counts. Each session's
transcript is scanned for the expected Skill invocation and fails closed on
a verified miss; a one-run resolution edge no longer promotes (noise
floor). Per-run timeouts and setup failures record an infra-error row
instead of aborting the sweep. Overlays touching skills no candidate arm
exercises are rejected up front.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fix skill routing paths, version headers, and skill rosters

Routing tables point at the tracked direct skill paths (matching the
post-#2434 generator output), AGENTS.md/CLAUDE.md headers match their
latest changelog rows, the 1.12.0 row describes what the migration actually
does, package/cursor READMEs list the full shipped skill roster, and the
swarm READMEs describe /gitnexus-review's expert lenses instead of calling
it single-agent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: drift-guard workflow for skill copies; pin eval pip deps; track docs/plans

ci.yml ignores '**.md', so an md-only skill edit would merge without the
shipped-skills-sync test running — skill-sync.yml triggers exactly on the
guarded trees. The eval job's pip install is version-pinned, and
docs/plans/ is unignored so gitnexus-plan output can be committed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep the runner-invocation literal in gitnexus-review; add concurrency block to skill-sync

skills-steering requires skills with a stale-index hint to carry the exact
'node .gitnexus/run.cjs analyze' form — restore it with the fallback ladder
as a parenthetical instead of replacing it. skill-sync.yml gains the
top-level concurrency block the workflow-convention check enforces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): token-economy guidance for expert lenses

Merge lenses that ground in the same material into one reviewer, and use
cheaper model/effort tiers for mechanical lenses where the harness offers
them, reserving the strongest engine for adversarial judgment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(eval): isolate transcript home on Windows

Ensure workflow_bench transcript tests set USERPROFILE alongside HOME so Path.home() resolves to the temporary test home on Windows.

* docs(skills): fold PR #2522 execution learnings into review/work/plan

Eight incident-backed hardenings from running the full skill cycle
(review -> plan -> work, 28-finding fix series) on PR #2522:

gitnexus-review:
- Expert lenses execute the code under review on candidate failing shapes
  (empirical probe outranks source reading — every HIGH the language
  lenses found came from a probe, not a read).
- Step 7 re-runs the exact CI check for refreshed baselines/fingerprints
  (a stale committed artifact is invisible in the diff; caught a red
  benchmarks arm).
- Step 8 treats version/invalidation constants as review surface
  (INCREMENTAL_SCHEMA_VERSION class recurred verbatim from #2494).

gitnexus-work:
- Step 4 proves regression tests discriminate against the pre-fix tree.
- Step 5 rebuilds executed build output before every verification run
  (parse workers load dist/; a correct fix 'failed' until rebuilt).
- Step 6 makes stage -> detect_changes -> commit one unbroken sequence.

gitnexus-plan:
- Phase 0 seeded-evidence mode: plan FROM a completed review's verified
  findings instead of re-running the graph ladder.
- Template §7: fingerprint/golden-guarded output rebaselines once, at the
  series tip.

All distribution copies resynced; shipped-skills-sync + skills-steering
24/24 locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): close the skill-evolution loop with an automated proposer driver

workflow_bench.evolve adds the three arrows the README described as manual:
a proposer session that turns loser trajectories (results.jsonl rows,
transcripts, patches, the learning queue) into ONE bounded candidate
overlay, a driver that iterates propose -> paired benchmark -> deterministic
gate up to --generations, and an --apply step that copies a promoted
overlay onto the canonical skills and shipped mirrors as a working-tree
diff. The trust boundary is unchanged: overlays re-validate through
candidate_overlay_files before any benchmark or apply consumes them, and
committing, CI, and the PR merge stay human.

learnings.jsonl is gitignored: it is machine-local evidence, like the
session transcripts it complements.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): route live-task friction into the evolution learning queue

Each family skill gains a short 'Skill feedback' section: on friction with
the skill's own instructions, append one JSON line to
eval/workflow_bench/learnings.jsonl (GitNexus repo only) — never self-edit
the skill from a live task. The proposer in workflow_bench.evolve consumes
the queue as hints; a learning reaches a shipped skill only by beating the
incumbent on the paired benchmark. All shipped mirrors re-copied byte-
identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(tests): run the evolve helper tests in the eval pytest job

test_evolve.py needs only pytest+pyyaml, same as the harness tests the job
already runs — without this line the new module had no CI coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): comment-triggered GitNexus review agent for PRs

'@gitnexus review' from a maintainer (OWNER/MEMBER/COLLABORATOR; the action
re-validates write access) runs the repo's gitnexus-review skill headlessly
against the PR and posts the review as a sticky comment — remote triggering
with no local setup. Read-only by construction: contents: read token,
Write/Edit and web tools disallowed, Bash allowlisted to git reads and the
gitnexus CLI; analyze parses PR code with tree-sitter, never executes it.
Requires the ANTHROPIC_API_KEY repository secret; activates once the file
is on the default branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): dispatch lane + existing OAuth secret for the review agent

Align with claude.yml: same action pin and the CLAUDE_CODE_OAUTH_TOKEN
secret the repo already carries — no new secret to configure. Add a
workflow_dispatch lane (PR number input) so the agent can be triggered from
the Actions UI and tested before the issue_comment trigger reaches the
default branch. Allowlist gh pr view/diff and gh api, which the review
skill uses to pin PR SHAs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): close a fork-PR RCE vector in the review agent's tool allowlist

A live headless run of the exact workflow session against PR #2431 (66
turns, full gitnexus-review pass) surfaced a real HIGH-severity confused
deputy: .gitnexus/ is gitignored, not blocked — a fork PR can commit its
own .gitnexus/run.cjs, issue_comment checks out PR-head content, and the
skill's runner ladder tries 'node .gitnexus/run.cjs analyze' first. That
would execute fork-controlled JS inside a job holding
CLAUDE_CODE_OAUTH_TOKEN and a write-scoped GITHUB_TOKEN — the opposite of
the 'PR code is read, never executed' claim in the workflow's own header.

Fix: drop the run.cjs allowlist entry so analyze always resolves through
npx gitnexus (npm registry, not the checked-out tree); the skill's
documented fallback mode covers the resulting graceful degradation. Also
drop 'gh api' (not read-only — accepts -X POST/PATCH/DELETE) and downgrade
pull-requests: write to read (comment posting only needs issues: write;
the prompt already forbids formal review submission).

Same session flagged a latent evolve.py bug: select_evidence's cost sort
used dict.get's missing-key default, which doesn't cover an explicit JSON
null in a foreign --seed-results row and crashes proposer setup with
TypeError. Guarded with 'or 0.0' and added a regression test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden PR review and evolution trust boundaries

* ci: follow workflow concurrency convention

* fix(eval): make terminating error paths explicit

* fix: unblock hardened review runtime checks

* test: make containment canaries deterministic

* test: expose Claude canary tool failures

* fix: adapt clean shell environment for Claude

* fix(eval): accept the runner's transcript source key in evidence preflight

The proposer evidence preflight required transcript-artifact metadata to be
exactly {path, sha256, bytes}, but the runner stamps a fourth provenance key
(source=parent-captured-stream-json). Any --seed-results or generation>=2 run
therefore aborted with SandboxError before proposing or promoting. Pin the
producer literal as PARENT_EVENT_STREAM_SOURCE and validate it in the metadata
check, and round-trip real producer output through sum_sessions into the
preflight so the schema can't drift again.

* fix(eval): treat an unmeasured session cost as unavailable, not $0

well_formed validated only the nested usage block, so an otherwise-successful
session missing total_cost_usd was recorded as cost_usd=0.0 — and cost_usd is
the default promotion metric (lower wins), so a cost-less session scored as
free and could win promotion it never earned. Extract cost via measured_cost()
(None on absent/garbage, a measured 0.0 preserved), propagate None through
sum_sessions/aggregate/savings/report, and have the gate refuse to rank on a
metric that was not measured on every run in both arms.

* fix(eval): warn when ranking on the main-loop-only num_turns metric

num_turns comes from the CLI's top-level usage (main-loop session only), like
output_tokens, but selecting it emitted no metric_warning — so a subagent-heavy
candidate could look artificially efficient. Add num_turns to
MAIN_LOOP_ONLY_METRICS and broaden the warning to cover turns.

* fix(eval): fail closed when an overlay adds a file with no committed base

An overlay adding a new .md under gitnexus-{plan,work} passes the structural
overlay checks but has no committed base for committed_destination_base_digests
to bind against, so it raised an uncaught ValueError that crashed the evolve
driver (and runner --candidate-overlay) mid-run. Catch it at both call sites:
evolve reports NOT PROMOTED and exits, runner routes it through parser.error.

* feat(eval): circuit-break the runner sweep on a systemic outage

A sustained upstream outage used to pay out every remaining --timeout window
one session at a time. Track consecutive session/infra/cleanup failures via a
pure systemic_outage_streak helper; after --outage-streak (default 5) in a row,
stop the sweep, still write report.md/promotion.json from partial evidence, and
exit non-zero so evolve.py halts instead of proposing from truncated evidence.
A task's own resolved=False never trips the breaker.

* fix(cli): report a dirty working tree as stale in gitnexus status

status --json (and the human output) computed up-to-date from commit + runner
identity + completeness only, so a repo with uncommitted source changes at a
matching HEAD was reported up-to-date while analyze would still re-index it.
A graph-backed agent gating on that JSON could skip re-analysis on a stale
graph. Extract analyze's dirty-tree check into a shared isWorkingTreeDirty()
in storage/git and fold it into the status freshness decision.

* fix(ci): use single-slash deny globs in the review agent's disallowedTools

github.workspace already expands to an absolute path, so Read(/${{ github.workspace }}/**)
and Read(//proc/**),(//sys/**),(//dev/**) produced double-slash patterns that a
normalizing matcher may not match — silently no-opping the deny layer. Not
exploitable (the allowlist is the primary control and never grants those
paths), but the globs should be well-formed. Update the pinned test strings.

* ci: install gitnexus-shared with npm ci from the committed lockfile

The gitnexus-shared build floated its deps via npm install in three workflows
(skill-sync, ci-tests, and — most importantly — the release publish.yml) while
every other install step uses npm ci. The lockfile is committed and in sync, so
switch all three to npm ci for reproducible, locked installs.

* test(cli): make the shipped-skills drift guard reject symlinks

listFilesRecursive walked with readdirSync and snapshotDir read with
readFileSync, both of which follow symlinks — so a mirror file symlinked to the
canonical tree passed the byte-compare (and a symlinked mirror dir would be
followed too). Reject a symlinked root via lstat and any symlinked entry via
Dirent.isSymbolicLink, with negative tests (skipped on Windows).

* test(eval): guard the candidate-skill vs mirror-root coverage invariant

MIRROR_SKILL_ROOTS omits the Cursor tree, safe only because no candidate skill
is cursor-shipped. Pin that invariant: every CANDIDATE_SKILLS entry must exist
under canonical + every mirror root and must not ship to Cursor, so adding a
cursor-shipped skill to the candidate set (the PR #2488 asymmetric-sync class)
fails loudly instead of syncing three of four trees.

* docs(ci): describe the review agent's staged post-merge rollout

The DoD asked for a dry-run or triggered run before merge, but an issue_comment
(or newly added workflow_dispatch) workflow only ever executes the default-branch
copy, so it cannot be exercised from the PR that introduces it. Reword the DoD
and the activation checklist to a staged rollout: merge registered-but-disabled,
validate same-repo and fork execution post-merge, then enable the variable.

* fix: pin plugin skill mcp.json to the release version via #2445 tooling

The ten plugin skill mcp.json launched `npx -y gitnexus@latest mcp` on every
skill connect — non-reproducible and a supply-chain surface, and (unlike the
persisted setup config) never pinned. Extend sync-plugin-manifests.mjs with an
mcp surface kind that stamps the gitnexus@<version> launch arg, pin all ten to
1.6.9 now, and keep them byte-identical so the drift guard stays green. The
release lifecycle + publish.yml --check now re-stamp them like the four manifest
surfaces; only READMEs stay on @latest as docs.

* test(eval): prove the proposer's built-in file tools are confined

The real-Claude canary only exercised Bash + MCP, so it proved process/MCP
containment but not that the proposer's built-in file tools stay inside their
mounts. Add a canary over the exact PROPOSER_ALLOWED_TOOLS surface and the same
read-only /evidence mount as run_proposer (allowlist extracted to a shared
constant so it can't drift): Read reaches /evidence, a Write into the read-only
evidence mount is denied, and a Write lands in the output tree.

* fix(eval): apply the candidate overlay after task setup for fair arms

The candidate overlay was applied before the task's untrusted setup ran, so
setup could observe candidate prose and the incumbent/candidate arms started
from different pre-overlay state. Reorder within the sandbox: capture the base
(pre-overlay) skill digest, run setup against the base skills, verify setup did
not tamper them, then apply the overlay and capture the post-overlay digest the
model must preserve. apply_candidate_overlay stages path-specific overlay files,
so setup's uncommitted changes stay out of the baseline and churn is unchanged.

Graph freshness for the review arm is handled by the status dirty-tree fix plus
the review skill's stale-triggered re-index, not by reordering the cached
per-task-sha graph materialization (which is mechanically blocked).

* test(eval): end-to-end containment proof of the autonomous proposer

Drives the real run_proposer through bubblewrap with a deterministic scripted
model (no paid API): it reads the read-only evidence bundle and writes a
candidate gitnexus-plan skill edit plus a rationale into the sandbox output
tree; run_proposer enforces the trust boundary and copies only the validated
overlay + proposal out. This exercises the autonomous-proposal stage of the
self-evolution loop end-to-end in the eval/containment CI job (the gate and
apply stages are covered by test_workflow_bench_evolution and
test_promotion_apply). Env-gated on GITNEXUS_REQUIRE_CLAUDE_CANARY, so it runs
only where the pinned Claude binary and user namespaces are available.

* fix(eval): let the proposer author its overlay via Bash

Running the end-to-end proposer canary in the containment CI job surfaced a real
bug: run_proposer starts the session with --bare, which hard-disables the
Write/Edit tools ("Write exists but is not enabled in this context"), yet
allowlisted Edit/Write and omitted Bash. The proposer therefore had no working
way to write its candidate overlay — the self-evolution loop could never produce
a candidate. The sandbox settings already pre-authorize Bash
(autoAllowBashIfSandboxed) and confine writes to workspace/tmp/home, so switch
PROPOSER_ALLOWED_TOOLS to Read/Grep/Glob/Bash and tell the proposer to author
files with Bash. The end-to-end test now drives the real run_proposer through
bubblewrap and asserts a validated overlay + proposal are produced (this also
replaces the earlier file-tool canary, whose Write/Edit premise was moot).

* test(eval): author the proposer overlay with newline-free Bash content

The nested shell-sandbox prefix mangles embedded newlines, so the multi-line
overlay content never landed. Use single-line content for the deterministic
proposer canary.

* test(eval): drop the unverifiable end-to-end proposer canary

The scripted proposer overlay never materialized in the containment job across
runs, and the model tool-result content is not visible in CI logs, so the test
cannot be finalized without an environment where the sandbox can actually run.
Keep the verified production fix (Bash-authoring in run_proposer); the proposer
sandbox/containment stays covered by the existing Bash+MCP and process-tree
canaries.

* test(cli): drop run-analyze.ts from the windowsHide spawn-family list

U7 moved run-analyze.ts's only child_process call (the git status --porcelain
dirty check) into storage/git.ts (already covered by this test, with
windowsHide). run-analyze.ts no longer imports a spawn-family function, so the
windowsHide-regression test's 'must have >=1 spawn call' invariant failed for
it. Remove it from SRC_FILES.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Zander Raycraft <zanderjraycraft@gmail.com>
Co-authored-by: Azizur Rahman <azizur100389@gmail.com>
2026-07-19 15:07:24 +01:00

826 lines
35 KiB
Python

"""Tests for evidence-bound, transactional promotion application."""
import json
import os
import stat
from pathlib import Path, PurePosixPath
import pytest
from workflow_bench import evolve, promotion_apply
from workflow_bench.evolution import (
CANDIDATE_SKILLS,
MAX_CANDIDATE_OVERLAY_BYTES,
candidate_overlay_digest,
candidate_overlay_payload,
)
from workflow_bench.promotion_apply import (
apply_promoted_overlay,
committed_destination_base_digests,
destination_base_digests,
freeze_overlay,
mirror_targets,
)
def _git(repo: Path, *arguments: str) -> str:
return (
__import__("subprocess")
.run(
["git", "-C", str(repo), *arguments],
check=True,
capture_output=True,
text=True,
)
.stdout.strip()
)
def test_evolve_reexports_public_promotion_helpers():
assert evolve.mirror_targets is mirror_targets
assert evolve.freeze_overlay is freeze_overlay
assert evolve.destination_base_digests is destination_base_digests
assert evolve.apply_promoted_overlay is apply_promoted_overlay
def test_mirror_targets_cover_canonical_and_shipped_copies():
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
assert targets == [
PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"),
PurePosixPath("gitnexus/skills/gitnexus-plan/SKILL.md"),
PurePosixPath("gitnexus-claude-plugin/skills/gitnexus-plan/SKILL.md"),
]
def test_apply_promoted_overlay_writes_all_mirrors(tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("evolved plan skill")
repo = tmp_path / "repo"
expected_targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in expected_targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
written = apply_promoted_overlay(overlay, repo_root=repo)
assert written == [
".claude/skills/gitnexus-plan/SKILL.md",
"gitnexus/skills/gitnexus-plan/SKILL.md",
"gitnexus-claude-plugin/skills/gitnexus-plan/SKILL.md",
]
contents = {(repo / path).read_text() for path in written}
assert contents == {"evolved plan skill"}
def test_apply_promoted_overlay_rejects_destination_drift(tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(f"old:{target}")
expected_bases = destination_base_digests(overlay, repo_root=repo)
drifted = repo / targets[1]
drifted.write_text("concurrent edit")
with pytest.raises(ValueError, match="drifted=.*gitnexus-plan/SKILL.md"):
apply_promoted_overlay(
overlay,
repo_root=repo,
expected_target_bases=expected_bases,
)
assert drifted.read_text() == "concurrent edit"
assert (repo / targets[0]).read_text() == f"old:{targets[0]}"
assert (repo / targets[2]).read_text() == f"old:{targets[2]}"
def test_apply_promoted_overlay_preserves_edit_racing_atomic_exchange(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_exchange = promotion_apply._exchange_at
raced = False
def edit_before_exchange(parent_descriptor, source, destination):
nonlocal raced
if not raced:
raced = True
(repo / targets[0]).write_text("concurrent edit")
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", edit_before_exchange)
with pytest.raises(RuntimeError, match="rolled back") as raised:
apply_promoted_overlay(overlay, repo_root=repo)
assert "atomic overlay exchange parity check failed" in str(raised.value.__cause__)
assert (repo / targets[0]).read_text() == "concurrent edit"
assert [(repo / target).read_text() for target in targets[1:]] == ["incumbent", "incumbent"]
assert list(repo.rglob(".wfevolve-*")) == []
def test_apply_promoted_overlay_rolls_back_raced_edit_when_exchange_then_raises(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_exchange = promotion_apply._exchange_at
exchanges = 0
def edit_exchange_then_raise(parent_descriptor, source, destination):
nonlocal exchanges
exchanges += 1
if exchanges == 1:
(repo / targets[0]).write_text("concurrent edit")
real_exchange(parent_descriptor, source, destination)
raise OSError("injected post-exchange failure")
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", edit_exchange_then_raise)
with pytest.raises(RuntimeError, match="rolled back"):
apply_promoted_overlay(overlay, repo_root=repo)
assert (repo / targets[0]).read_text() == "concurrent edit"
assert [(repo / target).read_text() for target in targets[1:]] == ["incumbent", "incumbent"]
assert list(repo.rglob(".wfevolve-*")) == []
def test_apply_promoted_overlay_preserves_second_edit_racing_rollback(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_exchange = promotion_apply._exchange_at
exchanges = 0
def edit_before_publication_and_rollback(parent_descriptor, source, destination):
nonlocal exchanges
exchanges += 1
if exchanges == 1:
(repo / targets[0]).write_text("first concurrent edit")
elif exchanges == 2:
(repo / targets[0]).write_text("second concurrent edit")
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", edit_before_publication_and_rollback)
with pytest.raises(RuntimeError, match="rollback was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["transaction_state"] == "rollback-incomplete"
raced_target = next(entry for entry in recovery["backups"] if entry["target"] == targets[0].as_posix())
assert Path(raced_target["candidate"]).read_text() == "second concurrent edit"
assert (repo / targets[0]).read_text() == "first concurrent edit"
def test_apply_promoted_overlay_preserves_mode_change_racing_exchange(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
destination.chmod(0o644)
real_exchange = promotion_apply._exchange_at
raced = False
def chmod_before_exchange(parent_descriptor, source, destination):
nonlocal raced
if not raced:
raced = True
(repo / targets[0]).chmod(0o600)
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", chmod_before_exchange)
with pytest.raises(RuntimeError, match="rolled back"):
apply_promoted_overlay(overlay, repo_root=repo)
assert (repo / targets[0]).read_text() == "incumbent"
assert stat.S_IMODE((repo / targets[0]).stat().st_mode) == 0o600
assert list(repo.rglob(".wfevolve-*")) == []
def test_apply_promoted_overlay_treats_candidate_hardlinked_at_both_names_as_incomplete(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_exchange = promotion_apply._exchange_at
linked = False
def hardlink_candidate_before_exchange(parent_descriptor, source, destination):
nonlocal linked
if not linked:
linked = True
os.unlink(destination, dir_fd=parent_descriptor)
os.link(
source,
destination,
src_dir_fd=parent_descriptor,
dst_dir_fd=parent_descriptor,
follow_symlinks=False,
)
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", hardlink_candidate_before_exchange)
with pytest.raises(RuntimeError, match="rollback was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
assert (repo / targets[0]).read_text() == "candidate"
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["transaction_state"] == "rollback-incomplete"
assert "linked at both" in recovery["rollback_failures"][0]
assert Path(recovery["backups"][0]["backup"]).read_text() == "incumbent"
def test_apply_promoted_overlay_rolls_back_every_completed_replace(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(f"old:{target}")
originals = {target: (repo / target).read_bytes() for target in targets}
real_exchange = promotion_apply._exchange_at
replacements = 0
def fail_second_exchange(parent_descriptor, source, destination):
nonlocal replacements
replacements += 1
if replacements == 2:
raise OSError("injected replacement failure")
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", fail_second_exchange)
with pytest.raises(RuntimeError, match="rolled back"):
apply_promoted_overlay(overlay, repo_root=repo)
assert {target: (repo / target).read_bytes() for target in targets} == originals
def test_apply_promoted_overlay_rolls_back_when_replace_lands_then_interrupts(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(f"old:{target}")
originals = {target: (repo / target).read_bytes() for target in targets}
expected_bases = destination_base_digests(overlay, repo_root=repo)
real_exchange = promotion_apply._exchange_at
apply_replacements = 0
def interrupt_after_second_landed_exchange(parent_descriptor, source, destination):
nonlocal apply_replacements
result = real_exchange(parent_descriptor, source, destination)
if destination == "SKILL.md":
apply_replacements += 1
if apply_replacements == 2:
raise KeyboardInterrupt("injected post-replace interruption")
return result
monkeypatch.setattr(promotion_apply, "_exchange_at", interrupt_after_second_landed_exchange)
with pytest.raises(KeyboardInterrupt, match="post-replace"):
apply_promoted_overlay(
overlay,
repo_root=repo,
expected_target_bases=expected_bases,
)
assert {target: (repo / target).read_bytes() for target in targets} == originals
def test_apply_promoted_overlay_prevalidates_all_targets_before_staging(tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
first = repo / mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))[0]
first.parent.mkdir(parents=True)
first.write_text("incumbent")
with pytest.raises(ValueError, match="destination (parent is unavailable|must already be a regular file)"):
apply_promoted_overlay(overlay, repo_root=repo)
assert first.read_text() == "incumbent"
assert list(repo.rglob(".wfevolve-*")) == []
def test_apply_promoted_overlay_rejects_internal_symlink_ancestor(tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in (targets[0], targets[2]):
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
redirected = repo / "redirected"
redirected.mkdir(parents=True)
(redirected / "SKILL.md").write_text("must-not-change")
symlink_parent = repo / targets[1].parent
symlink_parent.parent.mkdir(parents=True, exist_ok=True)
symlink_parent.symlink_to(redirected, target_is_directory=True)
with pytest.raises(ValueError, match="must not be a symlink"):
apply_promoted_overlay(overlay, repo_root=repo)
assert (redirected / "SKILL.md").read_text() == "must-not-change"
assert (repo / targets[0]).read_text() == "incumbent"
def test_apply_promoted_overlay_rejects_repository_swap_during_root_open(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
replacement_repo = tmp_path / "replacement-repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for root, content in ((repo, "incumbent"), (replacement_repo, "replacement")):
for target in targets:
destination = root / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(content)
detached_repo = tmp_path / "detached-repo"
real_open = promotion_apply.os.open
swapped = False
def swap_before_root_open(path, flags, mode=0o777, *, dir_fd=None):
nonlocal swapped
if not swapped and dir_fd is None and Path(path) == repo and flags & os.O_DIRECTORY:
swapped = True
repo.rename(detached_repo)
replacement_repo.rename(repo)
return real_open(path, flags, mode, dir_fd=dir_fd)
monkeypatch.setattr(promotion_apply.os, "open", swap_before_root_open)
with pytest.raises(ValueError, match="repository root changed while opening"):
apply_promoted_overlay(overlay, repo_root=repo)
assert [(detached_repo / target).read_text() for target in targets] == ["incumbent"] * 3
assert [(repo / target).read_text() for target in targets] == ["replacement"] * 3
assert list(tmp_path.rglob(".wfevolve-*")) == []
def test_apply_promoted_overlay_rejects_detached_parent_after_preparation(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
lexical_parent = (repo / targets[0]).parent
detached_parent = lexical_parent.with_name("gitnexus-plan-detached")
real_stage = promotion_apply._stage_replacement_at
replaced = False
def replace_parent_before_staging(parent_descriptor, content, mode):
nonlocal replaced
if not replaced:
replaced = True
lexical_parent.rename(detached_parent)
lexical_parent.mkdir()
(lexical_parent / "SKILL.md").write_text("incumbent")
return real_stage(parent_descriptor, content, mode)
monkeypatch.setattr(promotion_apply, "_stage_replacement_at", replace_parent_before_staging)
with pytest.raises(RuntimeError, match="rolled back") as raised:
apply_promoted_overlay(overlay, repo_root=repo)
assert "destination parent changed" in str(raised.value.__cause__)
assert (lexical_parent / "SKILL.md").read_text() == "incumbent"
assert (detached_parent / "SKILL.md").read_text() == "incumbent"
assert [(repo / target).read_text() for target in targets[1:]] == ["incumbent", "incumbent"]
assert list(repo.rglob(".wfevolve-*")) == []
def test_committed_destination_bases_ignore_and_reject_live_target_edits(tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-b", "main")
_git(repo, "config", "user.name", "Workflow Bench Test")
_git(repo, "config", "user.email", "workflow-bench@example.invalid")
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(f"committed:{target}")
_git(repo, "add", ".")
_git(repo, "commit", "-m", "incumbent")
expected = committed_destination_base_digests(overlay, repo_root=repo)
dirty = repo / targets[1]
dirty.write_text("user edit")
assert destination_base_digests(overlay, repo_root=repo) != expected
with pytest.raises(ValueError, match="drifted"):
apply_promoted_overlay(
overlay,
repo_root=repo,
expected_target_bases=expected,
)
assert dirty.read_text() == "user edit"
def test_mirror_roots_cover_every_candidate_skill_and_omit_none_that_ships_to_cursor():
# promotion_apply.mirror_targets writes canonical + MIRROR_SKILL_ROOTS, which
# today omits the Cursor tree. That is only safe because no candidate skill is
# cursor-shipped. If a future edit adds a cursor-shipped skill (e.g.
# gitnexus-review) to CANDIDATE_SKILLS, apply_promoted_overlay would rewrite
# the other trees and silently skip Cursor — the PR #2488 asymmetric-sync bug
# class. Pin the invariant to the filesystem, the source of truth the TS drift
# guard already enforces.
repo_root = Path(__file__).resolve().parents[2]
cursor_root = repo_root / "gitnexus-cursor-integration" / "skills"
for skill in sorted(CANDIDATE_SKILLS):
canonical = repo_root / ".claude" / "skills" / skill
assert canonical.is_dir(), f"candidate skill {skill} has no canonical .claude/skills dir"
for target in mirror_targets(PurePosixPath(".claude", "skills", skill, "SKILL.md")):
assert (repo_root / target).is_file(), f"candidate skill mirror missing on disk: {target}"
assert not (cursor_root / skill).exists(), (
f"candidate skill {skill} ships to Cursor, but MIRROR_SKILL_ROOTS does not cover "
"gitnexus-cursor-integration/skills — promotion would sync it asymmetrically"
)
def test_committed_destination_bases_reject_overlay_adding_uncommitted_target(tmp_path):
# An overlay that adds a file absent at HEAD has no committed base to bind
# against and raises ValueError — evolve.run / runner.main now catch that as
# NOT PROMOTED / a clean CLI error instead of an uncaught traceback.
overlay = tmp_path / "overlay"
new_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "NEW.md"
new_md.parent.mkdir(parents=True)
new_md.write_text("brand new candidate file")
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-b", "main")
_git(repo, "config", "user.name", "Workflow Bench Test")
_git(repo, "config", "user.email", "workflow-bench@example.invalid")
(repo / "README.md").write_text("seed")
_git(repo, "add", ".")
_git(repo, "commit", "-m", "seed")
with pytest.raises(ValueError, match="committed overlay destination is unavailable"):
committed_destination_base_digests(overlay, repo_root=repo)
def test_stage_replacement_removes_partial_file_when_fsync_fails(monkeypatch, tmp_path):
destination = tmp_path / "SKILL.md"
def fail_fsync(_descriptor):
raise OSError("injected fsync failure")
monkeypatch.setattr(promotion_apply.os, "fsync", fail_fsync)
with pytest.raises(OSError, match="injected fsync failure"):
promotion_apply._stage_replacement(destination, b"partial candidate", 0o644)
assert list(tmp_path.glob(".wfevolve-*")) == []
def test_apply_promoted_overlay_names_recovery_state_if_rollback_fails(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_exchange = promotion_apply._exchange_at
replacements = 0
def fail_apply_and_rollback(parent_descriptor, source, destination):
nonlocal replacements
replacements += 1
if replacements >= 2:
raise OSError("injected persistent replacement failure")
return real_exchange(parent_descriptor, source, destination)
monkeypatch.setattr(promotion_apply, "_exchange_at", fail_apply_and_rollback)
with pytest.raises(RuntimeError, match="rollback was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["rollback_failures"]
assert any(Path(entry["backup"]).exists() for entry in recovery["backups"])
def test_apply_promoted_overlay_writes_recovery_into_held_root_after_relocation(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
replacement_repo = tmp_path / "replacement-repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for root, content in ((repo, "incumbent"), (replacement_repo, "replacement")):
for target in targets:
destination = root / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(content)
detached_repo = tmp_path / "detached-repo"
real_exchange = promotion_apply._exchange_at
exchanges = 0
def relocate_after_exchange_then_fail(parent_descriptor, source, destination):
nonlocal exchanges
exchanges += 1
if exchanges == 1:
real_exchange(parent_descriptor, source, destination)
repo.rename(detached_repo)
replacement_repo.rename(repo)
raise OSError("injected post-exchange relocation")
raise OSError("injected rollback failure")
monkeypatch.setattr(promotion_apply, "_exchange_at", relocate_after_exchange_then_fail)
with pytest.raises(RuntimeError, match=r"recovery: .*detached-repo"):
apply_promoted_overlay(overlay, repo_root=repo)
assert list(repo.glob(".wfbench-overlay-recovery-*.json")) == []
recovery_files = list(detached_repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["rollback_failures"]
assert recovery["backups"]
assert all(str(detached_repo) in entry["backup"] for entry in recovery["backups"])
assert all(Path(entry["backup"]).exists() for entry in recovery["backups"])
assert [(repo / target).read_text() for target in targets] == ["replacement"] * 3
def test_apply_promoted_overlay_reports_published_state_and_closes_descriptors_on_cleanup_failure(
monkeypatch,
tmp_path,
):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
captured_descriptors: list[int] = []
real_prepare = promotion_apply._prepare_targets
def capture_descriptors(payload, repo_root):
root, root_descriptor, prepared = real_prepare(payload, repo_root)
captured_descriptors.extend([root_descriptor, *(item["parent_descriptor"] for item in prepared)])
return root, root_descriptor, prepared
def fail_cleanup(_parent_descriptor, _name):
raise OSError("injected cleanup failure")
monkeypatch.setattr(promotion_apply, "_prepare_targets", capture_descriptors)
monkeypatch.setattr(promotion_apply, "_unlink_temporary", fail_cleanup)
with pytest.raises(RuntimeError, match="transaction is published.*cleanup was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
assert [(repo / target).read_text() for target in targets] == ["candidate"] * 3
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["transaction_state"] == "published"
assert recovery["backups"]
for descriptor in captured_descriptors:
with pytest.raises(OSError):
os.fstat(descriptor)
def test_apply_promoted_overlay_tracks_candidate_when_backup_staging_and_cleanup_fail(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_stage = promotion_apply._stage_replacement_at
stages = 0
def fail_backup_stage(parent_descriptor, content, mode):
nonlocal stages
stages += 1
if stages == 2:
raise OSError("injected backup staging failure")
return real_stage(parent_descriptor, content, mode)
def fail_candidate_cleanup(_parent_descriptor, _name):
raise OSError("injected candidate cleanup failure")
monkeypatch.setattr(promotion_apply, "_stage_replacement_at", fail_backup_stage)
monkeypatch.setattr(promotion_apply, "_unlink_temporary", fail_candidate_cleanup)
with pytest.raises(RuntimeError, match="transaction is rolled-back.*cleanup was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["transaction_state"] == "rolled-back"
assert len(recovery["backups"]) == 1
assert recovery["backups"][0]["candidate_exists"]
assert recovery["backups"][0]["backup"] is None
assert Path(recovery["backups"][0]["candidate"]).exists()
assert [(repo / target).read_text() for target in targets] == ["incumbent"] * 3
def test_apply_promoted_overlay_tracks_candidate_before_identity_capture(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_identity = promotion_apply._entry_identity_at
identities = 0
def fail_candidate_identity(parent_descriptor, name):
nonlocal identities
identities += 1
if identities == 1:
raise OSError("injected candidate identity failure")
return real_identity(parent_descriptor, name)
monkeypatch.setattr(promotion_apply, "_entry_identity_at", fail_candidate_identity)
with pytest.raises(RuntimeError, match="rolled back"):
apply_promoted_overlay(overlay, repo_root=repo)
assert list(repo.rglob(".wfevolve-*")) == []
assert [(repo / target).read_text() for target in targets] == ["incumbent"] * 3
def test_apply_promoted_overlay_tracks_stage_name_when_parent_fsync_and_unlink_fail(monkeypatch, tmp_path):
overlay = tmp_path / "overlay"
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
skill_md.parent.mkdir(parents=True)
skill_md.write_text("candidate")
repo = tmp_path / "repo"
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-plan/SKILL.md"))
for target in targets:
destination = repo / target
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text("incumbent")
real_fsync = promotion_apply.os.fsync
real_unlink = promotion_apply.os.unlink
failed_parent_fsync = False
def fail_first_parent_fsync(descriptor):
nonlocal failed_parent_fsync
if not failed_parent_fsync and stat.S_ISDIR(os.fstat(descriptor).st_mode):
failed_parent_fsync = True
raise OSError("injected parent fsync failure")
return real_fsync(descriptor)
def fail_staging_unlink(path, *args, **kwargs):
if str(path).startswith(".wfevolve-"):
raise OSError("injected staging unlink failure")
return real_unlink(path, *args, **kwargs)
monkeypatch.setattr(promotion_apply.os, "fsync", fail_first_parent_fsync)
monkeypatch.setattr(promotion_apply.os, "unlink", fail_staging_unlink)
with pytest.raises(RuntimeError, match="transaction is rolled-back.*cleanup was incomplete; recovery:"):
apply_promoted_overlay(overlay, repo_root=repo)
recovery_files = list(repo.glob(".wfbench-overlay-recovery-*.json"))
assert len(recovery_files) == 1
recovery = json.loads(recovery_files[0].read_text())
assert recovery["transaction_state"] == "rolled-back"
assert len(recovery["backups"]) == 1
assert recovery["backups"][0]["candidate_exists"]
assert Path(recovery["backups"][0]["candidate"]).exists()
assert [(repo / target).read_text() for target in targets] == ["incumbent"] * 3
def test_freeze_overlay_detaches_authorized_bytes_from_mutable_input(tmp_path):
overlay = tmp_path / "overlay"
source = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
source.parent.mkdir(parents=True)
source.write_text("authorized")
frozen = tmp_path / "frozen"
digest = freeze_overlay(overlay, frozen)
source.write_text("mutated later")
assert candidate_overlay_digest(frozen) == digest
assert (frozen / source.relative_to(overlay)).read_text() == "authorized"
def test_freeze_overlay_matches_canonical_payload_digest_and_byte_boundary(tmp_path):
overlay = tmp_path / "overlay"
source = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
source.parent.mkdir(parents=True)
source.write_bytes(b"x" * MAX_CANDIDATE_OVERLAY_BYTES)
digest, payload = candidate_overlay_payload(overlay)
assert candidate_overlay_digest(overlay) == digest
frozen = tmp_path / "frozen"
assert freeze_overlay(overlay, frozen) == digest
assert candidate_overlay_payload(frozen) == (digest, payload)
source.write_bytes(b"x" * (MAX_CANDIDATE_OVERLAY_BYTES + 1))
with pytest.raises(ValueError, match="bounded evidence limit"):
candidate_overlay_digest(overlay)
rejected = tmp_path / "rejected"
with pytest.raises(ValueError, match="bounded evidence limit"):
freeze_overlay(overlay, rejected)
assert not rejected.exists()
def test_apply_promoted_overlay_rejects_out_of_boundary_files(tmp_path):
overlay = tmp_path / "overlay"
rogue = overlay / ".claude" / "skills" / "not-a-family-skill" / "SKILL.md"
rogue.parent.mkdir(parents=True)
rogue.write_text("smuggled")
with pytest.raises(ValueError, match="may only contain Markdown files"):
apply_promoted_overlay(overlay, repo_root=tmp_path / "repo")