GitNexus/gitnexus/test/unit/skill-evolution-workflow.test.ts
Gergő Magyar 4756e8a820
fix(release): gate publication on accuracy and paired evaluation (#3503)
* test(release): add fixed-answer native tool accuracy

* fix(release): gate stable artifacts on pinned paired evaluations

Apply review findings #1-4: gate Docker publication, grade in-flight retries, preserve repository context, and execute negative controls in CI.

* fix(eval): verify native context and resolve harness security findings

Inspect Claude's first startup reminder for ordinary repository guidance.
Bound launcher-option matching and escape Markdown backslashes and pipes.
Exclude only the parsed synthetic accuracy corpus from CodeQL; keep the
benchmark harness and evaluator under security analysis.

Validated focused Python and TypeScript regressions plus core typecheck.

* fix(eval): wake the EC2 runner and prove paired evaluator execution

Bootstrap the existing dedicated instance from a protected hosted job, bound
native runner pickup, and stop only instances started by this run. Keep paid
sessions inside the fixed stop window and add a runner-only dispatch.

Require a contained six-cell prepare/session/oracle/report canary in CI,
including a failed repair that remains in the measured denominator.

* fix(ci): reuse existing configuration and disambiguate accuracy cases

Remove the new AWS credential/settings path and release opt-in variable. Bound native runner pickup with the existing GitHub token and retain private EventBridge startup until its existing mechanism can be reused. Reuse current workflow environment names and evolution schedule controls.

Name malformed-observation cases explicitly so the execution audit distinguishes null from missing and actually tests invalid arrays.

* fix(review): isolate release builds and preserve benchmark evidence

* fix(release): gate RC publication on paired quality and manage EC2 lifecycle

* fix(ci): clarify release shell redirects and build script

* fix(eval): remove temporary patch sinks after capture

* fix(bench): retire tool-accuracy allowances repaired on main

Main at 50aa4be3b repaired three #3487 route checks (#3505) and all four
#3499 Python scope checks (#3502, #3504). The ratchet correctly failed the
merged head with stale allowances, so remove them and keep the fixed answers
as release protection: 18/31 pass, 13 recorded gaps remain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release): gate RCs on accuracy only and keep paid runs for stable

Follow the cost split proposed in discussion #3493: fixed-answer tool
accuracy on every RC, paired agent runs before stable releases and on a
weekly schedule against the latest RC.

RC publication no longer prepares a bundle, calls release-evaluation.yml
and waits for a paid run of up to 21 hours on the shared EC2 lock. Every
release attaches this run's accuracy report; stable npm publication still
requires exact-commit paired evidence. release-evaluation.yml drops its
workflow_call path and the now-unused candidate bundle module.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release): address PR #3503 review feedback

- Page through recent release-evaluation runs instead of reading only
  the newest 50, so valid stable evidence cannot be pushed off page one.
- Bound each runner-pickup jobs request by the remaining deadline.
- Replace non-finite floats at any depth in failure receipts so the
  incomplete report still serializes with allow_nan=False.
- Clarify that the summary.cap case has 501 names over 1001 sites.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(review): harden candidate build and close review gaps

- Mount the candidate checkout's .git read-only inside the build sandbox,
  so lifecycle scripts cannot plant git config (e.g. core.fsmonitor) that
  host git later runs in that checkout during actions/checkout cleanup.
  The real-Bubblewrap canary now also attempts that write.
- Type tool-accuracy observations instead of using explicit any.
- Cover the fail-closed branches reviewers found untested: unsafe or
  missing evidence artifacts, mismatched task pins, malformed repo/SHA,
  an instance that stops again during bootstrap, and unbalanced GitNexus
  guidance markers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(release): describe RC accuracy gate and stable evidence step

- CONTRIBUTING now states that RCs are gated by CI (including the
  fixed-answer tool-accuracy check) without a paid run, and that stable
  publication needs a passing Release evaluation for the exact commit
  within seven days.
- The oracle-control fixture passes tarfile's data filter only where it
  exists (3.11.4+), matching the declared Python 3.11 floor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(bench): pin tool-accuracy fixture line anchors in one place

run.ts bucketed rename edits and explain findings by line literals that
expectations.ts duplicated. Export FIXTURE_ANCHORS (file, line, pinned text)
from expectations.ts, build the fixed answers and run.ts classifiers from it,
and fail loudly (runner and unit test) when a fixture line no longer holds
the pinned text. Expected answers and observations are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(bench): fail closed when ripgrep is missing from the tool-accuracy run

rename's text-search pass shells out to rg and only logs a degraded warning
when it is absent, so a runner without ripgrep scored a degraded rename
(homonym.ts:1 missing) and still passed. Check rg --version before indexing,
record it in accuracy.json source.externalTools, and install ripgrep in the
ci-tests tool-accuracy job and the docker release-gate (same if: condition).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(eval): split baseline guidance, share bwrap preamble, drop dead bare option, harden patch cleanup

- Move the baseline (no-GitNexus) guidance scrubbing out of proposer_sandbox.py
  into baseline_guidance.py, with its unit tests.
- Expose real_directory, runtime_mount_args and bwrap_base_args publicly so
  release_build no longer imports private helpers or re-assembles the shared
  bubblewrap preamble; the candidate build command is byte-identical.
- Remove the unused run_claude `bare` parameter and its --bare/--tools branches.
- capture_patch: a failed temporary-directory removal no longer discards the
  read patch or replaces an earlier error; it is attached as a note, or warned
  on stderr when there was no earlier error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ec2-runner): keep stop observing through transient describe failures

- stop() treats a failed or timed-out describe-instances call as unknown
  state and keeps observing until the stop deadline; a persistent failure
  is still raised at the deadline with the last command error. Unusable
  states and identity mismatches still fail at once (new CommandError
  subclass separates command failures from semantic ones).
- Readiness CLI has one paid mode: --job-name NAME. --evolve is removed;
  skill evolution passes its job name, and tests tie the watchdog command
  to the paid job's actual name in both EC2 workflows.
- Jobs-API requests never get a timeout larger than the remaining pickup
  budget (previously rounded up to 1s past the deadline).
- Document that the shared concurrency group keeps only one pending run,
  so a queued release evaluation can be replaced and must be confirmed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(eval): record verified EC2 lifecycle configuration

The README still described the AWS settings as assumed and unverified.
The OIDC role, environment settings and runner_only lifecycle were
configured and verified live on 2026-10-09.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release-eval): grade both runtimes against one pinned task toolchain

Candidate and stable each staged task dependencies (node_modules,
gitnexus-shared/dist) from their own runtime checkout, so the same v1.6.12
task source was graded with different TypeScript, Vitest and LadybugDB
versions and a per-task solve delta also measured dependency drift.

- Check out and build the pinned task commit once (tasks-base) and point
  every task's repo at it for both runtimes; the runtime under test reaches
  sessions only through --gitnexus-root. prepare requires --task-repo at the
  single pinned task commit (new task-sha subcommand resolves it).
- Retain sandbox_dependency_content_digest per measured row and make
  release_gate fail when any task's cells, across candidate and stable, do not
  share one recorded dependency digest.
- Move the task-pin comparison into validate_report(task_pins=...) and use
  it from release_report check and release_gate.
- Cover the gate CLI with a missing and a corrupt candidate report.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(ci): warn that a pending release evaluation can be replaced

GitHub keeps one pending run per concurrency group, so a newer queued
EC2 run cancels a pending release evaluation. Operators must confirm the
startup job ran before relying on the evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(eval): prove hidden oracles against the grading toolchain

Release evaluation now grades every task with dependencies built at the
pinned task commit (TypeScript 5.9.3, Vitest 4.1.11), but the oracle
controls still ran against the current checkout's toolchain. The Ubuntu
containment job now builds the pinned task toolchain and points the
controls at it. All seven controls pass locally against that build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): pin the duplicated EC2 lifecycle jobs together

Both EC2 workflows carry their own copy of the live-verified start,
check, readiness and stop jobs. A reusable workflow would change that
verified job structure, so instead a test requires the copies to stay
identical apart from the job display name and the paid-job dependency.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release-eval): accept stable evidence only from main's current evaluator

Stable publish validated downloaded agent evidence against the release
commit's eval/ tree and never tied a report to the run that produced it.
Evidence from an older evaluator stayed acceptable for seven days after
graders changed, while task changes on main made every fresh evaluation
mismatch the tag. The first invalid artifact also hid older valid runs.

The composite action now extracts eval/ from main's head and runs the
validator there with locked base dependencies only (no dev extras or
project build). A report counts only when its harness_sha is the head_sha
of its trusted main run and the compare API shows no change under eval/,
the release-evaluation workflow or the pinned agent CLI between that
harness and the evaluator; a missing, diverged or 300-file comparison
fails closed. Runs are tried newest first, rejected runs are skipped
with a bounded reason list, every gh call has a timeout, and task pins
use validate_report's task_pins check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(eval): break the sandbox import cycle and resolve code-scanning alerts

- baseline_guidance is now pure text handling with its own GuidanceError;
  the baseline mount builder lives in proposer_sandbox again and converts
  that error to SandboxError, so imports run one way only.
- task_assets uses the public real_directory name; the unused private
  alias is gone.
- The oracle-control fixture always extracts with tarfile's data filter
  and skips on Pythons that lack it, instead of extracting unfiltered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release-eval): run candidate lifecycle scripts offline; reject smuggled report fields

- build_candidate now downloads locked dependencies with --ignore-scripts,
  then runs every lifecycle script (dependency installs, prepare, build)
  in a fresh network namespace, so candidate code cannot reach host-local
  services. Verified locally that the current head and the pinned task
  commit build this way; the real-Bubblewrap canary now asserts only
  loopback is visible to a lifecycle script.
- validate_report requires the evidence document to equal its
  field-whitelisted rebuild, so per-run or top-level fields outside the
  published schema are rejected before the publisher copies the file.
- The sandbox test suite also covers a no-GitNexus sandbox refusing
  unbalanced guidance markers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(eval): describe the offline candidate lifecycle

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(release-eval): pin downloads to the npm registry; require dependency binding

Address review feedback on PR #3503:
- The candidate's online install phase now refuses lockfile entries that
  resolve outside https://registry.npmjs.org/ (other than in-checkout
  workspace links) and refuses shipped .npmrc files, so a candidate
  cannot steer npm at loopback, private or metadata endpoints.
- Every measured row must carry a 64-hex task-dependency digest, so a
  single report cannot be accepted with its grading toolchain unbound.
- Oversized JSON integers mark a measurement untrustworthy instead of
  crashing report generation.
- Baseline guidance stripping recognises headings with up to three
  leading spaces (CommonMark).
- Remove a stale comment about the old private helper name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(eval): parse CommonMark ATX headings; compare evidence by exact JSON

- Baseline guidance stripping now parses headings as CommonMark ATX
  headings (0-3 leading spaces, optional closing # run, tabs), so a
  GitNexus section heading like '## GitNexus rules ##' is removed too.
- Evidence validation compares the serialized rebuild with the report,
  so a retyped value (true for 1, 2.0 for 2) no longer passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(eval): drop continuations of every CommonMark list item in baseline guidance

A removed GitNexus instruction left its indented continuation behind when
the item used '+' or an ordered marker (1. / 1)). The scrubber now
recognises every CommonMark list-item marker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(eval): isolate candidate MCP from agent task writes (#3503)

Run MCP in a nested Bubblewrap boundary with read-only task, graph,
registry and runtime mounts, private state and isolated processes/network.
Keep task writes on the agent's built-in tools and remove mutating MCP grants.

Add startup/tool mutation canaries and preserve explicit unsafe diagnostics.

Validation: focused pytest 138 passed, 16 skipped. Full locked evaluator
996 passed, 29 skipped; real containment canaries require Ubuntu CI.
Note: eight pre-existing comparator-reuse failures reproduce on the
unchanged head due to missing os.supports_dir_fd support. One process-control
timeout failed in the full run and passed alone and on the unchanged head.

* fix(eval): close baseline mount overlap and test symlink evidence (#3503)

Reject supplied no-MCP mounts that cover forbidden GitNexus paths, including ancestor mounts and lexical aliases. Keep ordinary dependency mounts available.

Exercise evidence symlink rejection with a real valid-report target and retain dangling-link coverage.

Validation: 1026 evaluator tests passed, 29 skipped; targeted mount tests 36 passed and evidence tests 34 passed. Removing the symlink guard in memory makes the repaired regression fail. Ruff and diff checks passed.

Note: eight pre-existing comparator-reuse failures remain in this environment because os.supports_dir_fd lacks the required os.lstat support.

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-10 08:26:08 +01:00

732 lines
31 KiB
TypeScript

import { execFileSync } from 'node:child_process';
import {
chmodSync,
existsSync,
mkdirSync,
mkdtempSync,
readFileSync,
realpathSync,
rmSync,
statSync,
writeFileSync,
} from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { load } from 'js-yaml';
import { describe, expect, it } from 'vitest';
// Contract guard for the online skill-evolution workflow. Both P1 blockers
// fixed here (a gate-passing run never applied its overlay; the benchmark
// could not resolve its task repo on a hosted runner) reached production
// because nothing exercised this workflow's path. Assert the structural
// contract so a regression fails loudly in CI instead of on the first real run.
const REPO_ROOT = path.resolve(__dirname, '../../..');
const WORKFLOW_PATH = path.resolve(REPO_ROOT, '.github/workflows/gitnexus-skill-evolution.yml');
const workflow = readFileSync(WORKFLOW_PATH, 'utf8');
const workflowDocument = load(workflow) as {
jobs?: Record<
string,
{
name?: string;
environment?: unknown;
needs?: string | string[];
permissions?: Record<string, string>;
'runs-on'?: string | string[];
env?: Record<string, string>;
if?: unknown;
'timeout-minutes'?: unknown;
steps?: Array<{
name?: string;
if?: unknown;
run?: unknown;
uses?: string;
'timeout-minutes'?: unknown;
with?: Record<string, unknown>;
env?: Record<string, string>;
'working-directory'?: string;
}>;
}
>;
};
const evolveJob = workflowDocument.jobs?.evolve;
type WorkflowStep = NonNullable<NonNullable<typeof evolveJob>['steps']>[number];
function findStep(stepName: string): WorkflowStep | undefined {
return evolveJob?.steps?.find(({ name }) => name === stepName);
}
function stepRun(stepName: string): string {
const step = findStep(stepName);
return typeof step?.run === 'string' ? step.run : '';
}
describe('dedicated runner readiness', () => {
it('bounds offline runner pickup from trusted hosted main before paid work', () => {
const check = workflowDocument.jobs?.['check-runner'];
expect(check?.['runs-on']).toBe('ubuntu-latest');
expect(check?.environment).toBe('gitnexus-evolution');
expect(check?.permissions).toEqual({ contents: 'read', actions: 'write' });
expect(check?.['timeout-minutes']).toBe(10);
expect(String(check?.if)).toContain("github.ref == 'refs/heads/main'");
expect(String(check?.if)).toContain("github.repository == 'abhigyanpatwari/GitNexus'");
expect(evolveJob?.needs).toEqual(['start-runner', 'check-runner', 'runner-ready']);
const pickup = check?.steps?.find(
(step) => step.name === "Verify the current run's native pickup probe",
);
expect(pickup?.run).toBe('python3 .github/scripts/evolution-runner-ready.py');
expect(pickup?.env).toEqual({ GH_TOKEN: '${{ github.token }}' });
const cancel = check?.steps?.at(-1);
expect(cancel?.if).toBe('failure()');
expect(cancel?.run).toBe(
'gh api --method POST "repos/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}/cancel"',
);
const ready = workflowDocument.jobs?.['runner-ready'];
expect(ready?.needs).toEqual(['start-runner']);
expect(ready?.if).toBe(check?.if);
expect(ready?.permissions).toEqual({});
expect(ready?.steps?.some((step) => String(step.run).includes('git --version'))).toBe(true);
expect(String(evolveJob?.if)).toContain('inputs.runner_only != true');
});
it('reuses the existing credentials and schedule variables with only the historical AWS configuration', () => {
const release = readFileSync(
path.join(REPO_ROOT, '.github/workflows/release-evaluation.yml'),
'utf8',
);
const existingSecrets = new Set([
'GITNEXUS_BENCH_ANTHROPIC_API_KEY',
'GITNEXUS_BENCH_AUTH_TOKEN',
'GITNEXUS_BENCH_OPENAI_API_KEY',
'GITNEXUS_EVOLUTION_AWS_ROLE_ARN',
'GITNEXUS_EVOLUTION_EC2_INSTANCE_ID',
'RELEASE_APP_ID',
'RELEASE_APP_PRIVATE_KEY',
]);
const existingVariables = new Set([
'GITNEXUS_EVOLUTION_ENABLED',
'GITNEXUS_EVOLUTION_WORKERS',
'GITNEXUS_EVOLUTION_AWS_REGION',
'GITNEXUS_EVOLUTION_STOP_SCHEDULE_UTC',
]);
for (const text of [workflow, release]) {
for (const match of text.matchAll(/secrets\.([A-Z_][A-Z_0-9]*)/g)) {
expect(existingSecrets.has(match[1]), match[1]).toBe(true);
}
for (const match of text.matchAll(/vars\.([A-Z_][A-Z_0-9]*)/g)) {
expect(existingVariables.has(match[1]), match[1]).toBe(true);
}
expect(text).not.toContain('configure-aws-credentials');
}
for (const job of ['check-runner', 'runner-ready', 'watch-evolve-pickup']) {
expect(workflowDocument.jobs?.[job]?.env).toBeUndefined();
for (const step of workflowDocument.jobs?.[job]?.steps ?? []) {
expect(Object.keys(step.env ?? {}).every((key) => key === 'GH_TOKEN')).toBe(true);
}
}
const releaseDocument = load(release) as typeof workflowDocument;
expect(Object.keys(releaseDocument.jobs?.evaluate?.env ?? {}).sort()).toEqual([
'EFFORT',
'INPUT_REF',
'MODEL',
]);
expect(String(releaseDocument.jobs?.evaluate?.if)).toContain(
"vars.GITNEXUS_EVOLUTION_ENABLED == 'true'",
);
expect(String(releaseDocument.jobs?.evaluate?.if)).toContain(
"vars.GITNEXUS_EVOLUTION_WORKERS == '3'",
);
});
it('watches actual paid pickup concurrently after the readiness gate', () => {
const watch = workflowDocument.jobs?.['watch-evolve-pickup'];
expect(watch?.needs).toEqual(evolveJob?.needs);
expect(watch?.['runs-on']).toBe('ubuntu-latest');
expect(watch?.environment).toBeUndefined();
expect(watch?.['timeout-minutes']).toBe(10);
expect(watch?.permissions).toEqual({ contents: 'read', actions: 'write' });
expect(watch?.if).toBe('inputs.runner_only != true');
expect(evolveJob?.needs).not.toContain('watch-evolve-pickup');
// The watchdog matches by job name, so a rename must update both together.
expect(evolveJob?.name).toBe('Propose, benchmark, and gate skill candidates');
expect(
watch?.steps?.some(
(step) =>
step.run ===
`python3 .github/scripts/evolution-runner-ready.py --job-name "${evolveJob?.name}"`,
),
).toBe(true);
expect(watch?.steps?.at(-1)?.if).toBe('failure()');
expect(watch?.steps?.at(-1)?.run).toBe(
'gh api --method POST "repos/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}/cancel"',
);
});
});
// The seed step's usability check is the proposer's OWN preflight
// (select_evidence + proposer_evidence_entries), invoked through uv. Stubbing
// uv would make these tests assert nothing about it: a stub accepts whatever
// fixture it is handed, so a fixture with a wrong digest, a wrong byte count,
// or world-readable transcripts would "pass" a check that rejects it in
// production — exactly backwards for a test whose subject is that rejection.
// So run the real thing, and skip rather than pretend when the eval project's
// environment is not provisioned (the node-only CI test jobs do not set up
// uv; `eval-tests` and this workflow's own runner do). UV_OFFLINE keeps the
// probe and the step itself from ever reaching the network mid-test.
const REAL_PREFLIGHT_AVAILABLE =
process.platform !== 'win32' &&
(() => {
try {
execFileSync(
'uv',
[
'run',
'--project',
'eval',
'--locked',
'--extra',
'dev',
'--offline',
'python',
'-c',
'import workflow_bench.evolve',
],
{ cwd: REPO_ROOT, stdio: 'ignore' },
);
return true;
} catch {
return false;
}
})();
// Provisioning uv is not free, and neither is the first `uv run` in a cold
// project, so give the two tests that shell out to it real headroom.
const PREFLIGHT_TEST_TIMEOUT_MS = 120_000;
// sha256 of the 3-byte transcript body the fixture writes. evolve.py re-hashes
// the file on disk and compares it against the results row, so this pair has
// to be genuinely consistent — and the wrong-but-well-formed digest below has
// to be 64 hex characters, or it would be rejected as malformed metadata
// before anything is ever hashed.
const TRANSCRIPT_DIGEST = 'ca3d163bab055381827226140568f3bef7eaac187cebd76878e0b63e9e442356';
const WRONG_TRANSCRIPT_DIGEST = '0'.repeat(64);
/** Bash that materializes one downloaded evidence artifact under `destination`. */
function artifactFixture({ generation, digest }: { generation: number; digest: string }): string {
const bench = `\${destination}/artifact/gen-${generation}/bench`;
const row = JSON.stringify({
task: 'demo',
arm: 'workflow',
run: 0,
resolved: false,
// A measured outcome, not a harness death: select_evidence keeps this and
// drops session-error/infra-error rows.
error_kind: 'oracle-failed',
transcript_artifacts: [
{
path: 'transcripts/session.jsonl',
sha256: digest,
bytes: 3,
source: 'parent-captured-stream-json',
},
],
});
return ` mkdir -p "${bench}/transcripts"
printf '%s\\n' '${row}' > "${bench}/results.jsonl"
printf '{}\\n' > "${bench}/transcripts/session.jsonl"
# upload-artifact normalizes to 0755/0644 on the way out; the step's
# chmod -R go-rwx is what has to restore the owner-only modes the real
# transcript reader requires, so hand it the un-restored modes.
chmod 0755 "${bench}/transcripts"
chmod 0644 "${bench}/transcripts/session.jsonl"`;
}
function runSeedStep(ghImplementation: string): {
output: string;
trace: string;
transcriptDirectoryMode?: number;
transcriptMode?: number;
} {
// realpath: _real_results_root() in evolve.py rejects a results directory
// whose path traverses a symlink, and macOS hands out $TMPDIR under one.
const root = realpathSync(mkdtempSync(path.join(os.tmpdir(), 'gitnexus-evolution-seed-')));
try {
const bin = path.join(root, 'bin');
const runnerTemp = path.join(root, 'runner-temp');
const githubOutput = path.join(root, 'github-output');
const trace = path.join(root, 'gh-trace');
mkdirSync(bin);
mkdirSync(runnerTemp);
writeFileSync(githubOutput, '');
const gh = path.join(bin, 'gh');
writeFileSync(gh, `#!/usr/bin/env bash\nset -euo pipefail\n${ghImplementation}\n`);
chmodSync(gh, 0o700);
// Only `gh` is stubbed — it is the step's input (which runs exist, what
// their artifacts contain). `uv` is deliberately NOT on the stub PATH, so
// the usability check below resolves the real uv and runs the real
// preflight against these fixtures. cwd is the repo root because that is
// where the workflow runs the step from, and `--project eval` is relative
// to it.
execFileSync(
'/bin/bash',
['-c', stepRun("Seed the proposer with the previous run's evidence")],
{
cwd: REPO_ROOT,
env: {
...process.env,
PATH: `${bin}:${process.env.PATH ?? ''}`,
GITHUB_OUTPUT: githubOutput,
GITHUB_REPOSITORY: 'abhigyanpatwari/GitNexus',
GITHUB_RUN_ID: '999',
RUNNER_TEMP: runnerTemp,
TRACE: trace,
UV_OFFLINE: '1',
},
stdio: 'pipe',
},
);
const output = readFileSync(githubOutput, 'utf8');
const seed = output.match(/^seed=(.+)$/m)?.[1];
const transcriptDirectory = seed ? path.join(seed, 'transcripts') : undefined;
const transcript = transcriptDirectory
? path.join(transcriptDirectory, 'session.jsonl')
: undefined;
return {
output,
trace: readFileSync(trace, 'utf8'),
transcriptDirectoryMode:
transcriptDirectory && existsSync(transcriptDirectory)
? statSync(transcriptDirectory).mode & 0o777
: undefined,
transcriptMode:
transcript && existsSync(transcript) ? statSync(transcript).mode & 0o777 : undefined,
};
} finally {
rmSync(root, { recursive: true, force: true });
}
}
describe('gitnexus skill-evolution workflow contract', () => {
it.each(['main', 'feature', 'detached'])(
'fetches review history while checked out on %s',
(branch) => {
const root = mkdtempSync(path.join(os.tmpdir(), 'evolution-baseline-'));
const remote = path.join(root, 'upstream.git');
const checkout = path.join(root, 'checkout');
const git = (cwd: string, ...args: string[]) =>
execFileSync(
'git',
['-C', cwd, '-c', 'user.name=Fixture', '-c', 'user.email=fixture@example.test', ...args],
{ encoding: 'utf8' },
).trim();
try {
mkdirSync(remote);
git(remote, 'init', '-q', '-b', 'main');
writeFileSync(path.join(remote, 'source'), 'before');
git(remote, 'add', 'source');
git(remote, 'commit', '-qm', 'base');
execFileSync('git', ['clone', '-q', remote, checkout]);
if (branch === 'feature') git(checkout, 'checkout', '-qb', 'feature');
if (branch === 'detached') git(checkout, 'checkout', '-q', '--detach');
const before = git(checkout, 'rev-parse', 'HEAD');
writeFileSync(path.join(remote, 'source'), 'after');
git(remote, 'commit', '-qam', 'advance');
const script = stepRun('Point the benchmark task repo at the checkout');
execFileSync('bash', ['-euc', script.slice(script.indexOf('git -C'))], {
env: {
...process.env,
GITHUB_WORKSPACE: checkout,
GITHUB_SERVER_URL: root,
GITHUB_REPOSITORY: 'upstream',
},
});
expect(git(checkout, 'rev-parse', 'refs/remotes/origin/main^{commit}')).toBe(
git(remote, 'rev-parse', 'HEAD'),
);
expect(git(checkout, 'rev-parse', 'HEAD')).toBe(before);
} finally {
rmSync(root, { recursive: true, force: true });
}
},
);
it('requires the real review canary before paid work and retains failed-sweep evidence', () => {
const steps = evolveJob?.steps ?? [];
const preflight = steps.findIndex(
(step) => step.name === 'Verify contained review execution before paid sessions',
);
const paid = steps.findIndex((step) => step.name === 'Run the propose → benchmark → gate loop');
expect(preflight).toBeGreaterThanOrEqual(0);
expect(preflight).toBeLessThan(paid);
expect(steps[preflight].if).toBeUndefined();
expect(steps[preflight].env?.GITNEXUS_REQUIRE_CLAUDE_CANARY).toBe('1');
expect(steps[preflight].env?.GITNEXUS_REQUIRE_BWRAP_CANARY).toBe('1');
expect(findStep('Upload benchmark evidence')?.if).toBe('always()');
expect(findStep('Detect and bound the applied promotion')?.if).toBeUndefined();
expect(stepRun('Open the promotion PR')).toContain(
'gitnexus-cursor-integration/skills/gitnexus-review',
);
});
it('applies gate-passing overlays so the promotion-PR path is reachable', () => {
const loop = stepRun('Run the propose → benchmark → gate loop');
expect(loop).toContain('./workflow_bench/run-evolution.sh --apply');
expect(loop).not.toContain('python -m workflow_bench.evolve');
});
it('passes the cell concurrency through to the benchmark', () => {
// Dispatch defaults to 3. Scheduled runs still fall back to serial unless
// GITNEXUS_EVOLUTION_WORKERS is set — a cell starved of CPU that hits the
// session ceiling is an excluded run the gate refuses.
expect(evolveJob?.env?.WORKERS).toBe(
"${{ inputs.workers || vars.GITNEXUS_EVOLUTION_WORKERS || '1' }}",
);
expect(workflow).toMatch(/workers:\n(?:[^\n]*\n){0,4} default: '3'/);
});
it('seeds from the newest usable completed main run, including failed runs', () => {
const seed = stepRun("Seed the proposer with the previous run's evidence");
// Failed sweeps deliberately upload partial evidence. Looking only at
// successful runs makes that evidence unreachable and leaves the weekly
// proposer memoryless once the last successful artifact expires.
expect(seed).toContain('--status completed');
expect(seed).not.toContain('--status success');
expect(seed).toContain('--limit 10');
expect(seed).toContain('for previous in ${previous_runs}');
expect(seed).toContain('continue');
expect(seed).toContain('gen-*/bench/results.jsonl');
expect(seed).toContain('chmod -R go-rwx');
expect(seed).toContain('select_evidence(load_jsonl');
// The usability check must stay the proposer's own preflight. Narrowing it
// to "the file has rows" would re-admit artifacts whose transcripts the
// proposer then refuses to read, costing the generation its evidence.
expect(seed).toContain('stage_proposer_evidence_bundle');
expect(seed).toContain('break');
});
it('lets a dispatch start from a blank slate while the schedule always seeds', () => {
// Evidence from a run whose harness leaked the hidden oracles cannot be
// trusted, and the staged prior proposal is what carries that taint into
// every later generation. Without an opt-out the only remedy is waiting
// for the tainted artifact to expire.
const triggers = (workflowDocument as { on?: Record<string, unknown> }).on;
const inputs = (triggers?.workflow_dispatch as { inputs?: Record<string, unknown> } | undefined)
?.inputs;
const input = inputs?.seed_from_previous as { type?: string; default?: unknown } | undefined;
expect(input?.type).toBe('boolean');
expect(input?.default).toBe(true);
const condition = findStep("Seed the proposer with the previous run's evidence")?.if;
expect(condition).toContain("github.event_name != 'workflow_dispatch'");
expect(condition).toContain('inputs.seed_from_previous');
});
it('bounds the best-effort seed walk well inside the job budget', () => {
// Every iteration blocks on a network download this job does not control,
// and the job-level timeout CANCELS rather than fails — which skips the
// `if: always()` upload and loses the sweep's evidence. So the walk needs
// its own budget: long enough to never trip on a healthy run, short
// enough that a wedged download is a fast, obvious failure.
const seedBudget = findStep("Seed the proposer with the previous run's evidence")?.[
'timeout-minutes'
];
expect(typeof seedBudget).toBe('number');
expect(seedBudget as number).toBeGreaterThanOrEqual(10);
expect(seedBudget as number).toBeLessThanOrEqual(30);
expect(seedBudget as number).toBeLessThan(evolveJob?.['timeout-minutes'] as number);
});
it.skipIf(!REAL_PREFLIGHT_AVAILABLE)(
'falls back past an empty newer artifact to an older usable run',
() => {
const result = runSeedStep(`
if [[ "$1 $2" == 'run list' ]]; then
printf '300\\n200\\n'
exit 0
fi
if [[ "$1 $2" == 'run download' ]]; then
run_id="$3"
shift 3
destination=''
while (( $# )); do
if [[ "$1" == '--dir' ]]; then destination="$2"; shift 2; else shift; fi
done
printf '%s\\n' "\${run_id}" >> "\${TRACE}"
if [[ "\${run_id}" == '300' ]]; then
mkdir -p "\${destination}/artifact/gen-3/bench"
printf '%s\\n' '{"error_kind":"session-error","resolved":false}' > "\${destination}/artifact/gen-3/bench/results.jsonl"
elif [[ "\${run_id}" == '200' ]]; then
${artifactFixture({ generation: 2, digest: TRANSCRIPT_DIGEST })}
fi
exit 0
fi
exit 1`);
// select_evidence drops session-error rows as unattributable, leaving
// gen-3 with nothing to propose from.
expect(result.trace).toBe('300\n200\n');
expect(result.output).toMatch(/seed=.*\/200\/artifact\/gen-2\/bench\n/);
expect(result.transcriptDirectoryMode).toBe(0o700);
expect(result.transcriptMode).toBe(0o600);
},
PREFLIGHT_TEST_TIMEOUT_MS,
);
it.skipIf(!REAL_PREFLIGHT_AVAILABLE)(
'falls back past a newer artifact whose transcript digest does not match',
() => {
// The sharp edge of running the real preflight: this artifact is
// non-empty and structurally well-formed, so every cheap check passes
// it. Only hashing the transcript and comparing against the row the
// proposer would trust rejects it — which is the whole reason the step
// shells out to the proposer's own code instead of grepping the JSONL.
const result = runSeedStep(`
if [[ "$1 $2" == 'run list' ]]; then
printf '400\\n200\\n'
exit 0
fi
if [[ "$1 $2" == 'run download' ]]; then
run_id="$3"
shift 3
destination=''
while (( $# )); do
if [[ "$1" == '--dir' ]]; then destination="$2"; shift 2; else shift; fi
done
printf '%s\\n' "\${run_id}" >> "\${TRACE}"
if [[ "\${run_id}" == '400' ]]; then
${artifactFixture({ generation: 4, digest: WRONG_TRANSCRIPT_DIGEST })}
elif [[ "\${run_id}" == '200' ]]; then
${artifactFixture({ generation: 2, digest: TRANSCRIPT_DIGEST })}
fi
exit 0
fi
exit 1`);
expect(result.trace).toBe('400\n200\n');
expect(result.output).toMatch(/seed=.*\/200\/artifact\/gen-2\/bench\n/);
},
PREFLIGHT_TEST_TIMEOUT_MS,
);
it.skipIf(process.platform === 'win32')(
'continues without a seed when every prior artifact is unavailable',
() => {
const result = runSeedStep(`
if [[ "$1 $2" == 'run list' ]]; then
printf '300\\n'
exit 0
fi
if [[ "$1 $2" == 'run download' ]]; then
printf '%s\\n' "$3" >> "\${TRACE}"
exit 1
fi
exit 1`);
expect(result.trace).toBe('300\n');
expect(result.output).toBe('');
},
);
it('accepts an OpenAI key as an alternative to the Anthropic token', () => {
const requireAuth = stepRun('Require the benchmark auth secret');
expect(requireAuth).toContain('HAS_ANTHROPIC');
expect(requireAuth).toContain('GITNEXUS_BENCH_ANTHROPIC_API_KEY');
expect(requireAuth).toContain('GITNEXUS_BENCH_OPENAI_API_KEY');
expect(requireAuth).toContain('provider=openai');
expect(evolveJob?.env?.PROVIDER).toBe("${{ inputs.provider || 'openai' }}");
const loop = findStep('Run the propose → benchmark → gate loop');
expect(loop?.env).toMatchObject({
GITNEXUS_BENCH_ANTHROPIC_API_KEY:
'${{ secrets.GITNEXUS_BENCH_ANTHROPIC_API_KEY || secrets.GITNEXUS_BENCH_AUTH_TOKEN }}',
GITNEXUS_BENCH_OPENAI_API_KEY: '${{ secrets.GITNEXUS_BENCH_OPENAI_API_KEY }}',
});
});
it('runs the proposer on its own model, separate from the benchmark arms', () => {
expect(evolveJob?.env?.MODEL).toBe("${{ inputs.model || 'gpt-5.6-sol' }}");
expect(evolveJob?.env?.PROPOSER_MODEL).toBe("${{ inputs.proposer_model || 'gpt-5.6-sol' }}");
expect(evolveJob?.env?.EFFORT).toBe("${{ inputs.effort || 'xhigh' }}");
});
it('runs only the read-only review profile with a pinned external comparator', () => {
const loop = findStep('Run the propose → benchmark → gate loop');
expect(loop?.env).toMatchObject({
EVOLUTION_PROFILE: 'review',
CE_PLUGIN_DIR: '${{ runner.temp }}/compound-engineering-plugin',
CE_PLUGIN_VERSION: '3.24.0',
});
const comparatorStep = findStep('Fetch pinned Compound Engineering review comparator');
const comparator = String(comparatorStep?.run);
expect(comparatorStep?.env).toMatchObject({
CE_COMMIT: '3ad9b51bceecf0158e590c882034d0398dbb9c5c',
});
expect(comparator).toContain('checkout --detach');
expect(comparator).not.toContain('main');
});
it('provisions the benchmark task repo at ~/GitNexus before the loop', () => {
const provision = stepRun('Point the benchmark task repo at the checkout');
expect(provision).toContain('[[ -e "${HOME}/GitNexus" && ! -L "${HOME}/GitNexus" ]]');
expect(provision).toContain('ln -sfn');
expect(provision).toContain('${GITHUB_WORKSPACE}');
expect(provision).toContain('${HOME}/GitNexus');
});
it('bounds promotion to gitnexus-review and all shipped mirrors', () => {
const containment = stepRun('Detect and bound the applied promotion');
expect(containment).toContain('.claude/skills/gitnexus-review/*');
expect(containment).toContain('gitnexus/skills/gitnexus-review/*');
expect(containment).toContain('gitnexus-claude-plugin/skills/gitnexus-review/*');
expect(containment).toContain('gitnexus-cursor-integration/skills/gitnexus-review/*');
});
it('installs node_modules for the monorepo root and gitnexus, and compiles shared from the parent', () => {
// The benchmark sandbox-copies node_modules from root and gitnexus
// (tasks.scenarios.yaml). Shared is compiled by gitnexus `npm run build`
// (scripts/build.js); a dedicated npm ci in gitnexus-shared stalls CI.
const rootStep = findStep('Install monorepo root dependencies');
expect(rootStep).toBeDefined();
expect(rootStep).not.toHaveProperty('working-directory'); // installs at the repo root
expect(String(rootStep?.run)).toContain('npm ci');
expect(findStep('Build pinned shared runtime')).toBeUndefined();
const gitnexusRun = stepRun('Install and build pinned GitNexus runtime');
expect(gitnexusRun).toContain('npm ci');
expect(gitnexusRun).toContain('npm run build');
expect(gitnexusRun).toContain('mkdir -p ../gitnexus-shared/node_modules');
expect(
evolveJob?.steps?.some(
(step) =>
step['working-directory'] === 'gitnexus-shared' &&
String(step.run ?? '').includes('npm ci'),
),
).toBe(false);
});
it('waits for the runner boot-time package lock before installing containment tools', () => {
const install = stepRun('Install sandbox runtime and pinned Claude CLI');
expect(install).toContain('DPkg::Lock::Timeout=600 update');
expect(install).toContain('DPkg::Lock::Timeout=600 install');
expect(install).toContain('ripgrep');
});
it('names the promotion branch with the run attempt for re-run recovery', () => {
const openPr = stepRun('Open the promotion PR');
expect(openPr).toContain('${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}');
});
it('emits only the promoted generation with a per-run random output delimiter', () => {
const detect = stepRun('Detect and bound the applied promotion');
// Random per-run delimiter, not a fixed heredoc marker that a summary
// value could close early.
expect(detect).toContain('openssl rand -hex');
expect(detect).not.toContain("echo 'summary<<PROMOTION_EOF'");
// Single promoted generation (highest-numbered gen-N), not a blind
// concatenation of every generation's promotion.json.
expect(detect).toContain('sort -V');
expect(detect).not.toContain('xargs -0 -r cat');
});
it('least-privileges the App token and gates the job on a protected Environment', () => {
expect(evolveJob?.environment).toBe('gitnexus-evolution');
const mint = findStep('Mint GitHub App token');
expect(mint?.with).toMatchObject({
'client-id': expect.any(String),
'permission-contents': 'write',
'permission-pull-requests': 'write',
});
expect(mint?.with).not.toHaveProperty('app-id');
});
it('keeps scheduled runs off until the three-worker proof is explicitly enabled', () => {
const condition = String(evolveJob?.if);
expect(condition).toContain("github.event_name == 'workflow_dispatch'");
expect(condition).toContain("vars.GITNEXUS_EVOLUTION_ENABLED == 'true'");
expect(condition).toContain("vars.GITNEXUS_EVOLUTION_WORKERS == '3'");
});
it('fails before paid work when runner survival protections are ineffective', () => {
const preflight = stepRun('Verify runner survival policy');
expect(preflight).toContain('/etc/needrestart/conf.d/90-gitnexus-evolution.conf');
expect(preflight).toContain("$nrconf{restart} = 'l';");
expect(preflight).toContain('/proc/self/oom_score_adj');
expect(preflight).toContain('oom_score_adjustment > -900');
expect(preflight).toContain('SECONDS + 5');
});
it('labels the upload-artifact pin with its real version', () => {
expect(workflow).toContain(
'actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1',
);
expect(workflow).not.toContain('# v6.0.0');
});
it('runs every multi-line shell step under strict mode', () => {
const runSteps = (evolveJob?.steps ?? []).filter(
(step): step is { name?: string; run: string } =>
typeof step.run === 'string' && step.run.includes('\n'),
);
expect(runSteps.length).toBeGreaterThan(0);
for (const step of runSteps) {
expect(step.run, `${step.name} must set -euo pipefail`).toContain('set -euo pipefail');
}
});
it('kills the benchmark with job time left to upload its evidence', () => {
// A job-level timeout cancels the job outright, so the upload step never
// runs and a multi-hour generation's evidence is lost. The sweep therefore
// needs its own, strictly shorter budget: a step timeout only fails that
// step, and the always() upload below still ships what it wrote.
const jobBudget = evolveJob?.['timeout-minutes'];
const loopStep = findStep('Run the propose → benchmark → gate loop');
const stepBudget = loopStep?.['timeout-minutes'];
expect(typeof jobBudget).toBe('number');
expect(typeof stepBudget).toBe('number');
expect(stepBudget as number).toBeLessThan(jobBudget as number);
// The runner is an EC2 box an EventBridge schedule stops 24h after it
// starts; when the box goes the runner vanishes mid-step and nothing
// uploads. The job must finish inside that window even when the schedule
// fires late (the 2026-08-01 run was queued 65 minutes after the cron).
expect(jobBudget as number).toBeLessThanOrEqual(21 * 60);
// A Friday dispatch inherits leftover uptime. The shared entrypoint — not
// the workflow YAML — must cap the sweep so it fails in-process and the
// always() upload still runs (run 33962002890).
const script = readFileSync(
path.join(REPO_ROOT, 'eval/workflow_bench/run-evolution.sh'),
'utf8',
);
// The flag, not a precomputed number: the CLI reads /proc/uptime in the
// same breath as it starts the clock the cap is measured against, so
// nothing between the two can be charged to the sweep.
expect(script).toContain('--max-runtime-from-instance-window');
expect(script).not.toContain('instance_window_budget_from_proc');
expect(script).toContain('export RUNTIME_DIGEST');
// The workflow's half of that contract is calling the entrypoint, not
// naming the flag. The YAML never mentions --max-runtime-from-instance-window
// at all — only the prose at gitnexus-skill-evolution.yml:179-184 describing
// the cap — so an assertion on flag text there would test a comment, and a
// loop step that had stopped invoking the script would still pass it.
expect(stepRun('Run the propose → benchmark → gate loop')).toContain(
'./workflow_bench/run-evolution.sh --apply',
);
});
it('uploads benchmark evidence unconditionally, on a path it addresses itself', () => {
const upload = findStep('Upload benchmark evidence');
// The sweep appends results.jsonl and transcripts as it goes, so a killed
// generation still holds the evidence explaining why — and a path taken
// from the killed step's outputs is exactly what would not be there.
expect(upload?.if).toBe('always()');
expect(upload?.with?.path).toBe('${{ runner.temp }}/wfevolve');
});
it('documents the App secrets and protected Environment on the activation checklist', () => {
expect(workflow).toContain('RELEASE_APP_ID');
expect(workflow).toContain('RELEASE_APP_PRIVATE_KEY');
expect(workflow).toContain('gitnexus-evolution');
expect(workflow).toContain('GITNEXUS_EVOLUTION_ENABLED=true for scheduled runs');
expect(workflow).toContain('GITNEXUS_EVOLUTION_WORKERS');
});
});