GitNexus/gitnexus/test/unit/analyzer-identity.test.ts
Gergő Magyar 8b5057f325
feat(skills): GitNexus Engineering Tool Kits (#2566)
* feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill

Adds .claude/skills/ce-plan: a planning-only skill that builds
implementation-ready plans from GitNexus graph navigation (query/context/
impact/trace), bounded statement-level PDG slices (pdg_query, impact
mode:pdg, explain), and targeted source verification, with a context
ledger to prevent repeated reads and a machine-readable implementation
context pack (stable contract for a future ce-implement). Whitelisted in
.gitignore and registered in AGENTS.md and CLAUDE.md outside the
auto-managed gitnexus block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions)

Tool contract: impact mode:'pdg' shape now includes the schema-required
direction param; CDG branch sense documented as the result 'label' field
(reason is cypher/raw-edge only); explain caveats corrected to its real
false-negative classes (cross-function TAINT_PATH is modeled).

Consistency: PDG slice homed in working memory (ledger keeps one-liners);
depth knob defined and category-overrides-baseline ordering stated;
call_depth (consumed by nothing) and content-hash bookkeeping dropped;
Never section folded into Hard rules; Phase 3 deduplicated to a pointer;
allowed-repeat escalations defined; budget/discard accounting clarified;
verification-commands gathering added to Phase 4; open_questions added to
the context pack.

From scenario runs: plans now pin the verified-at HEAD commit and index
freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed],
quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and
support an out:<path> destination override; output path defined as the
Phase 1 target repo root.

Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata
bumps; future ce-implement qualified as future.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): rename ce-plan → gitnexus-plan; add cross-CLI (Codex) entrypoints

Renames the skill dir, frontmatter, output filename convention, plan H1
(GitNexus Engineering Plan), the future executor handle
(gitnexus-implement), the .gitignore whitelist entry, and all
AGENTS.md/CLAUDE.md references. Follows the pr-swarm-review cross-CLI
pattern: SKILL.md is the canonical CLI-neutral spec, AGENTS.md § Engineering
planning is the Codex/any-agent entrypoint, and the README documents the
optional user-level ~/.codex/prompts/gitnexus-plan.md slash command plus an
invocation matrix. Skill prose de-branded from Claude Code (agent-neutral
verification layer).

Also fixes two post-review README contradictions: the anti-reread claim now
names the ledger's allowed escalations, and 'read-only by contract' is now
'planning-only' (the skill writes exactly one repo file — the plan); the
scope-creep rule and template §12 now agree on where deferred follow-ups
land. Drops the stale plugin-collision limitation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): document Codex user-level install path for gitnexus-plan

Codex discovers SKILL.md skills from ~/.agents/skills (same path the other
gitnexus-* skills install to); README now documents the cp install plus the
optional ~/.codex/prompts slash-command file, with the prompt body preferring
the repo copy and falling back to the user-level install.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan freshness gate + active PDG-layer refresh

Freshness is now a Phase 1 gate, not advisory: under the default
freshness:strict, a stale index is refreshed once per planning session via
node .gitnexus/run.cjs analyze --index-only (appending --pdg when the task
will reach the PDG phase), then the context resource is re-read. A missing
PDG layer likewise triggers the one permitted --index-only --pdg refresh
and re-probe instead of a passive recommendation. freshness:accept (or a
failed/impractical refresh) preserves the old behavior: plan on the stale
graph, source-weighted, labelled in the plan header. --index-only is the
load-bearing flag choice — it suppresses all file generation, so the
planning-only contract holds (only the .gitnexus store changes). Ledger
gains an index_refresh record; plan header states fresh / refreshed /
refresh-skipped-with-reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): gitnexus-plan runner build check before freshness refresh

When the target repo builds the analyzer from its own source (bin → dist/
mapping, as gitnexus/ does), the Phase 1 freshness gate now verifies dist/
is current before running the analyze refresh — rebuilding via the
package's build script when any analyzer source file is newer than the
built entrypoint — and prefers that freshly built CLI. Otherwise a stale
dist re-indexes with outdated extraction logic and the 'fresh' index lies.
Rebuilds are recorded in the ledger's index_refresh; the PDG-phase refresh
inherits the same check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): add gitnexus-work executor and gitnexus-lfg pipeline

gitnexus-work executes a gitnexus-plan as verified atomic commits: consumes
the §11 implementation_context pack, drift-checks the plan's evidence pin
against HEAD, re-verifies assumptions before relying on them, runs impact
before every symbol edit and detect_changes before every commit (repo
mandates), builds tests from the plan's scenarios, and routes structural
drift back to gitnexus-plan Deepen mode instead of coding around it.

gitnexus-lfg is a thin orchestrator: gitnexus-plan → blocking user gate
(deepen / proceed / stop, deepen loops allowed) → gitnexus-work → review
via the existing gitnexus-pr-review skill (open PR, else branch diff vs
default). One bounded fix cycle for review findings; never pushes or opens
a PR on its own.

gitnexus-plan gains a Deepen mode (re-run freshness gate, escalate to
depth:deep, re-verify graph/inferred/assumed claims toward verified,
rewrite the same file); its 'future gitnexus-implement' placeholder is
retired in favor of gitnexus-work. Registered via .gitignore whitelists,
AGENTS.md 1.10.0 (section renamed to Engineering planning & execution),
CLAUDE.md 1.5.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): apply cross-skill review findings to the gitnexus skill family

Two P1s: gitnexus-plan Deepen mode now re-anchors before re-pinning
(diffs the old evidence pin over every [verified]-claim file and re-reads
or downgrades before the header moves — moving the pin without this
laundered stale claims as verified); the index-refresh budget is stated
once in Phase 1 (one --index-only refresh plus at most one Phase 3 --pdg
upgrade per session, Deepen = its own session) with ledger and pdg-slice
deferring to it.

Contract fixes: gitnexus-work's drift check now covers every file the
pack cites (not just files_to_modify) and parses the full pack incl.
primary/related symbols and acceptance_criteria (walked in Phase 4
alongside §13); a pre-completed check skips §7 steps already landed and
Deepen gains a reconcile-execution-state step, closing the mid-execution
route-back loop; pack assumptions must name what to check and how.

lfg: Lane 4 passes the merge-base to detect_changes compare (two-dot
diff misattributes upstream commits when default advanced), branch-diff
is the stated normal case, oversized review findings route to the plan
gate instead of overflowing direct mode, the one-fix-cycle cap is
explicit on re-run, and headless runs end at the plan gate with the plan
as deliverable. work: blank mode narrowed to *gitnexus-plan*.md with a
re-execution guard, direct-mode discipline spelled out, branch
meaningfulness defined against the plan slug, and the plan document is
committed as the branch's docs commit (review diff includes it).
Planning-only contract now names the dist/ rebuild as the second
permitted state change; Phase 5.1 names the four claim tags; stale
AGENTS.md anchors fixed.

Known latent issue left untouched: gitnexus/gitnexus-pr-review pairs a
three-dot example with a two-dot detect_changes compare — that skill is
also shipped by the plugin, so fixing it here would drift the copies;
lfg compensates by passing the merge-base.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ship the engineering skill family with the gitnexus package

npm i -g gitnexus users now get gitnexus-plan / gitnexus-work / gitnexus-lfg:
the three skills are added to gitnexus/skills/ in directory form (SKILL.md +
references/), which installSkillsTo already enumerates dynamically and copies
recursively to every editor target (~/.agents/skills for Codex, Cursor,
OpenCode, Qoder, ...) on gitnexus setup — uninstall enumerates the same root,
so removal stays clean. The Claude Code plugin channel
(gitnexus-claude-plugin/skills/) carries the same copies plus the standard
per-skill mcp.json.

Global-install support in the skill text: gitnexus-plan Phase 1 now resolves
the analyzer runner explicitly — node .gitnexus/run.cjs analyze when the
project has a runner, else gitnexus analyze (installed CLI), else
npx gitnexus analyze — and all analyze mentions route through it, satisfying
the skills-steering policy (#1939/#1945) which sweeps the plugin copies.

New drift guard test/unit/shipped-skills-sync.test.ts asserts the npm and
plugin copies stay byte-identical to the canonical .claude/skills/ family
(plugin = canonical + mcp.json), same discipline as run.cjs ↔
resolve-invocation.ts. skills-steering + shipped-skills-sync: 11/11 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench — measure the skill workflow's token savings

Benchmarks gitnexus-plan → gitnexus-work against a baseline agent
(--disallowedTools Skill) on identical tasks, in fresh detached worktrees,
using real headless Claude Code sessions; every number comes from the CLI's
--output-format json usage report (field names validated against a live
2.1.207 session). Reports per-arm medians (input/cache/output tokens, cost,
wall time, turns), a savings row, and resolve status from a per-task verify
command — savings on failed tasks are flagged, not celebrated. Per-task
setup hook prepares fresh worktrees (deps); --permission-mode
bypassPermissions (default) lets sessions run unattended in the throwaway
trees.

Free-model support: --base-url/--auth-token/--model route headless sessions
through any Anthropic-compatible endpoint; free-model.litellm.yaml is a
ready litellm-proxy template for OpenRouter :free variants or local Ollama,
so benchmarking burns no paid tokens (README documents rate limits and the
small-model skill-following caveat).

Harness validated end-to-end with a stub CLI (worktree lifecycle, both
arms, plan→work chaining, verify, aggregation, report) and 4 pytest units
for the pure aggregation/savings/report helpers. AGENTS.md 1.11.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record first workflow_bench calibration run

Trivial-task calibration (add -V alias): both arms resolved; workflow arm
~4.3x baseline cost — the documented overhead-dominated regime, recorded so
the regime boundary is empirical rather than asserted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): workflow_bench scenario matrix — arm variants, task classes, churn

Ground-base measurement across scenarios: tasks.scenarios.yaml spans four
labeled classes (trivial → investigation-bug → investigation-feature →
cross-module) with deterministic verifies (prescribed test files). New arms:
workflow_direct (gitnexus-work direct mode — the middle option that locates
the routing boundary lfg's gate and work's triage encode) and baseline_nomcp
(no skills AND no graph tools — separates workflow-discipline value from
GitNexus-tool value; off by default). Records now carry task class and diff
churn (files/+ins/−del vs the starting commit) as an over-engineering proxy;
the report renders a class column and per-arm savings rows vs baseline.
5 pytest units + stub-CLI e2e of the full three-arm matrix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record workflow_bench ground base; fix churn measurement bias

Ground base (3 classes x 3 arms, n=1/cell): every arm resolved every task —
pass/fail quality saturates at this difficulty, making the comparison pure
cost. Full plan→work never amortized its ~$9-11 fixed cost on tasks a
baseline finishes in ≤35 turns (−211% to −333% cost); workflow_direct sits
near baseline (−15% to −55%, once faster wall) with more test coverage.
Routing implication recorded: direct mode/plain agent below this scale,
full workflow for cross-module / multi-session / plan-as-deliverable work.
The cross-module cell and multi-run variance are the next measurements.

Churn fix: git add --intent-to-add -A before diffing (arms that never
commit no longer undercount new files) and :(exclude)docs/plans (the
committed plan doc no longer inflates workflow churn); this run's churn
numbers predate the fix and are omitted from the recorded table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(skills): cost-optimize the workflow from measured ground base

Every optimization targets a measured fixed-cost component
(eval/workflow_bench ground base: workflow arm −211% to −333% vs baseline,
all tasks resolved):

- Plan form is category-priced: compact form (core sections w/ § anchors
  preserved, ≤80 lines excl. pack, mini-pack subset of the context pack)
  for narrow/default categories; the full 13 sections only for deep work
  (refactor/security/performance/concurrency/architecture). A compact plan
  outgrowing its cap reclassifies to full rather than overflowing.
- Freshness gate is category-priced: compact categories default to accept
  (source-weighted, refresh only when a graph claim becomes load-bearing);
  strict stays the default for full-plan categories — the rebuild+re-index
  was the largest single fixed cost.
- Turn economy: per-category tool-call budgets (~10 to ~45; architecture
  uncapped); budget exhaustion routes open questions to §12 instead of
  more digging.
- gitnexus-work fast path: HEAD == evidence pin → skip all citation
  re-reading (the pin's entire point); mini-pack fields tolerated.
- lfg Lane 1 boundary triage: tasks below the measured ~35-turn boundary
  get offered gitnexus-work direct mode before the plan lane is spent.

Copies re-synced (npm skills/, plugin, ~/.agents); steering + sync guards
green. Re-measurement of the workflow arm follows to verify the numbers
actually improve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): record optimization re-measurement — inv-bug workflow cell −20% cost

Same task, same conditions, post-830a0459 skills: $14.56→$11.70 (−20%),
83→72 turns, cache_read −24%; verified in-transcript that the compact form,
turn budget, and skipped rebuild/re-index all fired. Wall +15% from a work-
session test-debugging tail (n=1 variance). Regime unchanged (~3.5x baseline
on this class) — routing rule stands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): per-arm clone isolation — worktree ref-namespace leak contaminated an arm

The cross-module workflow_direct cell reported an impossible 28-turn solve
with churn byte-identical to the workflow arm: git worktree add shares the
repo's ref namespace, so the workflow arm's slug branch (created by
gitnexus-work Phase 2) survived worktree removal and the direct arm found
and adopted the completed work. Arms now get isolated git clone --shared
copies (object store via alternates, refs clone-local — agent branches and
stashes die with the clone; origin/<ref> fallback for non-default refs).
Leaked branch deleted; baseline arm verified clean (0 branch references in
its transcript); cell marked invalidated pending re-run.

Records the valid cross-module cells: workflow $18.32 vs baseline $18.03
(premium −1.6%, vs −211%..−333% on smaller classes) — fixed costs amortize
at this scale, with a less destructive diff and a plan artifact as bonus;
resolve rate still tied. Churn fingerprinting is what caught the
contamination — noted in the README as an integrity check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(eval): complete cross-module cell — direct mode wins 47% cost / 56% wall

Clean clone-isolated re-run: workflow_direct resolved the hardest class at
$9.53/52 turns/15m vs $18.03/98/34m baseline and $18.32/107/37m full
workflow. The measured story across all four classes: the execution
discipline (gitnexus-work) is the consistent sweet spot and delivers real
token savings on hard tasks; the planning pass buys its artifact, not
same-session savings. Resolve rate tied everywhere (n=1/cell caveat).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): add trajectory-gated skill evolution (#2431)

- Pair prompt candidates with incumbent workflow arms
- Gate promotions on pinned-model quality and efficiency
- Expire router evidence and document its lifecycle

* fix(eval): allow pr-review skill candidates

* feat(skills): rename and generalize GitNexus review

* feat(eval): external-comparator and review arms for workflow_bench

- ce_workflow / ce_workflow_direct: compound-engineering ce-plan/ce-work
  arms prompted with the same structure as the gitnexus arms
- review / ce_review: gitnexus-review vs ce-code-review on an identical
  diff applied by the task's setup
- plan handoff is snapshot-based: committed example plans in docs/plans/
  tie on clone mtimes and broke the name-glob pick (executed a stale plan)
- verify output tail is recorded per run and the final working-tree patch
  is kept, so failed rows are diagnosable after the clone is destroyed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills,eval): address #2431 review — data-safe rename migration, fail-closed bench evidence

- setup: never delete a legacy renamed skill dir — the installer cannot
  prove ownership (users customize or hand-write skills under these
  names); warn with the path instead, and the test now asserts survival
- workflow_bench: fail closed when a session's --output-format json
  report is empty, malformed, or missing usage fields — an exit-0 shell
  with no parseable usage no longer counts as measured evidence
  (5 parametrized regression tests)
- workflow_bench: document the trust model prominently (task setup/verify
  are shell-executed, sessions run bypassPermissions with the parent env,
  candidate overlays are prompt injection surface) in README + docstring
- free-model.litellm.yaml: master_key from LITELLM_MASTER_KEY env instead
  of a static token; loopback-binding warning
- ci: run the eval workflow_bench pytest suite on ubuntu (pytest+pyyaml
  only — no full eval stack)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): demand observed foreground verification in headless work-arm prompts

In a headless -p session there is no later turn: a work arm backgrounded
its slow test run, scheduled wakeups that can never fire, and reported
done while two of its tests failed. All four work-arm prompts (both
skill families, symmetric) now require verification output to be
observed inside the session.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): ask plan depth up front instead of offering deepen afterwards

gitnexus-plan Phase 0 now asks one blocking question in interactive
sessions — quick / standard / deep, mapped onto the existing depth/form/
freshness knobs — when the invocation carries no explicit depth signal.
Explicit knobs and headless runs skip the question (category posture
unchanged, so benchmarks and automation behave as before).

gitnexus-lfg's plan gate slims to proceed/stop: depth was already the
user's up-front choice, so deepening is no longer offered by default —
an explicit deepen request at the gate and executor route-backs still
run Deepen mode, which remains the mechanism for strengthening an
existing plan document.

All shipped copies resynced (npm skills/, Claude plugin); AGENTS.md
1.13.0 and CLAUDE.md 1.7.0 pointers updated, including the analyzer's
regenerated index-stats block at this branch's head.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): taint pass, expert lenses, and post-work index refresh

gitnexus-review gains a PDG-backed taint-and-dependence pass (explain +
pdg_query, --pdg folded into the stale refresh on trust-boundary diffs) and
an Expert lenses section: domain reviewers derived from the graph's
clusters plus four cross-cutting lenses (architectural fit, language
conformance per the repo's own contract, Definition of Done, simplicity),
dispatched once after the evidence-gathering steps and scaled to the diff.
gitnexus-work Phase 4 now refreshes the knowledge graph after the DoD walk
via the resolved-runner ladder with analyze --index-only, so the lfg review
lane and later sessions query the finished work without dirtying the tree.
lfg's threshold-governance paragraph moves to its README; eval citations
are tagged as measured in the GitNexus repo. All shipped copies re-synced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(cli): remove legacy gitnexus-pr-review on uninstall; cover the rename migration

uninstall's removal set now includes LEGACY_SKILL_DIR_NAMES derived from
RENAMED_SKILL_DIRS, so a pre-rename install is cleaned up instead of
orphaned. The rename warning gains behavioral coverage (fires with a legacy
dir present, silent without), and shipped-skills-sync asserts legacy names
stay absent from every shipped tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(eval): metric provenance, error-kind rows, skill-invocation verification, gate noise floor

The promotion gate defaults to cost_usd (the only metric that includes
subagent spend); token metrics carry an explicit main-loop-only warning in
the report and promotion.json. Rows are classified by error_kind
(session-error / verify-failed / infra-error), excluded from efficiency
medians, and the gate requires equal valid-run counts. Each session's
transcript is scanned for the expected Skill invocation and fails closed on
a verified miss; a one-run resolution edge no longer promotes (noise
floor). Per-run timeouts and setup failures record an infra-error row
instead of aborting the sweep. Overlays touching skills no candidate arm
exercises are rejected up front.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fix skill routing paths, version headers, and skill rosters

Routing tables point at the tracked direct skill paths (matching the
post-#2434 generator output), AGENTS.md/CLAUDE.md headers match their
latest changelog rows, the 1.12.0 row describes what the migration actually
does, package/cursor READMEs list the full shipped skill roster, and the
swarm READMEs describe /gitnexus-review's expert lenses instead of calling
it single-agent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: drift-guard workflow for skill copies; pin eval pip deps; track docs/plans

ci.yml ignores '**.md', so an md-only skill edit would merge without the
shipped-skills-sync test running — skill-sync.yml triggers exactly on the
guarded trees. The eval job's pip install is version-pinned, and
docs/plans/ is unignored so gitnexus-plan output can be committed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep the runner-invocation literal in gitnexus-review; add concurrency block to skill-sync

skills-steering requires skills with a stale-index hint to carry the exact
'node .gitnexus/run.cjs analyze' form — restore it with the fallback ladder
as a parenthetical instead of replacing it. skill-sync.yml gains the
top-level concurrency block the workflow-convention check enforces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(skills): token-economy guidance for expert lenses

Merge lenses that ground in the same material into one reviewer, and use
cheaper model/effort tiers for mechanical lenses where the harness offers
them, reserving the strongest engine for adversarial judgment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(eval): isolate transcript home on Windows

Ensure workflow_bench transcript tests set USERPROFILE alongside HOME so Path.home() resolves to the temporary test home on Windows.

* docs(skills): fold PR #2522 execution learnings into review/work/plan

Eight incident-backed hardenings from running the full skill cycle
(review -> plan -> work, 28-finding fix series) on PR #2522:

gitnexus-review:
- Expert lenses execute the code under review on candidate failing shapes
  (empirical probe outranks source reading — every HIGH the language
  lenses found came from a probe, not a read).
- Step 7 re-runs the exact CI check for refreshed baselines/fingerprints
  (a stale committed artifact is invisible in the diff; caught a red
  benchmarks arm).
- Step 8 treats version/invalidation constants as review surface
  (INCREMENTAL_SCHEMA_VERSION class recurred verbatim from #2494).

gitnexus-work:
- Step 4 proves regression tests discriminate against the pre-fix tree.
- Step 5 rebuilds executed build output before every verification run
  (parse workers load dist/; a correct fix 'failed' until rebuilt).
- Step 6 makes stage -> detect_changes -> commit one unbroken sequence.

gitnexus-plan:
- Phase 0 seeded-evidence mode: plan FROM a completed review's verified
  findings instead of re-running the graph ladder.
- Template §7: fingerprint/golden-guarded output rebaselines once, at the
  series tip.

All distribution copies resynced; shipped-skills-sync + skills-steering
24/24 locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): close the skill-evolution loop with an automated proposer driver

workflow_bench.evolve adds the three arrows the README described as manual:
a proposer session that turns loser trajectories (results.jsonl rows,
transcripts, patches, the learning queue) into ONE bounded candidate
overlay, a driver that iterates propose -> paired benchmark -> deterministic
gate up to --generations, and an --apply step that copies a promoted
overlay onto the canonical skills and shipped mirrors as a working-tree
diff. The trust boundary is unchanged: overlays re-validate through
candidate_overlay_files before any benchmark or apply consumes them, and
committing, CI, and the PR merge stay human.

learnings.jsonl is gitignored: it is machine-local evidence, like the
session transcripts it complements.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): route live-task friction into the evolution learning queue

Each family skill gains a short 'Skill feedback' section: on friction with
the skill's own instructions, append one JSON line to
eval/workflow_bench/learnings.jsonl (GitNexus repo only) — never self-edit
the skill from a live task. The proposer in workflow_bench.evolve consumes
the queue as hints; a learning reaches a shipped skill only by beating the
incumbent on the paired benchmark. All shipped mirrors re-copied byte-
identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(tests): run the evolve helper tests in the eval pytest job

test_evolve.py needs only pytest+pyyaml, same as the harness tests the job
already runs — without this line the new module had no CI coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): comment-triggered GitNexus review agent for PRs

'@gitnexus review' from a maintainer (OWNER/MEMBER/COLLABORATOR; the action
re-validates write access) runs the repo's gitnexus-review skill headlessly
against the PR and posts the review as a sticky comment — remote triggering
with no local setup. Read-only by construction: contents: read token,
Write/Edit and web tools disallowed, Bash allowlisted to git reads and the
gitnexus CLI; analyze parses PR code with tree-sitter, never executes it.
Requires the ANTHROPIC_API_KEY repository secret; activates once the file
is on the default branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): dispatch lane + existing OAuth secret for the review agent

Align with claude.yml: same action pin and the CLAUDE_CODE_OAUTH_TOKEN
secret the repo already carries — no new secret to configure. Add a
workflow_dispatch lane (PR number input) so the agent can be triggered from
the Actions UI and tested before the issue_comment trigger reaches the
default branch. Allowlist gh pr view/diff and gh api, which the review
skill uses to pin PR SHAs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): close a fork-PR RCE vector in the review agent's tool allowlist

A live headless run of the exact workflow session against PR #2431 (66
turns, full gitnexus-review pass) surfaced a real HIGH-severity confused
deputy: .gitnexus/ is gitignored, not blocked — a fork PR can commit its
own .gitnexus/run.cjs, issue_comment checks out PR-head content, and the
skill's runner ladder tries 'node .gitnexus/run.cjs analyze' first. That
would execute fork-controlled JS inside a job holding
CLAUDE_CODE_OAUTH_TOKEN and a write-scoped GITHUB_TOKEN — the opposite of
the 'PR code is read, never executed' claim in the workflow's own header.

Fix: drop the run.cjs allowlist entry so analyze always resolves through
npx gitnexus (npm registry, not the checked-out tree); the skill's
documented fallback mode covers the resulting graceful degradation. Also
drop 'gh api' (not read-only — accepts -X POST/PATCH/DELETE) and downgrade
pull-requests: write to read (comment posting only needs issues: write;
the prompt already forbids formal review submission).

Same session flagged a latent evolve.py bug: select_evidence's cost sort
used dict.get's missing-key default, which doesn't cover an explicit JSON
null in a foreign --seed-results row and crashes proposer setup with
TypeError. Guarded with 'or 0.0' and added a regression test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden PR review and evolution trust boundaries

* ci: follow workflow concurrency convention

* fix(eval): make terminating error paths explicit

* fix: unblock hardened review runtime checks

* test: make containment canaries deterministic

* test: expose Claude canary tool failures

* fix: adapt clean shell environment for Claude

* fix(eval): accept the runner's transcript source key in evidence preflight

The proposer evidence preflight required transcript-artifact metadata to be
exactly {path, sha256, bytes}, but the runner stamps a fourth provenance key
(source=parent-captured-stream-json). Any --seed-results or generation>=2 run
therefore aborted with SandboxError before proposing or promoting. Pin the
producer literal as PARENT_EVENT_STREAM_SOURCE and validate it in the metadata
check, and round-trip real producer output through sum_sessions into the
preflight so the schema can't drift again.

* fix(eval): treat an unmeasured session cost as unavailable, not $0

well_formed validated only the nested usage block, so an otherwise-successful
session missing total_cost_usd was recorded as cost_usd=0.0 — and cost_usd is
the default promotion metric (lower wins), so a cost-less session scored as
free and could win promotion it never earned. Extract cost via measured_cost()
(None on absent/garbage, a measured 0.0 preserved), propagate None through
sum_sessions/aggregate/savings/report, and have the gate refuse to rank on a
metric that was not measured on every run in both arms.

* fix(eval): warn when ranking on the main-loop-only num_turns metric

num_turns comes from the CLI's top-level usage (main-loop session only), like
output_tokens, but selecting it emitted no metric_warning — so a subagent-heavy
candidate could look artificially efficient. Add num_turns to
MAIN_LOOP_ONLY_METRICS and broaden the warning to cover turns.

* fix(eval): fail closed when an overlay adds a file with no committed base

An overlay adding a new .md under gitnexus-{plan,work} passes the structural
overlay checks but has no committed base for committed_destination_base_digests
to bind against, so it raised an uncaught ValueError that crashed the evolve
driver (and runner --candidate-overlay) mid-run. Catch it at both call sites:
evolve reports NOT PROMOTED and exits, runner routes it through parser.error.

* feat(eval): circuit-break the runner sweep on a systemic outage

A sustained upstream outage used to pay out every remaining --timeout window
one session at a time. Track consecutive session/infra/cleanup failures via a
pure systemic_outage_streak helper; after --outage-streak (default 5) in a row,
stop the sweep, still write report.md/promotion.json from partial evidence, and
exit non-zero so evolve.py halts instead of proposing from truncated evidence.
A task's own resolved=False never trips the breaker.

* fix(cli): report a dirty working tree as stale in gitnexus status

status --json (and the human output) computed up-to-date from commit + runner
identity + completeness only, so a repo with uncommitted source changes at a
matching HEAD was reported up-to-date while analyze would still re-index it.
A graph-backed agent gating on that JSON could skip re-analysis on a stale
graph. Extract analyze's dirty-tree check into a shared isWorkingTreeDirty()
in storage/git and fold it into the status freshness decision.

* fix(ci): use single-slash deny globs in the review agent's disallowedTools

github.workspace already expands to an absolute path, so Read(/${{ github.workspace }}/**)
and Read(//proc/**),(//sys/**),(//dev/**) produced double-slash patterns that a
normalizing matcher may not match — silently no-opping the deny layer. Not
exploitable (the allowlist is the primary control and never grants those
paths), but the globs should be well-formed. Update the pinned test strings.

* ci: install gitnexus-shared with npm ci from the committed lockfile

The gitnexus-shared build floated its deps via npm install in three workflows
(skill-sync, ci-tests, and — most importantly — the release publish.yml) while
every other install step uses npm ci. The lockfile is committed and in sync, so
switch all three to npm ci for reproducible, locked installs.

* test(cli): make the shipped-skills drift guard reject symlinks

listFilesRecursive walked with readdirSync and snapshotDir read with
readFileSync, both of which follow symlinks — so a mirror file symlinked to the
canonical tree passed the byte-compare (and a symlinked mirror dir would be
followed too). Reject a symlinked root via lstat and any symlinked entry via
Dirent.isSymbolicLink, with negative tests (skipped on Windows).

* test(eval): guard the candidate-skill vs mirror-root coverage invariant

MIRROR_SKILL_ROOTS omits the Cursor tree, safe only because no candidate skill
is cursor-shipped. Pin that invariant: every CANDIDATE_SKILLS entry must exist
under canonical + every mirror root and must not ship to Cursor, so adding a
cursor-shipped skill to the candidate set (the PR #2488 asymmetric-sync class)
fails loudly instead of syncing three of four trees.

* docs(ci): describe the review agent's staged post-merge rollout

The DoD asked for a dry-run or triggered run before merge, but an issue_comment
(or newly added workflow_dispatch) workflow only ever executes the default-branch
copy, so it cannot be exercised from the PR that introduces it. Reword the DoD
and the activation checklist to a staged rollout: merge registered-but-disabled,
validate same-repo and fork execution post-merge, then enable the variable.

* fix: pin plugin skill mcp.json to the release version via #2445 tooling

The ten plugin skill mcp.json launched `npx -y gitnexus@latest mcp` on every
skill connect — non-reproducible and a supply-chain surface, and (unlike the
persisted setup config) never pinned. Extend sync-plugin-manifests.mjs with an
mcp surface kind that stamps the gitnexus@<version> launch arg, pin all ten to
1.6.9 now, and keep them byte-identical so the drift guard stays green. The
release lifecycle + publish.yml --check now re-stamp them like the four manifest
surfaces; only READMEs stay on @latest as docs.

* test(eval): prove the proposer's built-in file tools are confined

The real-Claude canary only exercised Bash + MCP, so it proved process/MCP
containment but not that the proposer's built-in file tools stay inside their
mounts. Add a canary over the exact PROPOSER_ALLOWED_TOOLS surface and the same
read-only /evidence mount as run_proposer (allowlist extracted to a shared
constant so it can't drift): Read reaches /evidence, a Write into the read-only
evidence mount is denied, and a Write lands in the output tree.

* fix(eval): apply the candidate overlay after task setup for fair arms

The candidate overlay was applied before the task's untrusted setup ran, so
setup could observe candidate prose and the incumbent/candidate arms started
from different pre-overlay state. Reorder within the sandbox: capture the base
(pre-overlay) skill digest, run setup against the base skills, verify setup did
not tamper them, then apply the overlay and capture the post-overlay digest the
model must preserve. apply_candidate_overlay stages path-specific overlay files,
so setup's uncommitted changes stay out of the baseline and churn is unchanged.

Graph freshness for the review arm is handled by the status dirty-tree fix plus
the review skill's stale-triggered re-index, not by reordering the cached
per-task-sha graph materialization (which is mechanically blocked).

* test(eval): end-to-end containment proof of the autonomous proposer

Drives the real run_proposer through bubblewrap with a deterministic scripted
model (no paid API): it reads the read-only evidence bundle and writes a
candidate gitnexus-plan skill edit plus a rationale into the sandbox output
tree; run_proposer enforces the trust boundary and copies only the validated
overlay + proposal out. This exercises the autonomous-proposal stage of the
self-evolution loop end-to-end in the eval/containment CI job (the gate and
apply stages are covered by test_workflow_bench_evolution and
test_promotion_apply). Env-gated on GITNEXUS_REQUIRE_CLAUDE_CANARY, so it runs
only where the pinned Claude binary and user namespaces are available.

* fix(eval): let the proposer author its overlay via Bash

Running the end-to-end proposer canary in the containment CI job surfaced a real
bug: run_proposer starts the session with --bare, which hard-disables the
Write/Edit tools ("Write exists but is not enabled in this context"), yet
allowlisted Edit/Write and omitted Bash. The proposer therefore had no working
way to write its candidate overlay — the self-evolution loop could never produce
a candidate. The sandbox settings already pre-authorize Bash
(autoAllowBashIfSandboxed) and confine writes to workspace/tmp/home, so switch
PROPOSER_ALLOWED_TOOLS to Read/Grep/Glob/Bash and tell the proposer to author
files with Bash. The end-to-end test now drives the real run_proposer through
bubblewrap and asserts a validated overlay + proposal are produced (this also
replaces the earlier file-tool canary, whose Write/Edit premise was moot).

* test(eval): author the proposer overlay with newline-free Bash content

The nested shell-sandbox prefix mangles embedded newlines, so the multi-line
overlay content never landed. Use single-line content for the deterministic
proposer canary.

* test(eval): drop the unverifiable end-to-end proposer canary

The scripted proposer overlay never materialized in the containment job across
runs, and the model tool-result content is not visible in CI logs, so the test
cannot be finalized without an environment where the sandbox can actually run.
Keep the verified production fix (Bash-authoring in run_proposer); the proposer
sandbox/containment stays covered by the existing Bash+MCP and process-tree
canaries.

* test(cli): drop run-analyze.ts from the windowsHide spawn-family list

U7 moved run-analyze.ts's only child_process call (the git status --porcelain
dirty check) into storage/git.ts (already covered by this test, with
windowsHide). run-analyze.ts no longer imports a spawn-family function, so the
windowsHide-regression test's 'must have >=1 spawn call' invariant failed for
it. Remove it from SRC_FILES.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Zander Raycraft <zanderjraycraft@gmail.com>
Co-authored-by: Azizur Rahman <azizur100389@gmail.com>
2026-07-19 15:07:24 +01:00

1510 lines
60 KiB
TypeScript

import { createHash } from 'node:crypto';
import { writeFileSync } from 'node:fs';
import { link, mkdir, readFile, readdir, symlink, unlink, writeFile } from 'node:fs/promises';
import { performance } from 'node:perf_hooks';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { describe, expect, it } from 'vitest';
import {
_clearAnalyzerIdentityProcessCacheForTests,
_hashAnalyzerIdentityFramesForTests,
analyzerRunnerIdentitiesEqual,
captureAnalyzerIdentityBeforeLoad,
finalizeAnalyzerRunnerIdentity,
normalizeAnalyzerRunnerIdentityForComparison,
resolveAnalyzerRunnerIdentity,
} from '../../src/core/analyzer-identity.js';
import { getStoragePaths, loadMeta, saveMeta } from '../../src/storage/repo-manager.js';
import type { RepoMeta } from '../../src/storage/repo-manager.js';
import { setupMiniRepo } from '../helpers/mini-repo.js';
import { createTempDir } from '../helpers/test-db.js';
describe('analyzer runner identity', () => {
it('is versioned, resolved, and changes when the analyzer build tree changes', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(path.join(fixture.dbPath, 'package-lock.json'), '{"lockfileVersion":3}\n');
await writeFile(modulePath, 'export const analyzer = 1;\n');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first).toMatchObject({
schemaVersion: 4,
cliVersion: '9.8.7',
runtime: {
executablePath: expect.any(String),
version: process.version,
platform: process.platform,
architecture: process.arch,
modulesAbi: process.versions.modules ?? 'unknown',
libc: expect.any(String),
},
invokedArtifact: {
path: modulePath,
digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/),
},
build: {
kind: 'source',
rootPath: sourceRoot,
canonicalization: 'gitnexus-analyzer-build-v2',
digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/),
},
dependencyRuntime: {
manifestPath: path.join(fixture.dbPath, 'package.json'),
lockfilePath: path.join(fixture.dbPath, 'package-lock.json'),
canonicalization: 'gitnexus-analyzer-dependency-runtime-v4',
packageCount: 1,
artifactCount: 0,
digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/),
},
});
await writeFile(path.join(sourceRoot, 'new-module.ts'), 'export const changed = true;\n');
const second = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(second.invokedArtifact.digest).toBe(first.invokedArtifact.digest);
expect(second.build.digest).not.toBe(first.build.digest);
expect(second.dependencyRuntime.digest).toBe(first.dependencyRuntime.digest);
} finally {
await fixture.cleanup();
}
});
it('rejects build-tree symlinks instead of trusting unchanged link metadata', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const importedTarget = path.join(fixture.dbPath, 'outside-build-input.ts');
const importedLink = path.join(sourceRoot, 'linked-input.ts');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(importedTarget, 'export const imported = 1;\n');
try {
await symlink(importedTarget, importedLink, 'file');
} catch (error) {
if (['EPERM', 'EACCES'].includes((error as NodeJS.ErrnoException).code ?? '')) return;
throw error;
}
const resolve = () =>
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: path.join(fixture.dbPath, 'identity-cache'),
});
expect(resolve).toThrow(/build symbolic links are not supported/);
// Changing only target bytes leaves the symlink inode/text unchanged.
// The resolver must continue to fail closed, never return an old digest.
await writeFile(importedTarget, 'export const imported = 200;\n');
expect(resolve).toThrow(/build symbolic links are not supported/);
} finally {
await fixture.cleanup();
}
});
it('versions runtime semantics and rejects a cache from another runtime variant', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, Buffer.alloc(128 * 1024, 0x5a));
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.runtime).toMatchObject({
version: process.version,
platform: process.platform,
architecture: process.arch,
modulesAbi: process.versions.modules ?? 'unknown',
libc: expect.stringMatching(/\S/),
});
expect(
analyzerRunnerIdentitiesEqual(
{ ...first, runtime: { ...first.runtime, platform: `${first.runtime.platform}-other` } },
first,
),
).toBe(false);
expect(
analyzerRunnerIdentitiesEqual(
{
...first,
runtime: { ...first.runtime, architecture: `${first.runtime.architecture}-other` },
},
first,
),
).toBe(false);
expect(
analyzerRunnerIdentitiesEqual(
{
...first,
runtime: { ...first.runtime, modulesAbi: `${first.runtime.modulesAbi}-other` },
},
first,
),
).toBe(false);
expect(
analyzerRunnerIdentitiesEqual(
{ ...first, runtime: { ...first.runtime, libc: `${first.runtime.libc}-other` } },
first,
),
).toBe(false);
const [cacheFile] = await readdir(cacheDirectory);
const cachePath = path.join(cacheDirectory, cacheFile);
const envelope = JSON.parse(await readFile(cachePath, 'utf8')) as {
payload: {
schemaVersion: number;
runtimeVariant: {
nodeVersion: string;
platform: string;
architecture: string;
modulesAbi: string;
libc: string;
};
};
checksum: string;
};
expect(envelope.payload).toMatchObject({
schemaVersion: 6,
runtimeVariant: {
nodeVersion: process.version,
platform: process.platform,
architecture: process.arch,
modulesAbi: process.versions.modules ?? 'unknown',
libc: first.runtime.libc,
},
});
envelope.payload.runtimeVariant.platform = `${process.platform}-stale-cache`;
envelope.checksum = `sha256:${createHash('sha256')
.update(JSON.stringify(envelope.payload))
.digest('hex')}`;
await writeFile(cachePath, `${JSON.stringify(envelope)}\n`);
_clearAnalyzerIdentityProcessCacheForTests();
let hashedBytes = 0;
const afterIncompatibleCache = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onHashedInput: ({ bytes }) => {
hashedBytes += bytes;
},
});
expect(afterIncompatibleCache).toEqual(first);
expect(hashedBytes).toBeGreaterThanOrEqual(128 * 1024);
} finally {
await fixture.cleanup();
}
});
it('disables the default persistent cache without getuid but trusts an explicit override', async () => {
const fixture = await createTempDir();
const originalGetuid = Object.getOwnPropertyDescriptor(process, 'getuid');
const originalEnvironment = {
TMPDIR: process.env.TMPDIR,
XDG_RUNTIME_DIR: process.env.XDG_RUNTIME_DIR,
};
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const tempRoot = path.join(fixture.dbPath, 'tmp');
const explicitCache = path.join(fixture.dbPath, 'operator-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(tempRoot, { mode: 0o700 });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, Buffer.alloc(64 * 1024, 0x33));
process.env.TMPDIR = tempRoot;
delete process.env.XDG_RUNTIME_DIR;
Object.defineProperty(process, 'getuid', {
value: undefined,
configurable: true,
enumerable: true,
writable: true,
});
const hashedWith = (cacheDirectory?: string): number => {
let bytes = 0;
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
...(cacheDirectory ? { cacheDirectory } : {}),
onHashedInput: (input) => {
bytes += input.bytes;
},
});
return bytes;
};
expect(hashedWith()).toBeGreaterThanOrEqual(64 * 1024);
// Cross-platform process-local reuse remains available even when secure
// default persistence cannot be proved.
expect(hashedWith()).toBe(0);
expect(await readdir(tempRoot)).toEqual([]);
_clearAnalyzerIdentityProcessCacheForTests();
expect(hashedWith(explicitCache)).toBeGreaterThanOrEqual(64 * 1024);
_clearAnalyzerIdentityProcessCacheForTests();
expect(hashedWith(explicitCache)).toBe(0);
} finally {
if (originalGetuid) Object.defineProperty(process, 'getuid', originalGetuid);
else Reflect.deleteProperty(process, 'getuid');
if (originalEnvironment.TMPDIR === undefined) delete process.env.TMPDIR;
else process.env.TMPDIR = originalEnvironment.TMPDIR;
if (originalEnvironment.XDG_RUNTIME_DIR === undefined) delete process.env.XDG_RUNTIME_DIR;
else process.env.XDG_RUNTIME_DIR = originalEnvironment.XDG_RUNTIME_DIR;
await fixture.cleanup();
}
});
it('supports only an absolute, pre-provisioned, external operator-trusted cache directory', async () => {
const fixture = await createTempDir();
const protectedCache = await createTempDir();
const previous = process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR;
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, Buffer.alloc(64 * 1024, 0x37));
const url = pathToFileURL(modulePath).href;
const hashedBytes = (): number => {
let bytes = 0;
resolveAnalyzerRunnerIdentity(url, {
onHashedInput: (input) => {
bytes += input.bytes;
},
});
return bytes;
};
process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR = protectedCache.dbPath;
_clearAnalyzerIdentityProcessCacheForTests();
expect(hashedBytes()).toBeGreaterThanOrEqual(64 * 1024);
_clearAnalyzerIdentityProcessCacheForTests();
expect(hashedBytes()).toBe(0);
for (const invalid of [
'relative/cache',
path.join(fixture.dbPath, 'missing-cache'),
fixture.dbPath,
sourceRoot,
]) {
process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR = invalid;
_clearAnalyzerIdentityProcessCacheForTests();
expect(() => resolveAnalyzerRunnerIdentity(url)).toThrow(
/GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR/,
);
}
const linkedCache = path.join(fixture.dbPath, 'cache-link');
try {
await symlink(protectedCache.dbPath, linkedCache, 'dir');
process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR = linkedCache;
_clearAnalyzerIdentityProcessCacheForTests();
expect(() => resolveAnalyzerRunnerIdentity(url)).toThrow(/non-symlink|symbolic links/);
await unlink(linkedCache);
} catch (error) {
if (!['EPERM', 'EACCES'].includes((error as NodeJS.ErrnoException).code ?? '')) throw error;
}
} finally {
if (previous === undefined) delete process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR;
else process.env.GITNEXUS_ANALYZER_IDENTITY_CACHE_DIR = previous;
_clearAnalyzerIdentityProcessCacheForTests();
await protectedCache.cleanup();
await fixture.cleanup();
}
});
it('changes on lock and native/parser mutations while ignoring model caches', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const grammarRoot = path.join(fixture.dbPath, 'vendor', 'tree-sitter-fixture');
const nativePath = path.join(
grammarRoot,
'prebuilds',
`${process.platform}-${process.arch}`,
'tree-sitter-fixture.node',
);
const sharedLibraryPath = path.join(
path.dirname(nativePath),
process.platform === 'win32'
? 'tree-sitter-fixture.dll'
: process.platform === 'darwin'
? 'libtree-sitter-fixture.dylib'
: 'libtree-sitter-fixture.so.1',
);
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(path.dirname(nativePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
const originalLock = '{"name":"fixture-analyzer","lockfileVersion":3}\n';
await writeFile(path.join(fixture.dbPath, 'package-lock.json'), originalLock);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(grammarRoot, 'package.json'),
'{"name":"tree-sitter-fixture","version":"1.0.0"}\n',
);
await writeFile(nativePath, 'native-v1');
await writeFile(sharedLibraryPath, 'shared-v1');
await writeFile(path.join(grammarRoot, 'tree-sitter-fixture.wasm'), 'wasm-v1');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
let hashedBytes = 0;
let runtimeArtifactHashes = 0;
const resolve = () =>
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onHashedInput: (input) => {
hashedBytes += input.bytes;
if (input.kind === 'runtime-artifact') runtimeArtifactHashes += 1;
},
});
const first = resolve();
expect(first.dependencyRuntime.artifactCount).toBe(3);
expect(runtimeArtifactHashes).toBe(3);
expect(hashedBytes).toBeGreaterThan(0);
expect(analyzerRunnerIdentitiesEqual(structuredClone(first), first)).toBe(true);
expect(analyzerRunnerIdentitiesEqual({ ...first, schemaVersion: 1 }, first)).toBe(false);
// The second resolver call models a fresh status/analyze process: the
// persistent cache is reloaded from disk, and unchanged payload bytes are
// never read even though both build/dependency inventories are validated.
hashedBytes = 0;
runtimeArtifactHashes = 0;
expect(resolve()).toEqual(first);
expect(runtimeArtifactHashes).toBe(0);
expect(hashedBytes).toBe(0);
await writeFile(
path.join(fixture.dbPath, 'package-lock.json'),
'{"name":"fixture-analyzer","lockfileVersion":4}\n',
);
const lockChanged = resolve();
expect(lockChanged.build.digest).toBe(first.build.digest);
expect(lockChanged.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
expect(analyzerRunnerIdentitiesEqual(lockChanged, first)).toBe(false);
expect(runtimeArtifactHashes).toBe(0);
await writeFile(path.join(fixture.dbPath, 'package-lock.json'), originalLock);
await writeFile(nativePath, 'native-v2');
runtimeArtifactHashes = 0;
const nativeChanged = resolve();
expect(nativeChanged.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
expect(runtimeArtifactHashes).toBe(1);
await writeFile(nativePath, 'native-v1');
await writeFile(sharedLibraryPath, 'shared-v2');
const sharedLibraryChanged = resolve();
expect(sharedLibraryChanged.dependencyRuntime.digest).not.toBe(
first.dependencyRuntime.digest,
);
const modelCache = path.join(fixture.dbPath, '.cache', 'models');
await mkdir(modelCache, { recursive: true });
await writeFile(path.join(modelCache, 'weights.bin'), 'large-model-placeholder');
runtimeArtifactHashes = 0;
const cacheChanged = resolve();
expect(cacheChanged.dependencyRuntime).toEqual(sharedLibraryChanged.dependencyRuntime);
expect(runtimeArtifactHashes).toBe(0);
// A corrupt cache is fail-closed: valid inventory stats cannot rescue
// unverifiable cached digests, so every expensive artifact is rehashed.
const [cacheFile] = await readdir(cacheDirectory);
await writeFile(path.join(cacheDirectory, cacheFile), '{"payload":{},"checksum":"bad"}\n');
_clearAnalyzerIdentityProcessCacheForTests();
runtimeArtifactHashes = 0;
hashedBytes = 0;
resolve();
expect(runtimeArtifactHashes).toBe(3);
expect(hashedBytes).toBeGreaterThan(0);
} finally {
await fixture.cleanup();
}
});
it('reuses the secure runtime cache across isolated HOME and GITNEXUS_HOME values', async () => {
const fixture = await createTempDir();
const previous = {
HOME: process.env.HOME,
GITNEXUS_HOME: process.env.GITNEXUS_HOME,
TMPDIR: process.env.TMPDIR,
XDG_RUNTIME_DIR: process.env.XDG_RUNTIME_DIR,
};
const restore = (name: keyof typeof previous): void => {
const value = previous[name];
if (value === undefined) delete process.env[name];
else process.env[name] = value;
};
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const grammarRoot = path.join(fixture.dbPath, 'vendor', 'tree-sitter-fixture');
const nativePath = path.join(
grammarRoot,
'prebuilds',
`${process.platform}-${process.arch}`,
'tree-sitter-fixture.node',
);
const tempRoot = path.join(fixture.dbPath, 'runtime-cache-root');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(path.dirname(nativePath), { recursive: true });
await mkdir(tempRoot, { mode: 0o700 });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(grammarRoot, 'package.json'),
'{"name":"tree-sitter-fixture","version":"1.0.0"}\n',
);
await writeFile(nativePath, Buffer.alloc(2 * 1024 * 1024, 0x5a));
delete process.env.XDG_RUNTIME_DIR;
process.env.TMPDIR = tempRoot;
process.env.HOME = path.join(fixture.dbPath, 'home-a');
process.env.GITNEXUS_HOME = path.join(fixture.dbPath, 'gitnexus-home-a');
let runtimeHashes = 0;
let hashedBytes = 0;
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
onHashedInput: (input) => {
if (input.kind === 'runtime-artifact') runtimeHashes += 1;
hashedBytes += input.bytes;
},
});
expect(runtimeHashes).toBe(1);
expect(hashedBytes).toBeGreaterThanOrEqual(2 * 1024 * 1024);
process.env.HOME = path.join(fixture.dbPath, 'home-b');
process.env.GITNEXUS_HOME = path.join(fixture.dbPath, 'gitnexus-home-b');
_clearAnalyzerIdentityProcessCacheForTests();
runtimeHashes = 0;
hashedBytes = 0;
let cacheMissWalks = 0;
let cacheMissReads = 0;
const second = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
onHashedInput: (input) => {
if (input.kind === 'runtime-artifact') runtimeHashes += 1;
hashedBytes += input.bytes;
},
onCacheMissWork: (input) => {
if (input.kind === 'directory-walk') cacheMissWalks += 1;
else cacheMissReads += 1;
},
});
expect(second).toEqual(first);
expect(runtimeHashes).toBe(0);
expect(hashedBytes).toBe(0);
expect(cacheMissWalks).toBe(0);
expect(cacheMissReads).toBe(0);
} finally {
restore('HOME');
restore('GITNEXUS_HOME');
restore('TMPDIR');
restore('XDG_RUNTIME_DIR');
await fixture.cleanup();
}
});
it('keeps warm identities stable when unrelated siblings churn in a shared parent', async () => {
const fixture = await createTempDir();
let unrelated: Awaited<ReturnType<typeof createTempDir>> | null = null;
try {
const modulePath = path.join(fixture.dbPath, 'src', 'core', 'analyzer.ts');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
unrelated = await createTempDir();
_clearAnalyzerIdentityProcessCacheForTests();
let cacheMissWork = 0;
const second = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheMissWork: () => {
cacheMissWork += 1;
},
});
expect(second).toEqual(first);
expect(cacheMissWork).toBe(0);
} finally {
if (unrelated) await unrelated.cleanup();
await fixture.cleanup();
}
});
it('invalidates an absent path guard when a nearer ancestor package lock appears', async () => {
const fixture = await createTempDir();
try {
const packageRoot = path.join(fixture.dbPath, 'packages', 'fixture-analyzer');
const modulePath = path.join(packageRoot, 'src', 'core', 'analyzer.ts');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
const ancestorLock = path.join(fixture.dbPath, 'package-lock.json');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(packageRoot, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
const withoutLock = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(withoutLock.dependencyRuntime.lockfilePath).toBeNull();
await writeFile(ancestorLock, '{"lockfileVersion":3}\n');
const withLock = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(withLock.dependencyRuntime.lockfilePath).toBe(ancestorLock);
expect(withLock.dependencyRuntime.digest).not.toBe(withoutLock.dependencyRuntime.digest);
} finally {
await fixture.cleanup();
}
});
it('invalidates an absent path guard when a nearer dependency shadows a hoisted one', async () => {
const fixture = await createTempDir();
try {
const packageRoot = path.join(fixture.dbPath, 'packages', 'fixture-analyzer');
const modulePath = path.join(packageRoot, 'src', 'core', 'analyzer.ts');
const hoistedRoot = path.join(fixture.dbPath, 'node_modules', 'runtime-package');
const nearerRoot = path.join(fixture.dbPath, 'packages', 'node_modules', 'runtime-package');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(hoistedRoot, { recursive: true });
await writeFile(
path.join(packageRoot, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(hoistedRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
await writeFile(path.join(hoistedRoot, 'runtime.js'), 'export const source = "hoisted";\n');
const hoisted = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
await mkdir(nearerRoot, { recursive: true });
await writeFile(
path.join(nearerRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '2.0.0' }),
);
await writeFile(path.join(nearerRoot, 'runtime.js'), 'export const source = "nearer";\n');
const shadowed = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(shadowed.dependencyRuntime.digest).not.toBe(hoisted.dependencyRuntime.digest);
} finally {
await fixture.cleanup();
}
});
it('invalidates a warm identity when an intermediate package symlink is retargeted', async () => {
const fixture = await createTempDir();
try {
const packageRoot = path.join(fixture.dbPath, 'fixture-analyzer');
const modulePath = path.join(packageRoot, 'src', 'core', 'analyzer.ts');
const nodeModulesRoot = path.join(packageRoot, 'node_modules');
const packageLink = path.join(nodeModulesRoot, 'runtime-package');
const storeA = path.join(fixture.dbPath, 'store-a');
const storeB = path.join(fixture.dbPath, 'store-b');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(nodeModulesRoot, { recursive: true });
await mkdir(storeA);
await mkdir(storeB);
await writeFile(
path.join(packageRoot, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(storeA, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
// Keep the final candidate manifest's inode and stat state identical so
// only an exact lexical-component guard can observe the retarget.
await link(path.join(storeA, 'package.json'), path.join(storeB, 'package.json'));
await writeFile(path.join(storeA, 'runtime.js'), 'export const source = "a";\n');
await writeFile(path.join(storeB, 'runtime.js'), 'export const source = "b changed";\n');
try {
await symlink(storeA, packageLink, 'dir');
} catch (error) {
if (['EPERM', 'EACCES'].includes((error as NodeJS.ErrnoException).code ?? '')) return;
throw error;
}
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
await unlink(packageLink);
await symlink(storeB, packageLink, 'dir');
_clearAnalyzerIdentityProcessCacheForTests();
let cacheMissWork = 0;
const retargeted = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheMissWork: () => {
cacheMissWork += 1;
},
});
expect(retargeted.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
expect(cacheMissWork).toBeGreaterThan(0);
} finally {
await fixture.cleanup();
}
});
it('invalidates direct-stat guards for build, topology, and artifact inventory changes', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const grammarRoot = path.join(fixture.dbPath, 'vendor', 'tree-sitter-fixture');
const artifactDir = path.join(
grammarRoot,
'prebuilds',
`${process.platform}-${process.arch}`,
);
const nativePath = path.join(artifactDir, 'tree-sitter-fixture.node');
const addedNativePath = path.join(artifactDir, 'tree-sitter-extra.node');
const manifestPath = path.join(grammarRoot, 'package.json');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(artifactDir, { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(manifestPath, '{"name":"tree-sitter-fixture","version":"1.0.0"}\n');
await writeFile(nativePath, 'native-v1');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.dependencyRuntime.artifactCount).toBe(1);
let walks = 0;
let reads = 0;
let hashes = 0;
const resolve = () =>
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheMissWork: (input) => {
if (input.kind === 'directory-walk') walks += 1;
else reads += 1;
},
onHashedInput: () => {
hashes += 1;
},
});
const reset = () => {
walks = 0;
reads = 0;
hashes = 0;
};
expect(resolve()).toEqual(first);
expect({ walks, reads, hashes }).toEqual({ walks: 0, reads: 0, hashes: 0 });
await writeFile(path.join(sourceRoot, 'added.ts'), 'export const added = true;\n');
reset();
const buildAdded = resolve();
expect(buildAdded.build.digest).not.toBe(first.build.digest);
expect(walks).toBeGreaterThan(0);
await writeFile(addedNativePath, 'native-extra');
reset();
const artifactAdded = resolve();
expect(artifactAdded.dependencyRuntime.artifactCount).toBe(2);
expect(artifactAdded.dependencyRuntime.digest).not.toBe(buildAdded.dependencyRuntime.digest);
expect(walks).toBeGreaterThan(0);
expect(hashes).toBe(1);
await writeFile(manifestPath, '{"name":"tree-sitter-fixture","version":"2.0.0"}\n');
reset();
const manifestChanged = resolve();
expect(manifestChanged.dependencyRuntime.digest).not.toBe(
artifactAdded.dependencyRuntime.digest,
);
expect(reads).toBeGreaterThan(0);
await unlink(nativePath);
reset();
const artifactDeleted = resolve();
expect(artifactDeleted.dependencyRuntime.artifactCount).toBe(1);
expect(artifactDeleted.dependencyRuntime.digest).not.toBe(
manifestChanged.dependencyRuntime.digest,
);
expect(walks).toBeGreaterThan(0);
} finally {
await fixture.cleanup();
}
});
it('keeps same-name/version dependency instances distinct by package-root locator', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const nestedA = path.join(
fixture.dbPath,
'node_modules',
'parent-a',
'node_modules',
'duplicate',
);
const nestedB = path.join(
fixture.dbPath,
'node_modules',
'parent-b',
'node_modules',
'duplicate',
);
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(nestedA, { recursive: true });
await mkdir(nestedB, { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'parent-a': '1.0.0', 'parent-b': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
for (const parentName of ['parent-a', 'parent-b']) {
await writeFile(
path.join(fixture.dbPath, 'node_modules', parentName, 'package.json'),
JSON.stringify({
name: parentName,
version: '1.0.0',
dependencies: { duplicate: '1.0.0' },
}),
);
}
for (const nestedRoot of [nestedA, nestedB]) {
await writeFile(
path.join(nestedRoot, 'package.json'),
JSON.stringify({ name: 'duplicate', version: '1.0.0' }),
);
await writeFile(path.join(nestedRoot, 'runtime.wasm'), 'same-runtime-bytes');
}
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.dependencyRuntime.packageCount).toBe(5);
expect(first.dependencyRuntime.artifactCount).toBe(2);
const [cacheFile] = await readdir(cacheDirectory);
const envelope = JSON.parse(await readFile(path.join(cacheDirectory, cacheFile), 'utf8')) as {
payload: { artifactEntries: Array<{ canonicalPath: string }> };
};
const artifactLocators = envelope.payload.artifactEntries.map((entry) => entry.canonicalPath);
expect(artifactLocators).toEqual(
expect.arrayContaining([
expect.stringContaining('node_modules/parent-a/node_modules/duplicate/runtime.wasm'),
expect.stringContaining('node_modules/parent-b/node_modules/duplicate/runtime.wasm'),
]),
);
expect(new Set(artifactLocators).size).toBe(2);
await writeFile(
path.join(nestedB, 'package.json'),
JSON.stringify({ name: 'duplicate', version: '1.0.0', instance: 'parent-b' }),
);
const changedOneInstance = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(changedOneInstance.dependencyRuntime.packageCount).toBe(5);
expect(changedOneInstance.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
} finally {
await fixture.cleanup();
}
});
it('discovers runtime artifacts in every resolved package without an allowlist', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const dependencyRoot = path.join(fixture.dbPath, 'node_modules', 'ordinary-runtime');
const nativePath = path.join(dependencyRoot, 'build', 'addon.node');
const wasmPath = path.join(dependencyRoot, 'codec', 'runtime.wasm');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(path.dirname(nativePath), { recursive: true });
await mkdir(path.dirname(wasmPath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'ordinary-runtime': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(dependencyRoot, 'package.json'),
JSON.stringify({ name: 'ordinary-runtime', version: '1.0.0' }),
);
await writeFile(nativePath, 'native-v1');
await writeFile(wasmPath, 'wasm-v1');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.dependencyRuntime).toMatchObject({ packageCount: 2, artifactCount: 2 });
await writeFile(nativePath, 'native-v2-with-a-different-size');
const changed = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(changed.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
} finally {
await fixture.cleanup();
}
});
it('tracks generic runtime directories and filename-mismatched native payloads on cold and warm scans', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const dependencyRoot = path.join(fixture.dbPath, 'node_modules', 'runtime-package');
const foreignPlatform = process.platform === 'linux' ? 'darwin' : 'linux';
const foreignArchitecture = process.arch === 'x64' ? 'arm64' : 'x64';
const payloadPaths = [
path.join(dependencyRoot, '.cache', 'generated-loader.js'),
path.join(
dependencyRoot,
'cache',
`${foreignPlatform}-${foreignArchitecture}`,
'addon.node',
),
path.join(dependencyRoot, 'models', 'runtime-model.wasm'),
path.join(
dependencyRoot,
'prebuilds',
`${foreignPlatform}-${foreignArchitecture}`,
'foreign-target.node',
),
path.join(
dependencyRoot,
'codec',
`runtime-${foreignPlatform}-${foreignArchitecture}.wasm`,
),
];
await mkdir(path.dirname(modulePath), { recursive: true });
for (const payloadPath of payloadPaths) {
await mkdir(path.dirname(payloadPath), { recursive: true });
}
await mkdir(path.join(dependencyRoot, '.git'), { recursive: true });
await mkdir(path.join(dependencyRoot, '.hg'), { recursive: true });
await mkdir(path.join(dependencyRoot, '.svn'), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(dependencyRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
await writeFile(path.join(dependencyRoot, '.git', 'config'), 'ignored-vcs-state');
await writeFile(path.join(dependencyRoot, '.hg', 'dirstate'), 'ignored-vcs-state');
await writeFile(path.join(dependencyRoot, '.svn', 'wc.db'), 'ignored-vcs-state');
for (const payloadPath of payloadPaths) await writeFile(payloadPath, 'payload-v1');
const baseline = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: path.join(fixture.dbPath, 'baseline-cache'),
});
expect(baseline.dependencyRuntime.artifactCount).toBe(payloadPaths.length);
for (const payloadPath of payloadPaths) {
await writeFile(payloadPath, `payload-v2:${path.basename(payloadPath)}`);
}
const warmCacheDirectory = path.join(fixture.dbPath, 'warm-cache');
let runtimeHashes = 0;
const coldAfterMutation = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: warmCacheDirectory,
onHashedInput: (input) => {
if (input.kind === 'runtime-artifact') runtimeHashes += 1;
},
});
expect(coldAfterMutation.dependencyRuntime.digest).not.toBe(
baseline.dependencyRuntime.digest,
);
expect(runtimeHashes).toBe(payloadPaths.length);
runtimeHashes = 0;
expect(
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: warmCacheDirectory,
onHashedInput: (input) => {
if (input.kind === 'runtime-artifact') runtimeHashes += 1;
},
}),
).toEqual(coldAfterMutation);
expect(runtimeHashes).toBe(0);
for (const payloadPath of payloadPaths) {
await writeFile(payloadPath, `payload-v3-with-new-bytes:${path.basename(payloadPath)}`);
}
runtimeHashes = 0;
const warmAfterMutation = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: warmCacheDirectory,
onHashedInput: (input) => {
if (input.kind === 'runtime-artifact') runtimeHashes += 1;
},
});
expect(warmAfterMutation.dependencyRuntime.digest).not.toBe(
coldAfterMutation.dependencyRuntime.digest,
);
expect(runtimeHashes).toBe(payloadPaths.length);
} finally {
await fixture.cleanup();
}
});
it('fails closed when a resolved-package artifact walk exceeds its depth bound', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const dependencyRoot = path.join(fixture.dbPath, 'node_modules', 'deep-runtime');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(dependencyRoot, { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'deep-runtime': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(dependencyRoot, 'package.json'),
JSON.stringify({ name: 'deep-runtime', version: '1.0.0' }),
);
let cursor = dependencyRoot;
for (let depth = 0; depth < 66; depth += 1) {
cursor = path.join(cursor, 'd');
await mkdir(cursor);
}
expect(() =>
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory: path.join(fixture.dbPath, 'identity-cache'),
}),
).toThrow(/payload scan exceeded depth 64/);
} finally {
await fixture.cleanup();
}
});
it('stable-reads symlinked package locks and rejects broken lock links', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const lockTarget = path.join(fixture.dbPath, 'actual-package-lock.json');
const lockLink = path.join(fixture.dbPath, 'package-lock.json');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(lockTarget, '{"lockfileVersion":3}\n');
try {
await symlink(lockTarget, lockLink, 'file');
} catch (error) {
if (['EPERM', 'EACCES'].includes((error as NodeJS.ErrnoException).code ?? '')) return;
throw error;
}
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.dependencyRuntime.lockfilePath).toBe(lockLink);
await writeFile(lockTarget, '{"lockfileVersion":4,"changed":true}\n');
const targetChanged = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(targetChanged.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
await unlink(lockLink);
await symlink(path.join(fixture.dbPath, 'missing-lock-target.json'), lockLink, 'file');
expect(() =>
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, { cacheDirectory }),
).toThrow(/package lock symbolic link does not resolve to a file/);
} finally {
await fixture.cleanup();
}
});
it('uses one final warm validation pass and notices immediate file and topology changes', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const addedPath = path.join(sourceRoot, 'added-at-boundary.ts');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
let validationPasses = 0;
expect(
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheValidationPass: () => {
validationPasses += 1;
},
}),
).toEqual(first);
expect(validationPasses).toBe(1);
validationPasses = 0;
let changedFile = false;
const fileChanged = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheValidationPass: () => {
validationPasses += 1;
if (!changedFile) {
changedFile = true;
writeFileSync(modulePath, 'export const analyzer = 200;\n');
}
},
});
expect(fileChanged.build.digest).not.toBe(first.build.digest);
expect(validationPasses).toBe(2);
validationPasses = 0;
let changedTopology = false;
const topologyChanged = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onCacheValidationPass: () => {
validationPasses += 1;
if (!changedTopology) {
changedTopology = true;
writeFileSync(addedPath, 'export const added = true;\n');
}
},
});
expect(topologyChanged.build.digest).not.toBe(fileChanged.build.digest);
expect(validationPasses).toBe(2);
} finally {
await fixture.cleanup();
}
});
it('keeps warm-cache work at zero and materially below cold-path latency', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const payloadPath = path.join(sourceRoot, 'large-runtime-source.bin');
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
const payloadBytes = 32 * 1024 * 1024;
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(payloadPath, Buffer.alloc(payloadBytes, 0x61));
let coldHashedBytes = 0;
let coldTopologyWork = 0;
let coldValidationPasses = 0;
const coldStarted = performance.now();
const coldIdentity = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onHashedInput: ({ bytes }) => {
coldHashedBytes += bytes;
},
onCacheMissWork: () => {
coldTopologyWork += 1;
},
onCacheValidationPass: () => {
coldValidationPasses += 1;
},
});
const coldDurationMs = performance.now() - coldStarted;
expect(coldHashedBytes).toBeGreaterThanOrEqual(payloadBytes);
expect(coldTopologyWork).toBeGreaterThan(0);
expect(coldValidationPasses).toBe(1);
const warmDurationsMs: number[] = [];
for (let iteration = 0; iteration < 5; iteration += 1) {
let warmHashedBytes = 0;
let warmTopologyWork = 0;
let warmValidationPasses = 0;
const warmStarted = performance.now();
const warmIdentity = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onHashedInput: ({ bytes }) => {
warmHashedBytes += bytes;
},
onCacheMissWork: () => {
warmTopologyWork += 1;
},
onCacheValidationPass: () => {
warmValidationPasses += 1;
},
});
warmDurationsMs.push(performance.now() - warmStarted);
expect(warmIdentity).toEqual(coldIdentity);
expect(warmHashedBytes).toBe(0);
expect(warmTopologyWork).toBe(0);
expect(warmValidationPasses).toBe(1);
}
warmDurationsMs.sort((a, b) => a - b);
const medianWarmDurationMs = warmDurationsMs[Math.floor(warmDurationsMs.length / 2)];
expect(medianWarmDurationMs).toBeLessThan(coldDurationMs);
} finally {
await fixture.cleanup();
}
});
it('uses non-ambiguous framing and treats the invoked entrypoint as diagnostic', async () => {
const leftOldEncoding = Buffer.concat([
Buffer.from('a'),
Buffer.from([0]),
Buffer.from('b\0c'),
Buffer.from([0]),
]);
const rightOldEncoding = Buffer.concat([
Buffer.from('a\0b'),
Buffer.from([0]),
Buffer.from('c'),
Buffer.from([0]),
]);
expect(leftOldEncoding).toEqual(rightOldEncoding);
expect(_hashAnalyzerIdentityFramesForTests([['entry', 'a', 'b\0c']])).not.toBe(
_hashAnalyzerIdentityFramesForTests([['entry', 'a\0b', 'c']]),
);
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
const options = { cacheDirectory: path.join(fixture.dbPath, 'identity-cache') };
const identity = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, options);
const alternateEntrypoint = {
...identity,
invokedArtifact: {
path: path.join(sourceRoot, 'server', 'analyze-worker.ts'),
digest: `sha256:${'a'.repeat(64)}`,
},
};
expect(analyzerRunnerIdentitiesEqual(alternateEntrypoint, identity)).toBe(true);
expect(normalizeAnalyzerRunnerIdentityForComparison(alternateEntrypoint)).toEqual(
normalizeAnalyzerRunnerIdentityForComparison(identity),
);
expect(normalizeAnalyzerRunnerIdentityForComparison({ schemaVersion: 4 })).toBeNull();
expect(
analyzerRunnerIdentitiesEqual(
{ ...alternateEntrypoint, invokedArtifact: { path: '', digest: 'bad' } },
identity,
),
).toBe(false);
await writeFile(
path.join(sourceRoot, 'new-semantic-input.ts'),
'export const changed = 1;\n',
);
expect(() =>
finalizeAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, identity, options),
).toThrow(/changed during analysis/);
} finally {
await fixture.cleanup();
}
});
it('captures before loading and rejects a replacement that races module evaluation', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
await mkdir(path.dirname(modulePath), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
'{"name":"fixture-analyzer","version":"9.8.7"}\n',
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
const options = { cacheDirectory: path.join(fixture.dbPath, 'identity-cache') };
const prepared = await captureAnalyzerIdentityBeforeLoad(
pathToFileURL(modulePath).href,
async () => {
// Change the size as well as the bytes so filesystems with coarse
// timestamp granularity cannot make this race regression flaky.
await writeFile(modulePath, 'export const analyzer = 200;\n');
return 'loaded-after-replacement';
},
options,
);
expect(prepared.loaded).toBe('loaded-after-replacement');
expect(() =>
finalizeAnalyzerRunnerIdentity(
pathToFileURL(modulePath).href,
prepared.runnerIdentity,
options,
),
).toThrow(/changed during analysis/);
} finally {
await fixture.cleanup();
}
});
it('content-addresses every resolved package payload and reuses it without byte reads', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const dependencyRoot = path.join(fixture.dbPath, 'node_modules', 'runtime-package');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(dependencyRoot, { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(dependencyRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
const payloadNames = [
'index.js',
'legacy.cjs',
'module.mjs',
'data.json',
'addon.node',
'runtime.wasm',
'extensionless',
'runtime-config.txt',
];
for (const name of payloadNames)
await writeFile(path.join(dependencyRoot, name), `${name}:v1`);
const cacheDirectory = path.join(fixture.dbPath, 'identity-cache');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(first.dependencyRuntime.artifactCount).toBe(payloadNames.length);
let warmBytes = 0;
expect(
resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
onHashedInput: ({ bytes }) => {
warmBytes += bytes;
},
}),
).toEqual(first);
expect(warmBytes).toBe(0);
let priorDigest = first.dependencyRuntime.digest;
for (const name of payloadNames) {
await writeFile(path.join(dependencyRoot, name), `${name}:v2-with-new-bytes`);
const changed = resolveAnalyzerRunnerIdentity(pathToFileURL(modulePath).href, {
cacheDirectory,
});
expect(changed.dependencyRuntime.digest).not.toBe(priorDigest);
priorDigest = changed.dependencyRuntime.digest;
}
} finally {
await fixture.cleanup();
}
});
it('enforces iterative build, package, edge, entry, payload, byte, depth, and resolution bounds', async () => {
const fixture = await createTempDir();
try {
const sourceRoot = path.join(fixture.dbPath, 'src');
const modulePath = path.join(sourceRoot, 'core', 'analyzer.ts');
const dependencyRoot = path.join(fixture.dbPath, 'node_modules', 'runtime-package');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(path.join(dependencyRoot, 'deep', 'deeper'), { recursive: true });
await writeFile(
path.join(fixture.dbPath, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0', missing: '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(path.join(sourceRoot, 'extra.ts'), 'export const extra = true;\n');
await mkdir(path.join(sourceRoot, 'nested', 'deeper'), { recursive: true });
await writeFile(path.join(sourceRoot, 'nested', 'deeper', 'leaf.ts'), 'export {};\n');
await writeFile(
path.join(dependencyRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
await writeFile(path.join(dependencyRoot, 'a.js'), 'a');
await writeFile(path.join(dependencyRoot, 'b.json'), '{}');
await writeFile(path.join(dependencyRoot, 'deep', 'deeper', 'c.mjs'), 'c');
const url = pathToFileURL(modulePath).href;
let sequence = 0;
const bounded = (traversalLimits: Record<string, number>) => () =>
resolveAnalyzerRunnerIdentity(url, {
cacheDirectory: path.join(fixture.dbPath, `cache-${sequence++}`),
traversalLimits,
});
expect(bounded({ buildEntries: 1 })).toThrow(/build scan exceeded 1 entries/);
expect(bounded({ buildDepth: 1 })).toThrow(/build scan exceeded depth 1/);
expect(bounded({ buildBytes: 1 })).toThrow(/build scan exceeded 1 bytes/);
expect(bounded({ runtimePackages: 1 })).toThrow(/dependency graph exceeded 1 packages/);
expect(bounded({ runtimeEdges: 1 })).toThrow(/dependency graph exceeded 1 edges/);
expect(bounded({ runtimeEntries: 1 })).toThrow(/payload scan exceeded 1 entries/);
expect(bounded({ runtimeDepth: 1 })).toThrow(/payload scan exceeded depth 1/);
expect(bounded({ runtimePayloads: 1 })).toThrow(/payload scan exceeded 1 payloads/);
expect(bounded({ runtimeBytes: 1 })).toThrow(/runtime scan exceeded 1 bytes/);
expect(bounded({ resolutionAncestors: 1 })).toThrow(/exceeded 1 ancestors/);
} finally {
await fixture.cleanup();
}
});
it('persists the same receipt to both metadata mirrors on full and incremental runs', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, skipSkills: true },
{ onProgress: () => {} },
);
const { storagePath } = getStoragePaths(repo.dbPath);
const first = await loadMeta(storagePath);
const expectedIdentity = resolveAnalyzerRunnerIdentity(
pathToFileURL(path.resolve(__dirname, '../../src/core/run-analyze.ts')).href,
);
expect(first?.runnerIdentity).toEqual(expectedIdentity);
expect(first?.runnerIdentity).toMatchObject({
schemaVersion: 4,
cliVersion: expect.any(String),
invokedArtifact: { digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/) },
build: { digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/) },
dependencyRuntime: { digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/) },
});
if (!first?.runnerIdentity) throw new Error('analysis did not persist a runner identity');
const legacyMeta = {
...first,
runnerIdentity: { ...first.runnerIdentity, schemaVersion: 1 },
} as unknown as RepoMeta;
await saveMeta(storagePath, legacyMeta);
const upgraded = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, skipSkills: true },
{ onProgress: () => {} },
);
expect(upgraded.alreadyUpToDate).toBeUndefined();
expect((await loadMeta(storagePath))?.runnerIdentity).toEqual(first.runnerIdentity);
const changedPath = path.join(repo.dbPath, 'src', 'logger.ts');
const before = await readFile(changedPath, 'utf8');
await writeFile(changedPath, `${before}\n// force incremental identity restamp\n`, 'utf8');
const incremental = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, skipSkills: true },
{ onProgress: () => {} },
);
expect(incremental.alreadyUpToDate).toBeUndefined();
const second = await loadMeta(storagePath);
expect(second?.runnerIdentity).toEqual(first?.runnerIdentity);
const primary = JSON.parse(
await readFile(path.join(storagePath, 'gitnexus.json'), 'utf8'),
) as { runnerIdentity?: unknown };
const legacy = JSON.parse(await readFile(path.join(storagePath, 'meta.json'), 'utf8')) as {
runnerIdentity?: unknown;
};
expect(primary.runnerIdentity).toEqual(second?.runnerIdentity);
expect(legacy.runnerIdentity).toEqual(second?.runnerIdentity);
} finally {
await repo.cleanup();
}
}, 300_000);
});