The Run Claude Code Review step passed an invalid PR ref
(owner/repo/pull/N) which gh interprets as a branch name, causing
early gh pr view failures. More importantly, the prompt omitted
--comment, so the code-review plugin only displayed findings in
terminal output and never invoked gh pr comment to post to the PR.
Switch to a full PR URL and add --comment so the plugin posts the
review during the session, which also routes around upstream bugs
anthropics/claude-code-action#1061 and #1087 where the action's
post-step capture can silently drop output on issue_comment triggers.
Patch addressing two of the still-open changes-requested findings on PR
#1479, rebased onto the current feat/incremental-indexing head. F3
(parser fingerprint in the cache key), F5 (atomic saveMeta), and F6
(AGENTS.md phrasing) were already handled on the branch, so the
corresponding parts of the original patch were dropped as redundant.
F1 (Blocker) — Cross-file edges between unchanged files
Adds `computeEffectiveWriteSet(graph, toWriteSet)` to
subgraph-extract.ts: a single pass over the new graph's edges that
pulls the unchanged-side file of every writable-boundary-crossing
edge into the write set. run-analyze composes it ON TOP of the
existing importer-BFS expansion and feeds the combined set to BOTH
`deleteNodesForFile` and `extractChangedSubgraph`, so the delete
cascade and the writeback subgraph cover identical files (asymmetry
would leave stale rows or PK-conflict at COPY time). The BFS reads
IMPORTS from the pre-pipeline DB (catches files that *stopped*
importing a changed file); the edge walk reads the new graph
(catches refined CALLS edges the pre-run DB couldn't predict, e.g.
a barrel re-export shifting a symbol from B to D). `extractChangedSubgraph`
stays a pure filter — all expansion is the orchestrator's job.
F4 (Medium) — Restore alphabetical chunk sort
`parseableScanned` is sorted before chunking. Filesystem-scan order
isn't stable enough across runs/platforms (notably macOS APFS) to
keep chunk hashes consistent, so the parse cache thrashes without
it. The pre-existing Ruby cross-file resolution order-dependency the
old comment cited is independent — the sort surfaces it but doesn't
cause it; tracked separately rather than leaving the cache cold.
Tests — incremental-subgraph-extract.test.ts
Locks the F1 invariants: `extractChangedSubgraph` is a pure filter
(includes only the set it's given, plus graph-wide nodes; edges
fire on one writable endpoint), and `computeEffectiveWriteSet`
covers the barrel-re-export scenario, the symmetric edge-into-
changed-file case, the no-boundary-crossed no-op, graph-wide-node
edges, and input-immutability. Supersedes the prior
extractChangedSubgraph-only test file on the branch.
Co-authored-by: Val Vladescu <vvladescu-tb@users.noreply.github.com>
* fix(augment): add CONTAINS fallback when FTS indexes unavailable
When the MCP server holds the KuzuDB write lock, the augment CLI opens
the DB read-only. FTS indexes cannot be created in read-only mode, so
searchFTSFromLbug returns ftsAvailable=false and an empty results array.
The existing early-return path silently produced no enrichment.
Add a Cypher name CONTAINS fallback that fires only when ftsAvailable is
false and BM25 produced no symbol matches. This covers the read-only DB
case (concurrent MCP server) and the first-run case (indexes not yet
built). The fallback is wrapped in .catch(() => []) and cannot throw.
When FTS indexes exist, this branch is never reached — behaviour is
unchanged for users without a concurrent MCP server.
* fix(augment): guard against CONTAINS '' and add no-FTS test coverage
Blocker 1 — CONTAINS '' on whitespace-leading patterns:
pattern.split(/\s+/)[0] returns "" when the input has leading whitespace
(e.g. " ".split(/\s+/) → ["", ""]). In Kuzu, CONTAINS '' matches every
node with a name property, injecting arbitrary graph nodes into LLM context.
Fix: trim() before split, then guard on !firstWord || firstWord.length < 2.
No behaviour change for normal non-empty patterns.
Blocker 2 — zero test coverage on the FTS-unavailable code path:
The new CONTAINS fallback block (engine.ts lines 146-166) was exercised by
no existing test — all existing tests run with FTS indexes built. A second
withTestLbugDB fixture is added with no ftsIndexes, forcing searchFTSFromLbug
to return ftsAvailable: false, and asserts:
1. augment('login', ...) returns non-empty enrichment (fallback works)
2. augment(' ', ...) returns '' (CONTAINS '' guard holds)
3. augment('nxyz_notfound', ...) returns '' (no matching nodes)
4. executeQuery throwing returns '' (.catch(() => []) path)
* fix(augment): extend CONTAINS '' guard to FTS happy path and consolidate
The same split(/\s+/)[0] bug existed at line 125 (BM25 symbol filter,
FTS-available path) — a leading-whitespace pattern produced CONTAINS ''
there too, matching every node in BM25-matched files.
Fix: hoist patternFirstWord computation with trim() and the length guard
to the top of augment(), before any DB interaction. Both CONTAINS sites
(BM25 symbol filter and CONTAINS fallback) now use the single pre-validated
value. No behaviour change for normal patterns; the guard fires once for
all callers instead of being duplicated.
Also tighten the whitespace test in the no-FTS suite from 3 spaces to
4 spaces so it unambiguously exercises the patternFirstWord guard rather
than straddling the outer pattern.length < 3 boundary.
* test(augment): negative-safety test for ftsAvailable=true gate
Asserts the CONTAINS fallback does NOT fire when FTS is available but
BM25 returns zero results. Pins the safety property promised by the PR
description: behavior is unchanged for users without the read-only-DB
condition.
If anyone later loosens the gate to `symbolMatches.length === 0` alone,
this test fails.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* feat(cli): add --skip-skills and --index-only flags to analyze command
The `installSkills()` call in `generateAIContextFiles()` runs
unconditionally, injecting 6 skill files into `.claude/skills/gitnexus/`
even when `--skip-agents-md` is passed. This is problematic for bulk
indexing operations on read-only mirrors or third-party repos.
Add two new flags:
- `--skip-skills`: suppress standard GitNexus skill file injection
- `--index-only`: pure index mode that suppresses all file injection
(AGENTS.md, CLAUDE.md, and skills), writing only to `.gitnexus/`
This gives users three levels of control:
- `--skip-agents-md` — suppress only root context files
- `--skip-skills` — suppress only skill injection
- `--index-only` — suppress everything (pure indexing)
Discovery context: while bulk-indexing 176 repos with
`--skip-agents-md`, all 144 indexed repos were contaminated with
`.claude/skills/gitnexus/` files requiring manual cleanup.
* fix(cli): address PR #742 review — gate community skills, drop dangling refs, add tests
Bot review (#742) flagged three issues with the original commit:
1. `--index-only --skills` still wrote community-derived skill files
to `.claude/skills/generated/`. The `--skills` branch in analyze.ts
was not gated by `skipAll`, so the "skip all file injection" contract
was violated. Gate `generateSkillFiles()` with `!skipAll` so
`--index-only` truly wins over `--skills`.
2. `--skip-skills` without `--skip-agents-md` produced AGENTS.md /
CLAUDE.md that still referenced `.claude/skills/gitnexus/*/SKILL.md`
files that were never installed — every agent load incurred 6
failed reads. Pass `skipSkills` through to `generateGitNexusContent()`
and omit the standard-skill rows (and the entire `## CLI` heading
when the table is empty). Community skills, when present via
`--skills`, are unaffected.
3. No filesystem tests for `skipSkills` / `indexOnly`. Add three
regression guards to `test/unit/ai-context.test.ts`:
- `.claude/skills/gitnexus/` is NOT created when skipSkills=true
- Nothing is written when both skipAgentsMd and skipSkills are true
(the resolved-flag state from --index-only)
- AGENTS.md/CLAUDE.md routing table omits standard skill references
when skipSkills=true, but preserves the load-bearing imperative
sections (Always Do / Never Do / Resources)
* test(cli): PR 1485 review follow-ups (help text, gate test, --skip-skills docs)
- Assert --skip-skills and --index-only in analyze --help (skip-git-cli.test.ts).
- Export shouldGenerateCommunitySkillFiles; unit-test index-only+skills gate.
- Clarify --skip-skills does not suppress --skills community files; --index-only for full skip.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(cli): warn when --index-only silently overrides --skills
Address review findings on PR 1485 follow-ups:
- analyze.ts emits a one-line note when both --index-only and --skills
are set, so users see why a pipeline re-index ran with no skill files
written.
- index.ts --skills help text now flags the --index-only override.
- shouldGenerateCommunitySkillFiles JSDoc documents the dual role of
the gate (community skills + AGENTS.md/CLAUDE.md re-generation).
- skip-git-cli.test.ts pins the override-warning surface end-to-end.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(embeddings): forward GITNEXUS_EMBEDDING_DIMS as dimensions in HTTP request body
When GITNEXUS_EMBEDDING_DIMS is set, include it as the `dimensions` field
in the /v1/embeddings request body. This enables Matryoshka-capable models
(OpenAI text-embedding-3-*, Cohere embed-v3, Voyage) to return truncated
vectors at the requested size.
When the env var is unset, the request body remains `{ input, model }` —
no breaking change for backends that reject unknown fields.
Adds 4 unit tests covering both paths (with/without dimensions) on both
the batch embed and single-query embed code paths.
* fix(embeddings): address review findings — strict parseInt, multi-batch test, comment wording
1. Strict parseInt validation: reject non-numeric strings like '1024abc'
by checking /^\d+$/ before parseInt (Finding 1).
2. Add multi-batch test asserting dimensions is forwarded in every fetch
call when inputs exceed batch size (Finding 2).
3. Soften JSDoc comment: backends may ignore or reject the dimensions
field rather than universally ignoring it (Finding 3).
4. Add test for invalid GITNEXUS_EMBEDDING_DIMS values.
---------
Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Addresses the only remaining Claude production-readiness review finding
on PR #1479 (Low-Medium, test-quality only — Claude itself said it does
NOT block merge, but the central PR claim "incremental ≡ full rebuild"
deserves explicit CI coverage rather than implicit trust).
Changes to gitnexus/test/unit/incremental-orchestration.test.ts:
1) Tighten the existing "comment-only edit takes incremental path" test.
- Replace toBeGreaterThan(0) bounds assertions on stats.files and
stats.nodes with exact toBe(firstMeta) per-field equality across
files / nodes / edges / communities / processes. DoD §2.7 calls
out bounds-only assertions as masking regressions that drop half
the graph; this swap closes that gap.
- Rationale: a comment-only edit must change the file content hash
(driving the incremental path) without changing any graph data.
Therefore every stat MUST be identical to the first run. Anything
else is a regression.
2) New test: incremental output is byte-equivalent to a full rebuild.
- Run analyze → comment-only edit → analyze (incremental writeback)
→ analyze --force (full rebuild from same on-disk state).
- Assert files / nodes / edges / communities / processes are exactly
equal across the incremental and the --force passes.
- This is the PR's central correctness contract, now proven by a
test that exercises the real runtime path end-to-end against a
real on-disk LadybugDB.
All 5 orchestration tests pass locally (52s), including the new
equivalence test — every stat field matches exactly between incremental
and --force on the mini-repo fixture.
tsc --noEmit clean.
* fix(server): sanitize repo name to prevent argument injection
Sanitizes the extracted repository name to prevent argument injection during git clone operations and ensures compatibility with various file systems.
1. Strips leading dashes to prevent git command-line argument injection.
2. Replaces unsafe directory characters with underscores.
3. Blocks path traversal segments ('.' and '..') and Windows reserved names.
4. Fixes ReDoS vulnerability in parseRepoNameFromUrl regex.
5. Added unit tests for sanitization and path traversal edge cases.
* fix(server): expand Windows reserved name check to include extensions
- Updated sanitizeRepoName to block Windows reserved names (CON, NUL, etc.) even when they have extensions (e.g., CON.txt).
- Corrected regex and added unit tests for these edge cases to resolve CI failures on Windows.
- Ref: https://github.com/abhigyanpatwari/GitNexus/pull/1305#issuecomment-4407200914
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Bugbot review on commit e23e4400 surfaced two new findings against the
incremental writeback in run-analyze.ts:
HIGH — Incremental BFS misses importers of newly added files.
queryImporters() reads the pre-pipeline DB. For a NEWLY ADDED
file there are no IMPORTS rows pointing to it yet, so unchanged
files whose pre-existing import statements now resolve to the
newcomer keep stale CALLS edges pointing at the OLD resolution
target.
LOW — Deleted files double-counted in filesToDelete.
hashDiff.deleted entries can reappear in writableFiles via the
BFS expansion (queryImporters can return a now-deleted path),
so deleteNodesForFile() ran twice for the same file.
Fixes:
- Add gitnexus/src/core/incremental/shadow-candidates.ts: derive
the pre-existing file paths whose JS/TS module-resolution claim
an added file can steal. Pattern catalogue: same-basename/
different-extension, bare-file-beats-directory-index, and
directory-index-beats-bare-file. Emit both POSIX and Windows
separators because the prior fileHashes map may have been
written from either OS.
- In run-analyze.ts, seed the BFS frontier with shadow candidates
that exist in the prior meta.fileHashes. Their importers — found
via queryImporters — get pulled into the writable set so their
CALLS edges re-resolve against the new file.
- Dedupe filesToDelete via Set to avoid the double-call.
Tests: gitnexus/test/unit/incremental-shadow-candidates.test.ts —
8 cases covering each shadow pattern, separator handling, .d.ts as
a single extension token, deduplication, and the no-self-shadow
invariant. All 40 incremental tests (file-hash, parse-cache,
subgraph-extract, shadow-candidates, orchestration) pass locally.
Note on the third Bugbot finding ("Subgraph edges reference nodes
absent from subgraph"): re-anchored from a prior review pass — the
code at subgraph-extract.ts:48 is unchanged. Already verified as a
false positive: getNodeLabel parses labels from ID strings, CSV
write is by ID, and COPY resolves against the live DB.
Addresses remaining findings on PR #1479 from Claude's re-review of
commit ad7bd31 + verifies the outstanding Bugbot HIGH severity.
1. F1 — Transitive importer expansion (Claude, was Medium-but-noted).
Previous 1-hop importer expansion missed barrel re-export chains
(A imports C, C re-exports B; when B changes, only C was pulled in
— A was left with potentially-stale CALLS edges to refined targets).
Replaced the single pass with a bounded BFS over the IMPORTS graph
(depth ≤ 4). Catches nested barrel pyramids without ballooning into
a near-full rebuild on monorepos with deep re-export trees. `--force`
remains the escape hatch documented in GUARDRAILS.md for cases that
exceed the bound.
2. F2 — Integration test for incremental orchestration (Claude, BLOCKER,
DoD §2.7). The unit tests added in ad7bd31 covered `diffFileHashes`,
`extractChangedSubgraph`, `computeChunkHash`, `pruneCache`, and the
Map/Set JSON round-trip — but none of them exercised the real
`runFullAnalysis` orchestration. Added gitnexus/test/unit/
incremental-orchestration.test.ts with four end-to-end tests against
a real git-initialized fixture repo + real LadybugDB:
a. First run populates fileHashes + schemaVersion and clears
incrementalInProgress on success.
b. Second run on unchanged state takes the alreadyUpToDate fast
path (early-return).
c. Second run after a source edit takes the incremental path
(not full rebuild) and rotates fileHashes for the touched file
while keeping the dirty flag cleared.
d. A pre-set incrementalInProgress flag forces a full rebuild
that clears it (crash-recovery wire).
These would catch any regression that wires `isIncremental` from a
pre-pipeline prediction (the Bugbot finding from commit 5eb0597) or
accidentally re-gates the embedding re-insert on `!isIncremental`
(the Bugbot finding from commit 60c10f1).
3. F3 — GUARDRAILS.md docs accuracy (Claude, Low). Line 33 still said
"only changed files are re-parsed" — AGENTS.md was already corrected
in ad7bd31 but GUARDRAILS.md was missed. Reworded to match.
4. F5 — Atomic saveMeta (Claude, Medium; vvladescu-tb fork). The dirty
flag (`incrementalInProgress`) travels through meta.json. A crash
mid-write would leave a corrupt meta.json that `loadMeta` would
silently treat as "no prior index", losing the flag and skipping
recovery. Switched to tmp-file + rename matching saveParseCache.
5. Bugbot's "Subgraph edges reference nodes absent from subgraph"
(HIGH severity). Verified as FALSE POSITIVE: `getNodeLabel` in
lbug-adapter.ts derives labels from the node-ID string (parses
the table prefix), not from the in-memory graph. The CSV
generator writes (src_id, dst_id, type) rows without consulting
node objects; `splitRelCsvByLabelPair` routes by ID-derived label;
`COPY ... (from=X, to=Y)` resolves both endpoints against the live
LadybugDB where unchanged-file nodes still exist. No fix needed.
All 213 tests pass locally (including the 4 new integration tests
and the previously-failing CI tests).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses CHANGES_REQUESTED review on PR #1479:
1. Remove docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
per maintainer request.
2. BLOCKER (Claude Finding 1, Bugbot Round 3): Stale cross-file edges
between unchanged files. extractChangedSubgraph excluded edges where
both endpoints were unchanged-file nodes — when a barrel/re-export
file changes, cross-file resolution may update CALLS edges between
two unchanged files that would then be silently lost.
Fix: 1-hop importer-closure expansion of the writable set in
run-analyze.ts. Before deleting/rewriting rows, query DB for
importers of every changed/deleted file and add them to the writable
set. Their nodes get deleted+rewritten too, so cross-file's refined
edges land in the DB. Re-added queryImporters to lbug-adapter.ts.
3. BLOCKER (Claude Finding 3): Parse cache key omitted parser version.
After a GitNexus upgrade, the cache silently replays pre-upgrade
ParseWorkerResults against the new schema → wrong CALLS/IMPORTS/
scope edges with no visible signal.
Fix: PARSE_CACHE_VERSION now embeds the gitnexus npm package
version (read at module load via createRequire on package.json).
Format: `${SCHEMA_BUMP}+${PKG_VERSION}` e.g. "1+1.6.4". Any release
that bumps package.json automatically invalidates the on-disk cache.
Mismatched versions fall through to an empty cache (next save
overwrites with the new version baked in).
4. BLOCKER (Claude Finding 2): No automated tests for incremental
behavior. Added 28 unit tests across 3 files:
- incremental-file-hash.test.ts (10 tests)
diffFileHashes classification, computeFileHash determinism,
computeFileHashes batch / missing-file tolerance, sorted output.
- incremental-parse-cache.test.ts (12 tests)
computeChunkHash stability and order-independence, version
prefix format, pruneCache, load/save round-trip on empty /
missing / corrupt / version-mismatched files, AND a Map/Set
round-trip test that pins the JSON replacer/reviver behaviour
(without it, ParsedFile.scopes[*].typeBindings collapses to
{} and downstream `.get()` / iteration throws).
- incremental-subgraph-extract.test.ts (6 tests)
writable-set node inclusion, Community/Process always kept,
edge inclusion when at least one endpoint is writable, MEMBER_OF
edges via graph-wide endpoints, empty subgraph case.
5. Medium (Claude Finding 6): AGENTS.md "Keeping the Index Fresh"
said "only changed files are re-parsed." Imprecise — the pipeline
parses every file every run; the cache skips tree-sitter for chunks
whose contents haven't changed. Reworded to match the design doc.
Test plan still expects:
[x] Typecheck clean
[x] All 28 new unit tests pass
[x] All previously-failing tests still pass on the rebased branch
[x] Equivalence verified locally (incremental ≡ --force, byte-identical
stats on this repo)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV
tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:
- captures.ts (C# scope extraction)
- namespace-siblings.ts (extractFileStructure)
- parse-worker.ts (worker thread parse path)
- parsing-processor.ts (sequential parse fallback)
Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.
Additional C# scope fixes:
- scope-tree.ts: Module scopes may share the same range as a top-level
namespace_declaration (files with no leading `using` directives). The
rangeStrictlyContains check rejects equal ranges. Added
rangeNonStrictlyContains for Module parents.
- scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
same Module == Namespace range case caused orphaned scopes. Added
moduleAwareContains helper.
- scope-extractor-bridge.ts: empty captures from ERROR-root files still
called extractScope -> "no Module scope found" warning. Added early
return for empty/non-array captures.
- namespace-siblings.ts: three sites pushed onto binding arrays frozen by
finalize-algorithm. Fixed with spread-copy before mutation.
lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.
Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).
* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV
LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.
Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.
This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.
* fix: address codeql findings on PR #1433
The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.
Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).
* fix(windows): replace 32767-char truncation with chunked-input parsing
The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.
Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.
Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.
* fix(windows): cover all parse sites and correct vector-extension state
Address adversarial review on PR #1433:
1. Extend parseSourceSafe to all remaining parser.parse() call sites that
handle full file content. The first commit only converted the four
sites with active truncation hacks; cache-miss paths in
call-processor (x2), heritage-processor (x2), import-processor, and
the Go/Python/TypeScript captures + Go range-binding still called
parser.parse() directly. On Windows those would still SIGSEGV for
files > 32767 chars.
2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
in lbug-adapter.ts. The flag means "successfully loaded" and is
checked by an early-return at the top of loadVectorExtension; setting
it on the skip path made the second call return true and let
QUERY_VECTOR_INDEX run against a DB without the extension.
3. Drop the placeholder issues/... URL in the same comment.
4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
(direct/callback boundary), the 32 767 Windows crash boundary,
single-line > chunk size, CRLF near boundary, and large all-Chinese
source. Confirms the callback path is correct for non-ASCII content,
which is also exercised by the existing csharp-captures large-file
test.
Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.
* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement
Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.
Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
so group/ and embeddings/ can import without crossing into ingestion
internals. core/tree-sitter/ already houses parser-loader.ts and is
the natural shared facade. All 11 existing importers updated; no shim
left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
http-route, include, tree-sitter-scanner) and the embeddings
ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
module-load grammar smoke test parsing a 36-char literal. Trivially
safe by inspection, intentionally direct, filtered out by the lint
rule via the string-literal-arg skip.
Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
parseSourceSafe. The spy is what catches a regression: parser.parse
on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
inside each mock factory so vitest's hoister does not race the
static import binding.
Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
Skips JSON/URL/marked/Number/Math, string-literal first args
(smoke tests), test files, and the helper itself. Auto-fix rewrites
the call site only; the developer adds the import after tsc
surfaces the missing identifier — same tradeoff as
unused-imports/no-unused-imports.
Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md
* fix(test): use mkdtempSync in http-route-extractor regression test
Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Bugbot re-review caught: deleteNodesForFile cascades to the
CodeEmbedding table (DELETE WHERE e.nodeId STARTS WITH ...), so
changed-file embedding rows are wiped along with their nodes. The
previous fix gated re-insert on `!isIncremental`, which silently
dropped those embeddings — a regression versus the full-rebuild path's
"preserve embeddings by default" guarantee.
Remove the `!isIncremental` gate. The per-batch try/catch already
handles the unchanged-file PK-conflict case ("some may fail if node
was removed, that's fine") with the same semantics, so re-inserting
the full cached set on incremental works:
- changed-file rows: deleted, then re-inserted from cache (preserved)
- unchanged-file rows: still in DB, re-insert PK-conflicts and is
silently ignored (existing rows are correct)
Cost: re-inserting ~24K embeddings on incremental when only a few
files changed — most are no-op conflicts. Bounded by batch size of
200; ~3-5s overhead. Worth it for correctness.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Bugbot (PR #1479):
- Medium: pruneCache was exported but never called -> cache grew
unbounded. Wire pruneCache into run-analyze before saveParseCache,
using a transient usedKeys Set on ParseCache that the parse phase
populates as it processes chunks.
- Low: willTryIncremental (pre-pipeline) and isIncremental
(post-pipeline) could desync, silently dropping embeddings on
mispredicted runs. Removed the prediction; the embedding cache
now loads unconditionally when shouldLoadCache is true. The
re-insert step gates on the actual isIncremental value to avoid
PK-conflicts when the incremental-writeback path keeps DB rows.
CI test failures:
- cli-e2e #1169 + run-analyze.test.ts #1233: my dirty-tree gate on
the lastCommit==HEAD early-return saw GitNexus's own auto-generated
outputs (.claude/, .cursor/, AGENTS.md, CLAUDE.md) as dirty,
perpetually defeating the up-to-date fast path. Extended the
pathspec exclusion to cover all auto-gen outputs, not just
.gitnexus/.
- ruby field-type disambig: my chunk-stability sort exposed a
pre-existing order-dependency in Ruby cross-file resolution
(`user.address.save -> Address#save` only resolves correctly when
user.rb parses before address.rb in some configurations). Removed
the sort. Filesystem ordering is stable enough in practice that
the parse cache still hits the common case; the pre-existing
fragility is left for a separate fix.
- pipeline-graph-golden: regenerated. Seeded Leiden RNG produces a
partition different from the previous Math.random snapshot.
- staleness `parallel calls` was a CI timing flake; passes locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(cursor): upgrade hooks to Cursor 2.4 postToolUse for Read/Grep/Shell coverage
Cursor 2.4 (released 2026-01-22) shipped generic preToolUse/postToolUse hooks
matching `Shell|Read|Write|Grep|Delete|Task|MCP:<tool>`, replacing the
2.3-era beforeShellExecution hook that only fired on shell commands. The
existing integration only intercepted the shell path, so Cursor users got
graph augmentation roughly 10% as often as Claude Code users — only when
the agent dropped to rg/grep instead of using its native Read/Grep tools.
This swaps the integration over to postToolUse and ports the bash+jq
hook script to cross-platform Node:
- gitnexus-cursor-integration/hooks/hooks.json: registers a single
postToolUse hook matching Shell|Read|Grep that invokes the new
gitnexus-hook.cjs.
- gitnexus-cursor-integration/hooks/gitnexus-hook.cjs: new Node hook
mirroring the safety patterns from the Claude hook (absolute-cwd
validation, .gitnexus discovery with linked-worktree fallback,
npx.cmd on Windows, end-of-options `--` marker, debug truncation,
graceful failure). Extracts the search pattern per tool kind:
Grep -> toolInput.query; Read -> file basename stripped to identifier
chars; Shell -> existing rg/grep arg parser. Emits Cursor-shape
`{ "additional_context": "..." }` on stdout — no shell, no jq.
- gitnexus-cursor-integration/hooks/augment-shell.sh: removed (Windows
incompatible, narrower coverage).
- gitnexus/test/unit/cursor-hook.test.ts: 33 regression tests covering
manifest wiring, source-level invariants (no shell:true, npx.cmd,
isAbsolute, additional_context output shape, end-of-options marker),
extractPattern coverage per tool, and behavioral early-exit paths
(empty/invalid stdin, relative cwd, no .gitnexus, unknown tool name,
short patterns, non-search shell commands, case-insensitive matching).
- README.md / gitnexus/README.md: editor-support table now lists Cursor
as Full / hooks=Yes (postToolUse), matching reality.
- gitnexus/src/cli/augment.ts and gitnexus/src/core/augmentation/engine.ts:
doc-strings updated from `Cursor beforeShellExecution` to
`Cursor postToolUse`.
Closes#1466.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(cursor): hook timeout is in seconds, not milliseconds
Cursor's `timeout` field in hooks.json is in seconds (per
https://cursor.com/docs/agent/hooks and the original integration's
`"timeout": 5`). I'd written `10000` after blindly copying the issue
body's example — that resolves to ~2.8 hours, not 10 seconds. If the
script ever hangs before reaching its inner spawnSync timeouts (e.g.
during stdin read), Cursor would have waited that long before killing
it.
Drop to `10` (seconds), matching the Claude plugin's hooks.json and
giving plenty of headroom over the inner 7s augment-CLI timeout.
Add a regression-guard assertion in cursor-hook.test.ts so a future
ms/s mixup fails fast.
Reported by Cursor Bugbot on PR #1467.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(cursor): address Claude review findings — payload aliases, debug, install docs
Resolves three findings from Claude reviewer on PR #1467:
1. Cursor payload field-name uncertainty (SIGNIFICANT)
Claude flagged that the Grep `query` field is an unverified assumption
per Cursor 2.4 docs (https://cursor.com/docs/agent/hooks). Mitigated:
- Expanded Grep aliases: query | pattern | regex | q | search | searchQuery
- Added pickLongestStringValue() last-resort fallback so the hook
extracts *something* even if Cursor renames every documented field
- Added GITNEXUS_DEBUG=1 stderr logging of the raw stdin payload so
users can capture Cursor's actual contract when diagnosing silent
no-ops, and report it back if aliases drift
- Added Read alias `filePath` (camelCase variant alongside `file_path`)
- Inline comment block citing the docs URL and the uncertainty
2. Hook command path resolution + install docs (SIGNIFICANT)
Claude flagged `node ./hooks/gitnexus-hook.cjs` as relative without
documented install path. Added gitnexus-cursor-integration/README.md
with explicit install steps:
- .cursor/hooks.json + hooks/gitnexus-hook.cjs at project root
- Confirms Cursor's project-root CWD convention with doc link
- Verify steps including GITNEXUS_DEBUG capture
- Pattern-extraction contract table per tool
- Troubleshooting: not-firing, npx fallback, wrong-pattern diagnosis
3. README "Full" overclaim for Cursor (MODERATE)
Both README rows now read `Yes (postToolUse, manual install)` linking
to the new install README, accurately signaling that hooks aren't
automated by `gitnexus setup` like they are for Claude Code.
4. Shell quoted-pattern parser limitation (MINOR, documented)
Added inline comment in gitnexus-hook.cjs documenting the known
`rg "User Service"` -> `User` truncation, plus regression tests in
cursor-hook.test.ts pinning the behavior so a future change is
visible.
Test additions (33 -> 41):
- Wide-alias source coverage for Grep (query / pattern / regex / q /
search / searchQuery) plus pickLongestStringValue fallback
- Read alias coverage including camelCase filePath
- GITNEXUS_DEBUG behavioral test: stderr quiet by default, payload
echoed when env var set, stdout output contract preserved either way
- Shell quoted-pattern documented behavior tests
- Install README presence + content (.cursor/hooks.json, hooks/, debug
diagnostics)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
- Rewrite docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
to describe the architecture that actually shipped (parse cache +
incremental DB writeback + scope-resolution short-circuit), with the
v1 hydrate-phase post-mortem preserved as historical context.
- AGENTS.md "Keeping the Index Fresh" section: note that incremental
is the new default and --force is the explicit opt-out; mention
the parse-cache file location and that it's safe to delete.
- GUARDRAILS.md Signs: add an "Index seems corrupt or incremental is
misbehaving" entry pointing users to --force as the manual escape
hatch (the dirty flag handles automatic recovery).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two compounding optimizations that drop warm-cache analyze from
~134s to ~38s on a 1000-file repo (72% faster), and cold rebuild
from ~143s to ~86s (40% faster) by short-circuiting work that was
previously re-done.
1. SCOPE-RESOLUTION: REUSE WORKER PARSEDFILE
Previously, the scope-resolution phase re-parsed every file with
tree-sitter on the main thread (~58s on a 1000-file repo) because
worker-produced tree-sitter Trees can't cross the worker MessageChannel.
But the worker ALSO produces a artifact via
, which structured-clones fine — and it's exactly
what scope-resolution would re-derive. Threading those ParsedFiles
through the parse phase () into
( map) lets scope-
resolution skip its extract loop on a per-file basis.
The fast path is bounded only by per file (cheap
graph mutation). On this repo: scopeResolution went from 58s → 5s.
2. MAP-PRESERVING PARSE-CACHE SERIALIZATION
is a
which JSON.stringify collapses to . The first attempt at threading
parsedFiles through the parse cache crashed at runtime with
"importerModule.typeBindings is not iterable" because cached entries
came back as plain objects.
Added a JSON replacer/reviver pair in parse-cache.ts that round-trips
Map and Set instances through tagged plain objects (). Symmetric: save uses replacer, load uses reviver.
3. STABLE CHUNK ORDERING
The byte-budget chunker walked files in filesystem-scan order, which
on Windows isn't guaranteed to be stable across runs. Even with
identical source content, two scans could place files in different
chunks, shifting chunk hashes and causing 100% parse-cache misses.
Added a deterministic alphabetical sort on before
chunking. Chunk membership is now stable across runs, so a single-file
edit invalidates exactly one chunk, not all of them.
Measured on this repo (993 files, 24K nodes):
Cold rebuild: 86s (was 143s)
Warm cache, no source changes: 3s (early-return)
Warm cache + 1-file edit: 38s (was 134s)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The parse cache is keyed at chunk granularity. With the previous 20MB
budget, a typical mid-size repo (e.g. this worktree at 9MB total
parseable source) fits in a single chunk — meaning ANY file change
invalidates the whole chunk and re-parses every file.
2MB default produces ~5x more chunks on the same input, so a one-file
edit invalidates ~1/N of cached chunks instead of the whole thing.
Cold-run overhead from more chunks is <5% (one extra serialization
pass per chunk).
Override via GITNEXUS_CHUNK_BYTE_BUDGET env var for benchmarking.
Measured on this repo (~9MB / 887 parseable files):
Cold (no cache): 143s
Warm cache, no source changes: 2s (early-return)
Warm cache + 1-file edit: 81s (~43% off cold)
Speedup is bounded by the scopeResolution phase (~58s flat regardless
of parse cache) and by GitNexus's own auto-writes during analyze
(AGENTS.md / .claude/skills/ etc. mutate between runs and invalidate
chunks containing them). Both are addressable in follow-ups.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Composes with the incremental DB writeback (commit 27f3b49d) to deliver
the major-speedup half of incremental indexing. Previously, the parse
phase ran in full on every analyze; the speedup came purely from
selective DB rewriting. With this commit the parse phase also reuses
prior tree-sitter output for chunks whose contents haven't changed.
How it works:
* Cache layer (gitnexus/src/storage/parse-cache.ts):
- File: <repo>/.gitnexus/parse-cache.json. Versioned, atomic write.
- Key: chunk content hash = sha256(sorted(filePath:fileContentHash
for each file in chunk)).
- Value: ParseWorkerResult[] (raw worker output for the chunk,
pre-merge).
- Granularity: per chunk (~20MB byte-budget). A change to one file
invalidates only its chunk — typically 1 of ~50 on a 1000-file
repo (~98% cache hit ratio on a small edit).
* Worker contract (gitnexus/src/core/ingestion/parsing-processor.ts):
- Extracted the chunk-result merge loop into a public
mergeChunkResults() so the same logic applies to live worker
output AND replayed cache entries.
- processParsingWithWorkers / processParsing accept an optional
outRawResults out-parameter that captures worker output before
merging — used by parse-impl to populate the cache after a miss.
* Parse phase wiring (parse-impl.ts):
- For each chunk, compute its content hash (after reading file
contents). Cache hit → mergeChunkResults() on cached results,
skip the worker dispatch entirely. Cache miss → run workers
normally, capture raw results, store under the chunk hash.
- Cache mutations happen in-place on the ParseCache passed via
PipelineOptions.parseCache.
* Lifecycle (run-analyze.ts):
- loadParseCache() before pipeline runs.
- Cache passed via runPipelineFromRepo's PipelineOptions.
- saveParseCache() after the pipeline + DB writeback succeed.
Equivalence verified on this repo (993 files, 24K nodes):
Cold (no cache, full work): 141.1s
Warm cache + 1-file edit, incremental: 63.6s ← 55% speedup
Warm cache + 1-file edit, --force: 71.6s ← 49% speedup
All three runs produce byte-identical {nodes, edges, clusters,
flows}. The cache survives --force (content-addressed = always
correct), so even forced rebuilds get the parse-skip benefit.
Why chunk-level rather than per-file: workers process sub-batches and
emit aggregated ParseWorkerResults. Per-file granularity would require
restructuring the worker contract; chunk-level captures most of the
practical speedup with no worker-side changes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Equivalence-preserving incremental analyze. The pipeline still parses
every file (correctness invariant: cross-file resolution / scope
resolution / MRO / community detection all need full graph data); the
saving comes from selectively replacing only changed-file rows in
LadybugDB instead of wiping and reloading the whole graph.
How it works:
* On every analyze, we hash all source files (SHA-256 of content) and
store the map in meta.json.fileHashes alongside schemaVersion.
* The next run loads the prior map and diffs:
- changed: content hash differs → file's DB rows replaced.
- added: not in prior map → file's DB rows inserted.
- deleted: in prior map but not on disk → file's DB rows dropped.
* If the diff is non-empty AND no --force / no schema mismatch / no
dirty flag, take the incremental path:
- Set incrementalInProgress dirty flag (BEFORE any DB mutation).
- Open existing DB (no wipe).
- deleteNodesForFile() for each changed/added/deleted file.
- deleteAllCommunitiesAndProcesses() — Leiden regenerates these.
- extractChangedSubgraph() from the in-memory ctx.graph: nodes whose
filePath is in the writable set + Community + Process + edges with
at least one endpoint in the writable set (edges entirely between
hydrated unchanged nodes are skipped — already in DB).
- loadGraphToLbug() on the subgraph. Unchanged-file rows in DB
untouched.
- Recreate FTS indexes.
- Update meta with new fileHashes; clear dirty flag.
* Otherwise full-rebuild path runs as before.
Crash recovery: incrementalInProgress is the dirty flag. Set before
destructive ops; cleared on success. Set on next-run startup → forces
full rebuild (cheapest path back to known-good).
Other changes:
* Dirty-tree gate on the existing 'lastCommit==HEAD' early-return:
uncommitted edits no longer slip through as 'already up to date'.
* deleteAllCommunitiesAndProcesses helper in lbug-adapter.
* Skip the embedding cache+restore cycle when willTryIncremental is
true — embeddings stay in DB; re-inserting them would PK-conflict.
End-to-end equivalence verified on this repo (993 files, 24K nodes):
incremental run produces byte-identical {nodes, edges, clusters,
flows} to a full rebuild from the same edited state.
Speedup is currently modest (~5% on this repo) because the parse
phase still runs in full. Parse-cache integration is a separate
follow-up that composes cleanly on top of this work.
See docs/superpowers/specs/2026-05-10-incremental-indexing-design.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Reverts the v1 design that parsed only closure files into a fresh
graph and tried to hydrate the rest from DB. Real-repo equivalence
test failed: cross-file resolution operates on partial parse data
(closure files only), so CALLS edges that resolve through unchanged
files silently fall off. Diff against full rebuild on the same
edited state: -50 nodes, -425 edges, -5 communities, -48 processes.
Architecture pivot: switch to PR #533-style content-addressed parse
cache. Pipeline parses every file (cache-served when possible),
giving cross-file resolution full data, with DB writeback then
restricted to changed-file rows.
Reverts:
d4b9de47 fix(incremental): drop invalid --no-renames=false
f35f7634 feat(analyze): incremental orchestrator branch + meta schema
bc039686 feat(pipeline): hydrate phase + parse-filter
98bb893d feat(lbug): loadGraphFromLbug, queryImporters, ...
aa8d7ae3 feat(incremental): change-detection, surface signatures, closure
Kept:
d9e340b0 feat(communities): seed Leiden RNG (foundational)
8235ca36 docs: incremental indexing design spec (will be revised)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(release): skip rc build on release PRs
Suppress the auto-fired Release Candidate workflow when:
1. The HEAD commit subject matches `chore: release vX.Y.Z` (the canonical
release-PR title), or
2. The squash-merged PR carries the `release` label.
Either match short-circuits the guard to should_run=false. This prevents the
rc cycle from racing publish.yml on the v-tag (as happened on v1.6.4 where
we had to manually cancel the auto-fired RC run after merging PR #1473).
Adds pull-requests: read to the guard job for the label lookup. A failed
gh API call falls through to the existing dedup logic rather than silently
suppressing rc builds.
* ci(release): address PR #1474 review — anchor regex + sanitise log echo
Two minor follow-ups from Claude's review:
1. End-anchor the release-subject regex. The previous shape
^chore: release vX.Y.Z would match noisy variants like
chore: release v1.0.0 (something unrelated). The new shape
requires either the bare title or the canonical squash-merge
(#NNNN) suffix exactly.
2. Sanitise HEAD_SUBJECT before echoing to logs. git %s strips
newlines so LF injection is impossible, but a hypothetical
subject containing ::error:: or ::set-output:: could otherwise
forge GitHub Actions annotation entries. Defence-in-depth.
Both findings flagged minor / does not block merge — applying
anyway since they are trivial.
* test(u8): de-flake regex linearity assertions
The single-trial 2x input + 3x ratio bound was razor-thin: a real macOS
CI run failed at ratio 3.01x with small=7.41ms / large=22.31ms - both
above the 5ms noise floor but close enough that single-shot scheduler
jitter pushed the ratio over.
Replace the methodology with four stacked techniques:
1. Warmup runs before timing (let the JIT tier up)
2. Median of 5 trials per measurement (eliminates GC + jitter)
3. 4x input ratio (was 2x) - linear gives ~4x, O(n^2) gives ~16x
4. 8x ratio bound with a 20ms noise floor on the LARGE measurement
Headroom: linear is expected at ~4x, bound is 8x = 2x safety margin.
A real O(n^2) regression on a 4x input would clock 16x, well outside.
Catastrophic backtracking is still caught by the absolute <500ms cap.
Verified: 10 consecutive local runs all passed.
* test(u8): address PR #1475 review — tighten floor + rename for accuracy
Two follow-ups from Claude's review:
1. Floor semantics: revert to 'skip when BOTH measurements below floor'
(AND, not single-check) and lower threshold from 20ms back to 5ms.
Median-of-5 makes 5ms reliably resolvable above performance.now()'s
~10-100us band, so the higher floor was unnecessary defense.
Closes the gap where an O(n^2) regression on a fast runner could
stay under 500ms AND below 20ms-large to escape both detectors.
2. Rename assertSubLinearRatio -> assertNearLinearScaling. The bound
is SIZE_RATIO * 2 = 8x on a 4x input = sub-quadratic with 2x
headroom over linear, not strict sub-linearity. New name reflects
the actual semantics.
The flag --no-renames=false isn't valid git syntax (it's parsed as a
file path). Git's default rename detection is on; removing the flag
keeps that behavior.
Caught while running an end-to-end smoke test against a small fixture
repo: incremental setup failed with 'Command failed: git diff
--name-status -z --no-renames=false ...'. After the fix, the
incremental path runs cleanly: closure is computed, hydrate phase
loads unchanged-file state from DB, parse phase only re-parses files
in closure, and the writeback updates only changed nodes/edges.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires incremental indexing into runFullAnalysis. Highlights:
* RepoMeta schema extended: schemaVersion, surfaceSignatures, and
incrementalInProgress fields. INCREMENTAL_SCHEMA_VERSION = 1.
* core/incremental/file-hash.ts — v1 surface signature: SHA-256 of file
content. v2 will switch to a true surface-only signature (defined in
surface.ts) so body-only edits don't expand the closure. The plumbing
is signature-agnostic so the swap is local.
* core/incremental/orchestrator.ts — eligibility check, closure
computation (uses file-hash as the surface signal), dirty-flag
management, subgraph extraction, signature merge.
* run-analyze.ts adds:
- hasDirtyTree() check on the existing 'lastCommit==HEAD' early-exit
so an uncommitted edit triggers re-index (was a coarse equality
check before).
- incremental branch: try incremental first; fall through to full
rebuild on any setup failure or eligibility miss.
- runIncrementalBranch() — opens existing DB, deletes closure-file
rows + Community/Process, runs pipeline with filesToParse, writes
only the changed-subgraph back, refreshes FTS, updates meta with
new surfaceSignatures and clears the dirty flag.
- Full-rebuild path now populates surfaceSignatures + schemaVersion
in meta.json so the next run is eligible for incremental.
Crash recovery: incrementalInProgress is set BEFORE any DB mutation
and cleared on success by overwriting meta.json. A crash anywhere in
between leaves the flag set, and the next analyze run forces a full
rebuild (cheapest path back to a known-good index).
v1 limitation documented: body-only edits trigger 1-hop closure
expansion (content-hash signal). True surface-only optimization is
deferred to v2 — see design doc for the integration path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires the incremental-indexing infrastructure into the phase-based
pipeline. Three coordinated changes:
* New hydratePhase (deps: structure) — loads node/edge state for files
OUTSIDE ctx.options.filesToParse from the existing LadybugDB index.
Runs before parse so the parse phase can produce a partial graph
while downstream phases (mro, communities, processes) still see the
full graph. No-op in full-rebuild mode (filesToParse unset).
* PipelineOptions.filesToParse: optional ReadonlySet<string>. When
set, parse phase filters scanned files to this set; hydrate fills
the complement. Set by runFullAnalysis when it detects an eligible
incremental run; never set by callers directly.
* gitnexus-shared PipelinePhase enum: 'hydrate' added so progress
callbacks can report the new phase distinctly from 'structure'.
Phase order: scan → structure → hydrate → markdown,cobol → parse
→ routes,tools,orm → crossFile → scopeResolution → mro → communities
→ processes. Communities (Leiden) still runs on the full graph,
satisfying the correctness invariant.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three new primitives in lbug-adapter.ts to support incremental indexing:
* loadGraphFromLbug(graph, unchangedFilePaths) — streams all nodes for
files in the set across every hydratable node table (excludes
Community/Process — graph-wide, regenerated downstream). Then loads
edges where both endpoints belong to loaded nodes, excluding
MEMBER_OF / STEP_IN_PROCESS edges (also graph-wide).
FilePaths chunked at 200 per query to keep statement size bounded
on huge repos. Endpoint-level join filters by source-side filePath
in the query, target-side checked JS-side via the loadedNodeIds set.
* queryImporters(targetFilePath) — returns DISTINCT a.filePath where
a -[IMPORTS]-> b and b.filePath = target. Powers closure expansion:
when a changed file's surface signature changes, all its importers
must be re-parsed.
* deleteAllCommunitiesAndProcesses() — drops Community/Process nodes
(and their edges via DETACH DELETE) at the start of each incremental
run so the communities/processes phases regenerate them from the
fully-merged graph. Required for the 'Leiden runs on full graph'
correctness invariant.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three new modules supporting the incremental-indexing pipeline:
* core/incremental/git-diff.ts — getChangedFilesSinceCommit() unions
'git diff lastCommit HEAD' (committed) with 'git status --porcelain'
(dirty tree). Renames flattened to delete(orig) + add(new). Throws
LastCommitMissingError when lastCommit is gone (caller falls back to
full rebuild).
* core/incremental/surface.ts — extractSurfaceSignature() produces a
stable hash of a file's publicly-visible symbols (functions, classes,
methods, interfaces, types, heritage). Body-only edits → same hash.
Signature/heritage changes → different hash. Drives the closure
scoping optimization.
* core/incremental/closure.ts — computeImporterClosure() iterative
fixpoint: parse each closure file, extract surface, query DB
importers, expand. Uses a parseCache so each file is parsed once.
Generic over TParseResult so closure logic is decoupled from the
pipeline's parse representation.
32 unit tests across the three modules. Tests cover edge cases:
clean tree, dirty-only, mixed, renames, deletes, multi-hop cascade,
cycle termination, surface invariance, etc.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The vendored Leiden algorithm defaults to Math.random for tie-breaking
and randomized walks, which produces non-deterministic community
assignments and modularity values across runs on the same graph.
Pass a seeded mulberry32 RNG (LEIDEN_SEED=0xC0DE) so:
- The same graph always produces the same partition
- Modularity values are reproducible
- Equivalence tests for incremental indexing can compare community
assignments byte-for-byte
This is foundational for the upcoming incremental-indexing feature
(see docs/superpowers/specs/2026-05-10-incremental-indexing-design.md)
where the correctness contract is incremental output ≡ full rebuild
output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Captures the design agreed in brainstorming on 2026-05-10:
- Transitive importer closure with public-surface-change optimization
- Git-only change detection (non-git repos: full rebuild as today)
- New default behavior; --force opts out
- New hydratePhase + loadGraphFromLbug primitive
- Iterative closure expansion with parseCache reuse
- incrementalInProgress dirty flag for crash recovery
Prior art: PR #592 (zenprocess), PR #533 (davidbeesley),
PR #1146 (azeemshaik025) — referenced and credited.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>