* feat(storage): resolve the shared sibling-store identity and layout (#3352) Linked worktrees of one repository resolve to one store under GITNEXUS_HOME/stores/<key>, keyed by the canonical git common dir. The resolver reads the .git entry directly, so hot paths spawn no git. Slot naming moves to a leaf module so storage-resolver and shared-store do not import each other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): resolve a shared checkout's graph and existing store slot (#3352) getStoragePaths reads a flat slot's recorded graphPath only for checkout slots under the stores directory; other paths keep <storagePath>/lbug with no I/O. A recorded path outside the store's commit graphs is ignored. resolveStoragePath falls back to an existing store slot for an unregistered checkout, so reads never move to an empty slot. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): share one immutable commit graph across clean worktrees (#3352) A linked worktree writes its own slot in the shared store. A new slot is seeded with a pointer to the nearest commit graph, so a clean checkout at an indexed commit takes the up-to-date path and writes no graph. A checkout with local changes gets a copy-on-write private graph before its first write. After a successful run, a clean checkout at HEAD publishes its graph into commits/ under a store lock, or drops it when that commit graph already exists. Commit graphs are never written after publish. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): seed a worktree's graph from the nearest index (#3352) A new shared slot is seeded from the store's commit graph nearest to HEAD, else from this checkout's or the main checkout's repository-local index (copied under that index's lock, source left in place). The run that follows is up to date or incremental instead of a full build. If a pointed-at shared graph has been removed, analyze falls back to a full build instead of an incremental update over a missing baseline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): keep one parse cache per shared store (#3352) Linked worktrees read and write the parse cache and durable ParsedFile store under the store's caches/ directory. Before pruning, a run folds in the chunk keys recorded by every member slot and commit graph, and the fold, prune and save run under a store-wide cache lock so one member never evicts another's live chunks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(clean): remove only what no shared-store member references (#3352) Deleting a shared checkout slot (clean, clean --all, remove, and the server delete route) recounts references under the store's publish lock and deletes commit graphs no member points at, then the store itself once empty. A graph that cannot be deleted (open on Windows) is reported and kept for the next pass. clean --gc also drops member slots whose worktree is gone or no longer resolves to the store. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): let a clone opt in to a shared store with --share-with (#3352) An independent clone joins a linked worktree's shared store only with analyze --share-with <repo>, and only when its normalized origin URL matches that member's (credentials stripped, as #2054 compares). The registry remembers the choice. --no-share moves an opted-in clone back to its own .gitnexus and reclaims its old slot; linked worktrees always share and are pointed at GITNEXUS_SHARED_STORE=off instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): count a shared checkout's commit graph as its code index (#3352) A clean shared checkout reads a commit graph and owns no graph file, so the code-index presence check made status report it unindexed and registry validation skip it. The check now follows the slot's validated graphPath. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): adopt existing worktree indexes and report shared-store state (#3352) A shared checkout gets <repo>/.gitnexus/store.json pointing at its store slot; resolution follows it only when it names that checkout's own slot. An existing local index seeds the slot and is left in place. status (text and --json) and doctor report the store, whether the graph is shared or private, and any leftover local index, which clean --local-index removes while keeping the pointer. With GITNEXUS_SHARED_STORE=off a previously shared checkout indexes into its own .gitnexus again and never writes a commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(mcp): read the graph a shared checkout points at (#3352) MCP, the HTTP API, group sync, augmentation, and the Claude hook (all three byte-identical copies) resolve a flat slot's graph through resolveGraphPath instead of joining 'lbug' onto the storage path, so checkouts at one commit share one open database. The embeddings writers (embeddings sync and the server embed job) take a private copy first and never write an immutable commit graph. The post-analyze settle probe accepts fresh metadata that points at an existing commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: describe the shared worktree index store (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(storage): reuse helpers across the shared-store code (#3352) One graph-clone helper replaces two copy-then-rename blocks; clean and status reuse formatSlotSize; leaving a store reuses removeSharedStorePointer; withStoreLock is imported from its own module instead of a re-export. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review findings in the shared store (#3352) - Never publish a graph whose build saw dirty files, and never trust a local-index seed on the up-to-date path; it may hold reverted edits. - Record slot pointers and reclaim under the publish lock, and reclaim right after each publish, so a commit graph is never deleted between publish and pointer save and superseded graphs don't pile up. - clean --gc decides membership from the registry, so opted-in clones and a main checkout without worktrees are not dropped. - Leaving a store re-registers first, so an up-to-date run cannot leave the registry pointing at a deleted slot. - Drop pinned branch summaries when an entry moves into a store slot. - Keep run.cjs and the AGENTS.md runner path inside the checkout. - MCP handles follow the slot's current graph, and branch scoping reads the slot's own metadata. - Cache resolveGraphPath by metadata file identity; look up opted-in entries with canonical registry paths. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): localize help for the shared-store flags (#3352) --help renders option text from the i18n catalog, so --share-with, --no-share, clean --gc and clean --local-index need keys in both locales, not just index.ts literals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): lock the embed-job graph copy and skip it on forced rebuilds (#3352) The server embed job copies a shared checkout's graph under the slot's index lock, so a CLI analyze in another process cannot interleave. A forced rebuild with no embeddings to carry over drops the pointer without copying a graph it would discard unread. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(agents): bump AGENTS.md version for the shared store notes (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep shared-store file reads inside checked paths (#3374) Address CodeQL js/path-injection and js/file-system-race on shared-store.ts: every filesystem read rebuilds its path under a fixed parent with an inline path.relative barrier, the .git probe is one read (EISDIR marks a directory) instead of stat-then-read, and the graph pointer cache stats and reads through one file descriptor. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review feedback on the shared store (#3374) - Hold the slot index lock for the whole server embedding job, not just the graph copy. - --no-share re-registers and deletes the old slot under that slot's index lock, so a running shared analyze cannot re-register it. - clean --gc previews without --force, like every other destructive arm. - status reports a pinned branch index as private. - Slot names prefix Windows device names that carry an extension. - A slot or store named ..<name> is a legal direct child. - Shared-store suites run in the serialized lbug-db vitest project and clear an inherited GITNEXUS_SHARED_STORE. - Doc and test accuracy fixes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(storage): pin shared-store behavior for worktrees of a bare repository (#3374) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address second review round on the shared store (#3374) - Reject --no-share in a linked worktree before taking any lock or indexing anything. - clean --all and remove also delete each shared checkout's pointer. - Delete an empty store only while also holding its cache lock, and re-check emptiness under it. - A store pointer is trusted only when the slot's own metadata names this checkout; the editable pointer file just says where to look. - Device names with an extension are prefixed on Windows only, so POSIX slot names stay stable. - Test fixtures use the gitnexus-test- prefix the stale-sidecar sweep recognizes; help text and the private-graph label are accurate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address third review round on the shared store (#3374) - clean --gc never collects a slot whose index lock is held: an analyze holds it until it registers the checkout, so a seeded but not yet registered slot is busy, not orphaned. - Reclaim compares absolute paths, so a relative GITNEXUS_HOME does not make live commit graphs look unreferenced. - graphPath is followed only into a published <commit>-<featureKey> dir, never .publish-* staging (TS and all three hook copies). - clean --gc fails on an unreadable stores root instead of reporting nothing to collect. - --no-share help states it is for opted-in clones. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): land the English --no-share help and private-graph label (#3374) These two strings were meant fore2a9467andd746675but were left out of the commit; the zh-CN side landed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address fifth review round on the shared store (#3374) - analyze --no-share leaves the shared store even when routed to a branch sub-index (gate on sharedStore, not placement.branch) - clean --gc aborts a store whose member listing is unreadable instead of treating it as empty, and skips non-directory entries under stores/ - reusing an existing commit graph no longer fails a finished analyze when the redundant private graph cannot be wiped - a store pointer holding JSON null, a number, a string, or an array is treated as invalid - clean, remove, and DELETE /api/repo hold the checkout slot's index lock while deleting the slot - the server analyze launcher waits on the checkout slot the worker actually writes - sync the Factory hook copy; isolate hook tests from storage overrides; use a junction for the Windows worktree alias; the relative GITNEXUS_HOME test now sets a relative home Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address sixth review round on the shared store (#3374) - clean, remove, and analyze --no-share remove the checkout's store pointer while still holding the slot's index lock, so an analyze that takes the lock next cannot have its new pointer deleted - clean --gc does not follow a symlink under stores/ (lstat) - removing a store pointer keeps the checkout's .gitnexus directory when it cannot be listed, instead of treating it as empty Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address seventh review round on the shared store (#3374) - clean --gc aborts when the registry cannot be read instead of treating every member as orphaned; with no registry file it collects only members whose checkout directory is gone - seeding from a repository-local index re-reads its metadata under the index lock, so the copied graph and the saved metadata match - clean --gc skips a store another collector removed meanwhile, and reports "no shared stores" when stores/ holds only stray files Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address eighth review round on the shared store (#3374) - removing a checkout's storage unregisters it and removes its store pointer before deleting the slot directory, which the file lock backend uses for its lock file; withCheckoutSlotLock becomes removeCheckoutStorage, used by remove, clean, and DELETE /api/repo, and analyze --no-share follows the same order - the clean --gc preview skips a slot whose index lock is held, matching what --force would drop Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): share the index store between clones automatically (#3352) Clones of one repository now share like linked worktrees. A clone whose normalized origin URL matches another registered, still-present clone joins that clone's store, or founds one (keyed on its own path) that the sibling joins on its next analyze. A lone clone keeps its own .gitnexus. Graphs stay keyed by commit and feature key, so clones only ever share a graph built from the same commit with the same settings. `analyze --no-share` now records a lasting opt-out (`shareOptOut` on the registry entry, preserved across re-registration); `--share-with` clears it. The analyze worker reports the storage it wrote over IPC so the server settles a clone's first shared slot. A query-time base-plus-overlay graph stays out: LadybugDB reads one database per query. Instead, private graph copies record whether the filesystem cloned them copy-on-write (sharing unchanged pages on disk) or made a full copy, and `status` reports it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): never follow a symlinked checkout .gitnexus (#3374) A checkout whose `.gitnexus` is a symlink (e.g. `.gitnexus -> ..`) made the shared-store pointer helpers operate on whatever it pointed at: findLegacyLocalIndex listed the target as a "legacy index", so `clean --local-index --force` recursively deleted the checkout's siblings; writeSharedStorePointer wrote store.json/.gitignore there; and removeSharedStorePointer deleted store.json there and could rm -r the target directory. All four now go through probePointerDir, which lstats `.gitnexus` and accepts it only when it is a real directory whose realpath is realpath(checkout)/.gitnexus (mirroring stale-branch-slots.ts). Anything else is left untouched: no legacy index is reported or removed, and no pointer is written or removed. removeLegacyLocalIndex re-probes after sizing, just before its delete loop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): never share a graph from a checkout that hides files (#3374) Trigger: a sparse checkout, a skip-worktree/assume-unchanged entry, or an uninitialized submodule leaves `git status` clean, and the run's indexCoverage.dirtyPaths drops paths with no file hash, so publishSharedGraph published a graph missing those files as the commit graph. seedSharedSlot then copied the seed's lastCommit into the new pointer slot, so the next analyze of a sibling hit the up-to-date fast path without hashing. Fix: add isWorkingTreePristine (storage/git.ts): the unfiltered listWorkingTreeDirtyPaths must be empty (null fails closed) and every gitlink in the index must have a checked-out `.git`. publishSharedGraph requires it instead of isWorkingTreeDirty. seedSharedSlot keeps the seed's lastCommit only when the seed is at HEAD and the checkout is pristine, so a clean sibling still fast-paths onto the shared graph; any other seeded pointer gets an empty lastCommit, like seedFromLocalIndex, and its next run hash-diffs, then re-points at the commit graph on publish. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep graphs with pending embeddings out of the shared store (#3374) Trigger: an analyze that finished with embeddings still owed (meta carries embeddingCheckpoint) was published as the commit graph, because the feature key ignores the checkpoint and publish never checked it. The published meta kept the checkpoint while the graph could never change. The next run healed the embeddings privately, found the commit graph already there, wiped its own healed graph and pointed back at the partial one. Fix: publishSharedGraph treats a checkpointed graph as not shareable, so it stays private and no published commit graph carries a checkpoint. When the target commit graph exists but records a checkpoint, has fewer stats.embeddings than the private graph, or has unreadable metadata, the checkout keeps its private graph instead of re-pointing. The commit graph is left as is because other checkouts may read it. Embeddings sync and the server embed job already privatize a pointer slot before writing (ensurePrivateSharedGraph). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): rebuild a shared slot whose graph is missing (#3374) A publish interrupted between moving the checkout's graph into `.publish-<uuid>` staging and renaming staging onto the commit dir left slot metadata at HEAD with no graph and no graphPath; the next reclaim deleted the orphaned staging. The up-to-date fast path never checked that the graph exists, so every later analyze reported "Already up to date" over a checkout with no index. A failed restore in publishSharedGraph's catch had the same outcome, and also deleted the staging dir holding the only copy of the graph. - run-analyze: for a shared-store slot, stat the graph the checkout reads (resolveGraphPath for the flat slot, the branch slot's own lbug otherwise) before the fast path; if it is missing, force a full rebuild. An incremental run would diff nothing into a fresh, empty database. Private .gitnexus indexes are unchanged: they only lose their graph by hand, and metadata-only fast-path fixtures rely on that path. - publishSharedGraph: when moving the staged graph back fails and it is still in staging, keep the staging dir, log its path, and skip this run's reclaim, which would otherwise delete it. The next analyze finds no graph and rebuilds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): fail safe on unreadable slot metadata during reclaim (#3374) Trigger: reclaim counted commit-graph references with loadMeta, which returns null for a torn or unreadable gitnexus.json as well as for a missing one. A member whose metadata could not be read therefore "referenced nothing", and the graph it still used was deleted. Separately, in `clean --gc` an orphan slot that fs.rm could not delete threw out of reclaim, aborting the collection of every store after it. Fix: the reference loop reads slot metadata with a local loadMetaStrict. Only absent metadata (ENOENT/ENOTDIR, same legacy-mirror fallback as loadMeta) means no reference; any other read error or a parse failure throws with the file path, like listDirStrict. Every caller already treats a reclaim throw as best effort (analyze logs "skipped cleanup", slot removal ignores it) or surfaces it (clean --gc). A failed orphan slot delete is now kept: it stays a member, so the graph it names is kept too, and it is reported in ReclaimResult.keptMembers and by a new clean --gc line. loadMeta itself is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(server): unregister under the slot lock and answer 409 on DELETE /api/repo (#3374) DELETE /api/repo called removeCheckoutStorage(storagePath) with no unregister callback and no checkout path, swallowed every error, and unregistered later outside the slot lock. For a shared-store slot this left the checkout's store.json pointer behind, let an analyze re-register between the slot removal and the unregister, and reported {deleted} even when the slot lock could not be taken because an analyze held it. The handler now calls removeCheckoutStorage(storagePath, () => unregisterRepo(entry.path), entry.path), matching `gitnexus remove` and `gitnexus clean`, so the unregister and the pointer removal run under the slot lock; non-shared storage is still removed, then unregistered. The later standalone unregister is gone. An IndexLockTimeoutError answers 409 and leaves the entry registered; any other failure propagates to the handler's 500, also with the entry kept so the delete can be retried, instead of unregistering over leftover index files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): give concurrent sibling clones one deterministic founder key (#3374) Trigger: two registered clones of one origin, both still on local storage, analyzing at the same time each saw no sibling store yet and founded a store keyed on their own checkout path. registeredStore() then kept each clone in its own store for good, so they never shared a graph. Fix: when no sibling store exists, siblingCloneStore keys the new store on the canonical path that sorts first among this clone and its registered siblings (case-folded on Windows, like registryPathEquals), so every sibling computes the same key. An existing sibling store still wins. Clones that already diverged into two stores are not migrated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep subdirectories of a clone out of shared stores (#3374) Trigger: `analyze --skip-git <clone>/pkg` inside a clone with a registered sibling of the same origin joined, or founded, a clone store. getRemoteUrl answers from any subdirectory, so siblingCloneStore saw the enclosing clone's remote. That broke the shared-store invariant that only tree roots participate. Fix: resolveOptedInStore now returns undefined for a path with no `.git` entry. It uses hasGitDir, which accepts a directory or a linked-worktree file, the same as resolveSharedStore's readCommonDir gate and run-analyze's repoHasGit. `--share-with` from such a path throws a clear error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(hooks): prefix Windows device names with an extension in hook slot names (#3374) Trigger: on Windows, storage-slot.ts sanitizeSlotBasename prefixes a device-name basename that has an extension (`CON.txt` -> `repository-CON.txt`), but the hook's copy in registry-query.cjs only matched the bare device name. With GITNEXUS_STORAGE_ROOT set, a checkout named e.g. `con.txt` got a different slot from the hook than from the CLI, so the hook could not see its index. Fix: mirror the platform branch (extension form on win32, bare form elsewhere) in all four byte-identical registry-query.cjs copies, and point their header comments at storage-slot.ts. Add a hook-vs-TS parity test over device-name basenames on both stubbed platforms, and fix the byte-identity test title to say four copies. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(storage): tidy the shared-store fix series (#3374) Explain the keptStaging early return where it is checked, reuse the exported EmbeddingCheckpoint type in the publish test, and drop review-step labels from test comments. No behavior change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address ninth review round on the shared store (#3374) - removeCheckoutStorage: keep the lock-safe order (unregister and drop the pointer under the slot lock, delete the slot last), but when the final rm fails, throw an error that says the checkout was already unregistered, names the leftover slot, and points at `gitnexus clean --gc --force`. - clean --gc preview no longer sweeps staging files: acquireIndexLock takes `sweep: false`, which the dry-run reclaim passes. - listStoreMetaRoots reports completeness; a store directory that cannot be listed (other than missing) makes the parse-cache prune retain every chunk instead of evicting keys other members still use. - clean.shared.kept says "could not be removed" rather than "still open". - ARCHITECTURE.md describes the stores/<key>/ layout and store.json pointer; README lists every GITNEXUS_SHARED_STORE off value and that it also stops clone sharing. - Tests: close the direct Ladybug connection in finally; fix a fixture JSDoc. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(storage): pin the publish gate for cone and sparse-index checkouts (#3374) isWorkingTreePristine already rejects every sparse mode, because each one marks left-out entries skip-worktree and `git ls-files -v` expands a sparse index. Only no-cone was tested; cover cone and cone with a sparse index too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(storage): name the checks that answer recurring review findings (#3374) Comment-only. The review bot re-derives findings from the code on every push, so put the refuting fact next to each flagged line: the pristine check's hidden-path coverage, sparse checkouts being skip-worktree, the lock-safe unregister-then-delete order, the hook parity test for device names, and why the stores/ lstat is not a race guard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
66 KiB
Architecture — GitNexus
Monorepo: CLI/MCP (gitnexus/) + browser UI (gitnexus-web/).
Repository layout
| Path | Role |
|---|---|
gitnexus/ |
npm package gitnexus: CLI, MCP server (stdio), HTTP API, ingestion pipeline, LadybugDB graph, embeddings. |
gitnexus-web/ |
Vite + React thin client: graph explorer + AI chat. All queries via gitnexus serve HTTP API. |
gitnexus-shared/ |
Shared TypeScript types and constants (consumed by CLI and Web). |
.claude/, gitnexus-claude-plugin/, gitnexus-cursor-integration/ |
Agent skills and plugin metadata. |
eval/ |
Evaluation harnesses for benchmarking tool usage. |
.github/ |
CI workflows + composite actions (setup-gitnexus/, setup-gitnexus-web/). |
End-to-end flow: index → graph → tools
-
Ingestion —
analyze.ts→runFullAnalysis(run-analyze.ts) →runPipelineFromRepo(pipeline.ts). The default DAG of 19 phases builds aKnowledgeGraphin memory, then loads into LadybugDB under.gitnexus/. Repo registered in~/.gitnexus/registry.jsonfor MCP discovery. -
Persistence —
repo-manager.ts(paths, registry, LadybugDB cleanup).lbug-adapter.ts(graph load, queries, embedding batches). -
Query layer — three interfaces to the same backend:
- MCP (stdio):
mcp.ts→LocalBackend→ tools (tools.ts) + resources (resources.ts) - HTTP bridge:
serve.ts→ Express (api.ts,mcp-http.ts) for web UI - CLI direct:
gitnexus query|context|impact|cypherintool.ts
- MCP (stdio):
-
Staleness —
core/git-staleness.tscompares indexedlastCommittoHEADand classifies the result ascurrent,behind,diverged(HEAD moved off the indexed commit, gap uncountable) orunknown;core/staleness-status.tsbuilds the onestalenesspayload that MCPlist_repos, the read tools and theserverepo routes all emit.
MCP tools
| Tool | Purpose |
|---|---|
list_repos |
Discover indexed repos |
query |
Hybrid BM25 + vector search over the graph |
cypher |
Ad hoc Cypher against the schema |
context |
Callers, callees, processes for one symbol |
impact |
Blast radius (upstream/downstream) with risk summary |
detect_changes |
Map git diffs to affected symbols and processes |
rename |
Graph-assisted multi-file rename with dry_run preview |
api_impact |
Pre-change impact report for an API route handler |
trace |
Shortest directed path between two symbols (call + class-member edges); group-aware (repo: "@<group>") for cross-repo traces |
route_map |
API route → handler → consumer mappings |
tool_map |
MCP/RPC tool definitions and handlers |
shape_check |
Response shape vs consumer property access mismatches |
explain |
Persisted taint findings (source→sink data flows) — needs analyze --pdg |
pdg_query |
Control/data dependence — CDG (mode: controls) / REACHING_DEF (mode: flows) — needs analyze --pdg |
group_list |
List repo groups or details for one group |
group_sync |
Rebuild group Contract Registry (contracts.json) and bridge graph |
query, context, and impact are group-aware: pass repo: "@<groupName>" (or "@<groupName>/<memberPath>" to scope to one member) plus optional service: "<monorepo/path>". Group-mode query merges per-repo results via Reciprocal Rank Fusion; group-mode impact runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (gitnexus/src/core/group/cross-impact.ts). trace is also group-aware via repo: "@<groupName>" — but, unlike the others, it resolves from/to across all members (a @<groupName>/<memberPath> suffix is advisory for trace, not a scope); pass from_uid/to_uid to disambiguate a symbol name that occurs in more than one member.
Group-mode trace (gitnexus/src/core/group/cross-trace.ts) stitches a path that crosses repositories: it resolves from/to across all members, and when they live in different repos it joins the home-repo segment to the target-repo segment over a single ContractLink boundary (an HTTP consumer→provider link, joined on Contract.symbolUid), reported as a CONTRACT_LINK hop in crossings[]. The crossing is clamped to one boundary (MAX_SUPPORTED_CROSS_DEPTH, shared with cross-impact); deeper crossDepth is reported via notes[]. With pdg: true (experimental, opt-in), each boundary-adjacent segment is enriched with its intra-procedural REACHING_DEF data-flow when that repo was indexed with --pdg (reusing the same anchored flows query as pdg_query); data flow never crosses the repo boundary, and a missing PDG layer degrades to call-level hops with a note. Two stores meet only at the symbolUid grain — the per-repo PDG/call graph and the group bridge — so this is the documented join; full cross-program (SDG-like) data flow across the boundary remains deferred (see docs/plans/2026-06-18-002-feat-unified-pdg-impact-evaluation-plan.md). The previously-planned group_query, group_context, group_impact, group_contracts, group_status MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
| Resource URI | Purpose |
|---|---|
gitnexus://group/{name}/contracts |
Contract Registry (provider/consumer rows + cross-links) |
gitnexus://group/{name}/status |
Per-member index + Contract Registry staleness |
Where to change what
| Concern | Start in |
|---|---|
| CLI commands/flags | src/cli/ (index.ts, per-command modules) |
| Parsing/graph construction | src/core/ingestion/pipeline-phases/ + pipeline.ts |
| Graph schema/DB | src/core/lbug/ (schema.ts, lbug-adapter.ts) |
| MCP tools/resources | src/mcp/server.ts, tools.ts, resources.ts |
Cross-repo groups (sync, contracts, @<group> routing) |
src/core/group/ (service.ts, cross-impact.ts, sync.ts, bridge-db.ts) |
| Search ranking | src/core/search/ (BM25, hybrid fusion) |
| Embeddings | src/core/embeddings/ + src/core/run-analyze.ts |
| Wiki generation | src/core/wiki/ |
| Language support | src/core/ingestion/languages/ + tree-sitter-queries.ts + gitnexus-shared/src/languages.ts |
| Import resolution | src/core/ingestion/import-processor.ts + import-resolvers/configs/ + model/resolution-context.ts |
| Call resolution/inheritance/MRO | src/core/ingestion/scope-resolution/ (pipeline, passes, graph-bridge) |
| Type extraction | src/core/ingestion/type-extractors/ |
| Worker pool | src/core/ingestion/workers/ |
| Web UI | gitnexus-web/src/ |
| CI | .github/workflows/*.yml, .github/actions/ |
Paths above are relative to
gitnexus/unless they start withgitnexus-web/or.github/.
Pipeline Phase DAG
19 default phases are defined in gitnexus/src/core/ingestion/pipeline-phases/, each with explicit deps and typed output. --pdg adds taintSummaries and callSummaries (21 total).
scan → structure → [springConfig, markdown, cobol] → parse → [routes, tools, orm]
→ crossFile → scopeResolution → [springAutoConfiguration, springAop]
→ pruneLocalSymbols → mro → springAopInheritance → di → communities → processes
| Phase | File | Deps | Output |
|---|---|---|---|
scan |
scan.ts |
(root) | File paths + sizes |
structure |
structure.ts |
scan |
File/Folder nodes, CONTAINS edges, allPathSet |
springConfig |
spring-config.ts |
structure |
Spring configuration-property nodes and metadata |
markdown |
markdown.ts |
structure |
Section nodes, cross-link edges from .md/.mdx |
cobol |
cobol.ts |
structure |
COBOL program/paragraph/section nodes (regex, no tree-sitter) |
parse |
parse.ts + parse-impl.ts |
structure, markdown, cobol |
Symbol nodes, IMPORTS/CALLS/EXTENDS edges, extracted routes/tools/ORM queries |
routes |
routes.ts |
parse |
Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators, and JS/TS static route sources — see below) |
tools |
tools.ts |
parse |
Tool nodes + HANDLES_TOOL edges |
orm |
orm.ts |
parse |
QUERIES edges (Prisma, Supabase) |
crossFile |
cross-file.ts + cross-file-impl.ts |
parse, routes, tools, orm |
Cross-file type propagation in topological import order |
scopeResolution |
scope-resolution/pipeline/phase.ts |
parse, crossFile, structure |
Binding/reference + inheritance edges; disposes BindingAccumulator |
springAutoConfiguration |
spring-auto-configuration.ts |
structure, scopeResolution |
DECLARES and CONDITIONAL_ON metadata for Spring configuration candidates |
springAop |
spring-aop.ts |
scopeResolution |
Direct declarative/advice ADVISED_BY edges and pointcut evidence |
pruneLocalSymbols |
prune-local-symbols.ts |
scopeResolution |
Drops inert block-local Const/Variable/Static nodes (only a File→DEFINES edge) post-resolution |
mro |
mro.ts |
crossFile, scopeResolution, pruneLocalSymbols, structure |
METHOD_OVERRIDES + METHOD_IMPLEMENTS edges |
springAopInheritance |
spring-aop.ts |
springAop, mro |
Propagates declarative behavior through class/interface inheritance decisions |
di |
di.ts |
mro |
INJECTS edges from consumer Classes, factory Methods, or AST-captured programmatic lookup callables to provider Classes/declaration CodeElements (framework-neutral DI resolution; per-language matchers registered in di-extractors/) |
communities |
communities.ts |
mro, pruneLocalSymbols, structure |
Community nodes + MEMBER_OF edges (Leiden algorithm) |
processes |
processes.ts |
communities, routes, tools, pruneLocalSymbols, structure |
Process nodes + STEP_IN_PROCESS edges |
Non-phase files in the same directory: parse-impl.ts, cross-file-impl.ts (implementation), wildcard-synthesis.ts (whole-module import expansion), types.ts, runner.ts, index.ts.
DAG runner
runner.ts — static phase graph, no plugins, compile-time type safety.
-
Validation — Kahn's topological sort. Rejects on: duplicate names, missing deps, cycles (DFS traces the concrete cycle path, e.g.,
A -> B -> C -> A, plus count of transitively blocked dependents). -
Execution — sequential in topological order. Each phase receives:
ctx: PipelineContext— shared mutableKnowledgeGraph,repoPath, progress callback, optionsdeps: ReadonlyMap<string, PhaseResult>— declared deps only (runner filters the results map to prevent hidden coupling)
-
Error handling — wraps phase errors with the phase name, emits terminal
errorprogress event, swallows progress handler errors to preserve the original cause. -
Timing — per-phase
durationMsinPhaseResult, dev-mode console logging.
Design patterns:
- Single graph accumulator — all phases mutate the same
KnowledgeGraphinctx; the graph is the primary output. - Typed phase access —
getPhaseOutput<T>(deps, 'name')for type-safe upstream results. - Binding accumulator lifecycle — created in
parse, disposed bycrossFile(infinally). No other phase should take ownership. - Skippable phases —
skipGraphPhasesomits MRO/di/communities/processes (faster tests);pruneLocalSymbolsstill runs (it is graph cleanup, not analysis).skipWorkersis no longer a sequential escape hatch — it (like--workers 0/GITNEXUS_WORKER_POOL_SIZE=0) is rejected with an actionable error, since the worker pool is the sole parse path (§ Chunked parse-and-resolve). - Local-symbol pruning —
pruneLocalSymbolsremoves inert block-local value symbols after scope resolution has consumed them. Opt out per-call withPipelineOptions.keepLocalValueSymbolsor globally with theGITNEXUS_KEEP_LOCAL_VALUE_SYMBOLSenv var.
How to add a new phase
- Create
pipeline-phases/my-phase.tswith aPipelinePhase<MyOutput>(name, deps, execute) - Export from
pipeline-phases/index.ts - Add to
buildPhaseList()inpipeline.ts
import type { PipelinePhase, PhaseResult } from './types.js';
import { getPhaseOutput } from './types.js';
import type { ParseOutput } from './parse.js';
export interface MyPhaseOutput {
/* ... */
}
export const myPhase: PipelinePhase<MyPhaseOutput> = {
name: 'myPhase',
deps: ['parse'],
async execute(ctx, deps) {
const { allPaths } = getPhaseOutput<ParseOutput>(deps, 'parse');
// ... write to ctx.graph ...
return {
/* typed output */
};
},
};
Where routes come from
route-extractors/ holds four independent ways a route can be discovered, all
converging on the routes phase's (method, url) registry:
| Source | Shape | Examples |
|---|---|---|
| Filesystem convention | path → URL, no parsing | Next.js app/, Expo, PHP |
| Single-file framework route | isRouteFile + worker extraction |
Laravel routes/*.php |
| Cross-file framework route | discoverRootRouteFiles + extractRoutes |
Django urlpatterns |
| AST-level route in a normal file | extractDecoratorRoutes |
Spring, FastAPI, NestJS (@Controller + @Get/@Post/…; URLs are controller-relative — setGlobalPrefix and URI versioning live in the bootstrap file and are not applied), JS/TS dispatch guards and static data route tables |
The last row is the one whose name undersells it. A route is DECLARED by a
decorator, but it can also be inferred from a raw node:http server's own
dispatch — if (req.method === 'GET' && pathname === '/api/x') is a route with
a path, a verb and a handler, and nothing else in the pipeline could see it.
route-extractors/dispatch-guard.ts reads that shape; the transport, dedup and
handler resolution are shared with decorator routes, and
ExtractedDecoratorRoute.source carries the provenance difference through to
the HANDLES_ROUTE edge.
JS/TS data route tables share that transport when a route-named array contains
direct object literals with static path, method, and handler fields and a
same-scope for...of dispatcher positively compares the path and method before
directly invoking the handler. Dynamic values, computed keys, spreads,
inline/called handlers, unknown verbs, and ambiguous handler bindings are
suppressed. Bare import aliases and single-level member handlers are attributed
only through declared import and owner provenance; an unproven receiver never
falls back to a global name guess.
That extractor is deliberately precision-weighted: route_map presents its
output as fact, so a startsWith namespace test, a bare pathname === '/'
without a verb, and any regex it cannot translate exactly are all dropped rather
than guessed at. A missing route is a coverage limit; an invented one is a lie.
Two rules there need more than one comparison to decide, and are worth knowing about before changing either:
- Same-file constant folding.
pathname === `${basePath}/rules`is common enough that refusing it loses whole route modules — and loses them invisibly, since a module with unfoldable paths and a module with no routes produce the same empty answer. Folding is same-file, string literals only, one alias hop, and refuses on ambiguity (a name declared twice with different values is dropped, never guessed). - Whole-repo reconciliation (
reconcileDispatchGuardRoutes, applied in the routes phase). A split route table — one module listing every path it recognises so the dispatcher can 404 early, handlers in others — otherwise lists every route twice, once verb-less with the table as its "handler". It applies to dispatch-guard routes only: a framework route with no verb is method-agnostic by declaration, which is a fact, not a weaker observation.
Semantic model
SemanticModel (gitnexus/src/core/ingestion/model/semantic-model.ts) is the authoritative store for every symbol-indexed lookup (by nodeId, simpleName, qualifiedName, or filePath). The scope-resolution pipeline reads from here: findOwnedMember, pickOverload, and findExportedDefByName all consult model.methods / model.fields / model.symbols.
ParsedFile (gitnexus-shared/src/scope-resolution/parsed-file.ts) is the single per-file artifact the scope-resolution pipeline consumes. Scope-resolution passes MUST NOT build a parallel parse representation. If a per-language hook needs AST-level facts that ParsedFile doesn't expose, it should reuse the orchestrator's treeCache (RunScopeResolutionInput.treeCache) rather than re-invoking parser.parse(...) on its own — the C# populateNamespaceSiblings hook is the reference implementation of this pattern.
The scope-resolution pipeline additionally carries WorkspaceResolutionIndex for Scope-valued lookups (classScopeByDefId, moduleScopeByFile) that SemanticModel structurally cannot hold. No symbol-indexed duplicates exist outside SemanticModel.
Write / read phase contract. The model is mutable during three ordered phases and read-only afterward:
Phase 1: parse ──► symbolTable.add fans into types/methods/fields
Phase 2: scope-resolution ──► reconcileOwnership() registers corrected ownerIds
Phase 3: finalize ──► model.attachScopeIndexes(bundle) — one-shot freeze
─────────────────────────── phase boundary ───────────────────────────
Read phase: all resolution passes + MCP + HTTP + embeddings see
SemanticModel (read-only handle); writes are type-errors.
runScopeResolution narrows MutableSemanticModel → SemanticModel at the phase boundary so downstream passes physically cannot mutate the model even accidentally.
Reconciliation pass. reconcileOwnership (scope-resolution/pipeline/reconcile-ownership.ts) is a shim for languages whose parse-time extractor doesn't resolve enclosingClassId at parse time (Python class-body methods are the canonical case). It walks parsed.localDefs[i].ownerId after populateOwners and registers any missed methods/fields into the model. Idempotent — safe to re-run, safe alongside languages whose extractor already carries ownerId (C#).
The architectural end state is for every language's parse-time extractor to emit the correct ownerId directly, making reconciliation a no-op (tracked as a follow-up refactor). The dev-mode validator validateOwnershipParity surfaces any drift via onWarn under NODE_ENV !== 'production' && VALIDATE_SEMANTIC_MODEL !== '0'.
References: semantic-model.ts file-head (full write/read contract); contract/scope-resolver.ts Contract Invariant I9 (scope-resolution-side rule).
Scope-Resolution Pipeline (RFC #909 Ring 3)
Language-agnostic scope-resolution resolver. This is the resolution path for every language — it owns CALLS/ACCESSES/USES emission and inheritance edges. Adding a language is one interface implementation (ScopeResolver) plus one registration in the SCOPE_RESOLVERS map — no changes to shared code, no new pipeline phase. (RING4-1 #942 removed the legacy call-resolution DAG and the per-language MIGRATED_LANGUAGES flag, so SCOPE_RESOLVERS registration is all that's needed.)
Pipeline stages
ParsedFile[] (extractParsedFile per file)
│ finalizeScopeModel (+ provider hooks)
▼
ScopeResolutionIndexes
│ resolveReferenceSites (via MethodRegistry.lookup)
▼
ReferenceIndex
│ emitReceiverBoundCalls ── FIRST
│ emitFreeCallFallback ── THEN
│ emitReferencesViaLookup ── uses handledSites + deferred-site skip set
│ emitPropertyDispatchCalls ── registration USES + conservative CALLS
│ emitCallableValueFlow ── assigned/passed callable invocation CALLS
│ emitImportedValueReferences ── cross-file value reads via finalized imports
│ emitUniqueNamePropertyAccesses ── LAST-RESORT property reads by name,
│ narrowed same-file → direct-import, refusing to choose otherwise
│ emitImportEdges
▼
KnowledgeGraph (IMPORTS / CALLS / ACCESSES / INHERITS / USES)
Orchestrator: runScopeResolution(input, provider) in scope-resolution/pipeline/run.ts.
Pipeline phase: scopeResolutionPhase in scope-resolution/pipeline/phase.ts — iterates the registered SCOPE_RESOLVERS over the worker-serialized ParsedFiles. (Per-language emitScopeCaptures hooks may reuse a cached Tree via the orchestrator's treeCache, but in worker-pool runs that cache is empty — Trees can't cross MessageChannels — so they consume the pre-extracted ParsedFile instead; § Performance notes.)
Callable-value flow
First-class callable values use a language-neutral inclusion analysis in passes/callable-value-flow.ts. Providers recognize their own syntax and emit JSON-safe CallableFlowSite facts (seed, copy, alias, address, load, store, formal, argument, and invoke) into ParsedFile; shared ingestion never branches on a language name. These always-on facts cross workers and the durable parse store, whose schema is bumped whenever their semantic shape changes.
The emit stage defers only invocation sites proven to reference a flow cell. Ordinary receiver/free/reference passes still resolve direct callees first and record exact callee IDs by file/line/column. Property dispatch then runs before callable flow because a property-dispatched wrapper call can seed actual-to-formal propagation. The callable solver consumes those direct targets, propagates callable sets through lexical cells and formals, and emits CALLS at the real indirect invocation site with reason callable-value-flow (confidence 0.8 for a singleton, 0.7 for a bounded multi-target set).
The solver is flow-insensitive but bounded: dependency-indexed work items rerun only when a cell they read changes; target/address sets cap at 32; a hostile fact graph has a finite work budget; overflow or budget exhaustion emits no partial CALLS and produces a structured warning. Lexical shadowing is function/block aware, invocation/constructor results are not reinterpreted as callable designators, and overload selection uses provider-supplied signature metadata. C/C++ additionally associate visible prototypes with unique definitions so actual-to-formal flow crosses translation units; the provider-owned hasFileLocalCallableLinkage hook prevents static declarations or definitions from leaking across files. C++ member-function pointers preserve parameter/cv shape, keep non-virtual targets exact, and expand virtual targets through MethodDispatchIndex/MRO.
Property-key dispatch remains a separate conservative fallback. Its per-key fan-out cap is 32; capped keys synthesize no partial calls and are reported at warning level with language, skipped-key count, dropped key names (bounded), and cap; the count also travels in RunScopeResolutionStats.propertyDispatchSkippedKeys.
Interface-dispatch fan-out walks the subtype closure of the receiver's interface and is generic-instantiation aware (#2912): a call through IValidator<string> must not reach an implementor of IValidator<int>, which shares its declaration and therefore its subtype list. Each heritage clause's arguments reach resolution by one of three routes — read off the @reference.inherits anchor's own spelling where that anchor spans the whole base (most languages, no query change), through the @reference.type-arguments sub-tag where the anchor is the bare name and moving it would renumber inheritance edge ids (Rust impl T<A> for S, Dart extends), or on a heritage MARKER payload for clauses that never become reference sites (Dart implements/with). Whichever pass emits the edge records the pair through one sink: preEmitInheritanceEdges for heritage clauses, ScopeResolver.emitHeritageEdges for the rest.
The walk then carries a substitution: a subtype's own type parameters bind to the receiver's arguments, so class Wrapper<T> : IValidator<T> stays reachable from every instantiation while class IntValidator : IValidator<int> is pruned from the string one. Receiver arguments come from the declared type (Case 4), a class-level field's declared type (Case 6), or — for a compound receiver such as this._repo — the spelling the compound fold typed that position from, reported back through recordReceiverType and accepted only when it names the class the fold returned.
The filter prunes only on positive evidence: an unknown instantiation on either side, an argument list whose arity does not line up, a name that may be a type variable the language's captures never recorded, or an unresolved spelling whose simple name matches all keep the target. A type parameter of the declaration ENCLOSING either side is recognised as such and never compared — void Run<T>(IValidator<T> v) writes a receiver with no known instantiation, so it keeps the unfiltered fan-out. That recognition is what generic METHODS now carry @declaration.type-parameters for in C#, Java and Kotlin (TypeScript already did): without it an unbounded T grounds to nothing and a bounded one grounds to its BOUND, and both compare unequal to an implementor's concrete argument. Languages that capture neither type arguments nor type parameters therefore emit exactly the pre-#2912 fan-out. The fan-out cap (32, GITNEXUS_MAX_INTERFACE_DISPATCH_FANOUT) and its skipped-target reporting are unchanged and apply after filtering. Note the fan-out itself still fires only for a receiver whose folded type is an Interface symbol, so a Rust Trait or a Dart abstract Class receiver emits no secondary targets to filter in the first place.
Standalone (regex-based) providers such as COBOL participate via ScopeResolver.scopeResolutionEdgeMode: 'callable-flow-only': runScopeResolution runs for them, but every ordinary emission path — heritage, interface implementations, receiver-bound, free-call fallback, reference/import edges, post-resolution hooks — is gated off, so their legacy phase (e.g. cobolPhase) remains the sole owner of structural edges and the callable solver's CALLS are purely additive. A callable-flow-only provider whose files emitted no callable facts exits early, before finalize, keeping the opt-in proportional to source scanning.
Receiver chains and the drop census (#2766)
A compound receiver (svc.getUser().address.save()) is captured as a compact string on ReferenceSite.receiverChain. utils/receiver-chain-codec.ts is the ONE encoder/decoder — capture emitters, the scope-resolution fold, and the durable ParsedFile store all import it rather than hand-rolling the format.
Wire format is v2: 2|<base>|<step>|<step>…, one-character version prefix, then base-first steps, each a one-character kind sigil plus the member name (c = call, f = field). a (await) and i (index) are name-free and encode as a bare sigil — an awaited call's name already lives on its c step, and a subscript key is a value, not a lookup-able identifier. The version went 1 → 2 when those two kinds were added, and a decoder REFUSES a foreign version rather than decoding the prefix it understands: a chain missing its await/index hop decodes cleanly as a different, shorter chain and would type the receiver against the wrong member. The format is unescaped (| and ~ cannot occur in an identifier), so an unencodable name is refused rather than escaped, and the payload is capped at MAX_RECEIVER_CHAIN_BYTES / MAX_CHAIN_DEPTH steps. Because these strings live in the incremental parse cache and the durable ParsedFile store, a format change requires a PARSE_CACHE_VERSION schema bump — a stale cache would otherwise replay v1 chains this build discards.
Receivers the resolver could not type are not silently dropped. Each records a ResolutionOutcome (scope-resolution/resolution-outcome.ts) carrying the receiver's shape (classifyReceiverShape: chain-call / chain-field / chain-mixed / chain-unwrap / no-chain — the bench censuses these) and its origin (in-program / external / unknown). scope-resolution/unresolved-receivers.ts aggregates them per member name into the index-persisted unresolvedReceiverMembers summary, keeping in-program and external counts under separate keys. Only in-program drops make a count short: an external-rooted call (System.out.println, fetch(...)) has no in-graph node an edge could have reached, so it is reported but does not hedge. impact / context read that summary and publish epistemic: 'exact' | 'lower-bound', prose boundaries, and the machine-readable causes split (EpistemicCauses in mcp/local/local-backend.ts).
Optional CFG/PDG emission (--pdg, #2081–#2086)
On a --pdg run the parse worker builds a per-function control-flow graph from the tree-sitter AST (LanguageProvider.cfgVisitor; TypeScript/JavaScript today) and serializes it onto ParsedFile.cfgSideChannel as plain data. Scope-resolution then emits the program-dependence layers from that side-channel inside Phase 4 of runScopeResolution, while the disk-backed ParsedFile store is still live — the only window where the worker-built CFGs are loaded (the store is cleared right after the phase returns). A standalone post-mro phase would read an empty store, so the emit deliberately lives in-phase, mirroring the applyCaptureSideChannel pattern. The opt-in is off by default (graph byte-identical), folded into the parse-cache key (a pdg-off warm cache is never reused on a --pdg run), and each layer is bounded by a per-function edge cap that logs any dropped edges. All layers are BasicBlock → BasicBlock edges in the single CodeRelation table, keyed by type; there is no Function → BasicBlock edge — the symbol↔block join is reconstructed from the BasicBlock id prefix + line span. The layers build on each other:
- M1 — CFG (#2081):
BasicBlocknodes +CFGedges. Edge kind (seq/cond-true/loop-back/…) rides thereasoncolumn (CFG is oneCodeRelationtype, not one per kind). - M2 — REACHING_DEF (#2082): GEN/KILL def→use data dependence from a pure fixpoint solver; the variable name rides
reason. - M3/M4 — TAINTED / SANITIZES / TAINT_PATH (#2083–#2084): intra- and inter-procedural taint (source→sink) — the
explaintool's data. - M5 — CDG (#2085): Ferrante control dependence over a Cooper–Harvey–Kennedy post-dominator tree (the EXIT-rooted reverse CFG); branch sense (
'T'/'F') ridesreason. A CFG whose EXIT is unreachable from some block is skipped for CDG (post-dominance would be unsound) while its CFG/REACHING_DEF layers are kept. - M6 — read surface (#2086): the
pdg_queryMCP tool answers "what gates X?" (CDG,mode: controls) and "where does Y flow?" (REACHING_DEF,mode: flows);explainis the taint consumer. Both are always anchored +LIMIT-bounded (LadybugDB has no rel-property index) and share oneresolveBlockAnchorhelper. These PDG edge types are deliberately kept out of the defaultVALID_RELATION_TYPES/ web schema. - Cross-repo trace enrichment: group-mode
trace(pdg: true) reuses the same anchored REACHING_DEFflowsquery to annotate a boundary-adjacent segment with how a value reaches the cross-repo call — strictly intra-procedural (data flow never crosses the repo boundary). See the group-aware tools note above.
See core/ingestion/cfg/ (emit + the pure CFG / post-dominator / control-dependence / reaching-defs / taint passes) and mcp/local/local-backend.ts (_pdgQueryImpl, _explainImpl, the shared resolveBlockAnchor).
ScopeResolver contract
Single interface a language implements to plug into the pipeline. Contract fully documented in scope-resolution/contract/scope-resolver.ts.
| Hook | Purpose |
|---|---|
languageProvider |
Base LanguageProvider (tree-sitter query, emitScopeCaptures, import/binding interpreters, hooks) |
populateOwners(parsed) |
Fill deferred ownerId fields on method defs (captures can't always know the owning class at parse time) |
buildMro(graph, parsed, nodeLookup) |
Produce mroByClassDefId: Map<DefId, DefId[]> — C3, Ruby-mixin, or first-wins per language |
resolveImportTarget(target, fromFile, allFiles) |
(rawImportPath, sourceFile) → targetFilePath (PEP-328 for Python, etc.) |
isNamespaceImport(parsedImport, targetFile, fromFile) |
Optionally reclassify a resolved named import as a namespace handle when the imported symbol is itself a module |
mergeBindings(existing, incoming, scopeId) |
Shadowing / LEGB precedence |
arityCompatibility |
Provider consumed by registry during MethodRegistry.lookup Step 2 |
importEdgeReason |
Confidence-tier string for IMPORTS edge reason field |
propagatesReturnTypesAcrossImports? |
Opt out of cross-file return-type propagation (default on) |
fieldFallbackOnMethodLookup? |
Statically-typed languages turn this OFF — the heuristic over-connects (default on) |
elementTypeOf? |
(containerType, via: {kind:'index'} | {kind:'accessor',name}) → elementType | undefined — element type of a container, reached by subscript (repos[0]) or by a property-style collection view (data.Values). ONE hook for both routes (it replaced the split unwrapCollectionAccessor / unwrapCollectionElement, where implementing one silently answered nothing for the other). Consulted only where the source actually performed the access — never as a general type-name normalizer |
stripTypePreservingDecoration? |
(typeName) → strippedName | undefined — strip ONE layer of TYPE-PRESERVING decoration (pointer, reference, const, nullable, borrow, sigil) so a receiver declared *Host still finds the Host binding (#2766). Never a container: unwrapping Repo[] here would fold repos.find(x) to Repo.find — that is elementTypeOf's job, and only after a real subscript. Consulted only after every undecorated lookup fails, and only by receiver-chain base/step resolution — default off |
collapseMemberCallsByCallerTarget? |
One CALLS edge per (caller, target) instead of per-site — default off |
populateNamespaceSiblings? |
Cross-file implicit visibility (compiler-implicit namespace sharing) — default off; ctx carries treeCache |
hoistTypeBindingsToModule? |
Walk up to Module scope when looking up a method's return-type typeBinding — default off; enable only when bindings are stored at module level |
hasFileLocalCallableLinkage? |
Precise internal-linkage predicate used only when joining callable declarations/prototypes to cross-file definitions; C/C++ use it for static free functions |
constructorCallTargetsClass? |
A constructor-form call Type(...) links to the Class def rather than its explicit Constructor def — default off; Swift and Dart opt in |
constructionSyntax? |
How the language spells construction, so an INLINE constructor receiver (Service(db).m(), new Service(db).m(), Service.new.m()) can be typed — bare / keyword / selector; default off, opt in per language only where measured to be needed (#2708) |
Per-language registration
- Implement
ScopeResolverinlanguages/<lang>/scope-resolver.ts. - Add entry to
SCOPE_RESOLVERSinscope-resolution/pipeline/registry.ts.
CI auto-discovers the set via tsx. No workflow edit required.
Code references
| Module | Purpose |
|---|---|
scope-resolution/contract/scope-resolver.ts |
ScopeResolver interface + shared types |
scope-resolution/pipeline/run.ts |
Generic orchestrator |
scope-resolution/pipeline/phase.ts |
Pipeline-phase wrapper (deps: parse, structure) |
scope-resolution/pipeline/registry.ts |
SCOPE_RESOLVERS map |
scope-resolution/passes/*.ts |
Reference-resolution passes (receiver-bound, free-call fallback, compound-receiver, MRO, cross-file return-type propagation) |
scope-resolution/graph-bridge/*.ts |
CLI-local translation from resolved references → KnowledgeGraph edges |
scope-resolution/scope/*.ts |
Generic scope-chain walkers + namespace targets |
scope-resolution/workspace-index.ts |
Build-once O(1) lookup index |
languages/python/index.ts |
Python ScopeResolver hooks + known-limitation docs |
languages/python/captures.ts |
emitPythonScopeCaptures (honors cross-phase Tree cache) |
languages/csharp/index.ts |
C# ScopeResolver hooks + known-limitation docs |
languages/csharp/captures.ts |
emitCsharpScopeCaptures (honors cross-phase Tree cache) |
languages/csharp/namespace-siblings.ts |
Cross-file implicit-namespace visibility hook (reads treeCache) |
Performance notes
- Cross-phase Tree cache: the orchestrator's
treeCache(RunScopeResolutionInput.treeCache) lets a scope-resolution per-language hook (emitScopeCaptures) reuse a tree instead of re-parsing. Workers leave it empty — Trees can't cross MessageChannels — so in normal (worker-pool) runs scope-resolution does NOT rely on it: workers serialize each file'sParsedFile(+ capture side-channel) and stream them in, so scope-resolution consumes the pre-extracted artifact rather than re-parsing on the main thread (§ Chunked parse-and-resolve).PROF_SCOPE_RESOLUTION=1emits hit/miss counters and a worker-engaged warning. - Typed relationship iteration: heritage + MRO walk only the EXTENDS / IMPLEMENTS / HAS_METHOD edges via
iterRelationshipsByType, not the full relationship map. - Workspace-resolution-index: O(1)
findOwnedMember/findExportedDef/classScopeByDefIdbuilt once per run. - Callable-value worklist: dependency-indexed inclusion propagation is linear in a reverse-ordered copy-chain fixture; target/address sets cap at 32 and the whole worklist has a finite budget with no partial edge emission on exhaustion.
- SCC-ordered cross-file return-type propagation (PR #1050):
propagateImportedReturnTypeswalksindexes.sccsin reverse-topological order (leaves first), so multi-hop alias chains likemodels.User → service.user → app.usercollapse to the terminal class in a single linear pass. Within each importer, the source module'stypeBindingsis chain-followed BEFORE mirroring (so we mirror terminal types, not intermediate refs), and the importer's owntypeBindingsis chain-followed AFTER mirroring (so localconst x = importedFn()resolves before downstream importers run). Cyclic SCCs reach a partial fixpoint within a single pass without iterating to convergence — see thets-circularcross-file-binding fixture which only asserts pipeline-no-throw. PROF output (PROF_SCOPE_RESOLUTION=1) splitsfinalizefrompropagateso quadratic regressions in the chain-follow surface independently.
Language-agnostic graph feeding
18 languages → single unified graph. Four abstraction layers:
Unified Graph Schema (44 node types, 21 relationship types)
↑
Scope-Resolution Pipeline (registry lookup + 3-tier import resolution + MRO)
↑
Language Providers (import semantics, type config, export checker, MRO strategy)
↑
Tree-Sitter Queries (per-language S-expressions, unified capture tags)
Language providers
Each language implements LanguageProvider (language-provider.ts). Key fields:
| Field | Purpose |
|---|---|
id, extensions |
Language identity and file matching |
treeSitterQueries |
S-expression queries for AST extraction |
importSemantics |
named / wildcard-leaf / wildcard-transitive / namespace |
importResolver |
Language-specific path → file resolution |
exportChecker |
Public/exported symbol detection |
typeConfig |
Type annotation extraction rules |
mroStrategy |
first-wins / c3 / none |
descriptionExtractor |
Optional hook returning a symbol's doc-comment text as its description; feeds the embedding metadata header so doc-only terms are semantically searchable (issue #2270). Most languages register createLeadingDocDescriptionExtractor (shared, language-neutral; per-language comment/wrapper config passed at the call site) |
definitionPropertiesExtractor |
Optional language-owned hook for structured, clone-safe definition metadata. Shared ingestion persists these properties opaquely; the owning provider supplies the extraction semantics. |
18 providers in languages/index.ts via satisfies Record<SupportedLanguages, LanguageProvider> — missing a language is a compile error.
Unified capture tags
Per-language tree-sitter queries use different AST node names but produce the same semantic capture tags: @definition.class, @definition.function, @call.name, @import.source, @reference.inherits. Downstream extraction needs no language branching. Defined in tree-sitter-queries.ts.
Import resolution
Per-language import resolution uses the configs + factory pattern (like call/method/class extractors). Each language declares an ImportResolutionConfig in import-resolvers/configs/, listing an ordered chain of ImportResolverStrategy functions. createImportResolver() (in resolver-factory.ts) composes them: first non-null result wins. Low-level helpers shared across strategies live alongside the configs in import-resolvers/ (e.g. go.ts, rust.ts, python.ts).
Unified 3-tier algorithm (model/resolution-context.ts), per-language importSemantics controls which tier activates:
| Tier | Confidence | Mechanism |
|---|---|---|
| 1 — same-file | 0.95 | Symbol table for caller's file |
| 2 — import-scoped | 0.9 | NamedImportMap chains (named) or all files in importMap (wildcard) |
| 3 — global | 0.5 | O(1) index lookups: class, impl, callable. Fallback only |
| Import strategy | Languages | Behavior |
|---|---|---|
named |
TS, JS, Java, C#, Rust, PHP, Kotlin | Only explicitly imported names visible |
wildcard-leaf |
Go, Ruby, Swift, Dart | Whole-package import, no transitive re-exports |
wildcard-transitive |
C, C++ | #include closure chains through re-exports |
namespace |
Python | Module aliases resolved at call site |
Chunked parse-and-resolve
parse processes files in ~20 MB byte-budget chunks to bound memory. Per chunk:
- Worker pool dispatches files (the sole parse path — there is no sequential fallback;
skipWorkers,--workers 0, andGITNEXUS_WORKER_POOL_SIZE=0are rejected with an actionable error) - Each worker: detect language → load grammar → run queries → return unified
ParseWorkerResult - Synthesize wildcard bindings (
wildcard-synthesis.ts) - Resolve imports
- Collect
BindingAccumulatorentries for cross-file propagation
Inheritance edges are emitted later, by the scope-resolution phase (preEmitInheritanceEdges + emitHeritageEdges), not during parse.
Workers: workers/worker-pool.ts, workers/parse-worker.ts.
Worker-serialized ParsedFiles (#2038). To index very large repos (e.g. the Linux kernel) without OOM, the worker pool is the sole parse path and workers serialize each file's ParsedFile (plus its capture side-channel) in parallel, streaming them to scope-resolution through a disk-backed store. Scope-resolution consumes the pre-extracted artifact instead of re-parsing every file on the main thread — tree-sitter's native input buffers are not GC-reclaimable, so the former main-thread re-parse leaked native memory until the process died. Pool creation is lazy / cache-miss-gated, so a warm all-cache-hit run replays cached worker output without spawning a worker (hence usedWorkerPool can be false even when the repo has parseable files).
Inheritance and MRO
Inheritance is captured by the @reference.inherits tag and emitted by the scope-resolution phase: preEmitInheritanceEdges resolves each base in scope, then emitHeritageEdges writes the EXTENDS/IMPLEMENTS edges. The phase then computes method resolution order via each ScopeResolver's buildMro hook, feeding a MethodDispatchIndex used for owner-scoped lookups. Per-language strategy:
first-wins— Java, C#, C++, TS, Ruby, Goc3— Python (C3 linearization)ruby-mixin— Ruby (mixin-aware linearization)none— single-inheritance languages
Full analysis flow
runFullAnalysis in run-analyze.ts orchestrates everything around the pipeline:
CLI (analyze.ts) → runFullAnalysis(repoPath, options, callbacks)
1. Early exit if lastCommit == HEAD (unless --force) [0%]
2. Cache existing embeddings from prior index [0%]
3. runPipelineFromRepo() → KnowledgeGraph [0-60%]
4. Clean up legacy KuzuDB files [60%]
5. initLbug() → loadGraphToLbug() via CSV streaming [60-85%]
6. Create FTS indexes (File, Function, Class, Method...) [85-90%]
7. Restore cached embeddings (batch insert) [88%]
8. Generate new embeddings if --embeddings [90-98%]
9. Save metadata + register repo + update .gitignore [98-100%]
10. Generate AI context files (AGENTS.md, CLAUDE.md) [100%]
Options: --force (rebuild regardless), --embeddings (opt-in, skipped if >50k nodes), --skipGit, --noStats.
Storage
<repo>/.gitnexus/
├── lbug # LadybugDB database
├── lbug.wal # Write-ahead log
├── lbug.shadow # Shadow sidecar (checkpoint staging)
├── lbug.lock # Single-writer lock
├── lbug.wal.checkpoint, lbug.checkpoint.{intent,apply}.lock # checkpoint-in-flight artifacts; left behind only by an interrupted checkpoint, consumed by the next writable open
├── lbug.{wal,shadow}.dirty-recovery # parked sidecars from a crashed run; safe to delete
├── gitnexus.json # lastCommit, indexedAt, stats (primary metadata file)
└── meta.json # legacy mirror of gitnexus.json, kept in sync (see MIGRATION.md)
~/.gitnexus/
├── registry.json # Global repo registry (MCP discovery)
└── stores/<key>/ # Shared sibling index store (see below)
├── caches/ # parse cache + durable ParsedFile store
├── commits/<commit>-<featureKey>/ # one immutable graph per commit + settings
└── checkouts/<slot>/ # one checkout's metadata, membership, and
# private graph when it has local edits
The flat <repo>/.gitnexus/ layout applies to a standalone repository and
whenever GITNEXUS_STORAGE_PATH / GITNEXUS_STORAGE_ROOT is set. A repository
with linked worktrees, and clones with the same origin URL, share one
stores/<key>/ automatically (a clone opts out with analyze --no-share;
GITNEXUS_SHARED_STORE=off turns sharing off entirely). Each sharing checkout
keeps only a .gitnexus/store.json pointer to its store. Path resolution lives
in shared-store.ts.
Read-only opens self-heal an interrupted checkpoint: the refusal is
classified and cleared by one writable open (probe + CHECKPOINT) before
the read-only open is retried — see sidecar-recovery.ts
(isReadOnlyCheckpointInProgressError) and the
lbug-interrupted-checkpoint-recovery integration test.
Managed by repo-manager.ts.
LadybugDB schema
Defined in lbug/schema.ts. Separate node tables per type, single CodeRelation table.
Node tables: File, Folder, Function, Class, Interface, Method, Constructor, CodeElement, Struct, Enum, Macro, Typedef, Union, Namespace, Trait, Impl, TypeAlias, Const, Static, Property, Record, Delegate, Annotation, Template, Module, Community, Process, Route, Tool, Section, Embedding.
Relation types (CodeRelation.type): CONTAINS, DEFINES, CALLS, IMPORTS, INHERITS, EXTENDS, IMPLEMENTS, USES, DECORATES, HAS_METHOD, HAS_PROPERTY, ACCESSES, METHOD_OVERRIDES, METHOD_IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS, HANDLES_ROUTE, FETCHES, HANDLES_TOOL, ENTRY_POINT_OF, WRAPS, QUERIES, INJECTS, CONDITIONAL_ON, DECLARES, ADVISED_BY, BINDS_EVENT_HANDLER, EMITS_EVENT.
Optional --pdg additions (off by default, opt-in via gitnexus analyze --pdg; see Optional CFG/PDG emission above): a BasicBlock node table, plus the PDG relation types CFG, REACHING_DEF, CDG, TAINTED, SANITIZES, and TAINT_PATH on the same CodeRelation table. These are deliberately kept out of the default VALID_RELATION_TYPES / web graph schema — query them via cypher, explain, or pdg_query.
Embeddings and search
Embeddings (src/core/embeddings/): Snowflake arctic-embed-xs (384D). Embeddable: File, Function, Class, Method, Interface. Incremental via SHA1 content hash. Separate Embedding table.
Search (src/core/search/): Hybrid BM25 + semantic vector, merged via Reciprocal Rank Fusion (K=60).
Known limitations
Overloaded method resolution
Node IDs use arity suffix (#<paramCount>): Method:file:Class.method#1 vs #2.
Same-arity disambiguation: type-hash suffix ~type1,type2 when collision detected and type annotations present. Languages without types (Python, Ruby, JS) use arity-only. TS/JS overload signatures excluded (collapse to implementation body). See #651.
C++ const-qualified: $const suffix after type-hash when non-const collision exists: Method:file:Container.begin#0$const.
Generic/template types: type-hash uses rawType (full AST text including generics): ~vector<int> vs ~vector<std::string>.
ID stability: collision-only tags mean IDs change when overloads are added. save#1 becomes save#1~int when save(String) is added.
Variadic matching: confidence 0.7 when one side is variadic and the other has fixed count.
METHOD_IMPLEMENTS confidence tiering:
| Match quality | Confidence |
|---|---|
| Exact parameter types match | 1.0 |
| Arity match, types unavailable | 1.0 |
| Variadic vs fixed | 0.7 |
| Insufficient info | 0.7 |
Related docs
- MIGRATION.md — breaking changes and migration guidance
- RUNBOOK.md — operational commands and recovery
- GUARDRAILS.md — safety boundaries for humans and agents
- TESTING.md — how to run tests
- docs/languages/objective-c-provider.md — Objective-C provider behavior and limits
AGENTS.md/CLAUDE.md— agent workflows and tool usage