GitNexus/ARCHITECTURE.md
Gergő Magyar ad5c7364e1
feat(storage): share one index store across linked worktrees and sibling clones (#3374)
* feat(storage): resolve the shared sibling-store identity and layout (#3352)

Linked worktrees of one repository resolve to one store under
GITNEXUS_HOME/stores/<key>, keyed by the canonical git common dir. The
resolver reads the .git entry directly, so hot paths spawn no git. Slot
naming moves to a leaf module so storage-resolver and shared-store do not
import each other.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(storage): resolve a shared checkout's graph and existing store slot (#3352)

getStoragePaths reads a flat slot's recorded graphPath only for checkout
slots under the stores directory; other paths keep <storagePath>/lbug with
no I/O. A recorded path outside the store's commit graphs is ignored.
resolveStoragePath falls back to an existing store slot for an
unregistered checkout, so reads never move to an empty slot.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(analyze): share one immutable commit graph across clean worktrees (#3352)

A linked worktree writes its own slot in the shared store. A new slot
is seeded with a pointer to the nearest commit graph, so a clean
checkout at an indexed commit takes the up-to-date path and writes no
graph. A checkout with local changes gets a copy-on-write private
graph before its first write. After a successful run, a clean checkout
at HEAD publishes its graph into commits/ under a store lock, or drops
it when that commit graph already exists. Commit graphs are never
written after publish.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(analyze): seed a worktree's graph from the nearest index (#3352)

A new shared slot is seeded from the store's commit graph nearest to
HEAD, else from this checkout's or the main checkout's
repository-local index (copied under that index's lock, source left in
place). The run that follows is up to date or incremental instead of a
full build. If a pointed-at shared graph has been removed, analyze
falls back to a full build instead of an incremental update over a
missing baseline.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(analyze): keep one parse cache per shared store (#3352)

Linked worktrees read and write the parse cache and durable ParsedFile
store under the store's caches/ directory. Before pruning, a run folds
in the chunk keys recorded by every member slot and commit graph, and
the fold, prune and save run under a store-wide cache lock so one
member never evicts another's live chunks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(clean): remove only what no shared-store member references (#3352)

Deleting a shared checkout slot (clean, clean --all, remove, and the
server delete route) recounts references under the store's publish
lock and deletes commit graphs no member points at, then the store
itself once empty. A graph that cannot be deleted (open on Windows) is
reported and kept for the next pass. clean --gc also drops member slots
whose worktree is gone or no longer resolves to the store.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(analyze): let a clone opt in to a shared store with --share-with (#3352)

An independent clone joins a linked worktree's shared store only with
analyze --share-with <repo>, and only when its normalized origin URL
matches that member's (credentials stripped, as #2054 compares). The
registry remembers the choice. --no-share moves an opted-in clone back
to its own .gitnexus and reclaims its old slot; linked worktrees always
share and are pointed at GITNEXUS_SHARED_STORE=off instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): count a shared checkout's commit graph as its code index (#3352)

A clean shared checkout reads a commit graph and owns no graph file, so
the code-index presence check made status report it unindexed and
registry validation skip it. The check now follows the slot's
validated graphPath.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(storage): adopt existing worktree indexes and report shared-store state (#3352)

A shared checkout gets <repo>/.gitnexus/store.json pointing at its store
slot; resolution follows it only when it names that checkout's own slot.
An existing local index seeds the slot and is left in place. status
(text and --json) and doctor report the store, whether the graph is
shared or private, and any leftover local index, which
clean --local-index removes while keeping the pointer. With
GITNEXUS_SHARED_STORE=off a previously shared checkout indexes into its
own .gitnexus again and never writes a commit graph.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(mcp): read the graph a shared checkout points at (#3352)

MCP, the HTTP API, group sync, augmentation, and the Claude hook (all
three byte-identical copies) resolve a flat slot's graph through
resolveGraphPath instead of joining 'lbug' onto the storage path, so
checkouts at one commit share one open database. The embeddings writers
(embeddings sync and the server embed job) take a private copy first
and never write an immutable commit graph. The post-analyze settle
probe accepts fresh metadata that points at an existing commit graph.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: describe the shared worktree index store (#3352)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(storage): reuse helpers across the shared-store code (#3352)

One graph-clone helper replaces two copy-then-rename blocks; clean and
status reuse formatSlotSize; leaving a store reuses
removeSharedStorePointer; withStoreLock is imported from its own module
instead of a re-export.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address review findings in the shared store (#3352)

- Never publish a graph whose build saw dirty files, and never trust a
  local-index seed on the up-to-date path; it may hold reverted edits.
- Record slot pointers and reclaim under the publish lock, and reclaim
  right after each publish, so a commit graph is never deleted between
  publish and pointer save and superseded graphs don't pile up.
- clean --gc decides membership from the registry, so opted-in clones
  and a main checkout without worktrees are not dropped.
- Leaving a store re-registers first, so an up-to-date run cannot leave
  the registry pointing at a deleted slot.
- Drop pinned branch summaries when an entry moves into a store slot.
- Keep run.cjs and the AGENTS.md runner path inside the checkout.
- MCP handles follow the slot's current graph, and branch scoping reads
  the slot's own metadata.
- Cache resolveGraphPath by metadata file identity; look up opted-in
  entries with canonical registry paths.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cli): localize help for the shared-store flags (#3352)

--help renders option text from the i18n catalog, so --share-with,
--no-share, clean --gc and clean --local-index need keys in both
locales, not just index.ts literals.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): lock the embed-job graph copy and skip it on forced rebuilds (#3352)

The server embed job copies a shared checkout's graph under the slot's
index lock, so a CLI analyze in another process cannot interleave. A
forced rebuild with no embeddings to carry over drops the pointer
without copying a graph it would discard unread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(agents): bump AGENTS.md version for the shared store notes (#3352)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): keep shared-store file reads inside checked paths (#3374)

Address CodeQL js/path-injection and js/file-system-race on
shared-store.ts: every filesystem read rebuilds its path under a fixed
parent with an inline path.relative barrier, the .git probe is one read
(EISDIR marks a directory) instead of stat-then-read, and the graph
pointer cache stats and reads through one file descriptor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address review feedback on the shared store (#3374)

- Hold the slot index lock for the whole server embedding job, not
  just the graph copy.
- --no-share re-registers and deletes the old slot under that slot's
  index lock, so a running shared analyze cannot re-register it.
- clean --gc previews without --force, like every other destructive arm.
- status reports a pinned branch index as private.
- Slot names prefix Windows device names that carry an extension.
- A slot or store named ..<name> is a legal direct child.
- Shared-store suites run in the serialized lbug-db vitest project and
  clear an inherited GITNEXUS_SHARED_STORE.
- Doc and test accuracy fixes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(storage): pin shared-store behavior for worktrees of a bare repository (#3374)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address second review round on the shared store (#3374)

- Reject --no-share in a linked worktree before taking any lock or
  indexing anything.
- clean --all and remove also delete each shared checkout's pointer.
- Delete an empty store only while also holding its cache lock, and
  re-check emptiness under it.
- A store pointer is trusted only when the slot's own metadata names
  this checkout; the editable pointer file just says where to look.
- Device names with an extension are prefixed on Windows only, so POSIX
  slot names stay stable.
- Test fixtures use the gitnexus-test- prefix the stale-sidecar sweep
  recognizes; help text and the private-graph label are accurate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address third review round on the shared store (#3374)

- clean --gc never collects a slot whose index lock is held: an analyze
  holds it until it registers the checkout, so a seeded but not yet
  registered slot is busy, not orphaned.
- Reclaim compares absolute paths, so a relative GITNEXUS_HOME does not
  make live commit graphs look unreferenced.
- graphPath is followed only into a published <commit>-<featureKey>
  dir, never .publish-* staging (TS and all three hook copies).
- clean --gc fails on an unreadable stores root instead of reporting
  nothing to collect.
- --no-share help states it is for opted-in clones.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cli): land the English --no-share help and private-graph label (#3374)

These two strings were meant for e2a9467 and d746675 but were left out
of the commit; the zh-CN side landed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address fifth review round on the shared store (#3374)

- analyze --no-share leaves the shared store even when routed to a
  branch sub-index (gate on sharedStore, not placement.branch)
- clean --gc aborts a store whose member listing is unreadable instead
  of treating it as empty, and skips non-directory entries under stores/
- reusing an existing commit graph no longer fails a finished analyze
  when the redundant private graph cannot be wiped
- a store pointer holding JSON null, a number, a string, or an array is
  treated as invalid
- clean, remove, and DELETE /api/repo hold the checkout slot's index
  lock while deleting the slot
- the server analyze launcher waits on the checkout slot the worker
  actually writes
- sync the Factory hook copy; isolate hook tests from storage overrides;
  use a junction for the Windows worktree alias; the relative
  GITNEXUS_HOME test now sets a relative home

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address sixth review round on the shared store (#3374)

- clean, remove, and analyze --no-share remove the checkout's store
  pointer while still holding the slot's index lock, so an analyze that
  takes the lock next cannot have its new pointer deleted
- clean --gc does not follow a symlink under stores/ (lstat)
- removing a store pointer keeps the checkout's .gitnexus directory when
  it cannot be listed, instead of treating it as empty

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address seventh review round on the shared store (#3374)

- clean --gc aborts when the registry cannot be read instead of treating
  every member as orphaned; with no registry file it collects only
  members whose checkout directory is gone
- seeding from a repository-local index re-reads its metadata under the
  index lock, so the copied graph and the saved metadata match
- clean --gc skips a store another collector removed meanwhile, and
  reports "no shared stores" when stores/ holds only stray files

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address eighth review round on the shared store (#3374)

- removing a checkout's storage unregisters it and removes its store
  pointer before deleting the slot directory, which the file lock
  backend uses for its lock file; withCheckoutSlotLock becomes
  removeCheckoutStorage, used by remove, clean, and DELETE /api/repo,
  and analyze --no-share follows the same order
- the clean --gc preview skips a slot whose index lock is held, matching
  what --force would drop

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(storage): share the index store between clones automatically (#3352)

Clones of one repository now share like linked worktrees. A clone whose
normalized origin URL matches another registered, still-present clone
joins that clone's store, or founds one (keyed on its own path) that the
sibling joins on its next analyze. A lone clone keeps its own .gitnexus.
Graphs stay keyed by commit and feature key, so clones only ever share a
graph built from the same commit with the same settings.

`analyze --no-share` now records a lasting opt-out (`shareOptOut` on the
registry entry, preserved across re-registration); `--share-with` clears
it. The analyze worker reports the storage it wrote over IPC so the
server settles a clone's first shared slot.

A query-time base-plus-overlay graph stays out: LadybugDB reads one
database per query. Instead, private graph copies record whether the
filesystem cloned them copy-on-write (sharing unchanged pages on disk)
or made a full copy, and `status` reports it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): never follow a symlinked checkout .gitnexus (#3374)

A checkout whose `.gitnexus` is a symlink (e.g. `.gitnexus -> ..`) made
the shared-store pointer helpers operate on whatever it pointed at:
findLegacyLocalIndex listed the target as a "legacy index", so
`clean --local-index --force` recursively deleted the checkout's
siblings; writeSharedStorePointer wrote store.json/.gitignore there; and
removeSharedStorePointer deleted store.json there and could rm -r the
target directory.

All four now go through probePointerDir, which lstats `.gitnexus` and
accepts it only when it is a real directory whose realpath is
realpath(checkout)/.gitnexus (mirroring stale-branch-slots.ts). Anything
else is left untouched: no legacy index is reported or removed, and no
pointer is written or removed. removeLegacyLocalIndex re-probes after
sizing, just before its delete loop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): never share a graph from a checkout that hides files (#3374)

Trigger: a sparse checkout, a skip-worktree/assume-unchanged entry, or an
uninitialized submodule leaves `git status` clean, and the run's
indexCoverage.dirtyPaths drops paths with no file hash, so publishSharedGraph
published a graph missing those files as the commit graph. seedSharedSlot
then copied the seed's lastCommit into the new pointer slot, so the next
analyze of a sibling hit the up-to-date fast path without hashing.

Fix: add isWorkingTreePristine (storage/git.ts): the unfiltered
listWorkingTreeDirtyPaths must be empty (null fails closed) and every
gitlink in the index must have a checked-out `.git`. publishSharedGraph
requires it instead of isWorkingTreeDirty. seedSharedSlot keeps the seed's
lastCommit only when the seed is at HEAD and the checkout is pristine, so a
clean sibling still fast-paths onto the shared graph; any other seeded
pointer gets an empty lastCommit, like seedFromLocalIndex, and its next run
hash-diffs, then re-points at the commit graph on publish.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): keep graphs with pending embeddings out of the shared store (#3374)

Trigger: an analyze that finished with embeddings still owed (meta carries
embeddingCheckpoint) was published as the commit graph, because the
feature key ignores the checkpoint and publish never checked it. The
published meta kept the checkpoint while the graph could never change. The
next run healed the embeddings privately, found the commit graph already
there, wiped its own healed graph and pointed back at the partial one.

Fix: publishSharedGraph treats a checkpointed graph as not shareable, so it
stays private and no published commit graph carries a checkpoint. When the
target commit graph exists but records a checkpoint, has fewer
stats.embeddings than the private graph, or has unreadable metadata, the
checkout keeps its private graph instead of re-pointing. The commit graph
is left as is because other checkouts may read it. Embeddings sync and the
server embed job already privatize a pointer slot before writing
(ensurePrivateSharedGraph).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): rebuild a shared slot whose graph is missing (#3374)

A publish interrupted between moving the checkout's graph into
`.publish-<uuid>` staging and renaming staging onto the commit dir left
slot metadata at HEAD with no graph and no graphPath; the next reclaim
deleted the orphaned staging. The up-to-date fast path never checked that
the graph exists, so every later analyze reported "Already up to date"
over a checkout with no index. A failed restore in publishSharedGraph's
catch had the same outcome, and also deleted the staging dir holding the
only copy of the graph.

- run-analyze: for a shared-store slot, stat the graph the checkout reads
  (resolveGraphPath for the flat slot, the branch slot's own lbug
  otherwise) before the fast path; if it is missing, force a full
  rebuild. An incremental run would diff nothing into a fresh, empty
  database. Private .gitnexus indexes are unchanged: they only lose their
  graph by hand, and metadata-only fast-path fixtures rely on that path.
- publishSharedGraph: when moving the staged graph back fails and it is
  still in staging, keep the staging dir, log its path, and skip this
  run's reclaim, which would otherwise delete it. The next analyze finds
  no graph and rebuilds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): fail safe on unreadable slot metadata during reclaim (#3374)

Trigger: reclaim counted commit-graph references with loadMeta, which
returns null for a torn or unreadable gitnexus.json as well as for a
missing one. A member whose metadata could not be read therefore
"referenced nothing", and the graph it still used was deleted. Separately,
in `clean --gc` an orphan slot that fs.rm could not delete threw out of
reclaim, aborting the collection of every store after it.

Fix: the reference loop reads slot metadata with a local loadMetaStrict.
Only absent metadata (ENOENT/ENOTDIR, same legacy-mirror fallback as
loadMeta) means no reference; any other read error or a parse failure
throws with the file path, like listDirStrict. Every caller already
treats a reclaim throw as best effort (analyze logs "skipped cleanup",
slot removal ignores it) or surfaces it (clean --gc). A failed orphan
slot delete is now kept: it stays a member, so the graph it names is
kept too, and it is reported in ReclaimResult.keptMembers and by a new
clean --gc line. loadMeta itself is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(server): unregister under the slot lock and answer 409 on DELETE /api/repo (#3374)

DELETE /api/repo called removeCheckoutStorage(storagePath) with no
unregister callback and no checkout path, swallowed every error, and
unregistered later outside the slot lock. For a shared-store slot this
left the checkout's store.json pointer behind, let an analyze re-register
between the slot removal and the unregister, and reported {deleted} even
when the slot lock could not be taken because an analyze held it.

The handler now calls removeCheckoutStorage(storagePath,
() => unregisterRepo(entry.path), entry.path), matching `gitnexus remove`
and `gitnexus clean`, so the unregister and the pointer removal run under
the slot lock; non-shared storage is still removed, then unregistered.
The later standalone unregister is gone. An IndexLockTimeoutError answers
409 and leaves the entry registered; any other failure propagates to the
handler's 500, also with the entry kept so the delete can be retried,
instead of unregistering over leftover index files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): give concurrent sibling clones one deterministic founder key (#3374)

Trigger: two registered clones of one origin, both still on local storage,
analyzing at the same time each saw no sibling store yet and founded a store
keyed on their own checkout path. registeredStore() then kept each clone in
its own store for good, so they never shared a graph.

Fix: when no sibling store exists, siblingCloneStore keys the new store on
the canonical path that sorts first among this clone and its registered
siblings (case-folded on Windows, like registryPathEquals), so every sibling
computes the same key. An existing sibling store still wins. Clones that
already diverged into two stores are not migrated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): keep subdirectories of a clone out of shared stores (#3374)

Trigger: `analyze --skip-git <clone>/pkg` inside a clone with a registered
sibling of the same origin joined, or founded, a clone store. getRemoteUrl
answers from any subdirectory, so siblingCloneStore saw the enclosing
clone's remote. That broke the shared-store invariant that only tree roots
participate.

Fix: resolveOptedInStore now returns undefined for a path with no `.git`
entry. It uses hasGitDir, which accepts a directory or a linked-worktree
file, the same as resolveSharedStore's readCommonDir gate and run-analyze's
repoHasGit. `--share-with` from such a path throws a clear error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(hooks): prefix Windows device names with an extension in hook slot names (#3374)

Trigger: on Windows, storage-slot.ts sanitizeSlotBasename prefixes a
device-name basename that has an extension (`CON.txt` -> `repository-CON.txt`),
but the hook's copy in registry-query.cjs only matched the bare device name.
With GITNEXUS_STORAGE_ROOT set, a checkout named e.g. `con.txt` got a
different slot from the hook than from the CLI, so the hook could not see its
index.

Fix: mirror the platform branch (extension form on win32, bare form
elsewhere) in all four byte-identical registry-query.cjs copies, and point
their header comments at storage-slot.ts. Add a hook-vs-TS parity test over
device-name basenames on both stubbed platforms, and fix the byte-identity
test title to say four copies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(storage): tidy the shared-store fix series (#3374)

Explain the keptStaging early return where it is checked, reuse the
exported EmbeddingCheckpoint type in the publish test, and drop
review-step labels from test comments. No behavior change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address ninth review round on the shared store (#3374)

- removeCheckoutStorage: keep the lock-safe order (unregister and drop the
  pointer under the slot lock, delete the slot last), but when the final
  rm fails, throw an error that says the checkout was already unregistered,
  names the leftover slot, and points at `gitnexus clean --gc --force`.
- clean --gc preview no longer sweeps staging files: acquireIndexLock
  takes `sweep: false`, which the dry-run reclaim passes.
- listStoreMetaRoots reports completeness; a store directory that cannot
  be listed (other than missing) makes the parse-cache prune retain every
  chunk instead of evicting keys other members still use.
- clean.shared.kept says "could not be removed" rather than "still open".
- ARCHITECTURE.md describes the stores/<key>/ layout and store.json
  pointer; README lists every GITNEXUS_SHARED_STORE off value and that it
  also stops clone sharing.
- Tests: close the direct Ladybug connection in finally; fix a fixture
  JSDoc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(storage): pin the publish gate for cone and sparse-index checkouts (#3374)

isWorkingTreePristine already rejects every sparse mode, because each one
marks left-out entries skip-worktree and `git ls-files -v` expands a sparse
index. Only no-cone was tested; cover cone and cone with a sparse index too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(storage): name the checks that answer recurring review findings (#3374)

Comment-only. The review bot re-derives findings from the code on every
push, so put the refuting fact next to each flagged line: the pristine
check's hidden-path coverage, sparse checkouts being skip-worktree, the
lock-safe unregister-then-delete order, the hook parity test for device
names, and why the stores/ lstat is not a race guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 11:20:12 +01:00

66 KiB
Raw Blame History

Architecture — GitNexus

Monorepo: CLI/MCP (gitnexus/) + browser UI (gitnexus-web/).

Repository layout

Path Role
gitnexus/ npm package gitnexus: CLI, MCP server (stdio), HTTP API, ingestion pipeline, LadybugDB graph, embeddings.
gitnexus-web/ Vite + React thin client: graph explorer + AI chat. All queries via gitnexus serve HTTP API.
gitnexus-shared/ Shared TypeScript types and constants (consumed by CLI and Web).
.claude/, gitnexus-claude-plugin/, gitnexus-cursor-integration/ Agent skills and plugin metadata.
eval/ Evaluation harnesses for benchmarking tool usage.
.github/ CI workflows + composite actions (setup-gitnexus/, setup-gitnexus-web/).

End-to-end flow: index → graph → tools

  1. Ingestion — analyze.ts → runFullAnalysis (run-analyze.ts) → runPipelineFromRepo (pipeline.ts). The default DAG of 19 phases builds a KnowledgeGraph in memory, then loads into LadybugDB under .gitnexus/. Repo registered in ~/.gitnexus/registry.json for MCP discovery.

  2. Persistence — repo-manager.ts (paths, registry, LadybugDB cleanup). lbug-adapter.ts (graph load, queries, embedding batches).

  3. Query layer — three interfaces to the same backend:

    • MCP (stdio): mcp.ts → LocalBackend → tools (tools.ts) + resources (resources.ts)
    • HTTP bridge: serve.ts → Express (api.ts, mcp-http.ts) for web UI
    • CLI direct: gitnexus query|context|impact|cypher in tool.ts
  4. Staleness — core/git-staleness.ts compares indexed lastCommit to HEAD and classifies the result as current, behind, diverged (HEAD moved off the indexed commit, gap uncountable) or unknown; core/staleness-status.ts builds the one staleness payload that MCP list_repos, the read tools and the serve repo routes all emit.

MCP tools

Tool Purpose
list_repos Discover indexed repos
query Hybrid BM25 + vector search over the graph
cypher Ad hoc Cypher against the schema
context Callers, callees, processes for one symbol
impact Blast radius (upstream/downstream) with risk summary
detect_changes Map git diffs to affected symbols and processes
rename Graph-assisted multi-file rename with dry_run preview
api_impact Pre-change impact report for an API route handler
trace Shortest directed path between two symbols (call + class-member edges); group-aware (repo: "@<group>") for cross-repo traces
route_map API route → handler → consumer mappings
tool_map MCP/RPC tool definitions and handlers
shape_check Response shape vs consumer property access mismatches
explain Persisted taint findings (source→sink data flows) — needs analyze --pdg
pdg_query Control/data dependence — CDG (mode: controls) / REACHING_DEF (mode: flows) — needs analyze --pdg
group_list List repo groups or details for one group
group_sync Rebuild group Contract Registry (contracts.json) and bridge graph

query, context, and impact are group-aware: pass repo: "@<groupName>" (or "@<groupName>/<memberPath>" to scope to one member) plus optional service: "<monorepo/path>". Group-mode query merges per-repo results via Reciprocal Rank Fusion; group-mode impact runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (gitnexus/src/core/group/cross-impact.ts). trace is also group-aware via repo: "@<groupName>" — but, unlike the others, it resolves from/to across all members (a @<groupName>/<memberPath> suffix is advisory for trace, not a scope); pass from_uid/to_uid to disambiguate a symbol name that occurs in more than one member.

Group-mode trace (gitnexus/src/core/group/cross-trace.ts) stitches a path that crosses repositories: it resolves from/to across all members, and when they live in different repos it joins the home-repo segment to the target-repo segment over a single ContractLink boundary (an HTTP consumer→provider link, joined on Contract.symbolUid), reported as a CONTRACT_LINK hop in crossings[]. The crossing is clamped to one boundary (MAX_SUPPORTED_CROSS_DEPTH, shared with cross-impact); deeper crossDepth is reported via notes[]. With pdg: true (experimental, opt-in), each boundary-adjacent segment is enriched with its intra-procedural REACHING_DEF data-flow when that repo was indexed with --pdg (reusing the same anchored flows query as pdg_query); data flow never crosses the repo boundary, and a missing PDG layer degrades to call-level hops with a note. Two stores meet only at the symbolUid grain — the per-repo PDG/call graph and the group bridge — so this is the documented join; full cross-program (SDG-like) data flow across the boundary remains deferred (see docs/plans/2026-06-18-002-feat-unified-pdg-impact-evaluation-plan.md). The previously-planned group_query, group_context, group_impact, group_contracts, group_status MCP tools are intentionally not introduced — group-level state is exposed via resources instead:

Resource URI Purpose
gitnexus://group/{name}/contracts Contract Registry (provider/consumer rows + cross-links)
gitnexus://group/{name}/status Per-member index + Contract Registry staleness

Where to change what

Concern Start in
CLI commands/flags src/cli/ (index.ts, per-command modules)
Parsing/graph construction src/core/ingestion/pipeline-phases/ + pipeline.ts
Graph schema/DB src/core/lbug/ (schema.ts, lbug-adapter.ts)
MCP tools/resources src/mcp/server.ts, tools.ts, resources.ts
Cross-repo groups (sync, contracts, @<group> routing) src/core/group/ (service.ts, cross-impact.ts, sync.ts, bridge-db.ts)
Search ranking src/core/search/ (BM25, hybrid fusion)
Embeddings src/core/embeddings/ + src/core/run-analyze.ts
Wiki generation src/core/wiki/
Language support src/core/ingestion/languages/ + tree-sitter-queries.ts + gitnexus-shared/src/languages.ts
Import resolution src/core/ingestion/import-processor.ts + import-resolvers/configs/ + model/resolution-context.ts
Call resolution/inheritance/MRO src/core/ingestion/scope-resolution/ (pipeline, passes, graph-bridge)
Type extraction src/core/ingestion/type-extractors/
Worker pool src/core/ingestion/workers/
Web UI gitnexus-web/src/
CI .github/workflows/*.yml, .github/actions/

Paths above are relative to gitnexus/ unless they start with gitnexus-web/ or .github/.


Pipeline Phase DAG

19 default phases are defined in gitnexus/src/core/ingestion/pipeline-phases/, each with explicit deps and typed output. --pdg adds taintSummaries and callSummaries (21 total).

scan → structure → [springConfig, markdown, cobol] → parse → [routes, tools, orm]
  → crossFile → scopeResolution → [springAutoConfiguration, springAop]
  → pruneLocalSymbols → mro → springAopInheritance → di → communities → processes
Phase File Deps Output
scan scan.ts (root) File paths + sizes
structure structure.ts scan File/Folder nodes, CONTAINS edges, allPathSet
springConfig spring-config.ts structure Spring configuration-property nodes and metadata
markdown markdown.ts structure Section nodes, cross-link edges from .md/.mdx
cobol cobol.ts structure COBOL program/paragraph/section nodes (regex, no tree-sitter)
parse parse.ts + parse-impl.ts structure, markdown, cobol Symbol nodes, IMPORTS/CALLS/EXTENDS edges, extracted routes/tools/ORM queries
routes routes.ts parse Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators, and JS/TS static route sources — see below)
tools tools.ts parse Tool nodes + HANDLES_TOOL edges
orm orm.ts parse QUERIES edges (Prisma, Supabase)
crossFile cross-file.ts + cross-file-impl.ts parse, routes, tools, orm Cross-file type propagation in topological import order
scopeResolution scope-resolution/pipeline/phase.ts parse, crossFile, structure Binding/reference + inheritance edges; disposes BindingAccumulator
springAutoConfiguration spring-auto-configuration.ts structure, scopeResolution DECLARES and CONDITIONAL_ON metadata for Spring configuration candidates
springAop spring-aop.ts scopeResolution Direct declarative/advice ADVISED_BY edges and pointcut evidence
pruneLocalSymbols prune-local-symbols.ts scopeResolution Drops inert block-local Const/Variable/Static nodes (only a File→DEFINES edge) post-resolution
mro mro.ts crossFile, scopeResolution, pruneLocalSymbols, structure METHOD_OVERRIDES + METHOD_IMPLEMENTS edges
springAopInheritance spring-aop.ts springAop, mro Propagates declarative behavior through class/interface inheritance decisions
di di.ts mro INJECTS edges from consumer Classes, factory Methods, or AST-captured programmatic lookup callables to provider Classes/declaration CodeElements (framework-neutral DI resolution; per-language matchers registered in di-extractors/)
communities communities.ts mro, pruneLocalSymbols, structure Community nodes + MEMBER_OF edges (Leiden algorithm)
processes processes.ts communities, routes, tools, pruneLocalSymbols, structure Process nodes + STEP_IN_PROCESS edges

Non-phase files in the same directory: parse-impl.ts, cross-file-impl.ts (implementation), wildcard-synthesis.ts (whole-module import expansion), types.ts, runner.ts, index.ts.

DAG runner

runner.ts — static phase graph, no plugins, compile-time type safety.

  1. Validation — Kahn's topological sort. Rejects on: duplicate names, missing deps, cycles (DFS traces the concrete cycle path, e.g., A -> B -> C -> A, plus count of transitively blocked dependents).

  2. Execution — sequential in topological order. Each phase receives:

    • ctx: PipelineContext — shared mutable KnowledgeGraph, repoPath, progress callback, options
    • deps: ReadonlyMap<string, PhaseResult> — declared deps only (runner filters the results map to prevent hidden coupling)
  3. Error handling — wraps phase errors with the phase name, emits terminal error progress event, swallows progress handler errors to preserve the original cause.

  4. Timing — per-phase durationMs in PhaseResult, dev-mode console logging.

Design patterns:

  • Single graph accumulator — all phases mutate the same KnowledgeGraph in ctx; the graph is the primary output.
  • Typed phase access — getPhaseOutput<T>(deps, 'name') for type-safe upstream results.
  • Binding accumulator lifecycle — created in parse, disposed by crossFile (in finally). No other phase should take ownership.
  • Skippable phases — skipGraphPhases omits MRO/di/communities/processes (faster tests); pruneLocalSymbols still runs (it is graph cleanup, not analysis). skipWorkers is no longer a sequential escape hatch — it (like --workers 0 / GITNEXUS_WORKER_POOL_SIZE=0) is rejected with an actionable error, since the worker pool is the sole parse path (§ Chunked parse-and-resolve).
  • Local-symbol pruning — pruneLocalSymbols removes inert block-local value symbols after scope resolution has consumed them. Opt out per-call with PipelineOptions.keepLocalValueSymbols or globally with the GITNEXUS_KEEP_LOCAL_VALUE_SYMBOLS env var.

How to add a new phase

  1. Create pipeline-phases/my-phase.ts with a PipelinePhase<MyOutput> (name, deps, execute)
  2. Export from pipeline-phases/index.ts
  3. Add to buildPhaseList() in pipeline.ts
import type { PipelinePhase, PhaseResult } from './types.js';
import { getPhaseOutput } from './types.js';
import type { ParseOutput } from './parse.js';

export interface MyPhaseOutput {
  /* ... */
}

export const myPhase: PipelinePhase<MyPhaseOutput> = {
  name: 'myPhase',
  deps: ['parse'],
  async execute(ctx, deps) {
    const { allPaths } = getPhaseOutput<ParseOutput>(deps, 'parse');
    // ... write to ctx.graph ...
    return {
      /* typed output */
    };
  },
};

Where routes come from

route-extractors/ holds four independent ways a route can be discovered, all converging on the routes phase's (method, url) registry:

Source Shape Examples
Filesystem convention path → URL, no parsing Next.js app/, Expo, PHP
Single-file framework route isRouteFile + worker extraction Laravel routes/*.php
Cross-file framework route discoverRootRouteFiles + extractRoutes Django urlpatterns
AST-level route in a normal file extractDecoratorRoutes Spring, FastAPI, NestJS (@Controller + @Get/@Post/…; URLs are controller-relative — setGlobalPrefix and URI versioning live in the bootstrap file and are not applied), JS/TS dispatch guards and static data route tables

The last row is the one whose name undersells it. A route is DECLARED by a decorator, but it can also be inferred from a raw node:http server's own dispatch — if (req.method === 'GET' && pathname === '/api/x') is a route with a path, a verb and a handler, and nothing else in the pipeline could see it. route-extractors/dispatch-guard.ts reads that shape; the transport, dedup and handler resolution are shared with decorator routes, and ExtractedDecoratorRoute.source carries the provenance difference through to the HANDLES_ROUTE edge.

JS/TS data route tables share that transport when a route-named array contains direct object literals with static path, method, and handler fields and a same-scope for...of dispatcher positively compares the path and method before directly invoking the handler. Dynamic values, computed keys, spreads, inline/called handlers, unknown verbs, and ambiguous handler bindings are suppressed. Bare import aliases and single-level member handlers are attributed only through declared import and owner provenance; an unproven receiver never falls back to a global name guess.

That extractor is deliberately precision-weighted: route_map presents its output as fact, so a startsWith namespace test, a bare pathname === '/' without a verb, and any regex it cannot translate exactly are all dropped rather than guessed at. A missing route is a coverage limit; an invented one is a lie.

Two rules there need more than one comparison to decide, and are worth knowing about before changing either:

  • Same-file constant folding. pathname === `${basePath}/rules` is common enough that refusing it loses whole route modules — and loses them invisibly, since a module with unfoldable paths and a module with no routes produce the same empty answer. Folding is same-file, string literals only, one alias hop, and refuses on ambiguity (a name declared twice with different values is dropped, never guessed).
  • Whole-repo reconciliation (reconcileDispatchGuardRoutes, applied in the routes phase). A split route table — one module listing every path it recognises so the dispatcher can 404 early, handlers in others — otherwise lists every route twice, once verb-less with the table as its "handler". It applies to dispatch-guard routes only: a framework route with no verb is method-agnostic by declaration, which is a fact, not a weaker observation.

Semantic model

SemanticModel (gitnexus/src/core/ingestion/model/semantic-model.ts) is the authoritative store for every symbol-indexed lookup (by nodeId, simpleName, qualifiedName, or filePath). The scope-resolution pipeline reads from here: findOwnedMember, pickOverload, and findExportedDefByName all consult model.methods / model.fields / model.symbols.

ParsedFile (gitnexus-shared/src/scope-resolution/parsed-file.ts) is the single per-file artifact the scope-resolution pipeline consumes. Scope-resolution passes MUST NOT build a parallel parse representation. If a per-language hook needs AST-level facts that ParsedFile doesn't expose, it should reuse the orchestrator's treeCache (RunScopeResolutionInput.treeCache) rather than re-invoking parser.parse(...) on its own — the C# populateNamespaceSiblings hook is the reference implementation of this pattern.

The scope-resolution pipeline additionally carries WorkspaceResolutionIndex for Scope-valued lookups (classScopeByDefId, moduleScopeByFile) that SemanticModel structurally cannot hold. No symbol-indexed duplicates exist outside SemanticModel.

Write / read phase contract. The model is mutable during three ordered phases and read-only afterward:

 Phase 1: parse            ──► symbolTable.add fans into types/methods/fields
 Phase 2: scope-resolution ──► reconcileOwnership() registers corrected ownerIds
 Phase 3: finalize         ──► model.attachScopeIndexes(bundle) — one-shot freeze
 ─────────────────────────── phase boundary ───────────────────────────
 Read phase: all resolution passes + MCP + HTTP + embeddings see
             SemanticModel (read-only handle); writes are type-errors.

runScopeResolution narrows MutableSemanticModel → SemanticModel at the phase boundary so downstream passes physically cannot mutate the model even accidentally.

Reconciliation pass. reconcileOwnership (scope-resolution/pipeline/reconcile-ownership.ts) is a shim for languages whose parse-time extractor doesn't resolve enclosingClassId at parse time (Python class-body methods are the canonical case). It walks parsed.localDefs[i].ownerId after populateOwners and registers any missed methods/fields into the model. Idempotent — safe to re-run, safe alongside languages whose extractor already carries ownerId (C#).

The architectural end state is for every language's parse-time extractor to emit the correct ownerId directly, making reconciliation a no-op (tracked as a follow-up refactor). The dev-mode validator validateOwnershipParity surfaces any drift via onWarn under NODE_ENV !== 'production' && VALIDATE_SEMANTIC_MODEL !== '0'.

References: semantic-model.ts file-head (full write/read contract); contract/scope-resolver.ts Contract Invariant I9 (scope-resolution-side rule).


Scope-Resolution Pipeline (RFC #909 Ring 3)

Language-agnostic scope-resolution resolver. This is the resolution path for every language — it owns CALLS/ACCESSES/USES emission and inheritance edges. Adding a language is one interface implementation (ScopeResolver) plus one registration in the SCOPE_RESOLVERS map — no changes to shared code, no new pipeline phase. (RING4-1 #942 removed the legacy call-resolution DAG and the per-language MIGRATED_LANGUAGES flag, so SCOPE_RESOLVERS registration is all that's needed.)

Pipeline stages

 ParsedFile[]  (extractParsedFile per file)
    │  finalizeScopeModel (+ provider hooks)
    ▼
 ScopeResolutionIndexes
    │  resolveReferenceSites  (via MethodRegistry.lookup)
    ▼
 ReferenceIndex
    │  emitReceiverBoundCalls  ── FIRST
    │  emitFreeCallFallback    ── THEN
    │  emitReferencesViaLookup ── uses handledSites + deferred-site skip set
    │  emitPropertyDispatchCalls ── registration USES + conservative CALLS
    │  emitCallableValueFlow   ── assigned/passed callable invocation CALLS
    │  emitImportedValueReferences ── cross-file value reads via finalized imports
    │  emitUniqueNamePropertyAccesses ── LAST-RESORT property reads by name,
    │       narrowed same-file → direct-import, refusing to choose otherwise
    │  emitImportEdges
    ▼
 KnowledgeGraph  (IMPORTS / CALLS / ACCESSES / INHERITS / USES)

Orchestrator: runScopeResolution(input, provider) in scope-resolution/pipeline/run.ts. Pipeline phase: scopeResolutionPhase in scope-resolution/pipeline/phase.ts — iterates the registered SCOPE_RESOLVERS over the worker-serialized ParsedFiles. (Per-language emitScopeCaptures hooks may reuse a cached Tree via the orchestrator's treeCache, but in worker-pool runs that cache is empty — Trees can't cross MessageChannels — so they consume the pre-extracted ParsedFile instead; § Performance notes.)

Callable-value flow

First-class callable values use a language-neutral inclusion analysis in passes/callable-value-flow.ts. Providers recognize their own syntax and emit JSON-safe CallableFlowSite facts (seed, copy, alias, address, load, store, formal, argument, and invoke) into ParsedFile; shared ingestion never branches on a language name. These always-on facts cross workers and the durable parse store, whose schema is bumped whenever their semantic shape changes.

The emit stage defers only invocation sites proven to reference a flow cell. Ordinary receiver/free/reference passes still resolve direct callees first and record exact callee IDs by file/line/column. Property dispatch then runs before callable flow because a property-dispatched wrapper call can seed actual-to-formal propagation. The callable solver consumes those direct targets, propagates callable sets through lexical cells and formals, and emits CALLS at the real indirect invocation site with reason callable-value-flow (confidence 0.8 for a singleton, 0.7 for a bounded multi-target set).

The solver is flow-insensitive but bounded: dependency-indexed work items rerun only when a cell they read changes; target/address sets cap at 32; a hostile fact graph has a finite work budget; overflow or budget exhaustion emits no partial CALLS and produces a structured warning. Lexical shadowing is function/block aware, invocation/constructor results are not reinterpreted as callable designators, and overload selection uses provider-supplied signature metadata. C/C++ additionally associate visible prototypes with unique definitions so actual-to-formal flow crosses translation units; the provider-owned hasFileLocalCallableLinkage hook prevents static declarations or definitions from leaking across files. C++ member-function pointers preserve parameter/cv shape, keep non-virtual targets exact, and expand virtual targets through MethodDispatchIndex/MRO.

Property-key dispatch remains a separate conservative fallback. Its per-key fan-out cap is 32; capped keys synthesize no partial calls and are reported at warning level with language, skipped-key count, dropped key names (bounded), and cap; the count also travels in RunScopeResolutionStats.propertyDispatchSkippedKeys.

Interface-dispatch fan-out walks the subtype closure of the receiver's interface and is generic-instantiation aware (#2912): a call through IValidator<string> must not reach an implementor of IValidator<int>, which shares its declaration and therefore its subtype list. Each heritage clause's arguments reach resolution by one of three routes — read off the @reference.inherits anchor's own spelling where that anchor spans the whole base (most languages, no query change), through the @reference.type-arguments sub-tag where the anchor is the bare name and moving it would renumber inheritance edge ids (Rust impl T<A> for S, Dart extends), or on a heritage MARKER payload for clauses that never become reference sites (Dart implements/with). Whichever pass emits the edge records the pair through one sink: preEmitInheritanceEdges for heritage clauses, ScopeResolver.emitHeritageEdges for the rest.

The walk then carries a substitution: a subtype's own type parameters bind to the receiver's arguments, so class Wrapper<T> : IValidator<T> stays reachable from every instantiation while class IntValidator : IValidator<int> is pruned from the string one. Receiver arguments come from the declared type (Case 4), a class-level field's declared type (Case 6), or — for a compound receiver such as this._repo — the spelling the compound fold typed that position from, reported back through recordReceiverType and accepted only when it names the class the fold returned.

The filter prunes only on positive evidence: an unknown instantiation on either side, an argument list whose arity does not line up, a name that may be a type variable the language's captures never recorded, or an unresolved spelling whose simple name matches all keep the target. A type parameter of the declaration ENCLOSING either side is recognised as such and never compared — void Run<T>(IValidator<T> v) writes a receiver with no known instantiation, so it keeps the unfiltered fan-out. That recognition is what generic METHODS now carry @declaration.type-parameters for in C#, Java and Kotlin (TypeScript already did): without it an unbounded T grounds to nothing and a bounded one grounds to its BOUND, and both compare unequal to an implementor's concrete argument. Languages that capture neither type arguments nor type parameters therefore emit exactly the pre-#2912 fan-out. The fan-out cap (32, GITNEXUS_MAX_INTERFACE_DISPATCH_FANOUT) and its skipped-target reporting are unchanged and apply after filtering. Note the fan-out itself still fires only for a receiver whose folded type is an Interface symbol, so a Rust Trait or a Dart abstract Class receiver emits no secondary targets to filter in the first place.

Standalone (regex-based) providers such as COBOL participate via ScopeResolver.scopeResolutionEdgeMode: 'callable-flow-only': runScopeResolution runs for them, but every ordinary emission path — heritage, interface implementations, receiver-bound, free-call fallback, reference/import edges, post-resolution hooks — is gated off, so their legacy phase (e.g. cobolPhase) remains the sole owner of structural edges and the callable solver's CALLS are purely additive. A callable-flow-only provider whose files emitted no callable facts exits early, before finalize, keeping the opt-in proportional to source scanning.

Receiver chains and the drop census (#2766)

A compound receiver (svc.getUser().address.save()) is captured as a compact string on ReferenceSite.receiverChain. utils/receiver-chain-codec.ts is the ONE encoder/decoder — capture emitters, the scope-resolution fold, and the durable ParsedFile store all import it rather than hand-rolling the format.

Wire format is v2: 2|<base>|<step>|<step>…, one-character version prefix, then base-first steps, each a one-character kind sigil plus the member name (c = call, f = field). a (await) and i (index) are name-free and encode as a bare sigil — an awaited call's name already lives on its c step, and a subscript key is a value, not a lookup-able identifier. The version went 1 → 2 when those two kinds were added, and a decoder REFUSES a foreign version rather than decoding the prefix it understands: a chain missing its await/index hop decodes cleanly as a different, shorter chain and would type the receiver against the wrong member. The format is unescaped (| and ~ cannot occur in an identifier), so an unencodable name is refused rather than escaped, and the payload is capped at MAX_RECEIVER_CHAIN_BYTES / MAX_CHAIN_DEPTH steps. Because these strings live in the incremental parse cache and the durable ParsedFile store, a format change requires a PARSE_CACHE_VERSION schema bump — a stale cache would otherwise replay v1 chains this build discards.

Receivers the resolver could not type are not silently dropped. Each records a ResolutionOutcome (scope-resolution/resolution-outcome.ts) carrying the receiver's shape (classifyReceiverShape: chain-call / chain-field / chain-mixed / chain-unwrap / no-chain — the bench censuses these) and its origin (in-program / external / unknown). scope-resolution/unresolved-receivers.ts aggregates them per member name into the index-persisted unresolvedReceiverMembers summary, keeping in-program and external counts under separate keys. Only in-program drops make a count short: an external-rooted call (System.out.println, fetch(...)) has no in-graph node an edge could have reached, so it is reported but does not hedge. impact / context read that summary and publish epistemic: 'exact' | 'lower-bound', prose boundaries, and the machine-readable causes split (EpistemicCauses in mcp/local/local-backend.ts).

Optional CFG/PDG emission (--pdg, #2081–#2086)

On a --pdg run the parse worker builds a per-function control-flow graph from the tree-sitter AST (LanguageProvider.cfgVisitor; TypeScript/JavaScript today) and serializes it onto ParsedFile.cfgSideChannel as plain data. Scope-resolution then emits the program-dependence layers from that side-channel inside Phase 4 of runScopeResolution, while the disk-backed ParsedFile store is still live — the only window where the worker-built CFGs are loaded (the store is cleared right after the phase returns). A standalone post-mro phase would read an empty store, so the emit deliberately lives in-phase, mirroring the applyCaptureSideChannel pattern. The opt-in is off by default (graph byte-identical), folded into the parse-cache key (a pdg-off warm cache is never reused on a --pdg run), and each layer is bounded by a per-function edge cap that logs any dropped edges. All layers are BasicBlock → BasicBlock edges in the single CodeRelation table, keyed by type; there is no Function → BasicBlock edge — the symbol↔block join is reconstructed from the BasicBlock id prefix + line span. The layers build on each other:

  • M1 — CFG (#2081): BasicBlock nodes + CFG edges. Edge kind (seq/cond-true/loop-back/…) rides the reason column (CFG is one CodeRelation type, not one per kind).
  • M2 — REACHING_DEF (#2082): GEN/KILL def→use data dependence from a pure fixpoint solver; the variable name rides reason.
  • M3/M4 — TAINTED / SANITIZES / TAINT_PATH (#2083–#2084): intra- and inter-procedural taint (source→sink) — the explain tool's data.
  • M5 — CDG (#2085): Ferrante control dependence over a Cooper–Harvey–Kennedy post-dominator tree (the EXIT-rooted reverse CFG); branch sense ('T'/'F') rides reason. A CFG whose EXIT is unreachable from some block is skipped for CDG (post-dominance would be unsound) while its CFG/REACHING_DEF layers are kept.
  • M6 — read surface (#2086): the pdg_query MCP tool answers "what gates X?" (CDG, mode: controls) and "where does Y flow?" (REACHING_DEF, mode: flows); explain is the taint consumer. Both are always anchored + LIMIT-bounded (LadybugDB has no rel-property index) and share one resolveBlockAnchor helper. These PDG edge types are deliberately kept out of the default VALID_RELATION_TYPES / web schema.
  • Cross-repo trace enrichment: group-mode trace (pdg: true) reuses the same anchored REACHING_DEF flows query to annotate a boundary-adjacent segment with how a value reaches the cross-repo call — strictly intra-procedural (data flow never crosses the repo boundary). See the group-aware tools note above.

See core/ingestion/cfg/ (emit + the pure CFG / post-dominator / control-dependence / reaching-defs / taint passes) and mcp/local/local-backend.ts (_pdgQueryImpl, _explainImpl, the shared resolveBlockAnchor).

ScopeResolver contract

Single interface a language implements to plug into the pipeline. Contract fully documented in scope-resolution/contract/scope-resolver.ts.

Hook Purpose
languageProvider Base LanguageProvider (tree-sitter query, emitScopeCaptures, import/binding interpreters, hooks)
populateOwners(parsed) Fill deferred ownerId fields on method defs (captures can't always know the owning class at parse time)
buildMro(graph, parsed, nodeLookup) Produce mroByClassDefId: Map<DefId, DefId[]> — C3, Ruby-mixin, or first-wins per language
resolveImportTarget(target, fromFile, allFiles) (rawImportPath, sourceFile) → targetFilePath (PEP-328 for Python, etc.)
isNamespaceImport(parsedImport, targetFile, fromFile) Optionally reclassify a resolved named import as a namespace handle when the imported symbol is itself a module
mergeBindings(existing, incoming, scopeId) Shadowing / LEGB precedence
arityCompatibility Provider consumed by registry during MethodRegistry.lookup Step 2
importEdgeReason Confidence-tier string for IMPORTS edge reason field
propagatesReturnTypesAcrossImports? Opt out of cross-file return-type propagation (default on)
fieldFallbackOnMethodLookup? Statically-typed languages turn this OFF — the heuristic over-connects (default on)
elementTypeOf? (containerType, via: {kind:'index'} | {kind:'accessor',name}) → elementType | undefined — element type of a container, reached by subscript (repos[0]) or by a property-style collection view (data.Values). ONE hook for both routes (it replaced the split unwrapCollectionAccessor / unwrapCollectionElement, where implementing one silently answered nothing for the other). Consulted only where the source actually performed the access — never as a general type-name normalizer
stripTypePreservingDecoration? (typeName) → strippedName | undefined — strip ONE layer of TYPE-PRESERVING decoration (pointer, reference, const, nullable, borrow, sigil) so a receiver declared *Host still finds the Host binding (#2766). Never a container: unwrapping Repo[] here would fold repos.find(x) to Repo.find — that is elementTypeOf's job, and only after a real subscript. Consulted only after every undecorated lookup fails, and only by receiver-chain base/step resolution — default off
collapseMemberCallsByCallerTarget? One CALLS edge per (caller, target) instead of per-site — default off
populateNamespaceSiblings? Cross-file implicit visibility (compiler-implicit namespace sharing) — default off; ctx carries treeCache
hoistTypeBindingsToModule? Walk up to Module scope when looking up a method's return-type typeBinding — default off; enable only when bindings are stored at module level
hasFileLocalCallableLinkage? Precise internal-linkage predicate used only when joining callable declarations/prototypes to cross-file definitions; C/C++ use it for static free functions
constructorCallTargetsClass? A constructor-form call Type(...) links to the Class def rather than its explicit Constructor def — default off; Swift and Dart opt in
constructionSyntax? How the language spells construction, so an INLINE constructor receiver (Service(db).m(), new Service(db).m(), Service.new.m()) can be typed — bare / keyword / selector; default off, opt in per language only where measured to be needed (#2708)

Per-language registration

  1. Implement ScopeResolver in languages/<lang>/scope-resolver.ts.
  2. Add entry to SCOPE_RESOLVERS in scope-resolution/pipeline/registry.ts.

CI auto-discovers the set via tsx. No workflow edit required.

Code references

Module Purpose
scope-resolution/contract/scope-resolver.ts ScopeResolver interface + shared types
scope-resolution/pipeline/run.ts Generic orchestrator
scope-resolution/pipeline/phase.ts Pipeline-phase wrapper (deps: parse, structure)
scope-resolution/pipeline/registry.ts SCOPE_RESOLVERS map
scope-resolution/passes/*.ts Reference-resolution passes (receiver-bound, free-call fallback, compound-receiver, MRO, cross-file return-type propagation)
scope-resolution/graph-bridge/*.ts CLI-local translation from resolved references → KnowledgeGraph edges
scope-resolution/scope/*.ts Generic scope-chain walkers + namespace targets
scope-resolution/workspace-index.ts Build-once O(1) lookup index
languages/python/index.ts Python ScopeResolver hooks + known-limitation docs
languages/python/captures.ts emitPythonScopeCaptures (honors cross-phase Tree cache)
languages/csharp/index.ts C# ScopeResolver hooks + known-limitation docs
languages/csharp/captures.ts emitCsharpScopeCaptures (honors cross-phase Tree cache)
languages/csharp/namespace-siblings.ts Cross-file implicit-namespace visibility hook (reads treeCache)

Performance notes

  • Cross-phase Tree cache: the orchestrator's treeCache (RunScopeResolutionInput.treeCache) lets a scope-resolution per-language hook (emitScopeCaptures) reuse a tree instead of re-parsing. Workers leave it empty — Trees can't cross MessageChannels — so in normal (worker-pool) runs scope-resolution does NOT rely on it: workers serialize each file's ParsedFile (+ capture side-channel) and stream them in, so scope-resolution consumes the pre-extracted artifact rather than re-parsing on the main thread (§ Chunked parse-and-resolve). PROF_SCOPE_RESOLUTION=1 emits hit/miss counters and a worker-engaged warning.
  • Typed relationship iteration: heritage + MRO walk only the EXTENDS / IMPLEMENTS / HAS_METHOD edges via iterRelationshipsByType, not the full relationship map.
  • Workspace-resolution-index: O(1) findOwnedMember / findExportedDef / classScopeByDefId built once per run.
  • Callable-value worklist: dependency-indexed inclusion propagation is linear in a reverse-ordered copy-chain fixture; target/address sets cap at 32 and the whole worklist has a finite budget with no partial edge emission on exhaustion.
  • SCC-ordered cross-file return-type propagation (PR #1050): propagateImportedReturnTypes walks indexes.sccs in reverse-topological order (leaves first), so multi-hop alias chains like models.User → service.user → app.user collapse to the terminal class in a single linear pass. Within each importer, the source module's typeBindings is chain-followed BEFORE mirroring (so we mirror terminal types, not intermediate refs), and the importer's own typeBindings is chain-followed AFTER mirroring (so local const x = importedFn() resolves before downstream importers run). Cyclic SCCs reach a partial fixpoint within a single pass without iterating to convergence — see the ts-circular cross-file-binding fixture which only asserts pipeline-no-throw. PROF output (PROF_SCOPE_RESOLUTION=1) splits finalize from propagate so quadratic regressions in the chain-follow surface independently.

Language-agnostic graph feeding

18 languages → single unified graph. Four abstraction layers:

 Unified Graph Schema (44 node types, 21 relationship types)
           ↑
 Scope-Resolution Pipeline (registry lookup + 3-tier import resolution + MRO)
           ↑
 Language Providers (import semantics, type config, export checker, MRO strategy)
           ↑
 Tree-Sitter Queries (per-language S-expressions, unified capture tags)

Language providers

Each language implements LanguageProvider (language-provider.ts). Key fields:

Field Purpose
id, extensions Language identity and file matching
treeSitterQueries S-expression queries for AST extraction
importSemantics named / wildcard-leaf / wildcard-transitive / namespace
importResolver Language-specific path → file resolution
exportChecker Public/exported symbol detection
typeConfig Type annotation extraction rules
mroStrategy first-wins / c3 / none
descriptionExtractor Optional hook returning a symbol's doc-comment text as its description; feeds the embedding metadata header so doc-only terms are semantically searchable (issue #2270). Most languages register createLeadingDocDescriptionExtractor (shared, language-neutral; per-language comment/wrapper config passed at the call site)
definitionPropertiesExtractor Optional language-owned hook for structured, clone-safe definition metadata. Shared ingestion persists these properties opaquely; the owning provider supplies the extraction semantics.

18 providers in languages/index.ts via satisfies Record<SupportedLanguages, LanguageProvider> — missing a language is a compile error.

Unified capture tags

Per-language tree-sitter queries use different AST node names but produce the same semantic capture tags: @definition.class, @definition.function, @call.name, @import.source, @reference.inherits. Downstream extraction needs no language branching. Defined in tree-sitter-queries.ts.

Import resolution

Per-language import resolution uses the configs + factory pattern (like call/method/class extractors). Each language declares an ImportResolutionConfig in import-resolvers/configs/, listing an ordered chain of ImportResolverStrategy functions. createImportResolver() (in resolver-factory.ts) composes them: first non-null result wins. Low-level helpers shared across strategies live alongside the configs in import-resolvers/ (e.g. go.ts, rust.ts, python.ts).

Unified 3-tier algorithm (model/resolution-context.ts), per-language importSemantics controls which tier activates:

Tier Confidence Mechanism
1 — same-file 0.95 Symbol table for caller's file
2 — import-scoped 0.9 NamedImportMap chains (named) or all files in importMap (wildcard)
3 — global 0.5 O(1) index lookups: class, impl, callable. Fallback only
Import strategy Languages Behavior
named TS, JS, Java, C#, Rust, PHP, Kotlin Only explicitly imported names visible
wildcard-leaf Go, Ruby, Swift, Dart Whole-package import, no transitive re-exports
wildcard-transitive C, C++ #include closure chains through re-exports
namespace Python Module aliases resolved at call site

Chunked parse-and-resolve

parse processes files in ~20 MB byte-budget chunks to bound memory. Per chunk:

  1. Worker pool dispatches files (the sole parse path — there is no sequential fallback; skipWorkers, --workers 0, and GITNEXUS_WORKER_POOL_SIZE=0 are rejected with an actionable error)
  2. Each worker: detect language → load grammar → run queries → return unified ParseWorkerResult
  3. Synthesize wildcard bindings (wildcard-synthesis.ts)
  4. Resolve imports
  5. Collect BindingAccumulator entries for cross-file propagation

Inheritance edges are emitted later, by the scope-resolution phase (preEmitInheritanceEdges + emitHeritageEdges), not during parse.

Workers: workers/worker-pool.ts, workers/parse-worker.ts.

Worker-serialized ParsedFiles (#2038). To index very large repos (e.g. the Linux kernel) without OOM, the worker pool is the sole parse path and workers serialize each file's ParsedFile (plus its capture side-channel) in parallel, streaming them to scope-resolution through a disk-backed store. Scope-resolution consumes the pre-extracted artifact instead of re-parsing every file on the main thread — tree-sitter's native input buffers are not GC-reclaimable, so the former main-thread re-parse leaked native memory until the process died. Pool creation is lazy / cache-miss-gated, so a warm all-cache-hit run replays cached worker output without spawning a worker (hence usedWorkerPool can be false even when the repo has parseable files).

Inheritance and MRO

Inheritance is captured by the @reference.inherits tag and emitted by the scope-resolution phase: preEmitInheritanceEdges resolves each base in scope, then emitHeritageEdges writes the EXTENDS/IMPLEMENTS edges. The phase then computes method resolution order via each ScopeResolver's buildMro hook, feeding a MethodDispatchIndex used for owner-scoped lookups. Per-language strategy:

  • first-wins — Java, C#, C++, TS, Ruby, Go
  • c3 — Python (C3 linearization)
  • ruby-mixin — Ruby (mixin-aware linearization)
  • none — single-inheritance languages

Full analysis flow

runFullAnalysis in run-analyze.ts orchestrates everything around the pipeline:

CLI (analyze.ts) → runFullAnalysis(repoPath, options, callbacks)
  1. Early exit if lastCommit == HEAD (unless --force)     [0%]
  2. Cache existing embeddings from prior index             [0%]
  3. runPipelineFromRepo() → KnowledgeGraph                [0-60%]
  4. Clean up legacy KuzuDB files                          [60%]
  5. initLbug() → loadGraphToLbug() via CSV streaming      [60-85%]
  6. Create FTS indexes (File, Function, Class, Method...) [85-90%]
  7. Restore cached embeddings (batch insert)              [88%]
  8. Generate new embeddings if --embeddings               [90-98%]
  9. Save metadata + register repo + update .gitignore     [98-100%]
 10. Generate AI context files (AGENTS.md, CLAUDE.md)      [100%]

Options: --force (rebuild regardless), --embeddings (opt-in, skipped if >50k nodes), --skipGit, --noStats.

Storage

<repo>/.gitnexus/
  ├── lbug           # LadybugDB database
  ├── lbug.wal       # Write-ahead log
  ├── lbug.shadow    # Shadow sidecar (checkpoint staging)
  ├── lbug.lock      # Single-writer lock
  ├── lbug.wal.checkpoint, lbug.checkpoint.{intent,apply}.lock  # checkpoint-in-flight artifacts; left behind only by an interrupted checkpoint, consumed by the next writable open
  ├── lbug.{wal,shadow}.dirty-recovery  # parked sidecars from a crashed run; safe to delete
  ├── gitnexus.json  # lastCommit, indexedAt, stats (primary metadata file)
  └── meta.json      # legacy mirror of gitnexus.json, kept in sync (see MIGRATION.md)

~/.gitnexus/
  ├── registry.json  # Global repo registry (MCP discovery)
  └── stores/<key>/  # Shared sibling index store (see below)
      ├── caches/                         # parse cache + durable ParsedFile store
      ├── commits/<commit>-<featureKey>/  # one immutable graph per commit + settings
      └── checkouts/<slot>/               # one checkout's metadata, membership, and
                                          # private graph when it has local edits

The flat <repo>/.gitnexus/ layout applies to a standalone repository and whenever GITNEXUS_STORAGE_PATH / GITNEXUS_STORAGE_ROOT is set. A repository with linked worktrees, and clones with the same origin URL, share one stores/<key>/ automatically (a clone opts out with analyze --no-share; GITNEXUS_SHARED_STORE=off turns sharing off entirely). Each sharing checkout keeps only a .gitnexus/store.json pointer to its store. Path resolution lives in shared-store.ts.

Read-only opens self-heal an interrupted checkpoint: the refusal is classified and cleared by one writable open (probe + CHECKPOINT) before the read-only open is retried — see sidecar-recovery.ts (isReadOnlyCheckpointInProgressError) and the lbug-interrupted-checkpoint-recovery integration test.

Managed by repo-manager.ts.

LadybugDB schema

Defined in lbug/schema.ts. Separate node tables per type, single CodeRelation table.

Node tables: File, Folder, Function, Class, Interface, Method, Constructor, CodeElement, Struct, Enum, Macro, Typedef, Union, Namespace, Trait, Impl, TypeAlias, Const, Static, Property, Record, Delegate, Annotation, Template, Module, Community, Process, Route, Tool, Section, Embedding.

Relation types (CodeRelation.type): CONTAINS, DEFINES, CALLS, IMPORTS, INHERITS, EXTENDS, IMPLEMENTS, USES, DECORATES, HAS_METHOD, HAS_PROPERTY, ACCESSES, METHOD_OVERRIDES, METHOD_IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS, HANDLES_ROUTE, FETCHES, HANDLES_TOOL, ENTRY_POINT_OF, WRAPS, QUERIES, INJECTS, CONDITIONAL_ON, DECLARES, ADVISED_BY, BINDS_EVENT_HANDLER, EMITS_EVENT.

Optional --pdg additions (off by default, opt-in via gitnexus analyze --pdg; see Optional CFG/PDG emission above): a BasicBlock node table, plus the PDG relation types CFG, REACHING_DEF, CDG, TAINTED, SANITIZES, and TAINT_PATH on the same CodeRelation table. These are deliberately kept out of the default VALID_RELATION_TYPES / web graph schema — query them via cypher, explain, or pdg_query.

Embeddings (src/core/embeddings/): Snowflake arctic-embed-xs (384D). Embeddable: File, Function, Class, Method, Interface. Incremental via SHA1 content hash. Separate Embedding table.

Search (src/core/search/): Hybrid BM25 + semantic vector, merged via Reciprocal Rank Fusion (K=60).

Known limitations

Overloaded method resolution

Node IDs use arity suffix (#<paramCount>): Method:file:Class.method#1 vs #2.

Same-arity disambiguation: type-hash suffix ~type1,type2 when collision detected and type annotations present. Languages without types (Python, Ruby, JS) use arity-only. TS/JS overload signatures excluded (collapse to implementation body). See #651.

C++ const-qualified: $const suffix after type-hash when non-const collision exists: Method:file:Container.begin#0$const.

Generic/template types: type-hash uses rawType (full AST text including generics): ~vector<int> vs ~vector<std::string>.

ID stability: collision-only tags mean IDs change when overloads are added. save#1 becomes save#1~int when save(String) is added.

Variadic matching: confidence 0.7 when one side is variadic and the other has fixed count.

METHOD_IMPLEMENTS confidence tiering:

Match quality Confidence
Exact parameter types match 1.0
Arity match, types unavailable 1.0
Variadic vs fixed 0.7
Insufficient info 0.7