Commit graph

127 commits

Author SHA1 Message Date
Adam B.
7376347058
feat(ingestion): add tRPC procedure detection and MCP chain/route surfacing (#3339)
* feat(ingestion): tRPC pattern detection — procedures, curried calls, route extraction

Three changes to properly index tRPC router files:

1. HOC-in-pair patterns: detect procedures like
   'create: procedure.mutation(async ({ input }) => {...})'
   as named Function nodes (pair > call_expression > arguments > arrow_function)

2. Curried call detection: capture 'workflow(db)(input)' chained calls
   where call_expression.function is itself a call_expression

3. tRPC route extraction: new route-extractors/trpc.ts detects
   .query()/.mutation()/.subscription() procedures, maps to /trpc/* routes
   with prefix inference from router variable names

4. Function/Const dedup: structural check skips Const nodes when
   variable_declarator value is arrow_function/function_expression

5. Process-route linking: match by (filePath, methodName) instead of
   filePath only, preventing shared flows across procedures in same router

(cherry picked from commit eced15e60a39a181a48bff8c5bbe278d26aa53b8)

* fix(route-extractors): line-by-line scanner for chained tRPC procedures

Original regex only matched direct patterns (list: proc.query()) but not
chained patterns (create: proc.input(z.object({...})).mutation()). Only
41/256 routes were detected on Jurialis.

Rewrote extractTrpcRoutes() as a line-by-line scanner that:
- Detects procedure keys starting with *Procedure builders
- Tracks currentProcedure forward until terminal .query()/.mutation()
- Handles .input() chaining naturally
- Deduplicates via seen Set on procedurePath

Result: 245 routes from 30 routers (was 41), 256 total after re-index.
(cherry picked from commit 49edf953e93e6b67b77932e866fcbfb8d089b1f3)

* feat(mcp): expose tRPC chains via context/query tools

Eight fixes to make the tRPC route->procedure->workflow->sub-workflow chain
visible through the standard MCP tools (context, query) that AI agents use,
instead of requiring raw Cypher queries.

- A context: order incoming/outgoing CALLS test-last, raise LIMIT 30->100
- B query: batched ENTRY_POINT_OF lookup, surface route URL on processes
- C query: mark process_symbols entry point with is_entry_point=true
- D entry-point-scoring: skip UTILITY_PATTERNS (get*/set*) penalty for
  symbols in tRPC router files so framework boost (3.0x) is preserved
- E fts-schema: index Route nodes so /trpc/* URLs are keyword-searchable
- F tools: bump max_symbols default 10->25 to fit procedure->workflow chain
- G context: optional chain_depth (0-3) param walks CALLS edges in both
  directions and returns layered chain field
- H LadybugDB bug workaround: WHERE r.type IN [...] silently drops edges
  on relationship properties; replace with OR chains (7 occurrences in
  local-backend.ts, pdg-impact.ts, graph-queries.ts)

Validated on Jurialis: setProviderCap procedure now appears as caller of
setProviderCapWorkflow, /trpc/cabinet.setProviderCap Route is searchable,
chain_depth=3 returns the layered call graph.

(cherry picked from commit c24df0adce91e80f191a5c5d2e8b14e9da677366)

* feat(mcp): surface is_entry_point + routes in context, raise query limit

context() is the mandatory pre-edit tool (AGENTS.md). Until now an agent
had to issue a separate query() call just to learn whether its symbol is
a process entry point or which HTTP route it handles. These two fields
make context() self-sufficient for the bmad-dev flow.

- context: query STEP_IN_PROCESS now returns p.entryPointId; new
  ENTRY_POINT_OF lookup attributes routes only from processes where the
  symbol is the entry point (middle steps do not own the route) plus any
  direct Route->symbol handler edge (tRPC procedures outside any Process)
- context: new top-level is_entry_point (true only) and routes[] fields
  ({url, method?}), emitted only when non-empty to keep the diff additive
- query: raise default limit 5->10 so dense domains (tRPC action router
  with 20 chains, data-export with 9 flows) surface more of their flow
  set without an explicit param

(cherry picked from commit 004f5b0ba985fcef3dfac5e67a2f2f7262184215)

* docs(mcp): document chain_depth, routes, and entry-point flag; neutralize examples in comments

* fix(mcp): address code-review findings on tRPC entry points, queries, and dispatch

- entry-point scoring: optional-chain framework detection and normalize
  path separators before router-pattern matching (P1 crash on .js routers)
- tRPC extractor: emit controllerName as null (callers resolve via
  lookupClassByName; a router object stringified as name corrupted lookups)
- route dedup: key seenRoutes by method:url so GET/POST pairs survive
- legacy TS queries: drop curried-call patterns (arity-corrupting for
  overload resolution) and require non-array callee on pair member calls
- registry-primary TS query: mirror the non-array-callee predicate on the
  new pair member-expression patterns (fixes query compile error)
- group tool port: forward chain_depth into per-tool context args

* fix(mcp): nested tRPC router paths, controller-less route binding, query chain_depth

- trpc extractor: brace-depth nesting stack composes sibling router paths (bare router() import style included); merge-prefix dot normalization; strict publicProcedure allowlist gate; drop phantom router metadata

- call-processor: bind controller-less tRPC routes to same-file handlers via exact single-match symbol lookup; ambiguous or missing handlers skipped

- query/group query: optional chain_depth (0-3) enriches ranked processes with their context chain; fix _computeContextChain layer docstring

- tree-sitter TS/JS scope queries capture string-key function declarations; tRPC pattern scoring narrowed to server router files

- parse-cache schema bump to v103 for extractor and route-binding changes

- add trpc route extractor regression tests (9 cases incl. sibling nesting and merge prefix)

* Address PR review feedback (#3339)

Close remaining review threads: compact tRPC keys, comment-safe terminals,
quoted-key identifier HOC captures, group-query default alignment, and
LadybugDB label scalars on context chains.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test: pin PARSE_CACHE schema bump to 103 (PR #3339 review fixes)

* fix(ingestion): bind same-name tRPC handlers and surface HANDLES_ROUTE

Same-name procedures resolve by startLine, the extractor keeps nested paths
through multiline schemas, and context/query UNION HANDLES_ROUTE for leaf
procedures that never become Process entries.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Clamp group query bounds, drop HOC pair false positives, emit tRPC
terminal lines and hyphenated quoted keys, cap chain BFS concurrency,
and restrict the router utility exemption to accessor names.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Score JS/JSX tRPC routers like TypeScript, ignore inner db.query and
unrelated .merge calls, mask regex braces, and assert clamp tests invoke
the query mock.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(ingestion,mcp): emit-side HOC callback guards + MCP schema prettier (PR #3339 round 2)

* Address PR review feedback (#3339)

Treat `/` after return-style keywords as a regex, drop the synthetic
appRouter path prefix, and bind the create mutation via an inline callback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Qualify query() docs so is_entry_point is promised only when the entry
symbol is among the search hits.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Emit tRPC terminals only at procedure paren depth, drop the filename
prefix on bare appRouter = router(), mask regex after if (), and cap
groupQuery member fan-out.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

De-duplicate context-chain BFS nodes reached through multiple frontier
edges, and bind tRPC identifier callbacks (`.mutation(handler)`) to the
handler symbol instead of the procedure key.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Apply LIMIT 50 after DISTINCT neighbors in the context-chain BFS, and
treat the official lowercase `procedure` builder as a tRPC procedure key.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Keep the later duplicate tRPC object-literal key so route binding matches
the handler JavaScript actually ships.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Compose same-file identifier-mounted tRPC subrouters so admin: adminRouter
emits the live admin.list path instead of an unprefixed /trpc/list.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Memoize tRPC identifier-mount paths so a depth-N chain stays linear, and
gate it with the build-free measure.mjs / baselines.json harness.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Compose identifier mounts without phantom /trpc URLs, bind wrapped and
multiline handlers, fail-close missing bench budgets, and reject query
page bounds before search.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Reject advertised MCP query bounds in group mode and chain_depth
instead of clamping or accepting non-integers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Reject invalid groupContext chain_depth and ignore non-callable
same-file tRPC handler candidates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Memoize live tRPC mount paths so the depth-N chain bench stays linear,
and render invalid group/MCP bounds without throwing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Keep t.merge('prefix.', namedRouter) procedures live by recording a
zero-hop mount so the unmounted-router drop does not hide them.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3339)

Recognize type-annotated `const adminRouter: AppRouter = t.router(`
bindings so identifier mounts still compose.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 09:44:42 +01:00
Twisted_Arrow
dcb2eb5cb4
fix(swift): resolve imports from Package.swift targets, not path segments (#3105)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(swift): match repeated SPM target prefixes

* test(swift): cover repeated SPM target prefixes

* test(swift): cover valid prefix before later partial match

* docs(swift): clarify target grouping parity scope

* docs(swift): clarify target grouping parity scope

* docs(swift): clarify target grouping parity scope

* fix(swift): resolve imports from Package.swift targets, not path segments

Stop fabricating IMPORTS from import Foundation onto a same-named folder.
Declare modules from Package.swift when the manifest is usable; keep
Sources/* for grouping and fail-open folder resolve minus SDK names.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): keep empty Package.swift declarations and nested .target() deps external

An inferred Sources/* folder is grouping-only. A dependency .target(name:) is not a module. Treat both as unresolved so import Foundation cannot bind to a decoy folder.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): keep implicit IMPORTS intra-group and gate linear Package.swift resolve

@_exported must not paint sibling files as implicit imports. A dedicated
bench pins declaration-only resolve and (t_4n/t_n)/4 linearity.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Honor member-only @_exported imports, skip comments while scanning
Package.swift factories, fail-open mixed helper-built target lists, and
block CoreData/CoreGraphics decoy folders on the inferred path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Match path: "." as the package root, skip block-commented Package.swift
factories, read import kind from the clause only, and skip capture tests
when the optional Swift grammar is missing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): scan Swift import-kind without nested regex backtracking

CodeQL js/redos flagged the comment-skipping IMPORT_KIND_RE; a linear walk keeps the same kind tokens.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): linear Package.swift factory scan and gate Swift context

parseSwiftPackageManifest re-walked every prefix for comments (O(n²) in factory count). Resume the scan and cover nested factories in one pass. Wire Swift into the import-target context arm now that resolveImportTarget is 5-arg.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Tighten Package.swift and import-text scanners: skip comments/strings, reject escapes, treat ident + [ as incomplete, and drop the unused factory-comment wrapper.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): prettier the @_exported availability fixture

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Judge Package.swift completeness from Package(...)'s own targets: argument instead of raw-text regexes, nest block comments when reading an import kind, and strip leading ./ from declared target paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Collect Package.swift factories only from Package(targets: [...]), fail-open on computed array elements, and treat // after a label colon as a comment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Ignore stray factories when Package() exists but omits targets:.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3105)

Require the Package-scan seen box so the always-true undefined guard goes away.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-19 15:50:31 +01:00
azizur100389
a578747455
fix(swift): resolve nested constructors in extensions (#3308)
* fix(swift): resolve nested constructors in extensions

* fix(swift): preserve qualified extension owners

* Address PR review feedback (#3308)

Recover qualified Swift extension owners through public / attribute prefixes, and stop last-dot-guessing when source text is present.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): keep attribute text from stealing extension owners

Bound header recovery so @available messages cannot rekey a fragment, and still inject nested types when the Class scope has no bindings.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): nest comments and keep the public extension fixture valid

Review follow-up: skip nested /* */ in the header scan, put the Inner.Entry decoy in a parsed file, and mark Outer/Container/Entry public so the live fixture is valid Swift.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): recover Unicode identifiers as extension owners

The header regex was ASCII-only, so extension Café.Container keyed as Caf and dropped nested-type siblings. Match ID_Start/ID_Continue segments instead.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(swift): skip raw strings and decode UTF-8 scope columns

Header recovery treated #"..."# as an ordinary quote and sliced Tree-sitter byte columns as JS offsets, so a same-line Café prefix or a raw attribute message could steal or drop the extension owner.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-19 09:46:53 +01:00
Parafee41
795cf0e151
fix(swift): resolve inherited protocol extension calls (#3309)
* fix(swift): resolve inherited protocol extension calls

* fix(scope): gate inherited implicit receiver lookup

* fix(swift): resolve call result types by exact callee

* test(swift): align cache and local call expectations

* fix(scope): reconcile replay diagnostics

* fix(swift): preserve exact callable return types

* fix(swift): capture throwing async call results

* test(swift): refresh capture golden

* test(swift): refresh scope capture baseline

* fix(swift): require explicit callable returns

* fix(scope): preserve duplicate return metadata

* chore(scope): align index documentation

* Address PR review feedback (#3309)

Stamp only protocol/class-extension members (nested QN + SPM buckets),
arity-narrow implicit-this across MRO, and keep Swift type peeling out
of shared workspace-index via stripTypePreservingDecoration.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3309)

Stamp extension members even when the extension declares a nested type, keep inherited class members ahead of protocol-extension defaults, and report replay-only interface-dispatch fan-out drops.

Note: pre-existing failure in gitnexus tsc against an older gitnexus-shared dist not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3309)

Walk inherited implicit-this owners nearest-first so a nearer override wins, and tighten Swift owner-stamp tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-19 07:44:35 +01:00
Abhinav Pandey
0ede6ae501
fix(rust): respect Cargo target boundaries in name fallback (#3294)
* fix(rust): respect Cargo target boundaries in name fallback

* fix(rust): read Cargo sources through validated file descriptors

* fix(rust): require matching imports for crate-root guesses

* fix(rust): require Cargo root identity for cross-target imports

* fix(rust): enforce Cargo identity across fallback imports

* fix(rust): keep Cargo membership on typical derive/macros (#3294)

Abort only genuine unknown expansion, ignore Cargo artifact layouts instead of every `target` path segment, and require a covering import for unanswered crate-root name fallback so #3253 still holds on ordinary Rust sources.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(rust): gate Cargo target membership in CI (#3253)

Add a parse-dispatch-rounds-style bench so derive/macro abort, target-path globs, and superlinear membership walks fail in CI instead of staying graph-invisible.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(rust): verify Cargo macro and re-export evidence

* style(bench): format Cargo membership benchmark

* fix(rust): require imports for cross-file root guesses

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-16 11:36:07 +01:00
mengkaka
79543c8f83
feat(storage): add configurable index storage and content retention tiers (#3060)
* feat(storage): add configurable index storage and content retention tiers

Rebase #3060 onto current origin/main. Keep GITNEXUS_STORAGE_PATH,
GITNEXUS_STORAGE_ROOT, and GITNEXUS_CONTENT_RETENTION, and fold in
main's FTS skip, embed-session, and help-text updates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Keep legacy registry rows on the local storage fallback, resolve
symlinks before the destructive-path guard, and align hook lookup
with CLI branch slugs, branch-slot metadata, and longest-path match.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Only list swept upload directories after a successful removal so
callers cannot treat a permission or transient rm failure as gone.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Document that getStoragePath may consult registered storage while
this module still does not mutate the global registry.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(storage): close review findings for external indexes and retention

Re-inspect ownership under the analyze lock, fail-closed when the
registry file is missing, and keep skip-git hook discovery plus
retention fields on HTTP/MCP list surfaces. /api/file stays 410
unless contentRetention is full.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3060)

Treat lock-only index dirs as empty, honor HTTP --force storage policy, and prefer registered plus branch-aware slots in hooks and augment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Keep hook fallbacks inside the current worktree, compare foreign-local slots canonically, and make storage fixtures survive ownership validation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix macOS hook test expecting realpath'd registry paths.

resolveHookRepo returns the written registry path, not a filesystem realpath, so the assertion must match that.

* Address gitnexus-check warnings on hook install docs and slot tests.

The Cursor troubleshooting list omitted registry-query.cjs, and the writable-slot test only checked that isDirectory exists instead of that the path is a directory.

* Align the HTTP catalog source-scan with skippable resolveRepo validation.

resolveRepo lists fresh repos with validate: options.validateStorage !== false so DELETE can skip prune; the test still required a literal validate: true.

* Harden storage path sinks so CodeQL path-injection and ReDoS alerts clear.

Contain every filesystem probe inside the resolved storage slot with the inline path.relative idiom, reject filesystem-root slots, and trim slot basenames in linear time.

* Settle bridge stamps before writing so CI size/mtime matches stay stable.

LadybugDB can still flush into bridge.lbug after close+rename; persist whole-millisecond mtimes and wait for consecutive stats to agree so a freshly written pair matches.

* Type the settled bridge stat as fs.Stats so tsc does not see bigint.

Awaited<ReturnType<typeof fsp.stat>> collapsed the bigint overload and broke prepare/typecheck on CI.

* Keep the bridge mtime stamp exact so same-size swaps still fail the pair check.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Wrap the bridge stamp predicate so prettier --check stays green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Require a quiet interval before stamping a settled bridge file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Reuse shared storage and settle helpers instead of local copies.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-12 20:31:55 +00:00
Aakash Sharma
d1a3edd333
perf(mcp): avoid O(n) git spawns on tools/list with many repos (#3259)
* perf(mcp): avoid O(n) git spawns on tools/list with many repos

toolSchemaRepoRequirements called listAllowedRepos -> listRepos -> checkStalenessAsync for every registered repo. With ~200 repos, this spawned 200 parallel git rev-list processes on every tools/list discovery call, causing a ~30s delay.

Replaced with a lightweight countRepos() method that reads the registry file once without spawning git processes, preserving full staleness checks for list_repos.

* Address PR review feedback (#3259)

Count the validated registry in countRepos so tools/list cannot advertise a multi-repo schema for ENOENT ghosts, and update the unrestricted listTools mocks to that contract.

Note: pre-existing failure in update-notice.test.ts (missing dist/cli/mcp.js) not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Add a 200-repo tools/list bench for the countRepos path (#3259)

Pin the #1363 comparison (listRepos git fan-out vs validated countRepos / listTools) in-tree so the latency claim can be re-run. Also drop the change-history comments on the unrestricted schema arm.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Gate the tools/list bench with baselines and CI --check (#3259)

Exact registry/schema floors plus ratio timing, no millisecond ceiling, so restoring listRepos() on tools/list fails CI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3259)

Align unrestricted tools/list schema flags with the refreshed registry snapshot, and make the bench reject a non-positive BENCH_REPS and isolate fixtures by N.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3259)

Isolate the tools/list bench from GITNEXUS_MCP_READ_ONLY and create the default fixture under mkdtempSync so CodeQL is not looking at a predictable /tmp path.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-11 13:18:28 +01:00
Parafee41
c9b391357e
fix(group): detect function-local Python imports (#3254)
* fix(group): detect function-local Python imports

* fix(group): parse Python imports structurally

* fix(group): use safe Python parser wrapper

* fix(group): isolate Python parse timeouts

* fix(group): skip recovered indented Python from-imports

Error-recovered trees can promote unclosed-docstring lookalikes into
real import_from_statement nodes. Drop those indented imports, bound
file reads, log parse timeouts, and add fingerprint floors for the
tree-sitter scan.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-11 06:23:52 +00:00
Yahoo
0edf9ce0ff
fix(ruby): guard gem requires with dependency metadata (#3096)
* docs(plans): add ruby gem require boundary plan

* fix(ruby): guard gem requires with dependency metadata

* fix(ruby): scope gem sources by manifest

* test(ruby): model resolved lockfile specs

* fix(ruby): stop local gem suffix fallthrough

* docs: remove Ruby resolution plan

* test(ruby): gate gem resolution correctness and scaling

* test(ruby): align gem benchmark baseline with ratio gates

* Address PR review feedback (#3096)

- Strip a trailing .rb so local and external gem prefix matching follows Ruby's optional-suffix require.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-10 16:04:11 +01:00
Kevin Rajan
3236e2fbcd
fix(cobol): prefer copybook dirs so COPY EXTERNAL does not hit vendor decoys (#3240)
* fix(cobol): prefer copybook dirs so COPY EXTERNAL does not hit vendor decoys

COPY of an out-of-repo member first-won any same-named .cpy, so
vendor/EXTERNAL.cpy became a live cobol-copy IMPORTS edge. Share one
resolver between census and the regex processor: prefer copybooks/cpy/copy
plus the importer dir when present, else fail-open. Drop COBOL KNOWN_GAPS.

Fixes #2967

* bench(cobol): update depth budget for copybook-dir preference (#2967)

COBOL resolver now prefers well-known copybook directories (copybooks/,
cpy/, copy/, plus importer dir) over vendor paths when resolving COPY
statements. This intentional behavior change moves the resolver from
depth-free (prior measured ~0.885) to depth-sensitive (measured 1.751
on CI run 34394116972), because the new preferredCopybookDirs check
walks path components.

- Raise depth_budget from 1.6 to 2.4 (~1.37x the measured ratio)
- Update _measured.depth_ratio from 0.885 to 1.751
- Add _cobol_copybook_dir_preference_2967 note documenting the change

The COBOL fingerprints already reflect the new target set behavior
(vendor/EXTERNAL.cpy correctly returns null when a copybook dir is
present) per commit fc9e8270.

Co-authored-by: Kevin Rajan <kvnloo@users.noreply.github.com>

* bench(cobol): update baselines for copybook-dir preference (#2967)

COBOL preferred-dir filtering now affects resolution outcomes. The
unique-arm layouts (mixed copybooks/src dirs) drop from 1153 to 442
resolved as files outside preferred directories are correctly filtered.
The collide arm (all files in svc${d}/copybooks) keeps 1153 resolved
because ALL files remain in the preferred class.

- Update small/deep resolved: 1153 → 442
- Update small/large/deep fingerprints for new target set
- Raise heap_bound_bytes.cobol: 3500000 → 6200000 (1.5x measured 4112464 B)
- Add measure.mjs exception: collide legitimately differs from small
- Update _heap_bound_note with new cobol measurement context
- Expand _cobol_copybook_dir_preference_2967 note to explain collide delta

The collide arm now measures collision behavior within the preferred
class rather than across mixed layouts — an intentional outcome of the
preferred-dir semantics rather than a corpus defect.

Refs #2967

Co-authored-by: Kevin Rajan <kvnloo@users.noreply.github.com>

* fix(cobol): P1-A stem key with uppercase extensions, P1-B polyglot preferred-class latch

P1-A: Use raw extension for path.basename so CUSTREC.CPY keys as CUSTREC
  - Before: path.basename('CUSTREC.CPY', '.cpy') -> 'CUSTREC.CPY' (no strip)
  - After: path.basename('CUSTREC.CPY', '.CPY') -> 'CUSTREC' (stripped)
  - Processor used raw extension; census now matches

P1-B: Latch preferred-class only from copybook-tier paths
  - Before: docs/copy/README.md triggered hasPreferredDir = true
  - After: check preferred-dir only after extension filter
  - Processor receives polyglot allPathSet; test added

Test coverage:
  - cobol-copy-external-imports.test.ts: processor pins for both fixes
  - cobol-import-target-parity.test.ts: resolver parity updated to new behavior
  - All existing COBOL tests pass

Co-authored-by: Kevin Rajan <kvnloo@users.noreply.github.com>

* test(cobol): fix Mixed.CPY integration test for P1-A behavior

The integration test in cobol-import-index-reuse.test.ts had outdated
expectations from before P1-A. With P1-A, path.basename uses the raw
extension, so Mixed.CPY is now keyed as MIXED (extension stripped).

Before (pre-P1-A):
- Mixed.CPY → keyed as MIXED.CPY (uppercase ext not stripped)
- COPY MIXED.CPY → found, COPY MIXED → null

After (P1-A):
- Mixed.CPY → keyed as MIXED (raw ext stripped)
- COPY MIXED → found, COPY MIXED.CPY → null

Updated test expectations to match P1-A behavior. All three validation
tests now pass:
- cobol-import-index-reuse.test.ts (integration, index reuse)
- cobol-copy-external-imports.test.ts (processor P1-A/P1-B pins)
- cobol-import-target-parity.test.ts (resolver parity)

Fixes PER-994 CI failure on abhigyanpatwari/GitNexus#3240.

Co-authored-by: Kevin Rajan <kvnloo@users.noreply.github.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kevin Rajan <kvnloo@users.noreply.github.com>
2026-09-10 15:42:59 +01:00
Abhinav Pandey
2220f4d851
fix(resolution): label fallback guesses and preserve export visibility (#3190)
* fix(resolution): distinguish name guesses and preserve export visibility

* test(go): keep method enrichment fixture in one package

* fix(resolution): address split review edge cases and evidence reporting

* fix(exports): recognize imported and expression-local receivers

* fix(resolution): align export and target evidence with language scope

* fix(ingestion): preserve lexical import provenance through resolution

* test(ci): rebalance Windows shards from measured slow suites

* Address PR review feedback (#3190)

- Label constructor unique-name guesses as global-name-fallback and run language vetoes
- Tighten Go qualified, Rust crate::, and Ruby class-reopen fallback guards
- Ignore for-loop shadowed CommonJS receivers and exclude guesses from the resolved-call census
- Refresh FinalizeOutput and hook docs for lexical binding scopes

Co-authored-by: Cursor <cursoragent@cursor.com>

* Tighten review-feedback leftovers on fallback visibility.

Qualified Go calls still respect export and test-package rules, nested Rust src/ stays a module segment, and top-level conditional this.x is treated as CommonJS.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3190)

- Distinguish Swift package prefixes when comparing target modules
- Document lexical import binding and handledSites refusal marking
- Drop the stale ci-scope-parity workflow claim and prototype-safe export verdicts

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3190)

Supply caller source on Ruby visibility cases so they exercise the named allow branches instead of the missing-text bypass.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(resolution): label unique constructor types as name guesses

A workspace-unique class hit in findClassBindingInScope was treated as
an in-scope bind, so Go/JS constructor-form sites skipped the guess
label and the Go unexported veto.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(resolution): keep qualified constructors precise after unique-name split

Bare constructor unique-name hits stay guesses so Go can veto an
unexported type. A written qualifier is now carried as rawQualifiedName
so `new pkg.Foo()` and `models.Box[T]{}` can still recover the unique
class without that veto.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(bench): rebaseline Go/Java scope-capture fingerprints for constructor qualifiers

Generic Go composite literals and qualified Java `new pkg.Foo()` now
carry @reference.qualified-name on existing constructor matches.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(resolution): treat import-reached unique constructors as precise

C++ #include and Rust re-exports do not mint a lexical class binding.
A unique type in an imported file (or imported directory) is therefore
a real bind, not a name guess, so those CALLS edges stay import-resolved.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(resolution): require named or resolved imports for constructor precision

Bare Go Box[T]{} is not package-qualified, and a sibling-file import of a different name is not constructor visibility.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-10 11:38:19 +01:00
Navid EMAD
506432017f
fix(zig): model callable-value references, and stop reporting their absence as exact (#3219)
Some checks are pending
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
Scorecard / Scorecard analysis (push) Waiting to run
* fix(impact): stop reporting 'exact' over unmodelled callable-value references

A function named in VALUE position — `bridge.accessor(Element.getNamespaceUri,
null, .{})`, `{ onClick: handler }`, a comparator handed to a sort — is
registered somewhere rather than called. The registration is modelled (a
`value-ref` site becomes a USES edge; Kythe `ref` vs `ref/call`, Joern
METHOD_REF), but the invocation THROUGH the stored value is not: it happens
later via a struct field, a registry lookup or comptime reflection.

impact()/context() nonetheless reported such a target as `epistemic: 'exact'`.
For lightpanda-io/browser that meant the DOM `Element.namespaceURI` accessor
came back with two internal callers, LOW risk and a claim of completeness —
worse than no answer, because 'exact' tells the reader not to look further.
tools.ts defines 'lower-bound' as "the walk provably missed callers", which is
precisely this case.

computeEpistemicBoundary now probes for inbound USES edges stamped with the
value-ref reason and hedges when it finds any, contributing a boundary note and
a new `causes.callableValueReferences` (unit: distinct referrer symbols).

Read from the graph, not from index metadata: unlike a dropped receiver — which
leaves no edge to find and therefore needs a persisted summary — a value
reference IS in the graph. So the signal needs no re-index and no analyzer
change, and it works on indexes written before this commit.

The cause gets its own slot rather than joining `dispatchBoundary`: a value
referrer is neither an implementation nor an interface-level consumer, and the
two differ in what the reader should do about them — a dispatch boundary is
irreducible, a callable value usually becomes traceable once the provider
models the store/load that carries it.

Language-neutral: every provider that emits a `value-ref` capture participates.
The writer and the reader now share VALUE_REF_EDGE_REASON, because drift
between them would fail silently — the probe would match nothing and every
answer would go back to claiming certainty.

* fix(zig): emit `value-ref` for a callable named in value position

`mapReferenceKindToEdgeType` has handled the registration-vs-invocation case
since #2437, and TypeScript, JavaScript and C++ all emit `value-ref`. Zig
emitted zero — so a Zig function handed somewhere as a VALUE was absent from
the graph entirely.

That is not a corner: Zig's JS bridge is built out of this one shape,

    pub const namespaceURI = bridge.accessor(Element.getNamespaceUri, null, .{});

and `bridge.{accessor,function,indexed,…}` appears 2,047 times across 257 files
in lightpanda-io/browser — the project's whole JS<->Zig surface, none of it
reaching the graph. `impact` on `Element.getNamespaceUri`, which IS the DOM
`Element.namespaceURI` accessor, answered with its two in-file callers.

Three query rules, tagging `@reference.value-ref` on: a bare identifier in
argument position (`bridge.accessor(_tagName, …)`), a qualified one
(`bridge.accessor(Element.getNamespaceUri, …)`), and a const binding initialiser
(`pub const defaultHandler = onReset;`). Everything downstream already existed —
no new edge type, no schema change, no capture-machinery change, no baseline
edited.

One grammar detail is load-bearing: in tree-sitter-zig, call arguments are
DIRECT children of `call_expression` — there is no `arguments` node, only
builtins have one — so the callee has to be consumed explicitly by `function:`.
Without that binding the same rule also matches the callee of `foo(bar)` and
mints a USES edge shadowing the call's own CALLS edge. There is a test for it.

The rules are deliberately broad — `js.Bridge(Element)` and `register(count)`
match too — because the callable gate in the property-dispatch pass
(Function/Method/Constructor only) is the filter, the same design that keeps
TypeScript's `{ port: DEFAULT_PORT }` from registering anything. Measured on the
real corpus: all 3,169 emitted edges land on a Method (3,146) or Function (23).

Deliberately NOT modelled: the terminal invoke. The value reaches
`Accessor.init` -> a struct field -> `Factory.zig` reflection (`inline for`,
`@typeInfo`) -> `Caller.zig`'s `@call(.auto, func, args)` over `func: anytype`.
Resolving that needs comptime evaluation. The point is to stop dropping the
reference; the shortfall is now reported as `epistemic: "lower-bound"` by the
preceding commit instead of being papered over.

Measured on lightpanda-io/browser (698 .zig files), wiped index both times:
30,222 nodes / 71,070 edges -> 30,222 nodes / 74,229 edges. The delta is
entirely `USES` (3,169 after vs 10 before, and every one is a value-ref); CALLS,
ACCESSES, IMPORTS, HAS_METHOD, DEFINES and the rest are unchanged — purely
additive, no node invented.

The fixture joins the existing `zig-idioms` corpus rather than adding a new
one, because `bench/receiver-resolution` uses `test/fixtures/lang-resolution` as
its `--check` corpus. That gate, `scope-capture` (15 languages),
`zig-cross-file-resolution` and every other bench `--check` pass unchanged.

* fix(review): resolve qualified value references through their owner, and stop
hedging on registrations the analyzer already followed

Five findings from the review bot on #3219; four valid, all addressed.

1. QUALIFIED VALUE REFERENCES RESOLVED BY TAIL NAME (the serious one).
   `bridge.accessor(Element.getNamespaceUri, …)` was resolved with
   `findCallableBindingInScope(site.inScope, site.name, …)`, which never sees
   the receiver and gives LOCAL bindings precedence. Reproduced:

       const Element = @This();
       pub fn getThing(...)                 // main.getThing
       pub const JsApi = struct {
           fn getThing(...)                 // JsApi.getThing
           pub const thing = bridge.accessor(Element.getThing, null, .{});
       };

   emitted `USES JsApi → JsApi.getThing` — a WRONG edge, which is worse than
   the missing edge this PR set out to fix.

   `resolveValueRefTarget` now resolves a site carrying an explicit receiver
   through `findClassBindingInScope` → `findOwnedMember` (the machinery
   `receiver-bound-calls` already uses), gated on `CALL_TARGET_TYPES` because
   `findOwnedMember` also answers with fields. A bare site keeps the lexical
   walk, which is what an unqualified name means. When the owner cannot be
   resolved the site is DECLINED rather than falling back: declining costs a
   reference, falling back mints a confident edge to the wrong function, and
   the missing reference is now reported as `lower-bound` anyway.

   Cost on lightpanda-io/browser: value-ref edges 3,169 → 2,799 (−12%). Those
   370 were tail-name coincidences, not registrations — all 94 `Element.zig`
   `JsApi` entries survive, the cross-container case resolves and is correctly
   attributed (`IntersectionObserverEntry.JsApi` → `IntersectionObserverEntry.
   getTarget#0`, not the enclosing observer), and both acceptance probes are
   unchanged: `getNamespaceUri` 7/3/LOW/lower-bound, `getTagNameLower`
   31/10/HIGH/exact.

2. A FAILED PROBE READ AS "NO BOUNDARY". `.catch(() => [])` turned an
   unanswerable query into count 0 and no note, so `exact` could be published
   on the strength of a question that was never asked. It now returns `null`
   and emits a boundary note. This file's own `loadMeta` comment states the
   rule: a probe failing must never read as certainty.

3. NOT EVERY VALUE REFERENCE IS AN UNMODELLED INVOCATION. Where
   `emitPropertyDispatchCalls` sweep 2 synthesized the dispatch (reason
   `property-dispatch`), the walk did NOT provably miss the caller, so hedging
   was noise over an answer that was computed — and a signal that fires on
   every JS/TS hook table stops carrying information. A second probe excludes
   those targets. Zig never sets a property key, so the motivating case is
   untouched. The exclusion is symbol-level, not per-edge, because the graph
   does not record which registration produced which synthesized call; that
   residual is documented at the call site.

4. `LIMIT 50` SILENTLY UNDERSTATED THE COUNT. `rows.length` over a capped row
   set published a ceiling as the documented symbol count. Replaced with
   `COUNT(DISTINCT other.id)` — bounded work without a bounded answer, the
   shape `countByType` in the same file already uses.

5. THE TEST DID NOT PIN THE CALLER COUNT its own comment promised. Added
   `expect(result.impactedCount).toBe(2)`.

New regression tests: qualified references bind the written owner and not the
nearer lexical match; an unresolvable receiver emits nothing rather than a
wrong edge; a dispatch-modelled registration stays `exact`; a probe that cannot
run hedges instead of claiming certainty.

Known limitation, pinned by a test rather than left implicit: when a `@This()`
alias's NAME differs from its container's (`const Element = @This();` inside
`main.zig`), the receiver does not resolve and the reference is declined. The
provider has `rewriteZigThisAlias` for this, but it is applied to type nodes
and extending it to reference receivers would change existing CALL-site
behaviour. Lightpanda's `const Foo = @This();` in `Foo.zig` convention makes
the names coincide, which is why the corpus is unaffected.

* fix(review): bump the parse-cache schema and resolve module-owned value references

Review round 2 on #3219. Four inline items, all reproduced against the
worktree before deciding.

P1 — the new captures could stay INERT on a warm parse cache. Adding
`@reference.value-ref` rules to `ZIG_SCOPE_QUERY` changes
`ParsedFile.referenceSites`, which is a PARSE-TIME fact, but `SCHEMA_BUMP`
stayed at 93. A repo indexed before this branch and re-analyzed after it
replays the old, empty site list for every unchanged `.zig` file — `--force`
included, since shards are content-addressed — so no USES edge is emitted, the
boundary probe measures a real zero, and `impact` on a registered accessor goes
back to `epistemic: "exact"`. #3399 un-fixed on the incremental path most users
are on, with every cold-run test still green. DECISIONS D1-1's "zero changes to
… the schema" conflated the graph schema with the cache schema.

Bumped 93 -> 98, not 94: #3190 claims 94 and #3179 claims 94 through 97 in one
PR. `incremental-parse-cache.test.ts` re-pinned, with 93-97 added to the taken
list. Re-check against origin/main and open PRs immediately before merging.

P2 — a qualified value reference through a MODULE was declined with no hedge.
R1-2 resolves a written receiver with `findClassBindingInScope`, which requires
`isClassLike`; a namespace-only `@import` handle is not class-like, so

    const dom_utils = @import("dom_utils.zig");   // no `@This()` in that file
    pub const comparator = bridge.accessor(dom_utils.compare, null, .{});

resolved to nothing. That is not the conservative half of R1-2's trade-off: a
declined site emits NO edge, so there is nothing for the probe to read and
`impact` on `compare` reports `exact`. Silence, not a hedge — and the pass
comment claiming otherwise was wrong on this path.

`resolveValueRefTarget` now tries the second kind of owner a qualified name can
have. `findNamespaceValueRefTarget` reads the file's `namespace` import edges
for the handle and the target module's own `origin: 'local'` module-scope
bindings for the member — the same channel `receiver-bound-calls` Case 1
already trusts for `dom_utils.compare()`, with Case 1's three guards for Case
1's reasons: `isNamespaceNameShadowed`, local-origin bindings only, and two
distinct defs under one name resolve nothing. `CALL_TARGET_TYPES` gates module
owners exactly as it gates container owners. Language-neutral: it reads generic
namespace import edges, names no language.

Still declined, deliberately: a receiver this index knows under no name at all
(the `@This()`-alias case, owners outside the workspace). There the alternative
is a confident edge to a lexically-nearer function the source did not name.

P2 — `impact-callable-value-references.test.ts` was in neither vitest list.
It opens a real engine via `withTestLbugDB(poolAdapter: true)`, so TESTING.md
puts it in the `lbug-db` include list and the `default` exclude list; it was in
neither, so `default` also collected it into the parallel pool. Added next to
its `impact-epistemic-lower-bound` sibling in both arrays.

P3 — DECISIONS.md was stale against HEAD and embedded host paths. D2-3 still
described the `LIMIT 50` that R1-5 removed and D1-3 still described the
receiver-blind resolution that R1-2 replaced; both now carry explicit
"superseded by" pointers. The `~/code/...` and mise-node paths are replaced
with placeholders.

Two smaller corrections the new cause made necessary:

- `formatImpactResult`'s `lower-bound` header hard-coded "callers binding via
  DI / dynamic dispatch", which now contradicts the value-reference bullet
  printed directly under it. The bullets carry the cause; the header only
  states that the count is a floor.
- `tools.ts` said a `causes.callableValueReferences` of 0 means "nothing was
  missed". The dispatch exclusion is symbol-level, not edge-level, so a target
  with both a followed registration and an unfollowed escape also reads 0. The
  docs now say a 0 means "no unfollowed registration was proven".

Fixture: `src/webapi/dom_utils.zig` (namespace-only) plus three cases in
`Element.zig` — the module-qualified registration, a non-callable module member,
and a `u8` parameter shadowing the handle. The shadow case was verified to FAIL
with the guard disabled, so it is not passing for an unrelated reason.

Gates: `tsc --noEmit` clean, `npm run build` clean, prettier clean.
`resolvers/` 3,630 passed / 3 skipped; `unit/scope-resolution` 2,010 passed;
`impact-callable-value-references` 7 passed under `lbug-db`;
`incremental-parse-cache` 40 passed; eval formatters 104 passed. Bench --check
all PASS with no baseline edited: receiver-resolution, zig-cross-file-resolution,
scope-emission, callable-value-flow, scope-capture (15 languages), python-scope.

* fix(review): guard the class receiver against a shadowing binding, and pin the dispatchability partition

Review round 3 (`gitnexus-check` bot on `cf53bbaa`). Three findings, each
reproduced against the code before deciding.

R3-1 (Error, valid, REPRODUCED) — a CLASS receiver could resolve through a
shadowing value binding. `findClassBindingInScope` is a class-only walk: it
filters the scope chain by `isClassLike`, so it steps over a nearer binding that
is a value and keeps climbing — and past the chain entirely, into a
qualified-name fallback that answers with the unique workspace definition of the
name. A `u8` parameter named `Ticker`, in a file that neither declares nor
imports the `Ticker` container another file defines, therefore emitted
`register(Ticker.fire)` as a confident USES edge to that container's method.
That is exactly the wrong-edge failure R1-2 exists to prevent, arriving through
the class channel instead of the lexical one.

Fixed with `isOwnerNameShadowedBySomethingElse` — a sibling of
`isNamespaceNameShadowed` with one extra clause. The plain namespace guard could
NOT be reused: a container is often its own local declaration
(`fn make() { const Local = struct {…}; register(Local.go); }`), and reading that
binding as its own shadow suppresses precisely the resolutions this path exists
to make — the #2723 mistake, one channel over. So a scope that binds the name
answers immediately, and the answer is "not shadowed" only when one of that
scope's own bindings IS the def just resolved.

Both halves are pinned and both were verified to fail when the guard is
weakened: the parameter case fails with no guard at all, the local-container
case fails with the plain `isNamespaceNameShadowed`.

R3-2 (Warning; mechanism correct, unreachable today; fragility fixed instead) —
the dispatch exclusion could suppress an unfollowed registration. The bot is
right about the code: sweep 2 synthesizes CALLS only for a registration whose
site carried a `propertyKey`, while the exclusion zeroes the note on ANY inbound
`property-dispatch` CALLS edge. It is not reachable in the current rule set, and
the reason is measured rather than assumed: `@reference.value-ref` is emitted by
exactly three languages — JavaScript (2 rules), TypeScript (2), Zig (3) — every
JS/TS rule also captures `@reference.property-key` (both are object-literal
shapes) and no Zig rule does. A dispatchable registration is therefore always a
JS/TS one, an undispatchable one always a Zig one, and they cannot meet on one
symbol.

Rejected: splitting the edge `reason` into dispatchable / undispatchable. It is
the precise fix, but it is a graph-content change that churns whichever side
keeps the old literal — the Zig, TypeScript and probe suites all pin
'scope-resolution: value-ref' by hand as a drift canary — and it buys nothing
against a case no rule can produce.

What was actually wrong is that the exclusion's soundness rested on a
coincidence recorded nowhere, in files nobody reading `local-backend.ts` would
open. Fixed at both ends: the exclusion site now states the invariant, the three
facts it rests on and the two options for when it breaks; and
`value-ref-dispatchability.test.ts` fails the day it does — a JS/TS rule for a
bare callback argument, a Zig rule that grows a key, or a fourth language
emitting `value-ref` at all. Verified to fire (adding a property key to a Zig
value-ref rule fails the Zig case), and its rule splitter has its own guard test
so the suite cannot pass vacuously. `ZIG_SCOPE_QUERY` is exported for that test
only.

R3-3 (Nit, valid) — a test comment claimed the wrong epistemic result. The
`declines a qualified reference whose receiver cannot be resolved` case said the
shortfall shows up as `lower-bound`. It does not: with no edge there is no
evidence and the target stays `exact`. The pass docstring was corrected in round
2 and this comment was missed. It now says the decline costs the reference AND
the hedge, and why that is still the right trade.

Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,632 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission PASS with no baseline edited; callable-value-flow failed once on
its TIMING budget (2.006 > 1.9) with a byte-identical fingerprint, then passed
twice at 1.788 / 1.813 — machine load, not a regression.

* fix(review): let a file's own `@import` outrank the workspace-wide class fallback

Found by running the local preflight harness before pushing rather than after.
One of its five findings is a real hole in R2-2; the rest are documentation that
overclaimed.

R4-1 (valid, REPRODUCED) — the container channel preempted the file's own
`@import`. R2-2 added the module channel as a FALLBACK after
`findClassBindingInScope`, and that order is wrong: `findClassBindingInScope`
does not stop at the scope chain. When its `isClassLike` walk misses — and a
namespace handle binds a Module, so it always misses — it falls back to
`scopes.qualifiedNames`, a workspace-wide index, and answers with the unique def
of that name anywhere in the repo. A container named `dom_utils` in a file
`Element.zig` never imports therefore captured
`bridge.accessor(dom_utils.compare, …)`, binding
`Method:src/webapi/decoy.zig:dom_utils.compare#2` while the module channel that
would have answered correctly was never reached. R3-1's shadow guard cannot
catch it: the import binds at MODULE scope, which that guard treats as the floor.

Fixed by trying the module channel FIRST. An `@import` written in this file is
the strongest available statement about what the name means here and outranks a
global uniqueness guess; when the handle is not an import of this file the
channel answers nothing and the container path runs exactly as before. Pinned by
`decoy.zig` and a strengthened assertion on the existing module-owner test,
which fails on the old order.

R4-2 — the cause documentation named shapes nothing captures. `tools.ts`
illustrated `callableValueReferences` with "a callback argument", "a stored
function pointer" and `qsort(xs, n, sz, compareItems)`. Only Zig captures a call
argument or a const initialiser; JS/TS capture only object-literal property
values, and C has no value-ref rule, so the `qsort` example is counted in no
language. Both cause blocks now name the captured shapes and say that a bare
JS/TS callback argument is not among them, so a 0 does not rule it out.

R4-3 — the same block exempted itself from the re-index caveat this PR proves it
needs. "Read from the graph, so it needs no index-time metadata" is true of the
probe and false of the edges: an index built before a language emitted these
captures has none and reports 0 — the warm-cache failure R2-1 bumped
SCHEMA_BUMP for. It now says to re-analyze before reading a 0 as measured.

R4-4 — the dispatchability canary was narrower than its own promise. Its header
claimed it fails on "a fourth language emitting value-ref at all"; it reads
`languages/<dir>/query.ts`, so Vue — which owns no query and borrows
`emitTsScopeCaptures` / `emitJsScopeCaptures` — emits value-refs while the
assertion lists three languages, and a capture synthesized in code is invisible
to it. The case now asserts on query OWNERS, which is what it checks and is
sound because a delegating language inherits the rules it borrows; the header
states the synthesized-capture gap instead of letting a green tick imply it away.

R4-5 — the BARE docstring described a lexical walk that is not one.
`findCallableBindingInScope` applies the callable predicate WHILE walking, so a
nearer parameter or local is stepped over: the defect R3-1 fixed on the container
channel, unguarded here, pre-existing since #2437 and reachable in JS/TS. Out of
scope for #3399, so behaviour is unchanged and the sentence now says what the
walk does rather than implying a guarantee it does not give.

Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,632 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.

* fix(review): honour the hub opt-in for value references, so `hub.fn` means one thing

Review round 5 (`gitnexus-check` bot on `1c7c05ff`). One finding, valid and
reproduced before fixing.

`findNamespaceValueRefTarget` accepted only `ref.origin === 'local'`. R2-2
recorded that as deliberate — "the `namespaceExportsIncludeImportedNames` hub
opt-in is a provider decision this language-neutral pass does not make" — and
that reasoning was wrong twice over. Zig sets the flag
(`languages/zig/scope-resolver.ts:41`, measured on ghostty and tigerbeetle before
it landed), and the pass does have the provider in scope: `runScopeResolution`
takes one and already forwards several of its hooks to other passes.

The result was the exact asymmetry R2-2 argued against in its own first
paragraph. A Zig hub declares nothing — every name it publishes it imported — so
requiring a local declaration declines every member reached through one:

    // hub.zig
    pub const scale = @import("dom_utils.zig").scale;

    // Element.zig
    const hub = @import("hub.zig");
    pub fn callsThroughTheHub(v: u8) u8 { return hub.scale(v); }   // resolved
    pub const scaled = bridge.accessor(hub.scale, null, .{});      // declined

One name meaning two different things depending on whether a `(` follows it.

Fixed by forwarding `provider.namespaceExportsIncludeImportedNames` into the pass
and consulting the published channel when it is set — the same question
`receiver-bound-calls` Case 1 asks, through the same `lookupBindingsAt` read
`findExportedDefIncludingImportedNames` performs for the CALL form. The pass
still names no language; it asks the provider, which is the sanctioned hook.

Precedence is unchanged where it mattered: a locally declared member still wins
over a republished one, two distinct defs under one name still resolve nothing,
and `CALL_TARGET_TYPES` still gates the answer — pinned by `hub.DEFAULT_NS`, a
re-exported CONSTANT, which stays unregistered. Languages that do not opt in are
unaffected: the parameter defaults to `false`, and passing `false` was verified
to fail the new hub test, so it is load-bearing. The fixture republishes a member
(`scale`) that nothing else uses, so the hub assertion cannot be satisfied by an
edge another case emitted.

Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,634 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.

Also recorded in DECISIONS.md: `/autofix` returned "No successful autofix run"
because the `PR Autofix` run on `1c7c05ff` failed at `actions/upload-artifact`
with a 403 from GitHub's artifact storage, not because of anything in this
branch. This push triggers a fresh run.

* fix(review): inspect the module scope in the owner-shadow guard, and check the TSX suffix

The `gitnexus-check` pass on `bd6e577e` carried three findings of its own, and I
answered the later `1c7c05ff` pass without noticing them. Recording that as a
process failure too: bot reviews are per-head, and a later pass does not
necessarily repeat an earlier one's findings.

R6-1 (Error, valid, REPRODUCED) — the shadow guard stopped one rung short.
`isOwnerNameShadowedBySomethingElse` returned `false` on reaching the module
scope, justified as "a container declared there IS the binding, and the caller
already resolved it". That holds when the owner came from the scope chain and
fails when it came from the workspace-wide qualified-name fallback:

    // Gauge.zig — never imported by Element.zig
    const Gauge = @This();
    pub fn read(self: *Gauge) u8 { … }

    // Element.zig
    const Gauge = @import("dom_utils.zig").DEFAULT_NS;          // NOT a container
    pub const level = bridge.accessor(Gauge.read, null, .{});   // → Gauge.zig's read

`findClassBindingInScope` steps over the module-scope binding because it is not
class-like, answers from `scopes.qualifiedNames`, and the guard waved it through.

Worth recording: the first fixture attempt did NOT reproduce. A local
`const Gauge: u8 = 3;` also claims the workspace qualified name `Gauge`, leaving
two candidates, and the fallback refuses to guess between two — so the shape
defeated itself. Binding the name by IMPORT claims no qualified name, the
fallback stays unique, and it fires. A negative result on the first shape was not
evidence the finding was wrong.

Fixed by inspecting the module scope as the last rung instead of skipping it.
The identity exemption is what makes that safe where `isNamespaceNameShadowed`
cannot do it (#2723: a namespace import writes its own name into the module scope
and would read as its own shadow) — the binding that IS the owner exempts itself,
and only a binding to something else answers `true`. `lookupBindingsAt` is
consulted at that scope and only there, because an imported alias lives in the
finalized channel rather than in `scope.bindings`.

R6-2 — the hub re-export finding, already fixed in `5d8fe9d8`; the same defect
restated on the later head.

R6-3 (valid, fixed) — the dispatchability canary omitted the TSX suffix.
`getTsScopeQuery` analyzes a `.tsx` file with `TYPESCRIPT_SCOPE_QUERY +
TSX_JSX_QUERY_SUFFIX`, and the test read only the base, so a `value-ref` rule
added to the suffix would be emitted in TSX analysis with the canary green. The
suffix is now exported and concatenated into the check; verified load-bearing by
adding an unkeyed `jsx_expression` value-ref rule to it, which fails the
TypeScript case.

Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,635 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.

* perf(bench): gate callable-value reference resolution on linear scaling

`resolveValueRefTarget` (#3399) replaced one lexical walk with four
channels, and the last of them — a qualified receiver resolved through
`findClassBindingInScope` — falls back to `scopes.qualifiedNames`, a
WORKSPACE-WIDE index consulted once per site. Keyed that is O(1); scanned
it is O(files) per site, and a registration table that costs O(sites)
today costs O(sites x files) tomorrow. Nothing in the suite can see that:
the fixtures are single-file, and the corpus it actually matters on is
lightpanda-io/browser, where `bridge.{accessor,function,…}` appears 2,047
times across 257 files.

`bench/value-ref-resolution/measure.mjs` builds two synthetic Zig corpora
of identical shape 4x apart in file count, and times ONLY the per-site
resolution loop — extraction, `reconcileOwnership` and finalize are setup.
`linear_factor` is `(t_large/t_small)/(N_large/N_small)`: measured
1.00-1.09 across runs, with `us_per_site` flat at 1.4-1.5 between the arms.

Correctness comes first, because a timing gate alone is satisfied by a fast
wrong answer: exact site/resolved/declined counts per arm plus an
order-independent sha256 over every (site -> resolved target) pair. The
corpus exercises all four channels — CONTAINER, NAMESPACE, HUB and BARE —
and carries two DECLINE controls per module (a non-callable namespace
member, a non-callable bare argument), so a widened callable gate moves
`declined` instead of hiding inside the timing.

Verified load-bearing rather than assumed: making the qualified lookup scan
a workspace-sized collection leaves the fingerprint IDENTICAL and takes
`linear_factor` to 3.752 against a slack of 1.375 — the regression class
this exists for is exactly the one no correctness gate can see.

Zig is the corpus because it is the only language whose provider sets
`namespaceExportsIncludeImportedNames`, so it is the only one that can
exercise the hub channel at all; the pass itself names no language.

`resolveValueRefTarget` is exported for the bench. Timing
`emitPropertyDispatchCalls` instead would fold the signal into edge
emission, and re-implementing the channel order in the bench would pin the
bench's idea of the function rather than the function.

* feat(zig): index every build package in the repo, not only the root one

Zig already had workspace setup — `loadZigBuildConfig` has parsed
`build.zig.zon` `.path` deps, the root `build.zig`'s named modules and each
build module's own `addImport` table since #1432, threaded through
`ScopeResolver.loadResolutionConfig` exactly as tsconfig is for TypeScript.
What it did not have is the part `tsconfigFor` supplies: per-package scope.

`loadZigBuildConfig` reads `<repoRoot>/build.zig{,.zon}` and nothing else, so
a repo laying its packages out as `packages/<name>/build.zig` — no root build
files at all — got `null`, and EVERY bare `@import("<module>")` in it went
unresolved. Cross-file resolution silently degraded to relative imports.
Measured on a two-package probe before this change: `config = null`,
`@import("core")` from `packages/app/src/main.zig` → `null`.

`loadZigWorkspaceIndex` discovers packages the way `findTsconfigFiles`
discovers configs — one bounded breadth-first walk that skips the hardcoded
ignore set — and `zigPackageFor` is the `tsconfigFor` analogue: the nearest
enclosing package governs a file, deepest-first, with no fall-through to an
outer package. Fall-through is what makes a vendored dependency's
`@import("config")` resolve to the outer repo's `config` module, the same
failure `loadTsconfigIndex` documents for a package declaring no `baseUrl`.

`resolveZigImportInternal` is UNCHANGED — impact analysis puts it at HIGH risk
with 6 dependents, and it does not need to move: it is handed one package's
config, and which config it receives is the only difference. Its 48 existing
tests pass untouched. A nested package's paths are rebased to repo-relative at
load time (`.path = "../core"` → `packages/core`, which the root-relative
reading rejects outright as an escape); the ROOT package keeps its raw
spelling, which is what `parseZigBuildZon` promises and its tests pin, so a
single-package repo is byte-identical.

`loadImportConfigs` — which runs unconditionally for every repo, Zig or not —
keeps calling the root-only loader, so no non-Zig repo pays for the walk. That
is the split TypeScript already has between the cheap `loadTsconfigPaths` and
the repo-walking `loadTsconfigIndex`, which `loadResolutionConfig` reaches only
during a language pass.

The `zig-monorepo` fixture carries the discriminating case rather than only the
happy path: `tool` binds the alias `core` to its OWN `src/core.zig`, so a
repo-wide flattened module map — the shape a workspace index invites — would
point `measure` at `packages/core`, a confident edge into a package `tool` does
not depend on. That is worse than the unresolved import this fixes, and only
per-package scoping keeps them apart.

Tests: 7 unit cases pinning the index (including `loadZigBuildConfig` answering
`null` on the same fixture, side by side) and 3 integration cases pinning the
EDGES — cross-package CALLS from the module root AND from a non-root file, and
`tool` not crossing over. Verified load-bearing: restricting the index to the
root package fails all three.

Second consumer caught by the integration test and fixed with it:
`populateZigWorkspaceStaticGating` reads the same `resolutionConfig`, per file
because two files of that pass can belong to different packages.

Gates: build clean, `tsc --noEmit` clean, prettier clean, eslint 0 errors.
`test/integration/resolvers` 3,768 passed / 4 skipped, `test/unit/scope-resolution`
2,015 passed. Bench `--check` with NO baseline edited: `receiver-resolution`
(whose corpus is `test/fixtures/lang-resolution`, where the fixture lands),
`zig-cross-file-resolution`, `value-ref-resolution`, `scope-capture` (15
languages), `import-target`, `callable-value-flow`, `scope-emission`.

* perf(zig): dequeue the package walk by head index, not `shift()`

`findZigPackageDirs` pushes children while it drains the queue, which keeps
the array in a mode where `Array.prototype.shift()` memmoves the whole
remainder rather than taking V8's left-trimming fast path — so the walk is
quadratic in the frontier, bounded only by `ZIG_SCAN_MAX_DIRS` (20,000).

Measured at that bound rather than estimated from the move count: 53 ms at
fan-out 4 and 81 ms at fan-out 20, against 0.8 ms with a head index — 66-106x,
and paid before a single config is read.

FIFO order is unchanged, and the package ordering never depended on the walk
anyway: `loadZigWorkspaceIndex` sorts by (dir length desc, localeCompare).
Memory is unchanged too — the entries were already retained by the pushes;
`shift()` only dropped the head.

`findTsconfigFiles` and `loadNodeWorkspacePackages` carry the same walk with
the same bound and are equally affected. Deliberately NOT fixed here: they are
pre-existing, outside this PR's diff, and reach TypeScript/Node import
resolution, which nothing in this change set covers. Filed separately.

Reported by gitnexus-check on aea0ab06.

* chore: drop DECISIONS.md from the branch

Review feedback: the working log does not belong in the repository. The
reasoning it carried that is still load-bearing lives in the code comments
and in the PR description.

* perf(bench): gate value-ref resolution on scaling alone, not wall clock

A millisecond ceiling measures the runner. This repo has been bitten by that
twice already — bench/callable-value-flow's widening_overhead failed at 2.07
and 1.975 against a 1.9 budget on a shared runner while the code was correct,
both times on a sub-11ms measurement — so the arm is dropped rather than
loosened. `ms_budget` is gone from both arms and from `--check`; `min_ms`
and `us_per_site` are still reported, and nothing compares them to anything.

The remaining timing gate is the ratio, reshaped after
bench/parse-dispatch-rounds/baselines.json: `_what` / `_triage` notes, a
`_measured` block recording samples for context, and min-of-15 reps instead
of 7 (bench/import-target measured N=5 tripping its own budget about one run
in twenty, N=15 holding).

Budget 1.6 — 1.40x the measured maximum over 12 runs (0.907 .. 1.146), the
~1.5x headroom its siblings use on ratios.

Re-verified rather than re-quoted, and the previous note was wrong: replacing
QualifiedNameIndex.get with a full scan moves linear_factor from ~1.0 to 2.07,
not the ~3.7 recorded. Only 1 of the 17 value-ref sites per module reaches
that workspace-wide fallback. 1.6 sits clear of both ends. The note also
records the trap the first attempt fell into: patching gitnexus-shared/src
changes nothing, because the bench resolves the built package.

* fix(review): settle namespace precedence on the name, before the type gate

findNamespaceValueRefTarget's local lookup applied CALL_TARGET_TYPES while
selecting, so a target module declaring a NON-callable under the name answered
nothing and fell through to the published channel — binding a re-exported
callable under a name the module's own declaration owns. findExportedDef does
not do that: it returns any local def and lets its caller's type gate reject
it, so findExportedDefIncludingImportedNames never reaches the imported names
for a name the file declares. `x.f` and `x.f()` must not disagree about which
module owns f.

Not reachable through valid Zig today (a container cannot declare a name
twice, and Zig is the only provider setting namespaceExportsIncludeImportedNames),
which is why the regression test builds the indexes directly instead of adding
a fixture — there is no valid source to write. Its second case fails with the
guard removed.

Three smaller corrections from the same review:

- ZigBuildZonConfig.pathDeps promised the raw `.path` string; a nested package
  stores the normalized repo-relative value. The interface now documents both
  spellings and why either is safe to hand to normalizeZigDepPath. Comment
  only — no behaviour change.
- The 'single-package repo' compatibility test ran against zig-idioms, which
  declares libs/geo as a path dep and that directory has its own build.zig, so
  the walk finds two packages and the name was a claim rather than a check. It
  now runs against libs/geo itself and asserts the package list is exactly
  ['']; a second test pins the multi-package case, including that a file inside
  libs/geo is governed by that package and not by the root.
- The lower-bound header asserted 'some callers are not traced'.
  callableValueReferenceBoundaries hedges when its probe could not RUN and says
  whether the symbol is registered is unknown, so the header claimed an
  omission nothing established. It now says the count may be incomplete and
  names no cause; the per-cause bullets underneath carry that.

* feat(zig): bind a `@This()` alias to the container it names

`@This()` IS the enclosing container, and `const Self = @This();` is how most
Zig files say so. The container is minted under the FILE STEM and the alias
bound nothing class-like: a file-level alias mints no Const at all
(isZigFileThisAlias suppresses it so it cannot shadow the type for `w: *Widget`),
a container-level one mints a Variable every isClassLike walk steps over. So
`Self.member` resolved to nothing — not a wrong edge, no edge, and a caller
list missing it is the false confidence #3399 is about. This was the PR's
declared known limitation.

bindZigThisAliases binds the alias name to its container definition in
indexes.bindingAugmentations — the sanctioned post-finalize channel (I8),
consulted only AFTER a scope's own bindings, so it can never outrank a real
local declaration and can only answer where nothing answered before. Nothing
is replaced or removed. No query, capture or SCHEMA_BUMP change.

It runs inside populateZigRangeBindings, sharing that pass's parsed tree: a
pass of its own would re-parse every Zig file on a cold tree cache. Only
aliases declared directly in a container body or at file level are bound — the
same set collectZigThisAliases recognizes — because a function-local alias
belongs in that function's scope.

Measured cold-index before/after, same command, same build:

  tigerbeetle (246 .zig)  CALLS 17,066 -> 17,140 (+74), USES 437 -> 443,
                          MEMBER_OF 5,131 -> 5,143; every other edge type
                          unchanged, all 24 node-label counts identical
  mach (132 .zig)         CALLS 7,600 -> 7,615 (+15), MEMBER_OF +1, USES flat

The +74 matches a source census of 72 `Alias.member(` call sites in
tigerbeetle files whose alias differs from the stem. Spot-checked end to end:
src/aof.zig:636 writes `try AOF.init(io, output_path)` inside AOFType, and the
edge AOFType.merge -> AOFType.init#2 is present after and absent before. Node
totals move only through the derived layers (Community 524->527, Process
774->763) — no source symbol added or removed.

Scale of what was dropped: 73 of ghostty's 185 `@This()` files, 93 of
tigerbeetle's 94 and 8 of mach's 42 spell the alias differently from the stem,
carrying 302 `Alias.member` references between them, 96 of those calls.

Fixture zig-idioms/src/webapi/Widget.zig exercises both paths — a file-level
`Self` and a container-level `Me` in Metrics — through a call and a
registration. Five cases; three fail with the binding disabled, and the two
that do not are the controls: Element.zig's stem-spelled alias must keep
resolving byte-identically, and the alias name must not become resolvable from
another file.

* docs(review): stop claiming what the alias pass does not do

Two comments asserted things the adjacent code does not support.

The alias pass's call site said it ran before the payload walk "so a subject
spelled through the alias resolves here too". It does not: the payload walk
types subjects through findReceiverTypeBinding, which reads typeBindings plus
the namespace/workspace type channels and never bindingAugmentations, where
bindZigThisAliases writes. Measured on `for (Self.items) |it|` — `it` is bound
neither before nor after. The pass is in that loop for the tree and nothing
else, which is now what the comment says.

Nor is that a gap to close by also writing a typeBinding: no container name has
one, the file stem included, so a payload subject written `Type.member`
resolves for no spelling at all. Giving the alias an entry would make it behave
unlike the container it names. Recorded at the call site so the next reader
does not re-derive it.

The module-shadow test said Element.zig declares `const Gauge: u8 = 3;`. It
binds `Gauge` by IMPORT, and the distinction is the point of the fixture: a
local declaration would also claim the workspace qualified name `Gauge`,
leaving two candidates, and the fallback refuses to guess between two — so the
case the test exists for would never be reached. The comment now says which
binding it is and why the other shape would be self-defeating.

Comment-only: node and edge counts are byte-identical across a full re-index
(52,083 / 165,261 both sides).

* docs(zig): stop describing loadZigBuildConfig as root-only in the present tense

It takes a `packageDir` since this branch and reads
`path.join(repoRoot, packageDir, name)`; `loadZigWorkspaceIndex` is what
supplies it one, per package. Two comments still described the historical
root-only invocation as the function's current behaviour:

- the monorepo integration-test header, flagged by review;
- loadZigWorkspaceIndex's own docstring — the same sentence, in the function
  that calls it WITH a packageDir a few lines below, so fixing only the test
  copy would have left the worse of the two.

Both now attribute root-only reading to the CALL (no `packageDir`), which is
what the argument actually rests on: a monorepo has no root build files, so
that call answers null and every bare @import goes unresolved. The trailing
note about `loadImportConfigs` is reworded the same way — it calls the loader
for the root package alone; the loader is not root-bound.

Comment-only: node and edge counts byte-identical across a full re-index
(52,083 / 165,261 both sides), and detect-changes reports the two hunks
overlap no indexed symbol.

* fix(zig): a written namespace handle owns its own decline

`findNamespaceValueRefTarget` returning `undefined` conflated two different
answers: "no namespace import named this receiver" and "the module this file
named does not expose that member as a callable". Only the first should fall
through to the container channel. The second did too, and
`findClassBindingInScope`'s miss path answers from the WORKSPACE-wide
qualified-name index — so a same-named container in a file this one never
imported supplied the member the written module does not have.

The owner-shadow guard does not stop it, which is the part that is not obvious:
a plain `const utils = @import("utils.zig");` records a namespace IMPORT EDGE,
not a module-scope binding, so the guard finds nothing bound under the name and
reads the container as unshadowed. It catches `const Gauge = @import(x).MEMBER`
(a real binding) and misses the handle form.

Reproduced before fixing, not argued: `decoy.zig`'s `dom_utils` struct gains
`onlyOnDecoy`, a callable `dom_utils.zig` does not have, and `Element.zig`
registers `dom_utils.onlyOnDecoy`. That minted `JsApi -> onlyOnDecoy` — a
confident USES edge into a file `Element.zig` never imports, the wrong-edge
failure this PR exists to avoid, arriving through the container channel after
the namespace channel said no.

The channel now returns 'owned' for every outcome reached once the receiver is
established as this file's unshadowed namespace handle, and the caller declines
on it. A locally shadowed handle still falls through, because there the name
does not mean the import at that site and the container channel's guard is the
right decider.

Bench fingerprint and counts unchanged; no baseline edited.

* fix(zig): reject an absolute nested-package `.path` before rebasing it

The nested-package branch prefixes the package directory and THEN normalizes,
so an absolute `.path` stops looking absolute on the way: `packages/app/` plus
`/src` is `packages/app//src`, which is relative by inspection.
`normalizeZigDepPath` drops the empty segment and the dep lands on
`packages/app/src` — a directory that really exists — so a dependency pointing
outside the repository is fabricated into an in-repo resolution. The root
package was never affected: its prefix is empty, so the value reached the check
as written.

`isAbsoluteZigDepPath` is now asked of the value AS WRITTEN, before any
prefixing, and `normalizeZigDepPath` asks the same helper so the two cannot
drift. `..` is deliberately not handled there: `../core` escapes the package
but not the repo, and rebasing it is what the branch exists to do —
`normalizeZigDepPath` still rejects what escapes the ROOT afterwards.

Note for the reviewer: `path.posix.join(pkg, depPath)` does NOT fix this.
`join('packages/app/', '/dep')` is `packages/app/dep` — it strips the leading
slash too, producing the same fabricated path without rejecting anything.
Measured before writing the fix.

Regression test pins it through the real loader: `packages/app` declares
`.escapes = .{ .path = "/src" }`, and the test fails without the guard.

`impact normalizeZigDepPath` is HIGH (14 impacted, 4 direct, exact). The edit
is behaviour-preserving for that function — the same two conditions moved into
a named helper it calls — and its existing absolute-path suite, POSIX, Windows
drive and UNC spellings included, passes unchanged.

* docs(mcp): stop defining lower-bound as proof that callers were missed

The CLI header was corrected in round 8; the MCP tool contract still made the
assertion the CLI stopped making. `context` said lower-bound "means callers
exist that this view provably does not list" and `impact` said "the walk
provably missed callers" — but `callableValueReferenceBoundaries` also
publishes lower-bound when its probe could not RUN, and says in its own note
that whether the symbol is registered is unknown. A client following the
contract would read an unanswered question as evidence of an omission.

Both now define it as a FLOOR with two possible causes — the walk provably
missed callers, or a probe that would have established completeness could not
run — and point at `boundaries` for which. That keeps the common case exactly
as strong as it was; it only stops the contract asserting the one case it
cannot support. The `causes.callableValueReferences` bullet already documented
the probe-failure branch, so the headline was contradicting the body.

`local-backend.ts` quotes that definition to justify hedging; the quote is
updated to name which half it relies on.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-09-09 20:12:17 +00:00
mengkaka
4154b63131
feat(indexing): add Objective-C semantic indexing support (#3179)
* docs: add Objective-C fork provider notes

* feat(objective-c): add deterministic provider and grammar

* feat(objective-c): finalize provider MVP

* fix(objective-c): harden provider integration

* fix(objective-c): normalize bare macro markers

* docs(objective-c): integrate provider documentation

* fix(objective-c): harden resolution and header classification

* fix(objective-c): complete provider follow-ups

* fix: address Objective-C review follow-ups

* chore: format Objective-C grammar sources

* fix(objective-c): harden review follow-ups

* Address PR review feedback (#3179)

Keep Objective-C chunking and macro recovery aligned with the grammar, and stop Community MEMBER_OF edges from leaking into symbol context.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address follow-up review on ObjC chunking and language fallback.

Keep preprocessor directive text from changing file-scope brace depth, group real ivar nodes, skip header modifiers, and restore Rakefile/Gemfile detection through getLanguageFromFilename.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Parse Objective-C headers with the objc grammar in embeddings.

ensureAndParse and structural extraction now use the same content classifier as ingest, including method snippets from .h files, so Protocol/Category/Class chunks are not re-parsed as C++.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3179)

Keep file-scope macro elision off C line splices and @interface/@protocol/@implementation bodies, and attach ivar attributes to the following instance variable when chunking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(bench): rebaseline Objective-C CSV emit

* feat(objective-c): add workspace resolution and linear emit benches

Plain .h files are classified as C++, so the ObjC pass could not
resolve #import of those headers. Load a C/C#-style workspace once
per pass, and keep protocol-candidate USES linear.

Refs #3179

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3179)

- Compare LadybugDB labels() as a scalar when excluding Community MEMBER_OF edges.
- Walk superclass members, skip file-static C sibling defs, and ignore comments in ObjC header/macro scans.

Note: pre-existing failure in objective-c-provider integration (worker-pool ready timeout) not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3179)

Emit Objective-C declaration captures so compilation-unit siblings can share
header/implementation bindings, and keep class vs protocol visibility groups
distinct.

Note: pre-existing failure in worker-pool startup (GITNEXUS_WORKER_READY_TIMEOUT_MS) not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3179)

Emit every comma-separated property/ivar declarator, and count @interface
after a multiline block comment closes so in-declaration macros stay intact.

Note: pre-existing failure in worker-pool startup (GITNEXUS_WORKER_READY_TIMEOUT_MS) not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: ximengkai <ximengkai@soyoung.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-09 09:40:21 +00:00
Gergő Magyar
96132bd13a
perf(scope-resolution): stop re-scanning the ParsedFile store once per language (#3211)
* perf(scope-resolution): stop re-scanning the ParsedFile store once per language

Scope resolution calls `loadParsedFilesForPaths` once per language, and every
call walks every shard in the store. The skip decision needs the envelope's
path listing, and that listing is only trustworthy after the payload digest
has been checked -- so a pass that wants 50 Python files still opens and
SHA-256s all 413 shards / 301MB of a TypeScript-dominated store to prove it
can skip them. A pass wanting a SINGLE file costs 335ms. The cost scales with
language count, not with the files that language has, so a polyglot repo pays
it worst.

`tryLoadV8Cache` now returns the listing it already parsed for that skip
decision, and the store memoizes it per run. Later passes skip on the
memoized listing without reopening the file.

Measured on a 2234-file, 3-language repo, min-of-5:

  before   python 411ms   typescript 2872ms   javascript 507ms  = 3834ms
  after    python 415ms   typescript 2805ms   javascript 248ms  = 3484ms

-350ms here, roughly -250ms per additional language elsewhere. The first pass
is unchanged by construction -- it is what populates the memo. End-to-end the
graph is byte-identical: 51,288 nodes / 163,094 edges / 2106 clusters /
759 flows on a true incremental run.

Keyed on size+mtime as well as name. Shard names are content-addressed, so a
name collision across different content should be impossible, but that
invariant lives in the parse-cache keying rather than here and one stat per
shard is a few ms against the hundreds this saves. The memo holds one store
directory at a time, so a new repo in a long-lived MCP process drops the
previous set instead of accumulating.

The failure mode a listing memo introduces is a FALSE SKIP: a pass concludes a
shard holds nothing it wants and those files silently never reach the graph --
an exit-0 wrong answer, not a crash. The new test walks four passes with
disjoint wants over one store, plus a shard written after the memo is warm;
it fails when the skip is forced.

Also records the full scopeResolution breakdown in bench/. The headline is
that `emit` is 7161ms of the 14.7s phase and ~21% of the edit loop, spread
across a fan of passes with no hot inner loop -- so the win there is not
running them for unchanged files, which is a design rather than a patch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3211)

Strengthen the shard-listing memo test so a later miss asserts fs.open and
v8.deserialize never run for the skipped shard. Key-set checks alone still
passed if the memo never skipped.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-08 13:48:30 +01:00
Gergő Magyar
0d1aed942f
docs(bench): close FTS as an optimization target with measured evidence (#3209)
PR #3208 landed two claims that further measurement disproved.

The narrowing is not inconsistent. A true incremental leaf edit rebuilds
exactly 8 of the 20 configured indexes -- the tables the writeback DMLs --
and the 20-index run I compared it against was a forced full rebuild
(runner-identity trap). Per-index costs for those 8 sum to 7769ms against
the 7525ms measured in-analyze. There is nothing to fix in `touchedFts`.

The 845ms was not `import('./platform/capabilities.js')`. The CLI already
imports that module statically; a cached dynamic import measures 0.035ms.
The `await` is the first yield after the native FTS build and absorbs the
libuv work still queued behind it. Recorded as a fourth measurement trap,
since it invalidates any mark placed on an await that follows native work.

What replaces them is a floor, established by probing a copy of the corpus
index directly:

  - narrowing further: nothing left, the 8 tables are exactly the DML'd set
  - concurrent builds: hard error, one write transaction at a time
  - connection thread count: flat at 4/8/16/24 (7298/7133/7109/7345ms
    min-of-3), though the default burns ~60% more CPU for it
  - dropping `content` from File: 3541ms -> 241ms, but that deletes
    full-file keyword search (#2317/#2323); capping is a bad trade because
    the size distribution is flat

The one lever left is overlap: the build runs on a libuv thread and hides
behind main-thread JS (3337ms for the index plus a 3000ms JS burn, against
~6859ms serial). File rows are `{ name, filePath }` from `processStructure`
with content lazy-read at CSV time, so they are known before parsing. What
blocks it is that the DB is closed for the whole pipeline and that an early
write moves `liveIndexMutationStarted` ahead of it.

The `ponytail:` comment in `createSearchFTSIndexes` invited exactly the fix
that cannot work -- every caller has already dropped the indexes it passes,
so a presence gate would never fire. Replaced with the measured reason.

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:22:24 +01:00
Gergő Magyar
eba42d2994
docs(bench): correct the edit-loop numbers and add the FTS per-index breakdown (#3208)
The numbers this file shipped with were full rebuilds labelled as incremental.
Rebuilding or re-copying `dist` between runs changes the analyzer runner
identity, and the tool then forces a full rebuild — every measurement taken
that way is a full run wearing an incremental label. That is the third
"silently fall back to full work" guard in this pipeline, after the non-git
corpus and the escalation gate, so the method section now says to read the
banner on every run.

Corrected, measured on a leaf file with a stable runner identity so the
incremental path is genuinely taken: 31.7s, not 36.5s. The graph write is a
3,980-node subgraph rather than the full 51,288, and parse is ~9% of the loop.

Adds the per-index FTS breakdown, which is the actionable finding: 10.2s across
20 indexes, of which File.file_fts alone is 3.5s because File nodes carry file
content and that index re-tokenizes ~30MB of source to reflect one changed row.
Also records that a two-importer leaf edit rebuilt 20 indexes while another run
rebuilt 8 — `touchedFts` narrowing is at least inconsistent, and it has a
withdrawal path when the index catalog cannot be read. That needs pinning
before anyone optimizes against the narrowed set.

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 17:35:31 +01:00
Gergő Magyar
f48bf81256
perf(parse): tighten the dispatch-round memory bound and unclamp the worker-pool override (#3200)
* docs(parse): record why dispatchGroups is a required interface member

Review finding #10 argued dispatchGroups should be optional to match
`getQuarantinedPaths?` / `getStats?`. Those are compatibility accommodation
for WorkerPool shapes that predate them, not a convention for new members;
optional here would force a `?.` plus an unreachable fallback at the single
production call site. Documenting the decision so the next reader does not
re-litigate it from the neighbouring optional markers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit addaab647377f3c4553f752fa3ca1388bcb9ca81)

* refactor(parse): simplify round accounting and dispatch setup

Simplification pass over the dispatch-rounds change. Behavior preserved:
identical graph on a full analyze (51,286 nodes / 163,092 edges).

- Drop `roundMissBytes`. `roundBufferedBytes` counts the same bytes plus the
  cache hits, so it is always the greater of the two and the first disjunct of
  the close condition could never fire on its own. One counter, one reset, one
  check.
- Measure round bytes with `Buffer.byteLength(content, 'utf8')` instead of
  `String.length`. UTF-16 code units undercount non-ASCII source by up to 3x,
  so the cap meant to bound main-thread retention was letting a CJK-heavy repo
  hold well past its nominal budget. Matches `estimateItemBytes` in the pool.
- Reset the durable ParsedFile directories for a round's chunks concurrently.
  Each targets its own chunk-hash directory, and running them serially put N
  round trips of fs work on the critical path the round exists to shorten.
  The try/catch stays inside the mapped callback, so one failure still
  degrades that chunk alone.
- Skip the quarantine filter entirely when nothing is quarantined, which is
  every run without a worker death. It was an identity copy of every group.
- `dispatchChunkParseRound` takes `DispatchGroup<...>` rather than re-declaring
  that shape inline; the type was already imported and used in its body.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 527d5b6e0ca8ae7bbc6a414c5ac7e27fd85e9995)

* refactor(parse): count round misses with the same idiom startRound uses

`drainRound` hand-rolled a reduce to count 'miss' entries while `startRound`,
one function above, filters the same predicate over the same union. Same
integer, one idiom.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 7eaa193b0cb5fa515844f36ae1401d6fb2fed7b8)

* fix(parse): honor GITNEXUS_WORKER_POOL_SIZE above the auto sizing cap

The auto pool size is bounded by source bytes so a tiny repo does not spawn a
full idle pool. That bound was also clamping the operator's env override,
because the env value is read inside `resolveAutoPoolSize()` and the result
went through `Math.min(..., workProportionalCap)`.

`DEFAULT_POOL_SIZE_CAP`'s own comment offers `GITNEXUS_WORKER_POOL_SIZE` and
`--workers <N>` as equivalent escape hatches for operators on bigger machines.
They were not. Measured on a 30MB corpus, where the byte-derived cap is 16:

  --workers 24                  -> pool: 24/24 active
  GITNEXUS_WORKER_POOL_SIZE=24  -> pool: 16/16 active   (silently ignored)

Both are deliberate operator input, so both now bypass the work-proportional
cap, which goes back to bounding only the auto default. After the fix, on the
same corpus, with identical graph output (51,286 nodes / 163,092 edges):

  GITNEXUS_WORKER_POOL_SIZE=24  -> pool: 24/24 active
  GITNEXUS_WORKER_POOL_SIZE=4   -> pool: 4/4 active
  unset                         -> pool: 16/16 active

Verified by hand against the pool's own throughput log; not covered by an
automated regression test, since the pool size is only observable through
that log line and not through the progress stream a test can read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 17ed08608c878079b2927da25cfd39c1608a02a2)

* fix(parse): bound the durable-reset fan-out and pin the pool-size override

Review follow-ups on #3200.

The round's durable ParsedFile directory resets went out as one unbounded
`Promise.all` — one recursive rm + mkdir per miss chunk, all at once. A round
can hold hundreds of small packs, and those resets compete for descriptors with
the chunk prefetch this loop already has in flight. `readFileContents` degrades
a losing read SILENTLY by documented contract, so a dropped file would vanish
from the chunk, from the graph, and from the chunk hash — shipping a narrowed
index with exit 0. Now routed through `mapConcurrent` at the same width the file
reads use, which keeps the pipelining win and caps in-flight descriptors.

An operator's pool size is now also bounded by the number of files there are to
parse, so `GITNEXUS_WORKER_POOL_SIZE=100000` on a five-file repo cannot become
the literal thread count. This applies to `--workers` and the env var alike, so
the parity the previous commit established is intact. It does NOT shrink an
incremental re-analyze: `totalParseable` counts every parseable file in the
scan, not the changed ones.

Adds the regression test a reviewer asked for. The existing coverage
(`worker-pool-resilience` calling `resolveAutoPoolSize` directly,
`analyze-worker-pool-size` mocking `runFullAnalysis`) never reaches
`runChunkedParseAndResolve`'s `effectivePoolSize`, so both stayed green through
a revert of the fix. The new test drives the real parse phase with a worker
double that writes a per-`threadId` marker, and counts them: verified it fails
on the reverted line with `expected [ 'worker-1' ] to have a length of 3 but
got 1`, and passes on HEAD.

Also corrects the `GITNEXUS_PARSE_ROUND_BYTES` docstring, which still described
the cache-miss counter deleted two commits ago.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(parse): skip caching a chunk with a stale durable generation; warn on over-subscription

Closes the two findings left open by the review of #3200.

When `prepareDurableParsedFileChunk` fails, the previous generation's shards
are still on disk, so a later warm hit would union them with the new ones. The
chunk is now recorded and its parse-cache write skipped -- the same posture
`finalizeWorkerChunk` already takes for a quarantined chunk, and for the same
reason: do not cache what we cannot vouch for. The next run re-dispatches into
a directory it can actually clear. Bounding the reset fan-out removed the
correlated trigger; this closes the individual case.

Pool size over-subscription now warns rather than caps. Silently capping is
precisely what the override exists to prevent, so an operator's number is still
honored -- but an exported GITNEXUS_WORKER_POOL_SIZE applies to every analyze
in a long-lived caller (watch auto-sync, the MCP server), including small
incremental ones, and that is easy to set once and forget. The warning names
the host's usable core count, so it is a hardware fact rather than an invented
threshold. `resolveHostParallelism` is extracted from `resolveAutoPoolSize`
rather than re-deriving the cgroup-aware fallback at the new call site.

Tests: the stale-generation skip is pinned by a new case asserting nothing is
written under any key; verified it fails without the guard with
`expected 1 to be +0`. 60 unit and 49 integration tests pass across the
affected suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(parse): guard dispatch-round cadence with a bench, not a wall-clock budget

Round boundaries are deliberately invisible to graph output — batching that
changed output would be a bug — so nothing in the repo could see the #3196 win
regress. It would have come back as a silent ~1.5x on every cold analyze. Two
earlier attempts to pin it as a unit test failed for that exact reason: one
scraped a logger line the progress stream does not carry, the other asserted
graph content that is identical either way.

Extracts the round-close fold into `createRoundBudget`, so the decision is a
shared unit the bench measures rather than a copy that drifts. The parse loop
is streaming and cannot know chunk sizes up front, so an accumulator is the
honest shape — not a planner.

Four deterministic arms, one ratio, no millisecond gate:
- layout_fingerprint — pack membership. Every cache key derives from it, so
  drift needs a SCHEMA_BUMP, never a lone re-baseline.
- packs / single_file_packs — the FLOOR. `rounds` only asserts something while
  the corpus over-splits (774 packs where the byte budget needs 5). This is
  bench/import-target's lesson, where four heap arms read 0 B and passed every
  ceiling: a ceiling says "not too big", nothing said "still measuring".
- rounds — the regression signal, both directions.
- cjk_rounds vs ascii_rounds — pins UTF-8 byte accounting. The two corpora
  share a UTF-16 length and differ only in encoded size, so String.length
  collapses them to equal. This is the arm no unit test could be.
- pack_scaling_ratio — (t_4n/t_n)/4, min-of-15. A ratio because wall-clock is
  runner-speed-dependent and this repo has the scar: callable-value-flow's ms
  gate failed twice at 2.07 and 1.975 against 1.9 with correct code, on a
  sub-11ms measurement.

Every arm verified to fail before being recorded: close-every-chunk reads 774
rounds, disabling the close reads 1, reverting roundFileBytes to String.length
takes cjk_rounds 8 -> 3, and shrinking the corpus trips the shape floor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(bench): record the analyze phase breakdown and the rejected optimizations

Where analyze time actually goes, measured while landing #3194/#3196/#3200,
plus the two optimizations that looked compelling and were measured away.

The headline is that the parse work is done: a one-file-edit re-analyze is
36.5s, of which parse is 2.8s (8%). scopeResolution is 40% and the unlogged
graph emit + FTS rebuild is 49% — neither is incremental, and the ~18s sits
outside the phase runner so every phase log is blind to it.

Also records the trap that invalidated an earlier measurement: a non-git
corpus never records a schema fingerprint, so every run is a forced rebuild
and any "warm" number taken that way is fiction.

Rejected, with numbers: more workers (16/20/24 land inside run-to-run spread)
and bundling the worker entry (~250ms on a normal filesystem; the 8.6s that
motivated it was a 9p-mount artifact).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:53:31 +01:00
Parafee41
1c1cbf111e
fix(zig): resolve cross-file static gates (#3185)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
Gitleaks / gitleaks (push) Has been cancelled
Publish / Classify release event (push) Has been cancelled
Scorecard / Scorecard analysis (push) Has been cancelled
Trivy Image Scan / Trivy (gitnexus-cli) (push) Has been cancelled
Trivy Image Scan / Trivy (gitnexus-web) (push) Has been cancelled
Publish / Build & Push RC Docker images (push) Has been cancelled
Publish / RC guard (marker + release-PR skip) (push) Has been cancelled
Publish / ci (push) Has been cancelled
Publish / Publish to npm (push) Has been cancelled
* fix(zig): resolve cross-file static gates

* Fix Zig workspace import alias enrichment

* Handle extensionless Zig workspace imports

* Reuse Zig import resolution for static gates

* Document Zig workspace static gating

* Harden Zig workspace reference enrichment

* Benchmark Zig cross-file static gating

* Enforce linear Zig benchmark scaling

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-09-07 07:19:30 +01:00
Garrett Griffin
3c2b14aff5
feat(zig): mark CALLS edges inside comptime-false branches as staticGated (#3161)
* feat(zig): static-gating analysis module + fixture (ported from feat/zig-static-gated-edges-v2)

Squashes c6fe922c, 2f3c8e9e, fab088f4, 9b58af74, aef2ae83, 86b892ef, d5657861,
f3780b3a: file-local comptime bool constants, cross-file flag resolution via the
@import alias map, re-aliased const chains, == / != against known bools,
else / else-if branch awareness. The module is self-contained; the hooks that
call it land in the next commit.

* feat(zig): stamp static-gated call sites through the scope pipeline

Wires the ported gating module into the scope-resolution pipeline that
now emits every Zig CALLS edge (PR #1432), replacing the parse-worker /
call-processor hooks of the original branch, which targeted the legacy
DAG path the merged provider no longer uses.

Data flow, one new fact carried end to end:

  emitZigScopeCaptures     stamps `@reference.static-gated` on a call
                           capture whose anchor lies in a statically dead
                           range (body of `if (CONST_FALSE)`, else of
                           `if (CONST_TRUE)`), via the module's new
                           `collectZigStaticGatedRanges` (line/col ranges,
                           because a Capture keeps no node)
  scope-extractor          marker -> `ReferenceSite.staticGated`
  buildReference           -> `Reference.staticGated`
  references-to-edges,     -> `GraphRelationship.staticGated` on the
  free-call-fallback,         emitted CALLS edge (both emit paths, plus
  edges.ts (tryEmitEdge*)     the generic bridge)
  local-backend impact     -> `staticGated` on impact frontier edges

Same marker idiom as Go's `@reference.callee-position` / `embedded-pointer`:
zero-range, present or absent, so every ungated site's capture set is
byte-identical and no other language changes.

SCHEMA_BUMP 92 -> 93: parse-time captures changed.

Cross-file constants (`if (cfg.FOO)` with `cfg = @import("cfg.zig")`) are
NOT stamped yet: the module resolves them through `lookupBoolsForPath`, but
the capture emitter runs per file in the parse worker with only
`{ path, content }`, so it cannot see the sibling source. The two positive
cross-file cases in zig-static-gating.test.ts are `it.skip` with that
reason; the negative ones pass unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* feat(graph): add staticGated edge property

Adds an optional `staticGated?: boolean` field to `GraphRelationship`
that flags edges originating in code branches known at index time to
be unreachable in production — e.g. `if (CONST_FALSE)` blocks where
the condition reduces to a comptime-known `false`.

Schema + persistence wiring:

- `gitnexus-shared/src/graph/types.ts` — additive optional field on
  `GraphRelationship`; absent edges read identically to live ones.
- `gitnexus/src/core/lbug/schema.ts` — `staticGated BOOLEAN` column
  on the `CodeRelation` REL table.
- `gitnexus/src/core/lbug/csv-generator.ts` — appends a
  `staticGated` column (0/1) to the `relations.csv` written for
  bulk COPY ingest.
- `gitnexus/src/core/lbug/lbug-adapter.ts` — fallback per-row
  `MATCH ... CREATE` insert reads the optional column and threads
  it into the relationship properties.

No language has populated this field yet — the Zig hookup lands in
the next commit. Existing DBs need a re-index for the new column to
appear; existing readers are unchanged because the field is
optional and absent on every other language's edges.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
(cherry picked from commit c8cd5efe27a77a5d1b3f05e3f4e89069b5c6b19e)

* fix(zig): gate bare literal branches; AND the flag over deduplicated free-call sites

PR #3161 review, two findings:

1. `stampZigStaticGating` returned early when the file declared no boolean
   constants, but `collectZigStaticGatedRanges` also folds bare literals, so
   `if (false) { foo(); }` in a constant-free file went unstamped. The early
   return is gone; the range walk runs for every file.

2. `emitFreeCallFallback` deduplicates CALLS edges per (caller, callee) and
   wrote `staticGated` from whichever site it met first, so a callee reached
   from one live site and one dead site was gated or not by traversal order.
   Emission is now deferred to the end of each file's sites and the flag is
   the AND over every site that collapsed into the edge: one live site keeps
   the edge live. The other emit path keys its dedup on the site range and
   was not affected; `collapseByCallerTarget` in the generic bridge would
   have the same shape but no language that sets the marker opts into it.

Fixture + tests: `gated_bare_literal`, `live_and_gated_same_callee` (live
site first) and `gated_then_live_same_callee` (dead site first) in
zig-static-gating.test.ts. All 70 resolver suites (3,603 tests) pass with
the shared emitter change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* test(zig): move the SCHEMA_BUMP pin to 93; rebaseline emit fingerprints for the staticGated column

Three CI failures on f4954963, all consequences of this PR:

- test/unit/incremental-parse-cache.test.ts pins SCHEMA_BUMP so concurrent
  bumps cannot collide; 92 -> 93 for #3161 (parse-time call captures gain
  `@reference.static-gated`), 92 added to the taken list.
- bench/emit-persistence `measure.mjs --check` and `measure-streaming.mjs
  --check`: byte-identity fingerprints drift because every relationship row
  now ends in a `staticGated` cell. Regenerated with the inverse-operation
  evidence recorded under `_rebaselined_3161_static_gated_column` in both
  baseline files: stripping ONLY the new column from the emitted CSVs
  reproduces the prior fingerprints exactly (36 files, 3 rel_* files differ,
  33 byte-identical; 36,000 PDG rows each +2 bytes), so no row moved between
  pair files or reordered. Timing and retention gates passed throughout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* docs(graph): state the CALLS contract on staticGated; say "provably unreachable at compile time"

Review on #3161 (magyargergo): the flag must not redefine what a CALLS
edge means. The field's doc now says so explicitly: CALLS still means
"there is a resolved call site from A to B", never "B is reachable from
A"; `staticGated` is additional, statically provable path-feasibility
metadata, an opt-in analysis layer that no core pass acts on. The edge is
emitted, persisted, traversed and counted exactly as before.

Wording: "unreachable in production" -> "provably unreachable from the
indexed source at compile time" on GraphRelationship, ReferenceSite and
Reference. The index has no production build configuration and should not
claim one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* feat(zig): surface staticGated on impact byDepth items; gate if-expressions and negated/parenthesized conditions

Addresses the tri-review on #3161.

- impact: the depth traversal already selected r.staticGated but dropped it
  when building the byDepth item. Forward it (present only when true) and
  document the field on the impact tool's byDepth contract. Traversal and
  ranking still do not act on it; that stays opt-in for consumers.
- zig-static-gating: walk `if_expression` (`const x = if (c) a() else b();`)
  in addition to `if_statement`. The expression form has no field names and
  no else_clause wrapper, so the arms are located positionally
  (`ifExpressionArms`). Labeled-block arms are covered.
- evalCond: `parenthesized_expression` is transparent, so `!(A and B)` and
  `((FLAG))` fold. Prefix `!` has no unary node in tree-sitter-zig; the
  header now says exactly which shapes fold instead of "simple negation".
- fixture + tests: nine new cases (negation x2, parentheses x2,
  if-expression then/else/labeled-block x5). Cross-file `@import` constants
  remain skipped and now cite the tracking issue #3162.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* refactor(zig): drop the unreachable cross-file gating builders; document the real wiring

gitnexus-check on b77cb4ed: `buildZigImportAliasMap` / `buildZigRawImportAliasMap`
and the per-call ancestor walk (`isCallStaticGated`, `ifBranchDirection`,
`nodesEqual`) had no caller anywhere. They were ported from the legacy
call-processor design; the scope-resolution provider stamps ranges via
`collectZigStaticGatedRanges` instead, and nothing populates the cross-file
seam yet (#3162). Remove them so the module exports only what runs.

The evaluator keeps `importAliases` + `lookupBoolsForPath` (the seam #3162
will fill); the header now says so explicitly and points at the actual
wire-up (`stampZigStaticGating` in languages/zig/captures.ts) instead of the
retired `configs/zig.ts` hook.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015ciQ3MTjpQXkCoNntZR9zG

* refactor(zig): reuse descendantsOfType and tighten static-gated capture matching

Walk if-nodes through tree-sitter instead of a hand-rolled stack, return the original capture array when nothing is gated, and register the marker as a known sub-tag so it cannot be mistaken for an anchor.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style(zig): wrap a long else-clause assignment to satisfy prettier

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-03 19:49:16 +01:00
Navid EMAD
932d937085
feat: add Zig language support (#1432)
* feat: add Zig language support

Adds a Zig LanguageProvider grounded in @tree-sitter-grammars/tree-sitter-zig 1.1.2.
The grammar's published peerOptional `tree-sitter@^0.22.1` is suppressed via an
npm `overrides` entry that aliases the peer to the bundled `tree-sitter@0.21.x`;
load-time smoke testing confirmed ABI compatibility.

v1 capabilities:
  - .zig file detection + Prism syntax mapping
  - Top-level + nested function_declaration as Function/Method
  - struct/enum/union (anonymous in the grammar — owner name resolved from the
    enclosing variable_declaration in class/field/method extractors)
  - container_field as struct/union fields and enum variants
  - top-level const/var as Variable nodes
  - free + member call_expression as @call edges
  - @import("./foo.zig") local-file resolution; std and external packages
    return empty (no ghost edges)
  - pub keyword detection for export checking
  - no heritage hooks (Zig has no inheritance; queries never emit @heritage.*)

Generic extractor changes (backward-compatible):
  - field-extractors/generic.ts and method-extractors/generic.ts: empty
    `bodyNodeTypes` falls back to the type declaration node itself as its own
    body container — needed because Zig's struct_declaration directly contains
    its container_field children. (The `extractOwnerName` hook this commit
    originally introduced now exists upstream; Zig just configures it.)

Out of scope (deferred):
  - usingnamespace, build.zig.zon package graph, comptime/anytype
  - scope-resolution hooks (emitScopeCaptures, interpretImport, …): Zig is
    classified `experimental` and uses the generic fallback resolution path
  - cross-package imports (std, deps)

Tests:
  - new fixture test/fixtures/sample-code/simple.zig
  - Zig describe block in tree-sitter-languages integration test
  - simple.zig added to parsing.test.ts fixture-existence list
  - Zig added to ingestion-utils detection unit test
  - Zig smoke case in parser-loader-abi.test.ts
  - Zig grammar registered in the grammar-literal validation gate

* feat(zig): integrate build.zig.zon resolution + Union label from PR 1096

Ports the additive pieces of grgisme's standalone Zig provider PR
(https://github.com/abhigyanpatwari/GitNexus/pull/1096) onto the rebased
provider:

  - build.zig.zon `.path` dependency resolution for bare-name
    @import("pkg") (parseZigBuildZon / loadZigBuildZon in
    language-config.ts, resolveZigImportInternal in import-resolvers/zig.ts,
    wired through ImportConfigs.zigBuildZon). `.url` deps and
    repo-escaping paths return null cleanly. 13 unit tests.
  - `union(enum)` containers now produce `Union` nodes (not Struct):
    'Union' added to ClassLikeNodeLabel + CLASS_LIKE_LABELS, the Zig query
    tags @definition.union, CONTAINER_TYPE_TO_LABEL maps union_declaration
    to 'Union'. The label was already plumbed graph-wide on main.
  - /^build$/ entry-point pattern (build.zig).
  - zig-basic lang-resolution fixture + resolvers integration test.

Adapted to current main while porting:

  - labelOverride relabels container-nested fns Function → Method
    (mirrors isKotlinClassMethod); the structure phase no longer derives
    Method from the legacy method-extraction path for plain
    @definition.function captures.
  - IMPORTS/CALLS edges require scope-resolution hooks
    (emitScopeCaptures / interpretImport) since the legacy DAG removal;
    Zig does not implement them yet, so the integration test documents
    that with a skipped import-edge case. The resolver itself is wired
    into the resolver factory and becomes live when the hooks land.

Not ported: named-bindings extractor (the legacy namedBindingExtractor
API no longer exists) and the bespoke field extractor (the generic
factory's extractOwnerName / empty-bodyNodeTypes hooks cover Zig).

Co-authored-by: Garrett Griffin-Morales <grgisme@gmail.com>

* feat(zig): scope-resolution hooks — IMPORTS and CALLS edges (Ring 3)

Implements the registry-primary scope-resolution path for Zig, the
prerequisite for cross-file edges since the legacy DAG removal. Adds the
standard per-language stack under languages/zig/:

  - query.ts: scope query (containers as Class scopes, blocks, functions),
    declarations (container anchors placed on the container node itself so
    the def lands in its own Class scope and the name binding auto-hoists
    to the parent — populateClassOwnedMembers needs the class-like def
    among the class scope's ownedDefs), @import statements (#eq?-gated
    builtin), parameter/constructor type bindings, and call/constructor
    reference sites. The grammar is required lazily (optionalDependency).
  - captures.ts: emitZigScopeCaptures — groups query matches, drops the
    plain-variable group for container/import bindings (their dedicated
    rules bind the name), and relabels container-nested fns
    @declaration.function → @declaration.method (labelOverride parity).
  - interpret.ts: namespace-kind imports (const x = @import("…")) and
    type bindings — self-parameter convention marks the receiver, Zig
    sigils (*, ?, [], error unions, const) stripped from type names while
    dotted qualifiers (mod.T) are preserved for Case-3 namespace-prefix
    receiver dispatch.
  - simple-hooks.ts: parameter bindings stay function-local (Go
    rationale), local-over-import merge precedence, bounds-check arity
    (always 'unknown' today — no synthesized arity metadata).
  - scope-resolver.ts: emit-side wiring; build.zig.zon threads through
    loadResolutionConfig into the same resolveZigImportInternal the legacy
    resolver config wraps. fieldFallbackOnMethodLookup off (statically
    typed). Registered in SCOPE_RESOLVERS.

The resolvers integration test un-skips the import-edge case and gains
CALLS assertions: free call (main → helper) and receiver-bound method
dispatch through a namespace-qualified constructor
(var p = pioneer.Pioneer{…}; p.tick() → main → tick).

* fix(zig): Union is class-like + missing-grammar warning (review pass)

Self-review findings on the Zig branch:

  - scope/walkers.ts `isClassLike` and finalize-algorithm's
    CALLABLE_OR_TYPE_LIKE did not include 'Union': a `union(enum)`
    container's methods got no ownerId from populateClassOwnedMembers, so
    method dispatch on union receivers silently dropped. Widened both
    sets; the zig-basic fixture gains a Tag method + a CALLS assertion
    (main → isEnergy) that fails without the widening (verified by
    reverting).
  - optional-grammars.ts now lists tree-sitter-zig with an npm `probe`
    (it is an optionalDependency, not vendored): users with .zig files
    and no prebuild get the standard one-line stderr warning instead of
    a silently degraded index.
  - Deduplicated the container-method predicate: `isZigContainerMethod`
    + ZIG_CONTAINER_TYPES now live once in languages/zig/captures.ts and
    feed both the provider labelOverride and the scope-capture relabel.
  - README language matrices: Zig row now claims Type Annotations,
    Constructor Inference, and Config (build.zig.zon) — all true since
    the scope-resolution hooks landed.

* fix(zig): anchor @declaration.variable to the binding identifier

`(variable_declaration (identifier) @declaration.name)` matched EVERY
identifier child of the node, so `const first = target;` also declared a
phantom local named `target` in the enclosing block. That phantom shadowed
the real function for later references (and starved callable-value-flow
seeds of a target). The `.` anchor pins the pattern to the first named
child — the bound name.

Regression test in resolvers/zig.test.ts pins the capture set.

* feat(zig): callable-value-flow captures + main's per-language conformance gates

Post-rebase catch-up: since this branch forked, main added three "every
registered language must appear here" tests. Each needs a Zig entry:

- callable-value-flow (#2522): Zig now emits `@callable-flow.*` facts via
  `synthesizeCallableFlowCaptures` (ZIG_CALLABLE_CAPTURE_OPTIONS in
  zig/captures.ts). tree-sitter-zig's `call_expression` carries arguments
  as direct children with no wrapper node, which the shared helper could
  not decompose, so this adds a language-neutral `extractCallArguments`
  hook (mirror of `extractFunctionParameters`; `undefined` = shared path).
  Zig joins the provider matrix as 'matrix' with a real assign→copy→
  argument→invoke case.
- external-import-conformance (#2953): `@import("std")` beside a decoy
  `src/std.zig` resolves to nothing; the decoy stays reachable via the
  relative spelling. Zig holds the property (no suffix fallback), so it is
  a case, not a KNOWN_GAPS entry.
- import-target-index-reuse contract (#2909): membership-probe-only
  fixture (minimumScans: 0, same shape as Rust).

* fix(zig): address gitnexus-check review findings

One commit per the bot's list so each item is easy to check off:

- parser-loader: Zig row gains `userSkippable: true`, so
  `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` (=1 or a list naming `zig`) disables it
  at analyze time like swift/dart/kotlin — as `optional-grammars.ts` already
  documented. Covered in parser-loader-skip-optional.test.ts.
- export-detection: `zigExportChecker` stops at the first declaration it
  reaches. A non-`pub` fn inside `pub const T = struct {…}` was reported
  exported because the walk continued up to the wrapper.
- import-resolvers/zig: a `.path = "."` dep normalizes to '' and no longer
  grows a leading slash (`/src/main.zig` could never match).
- language-config: build.zig.zon parsing strips `//` comments string-aware
  (a `//` inside `.url = "https://…"` survives) and matches braces while
  skipping string literals, so a commented-out `.path` cannot declare a dep
  and a `}` in a comment/string cannot truncate the block.
- method-extractors/configs/zig: the leading `self` receiver is excluded
  from `parameters` (Rust parity). Fixing that exposed a worse bug: the
  `parameters` node is a plain child of `function_declaration`, not a
  `parameters:` field, so `childForFieldName('parameters')` was always null
  and every Zig method had no parameters, no receiver and `isStatic: true`.
  One `zigParameterList` helper now feeds all three readers.
- variable-extractors/configs/zig: container (`struct`/`enum`/`union`) and
  `@import` bindings are skipped via the same predicate the scope captures
  use (`isZigContainerOrImportBinding`), instead of the comment merely
  claiming they were.
- tree-sitter-languages.test.ts: the "missing grammar" case now forces the
  absent-binding path through the loader's runtime opt-out on a fresh module
  instead of passing vacuously when the package is installed.

New: test/unit/zig-extractors.test.ts (exports, receiver/parameters,
variable guard); zig-import-resolver.test.ts gains the `.` dep, comment and
brace cases.

The `createFieldExtractor` heads-up needs no change: the added branch is
unreachable for every existing config (none has empty `bodyNodeTypes`).

* fix(zig): address second gitnexus-check review pass

- import-resolvers/zig: `..` above the repository root now returns null
  instead of aliasing a same-named root file (`../bar.zig` from `main.zig`
  is not `bar.zig`); the stale "extension is stripped and re-added" comment
  is corrected to what the code does.
- variable-extractors/configs/zig `extractType`: read the `type:` field only.
  The positional fallback returned the INITIALIZER of `const f = target;`
  as its type and gave up on compound annotations (`*Foo`, `?[]const u8`).
  The comment claiming 1.1.2 has no `type` field on variable_declaration
  was wrong (verified by AST dump) — and it is what led the review to
  suspect the callable-flow `extractAssignment` callback, which was
  already correct for `extern var f: T;`.
- receiver detection: only a FIRST parameter named `self` is the receiver.
  `emitZigScopeCaptures` tags first-position parameters
  (`@type-binding.first-parameter`), `interpretZigTypeBinding` requires the
  tag as well as the name, so `zigReceiverBinding` no longer turns
  `fn f(a: u32, self: T)` into an instance method.
- resolvers/zig.test.ts: both suites `describe.skipIf(!zigAvailable)`
  (Swift/Dart pattern) — the grammar is an optionalDependency.
- tree-sitter-languages.test.ts: the Zig parsing case gates on
  `isLanguageAvailable` instead of a catch-all `return`, so an installed
  grammar that fails to load fails the test; comment no longer calls Dart
  and Swift npm optionalDependencies (they are vendored).
- test/helpers/literal-collectors: `DIR_LANG` gains `zig`, so literals under
  `languages/zig/**` are validated against the Zig grammar alone rather than
  against every grammar.
- walkers.ts `isShapeLike` doc: Union IS included (via isClassLike, wired by
  Zig's union member container); Typedef remains the only deferred one.
- language-config `ZigBuildZonConfig.pathDeps` doc: values are the raw
  `.path` strings; the resolver normalizes.

Tests: zig-import-resolver (+1), zig-extractors (+3).

* fix(zig): address third gitnexus-check review pass

- language-config `parseZigBuildZon`: the `.dependencies = .{` header and
  the per-entry `.<name> = .{` headers are matched only OUTSIDE string
  literals (per-offset string mask + `matchZonHeader`). A `.name` or
  `.description` value spelling `.dependencies = .{ .fake = .{ .path = … } }`
  used to be taken as the block and returned the fake dep instead of the
  real top-level one.
- import-resolvers/zig: an absolute import (`@import("/foo.zig")`) returns
  null. The path walker skipped every empty component, so the leading `/`
  vanished and `/foo.zig` resolved as importer-relative `src/foo.zig` — an
  in-repo edge for an import Zig rejects as outside the module path.
- tree-sitter-languages.test.ts: the Zig parsing case gates on the PACKAGE
  being installed (`createRequire().resolve`, minus a deliberate
  `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` opt-out), not on `isLanguageAvailable`,
  which is false for absent AND for installed-but-broken bindings — so a
  load failure (ABI mismatch, bad export) now fails the test instead of
  skipping it, as the comment already claimed.
- walkers.ts `isShapeLike` doc: `Union` sits in `isClassLike` because that
  is the label set the ownership walkers consult, not because unions
  inherit — Zig has no inheritance and no heritage hooks. The previous
  wording ("inheritance-capable owner") said otherwise.
- Not re-fixed (already addressed in the second pass, findings carried
  over): "receiver = any parameter named self" — `interpretZigTypeBinding`
  only sources a first-position parameter as `self`; `zigReceiverBinding`'s
  doc now states that invariant. "`DIR_LANG` has no zig entry" — it does.
  Extended one level out: `BASENAME_LANGS` (`zig.ts`) and `PREFIX_LANGS`
  (`ZIG_`) map to the Zig grammar too, so extractor configs and the
  export-detection set are validated against Zig alone. That immediately
  caught a dead `childForFieldName('parameters')` in
  method-extractors/configs/zig `zigParameterList` (there is no such field;
  the named-child lookup was already the one doing the work) — removed.

Tests: zig-import-resolver (+2: absolute path, header inside a string);
both fail on the previous code.

* fix(zig): address fourth gitnexus-check review pass

- language-config: `parseZigBuildZon` only accepts a `.path` that is a
  DIRECT field of a dependency entry. Nested blocks inside the entry body
  are blanked (string-aware, offsets preserved) before the `.path` regex
  runs, and a match starting inside a string literal is rejected, so
  `.dep = .{ .url = "…", .meta = .{ .path = "x" } }` no longer becomes a
  path dep. Regression test in zig-import-resolver.test.ts (fails on the
  previous code).
- tree-sitter-languages test: the "grammar is absent" case now drives the
  loader's real `source.load()` catch branch — `node:module` is stood in
  with a `createRequire` whose require throws MODULE_NOT_FOUND for
  `@tree-sitter-grammars/tree-sitter-zig` and delegates everything else —
  instead of the `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` opt-out, which has its
  own test. It also asserts the opt-out flag is NOT set and that other
  grammars still load (non-fatal optional failure).

Not re-fixed:
- "Return type is read from the wrong tree-sitter field": tree-sitter-zig
  1.1.2 has NO `return_type` field on function_declaration — the type after
  `)` is the `type` field (AST dump: `builtin_type "i32" field=type`; the
  proposed `childForFieldName('return_type')` is null for every function).
  A pin test in zig-extractors.test.ts asserts both the grammar fact and
  that `returnType` is extracted (`void`, `!*Counter`).
- "`DIR_LANG` has no zig entry": it does (added in the first pass and
  answered again in the third); the finding is carried over unchanged.

* fix(zig): address fifth gitnexus-check review pass

- tree-sitter-languages test: the Zig "functions, structs, enums, and
  imports" case now asserts the `import.source` capture for
  `const std = @import("std");` (the fixture's only import), so a query
  change that drops Zig import matching fails it instead of passing
  unchanged.

Not re-fixed:
- "Return type is read from the wrong tree-sitter field": carried over from
  the fourth pass unchanged. tree-sitter-zig 1.1.2 has no `return_type`
  field on function_declaration; the return type IS the `type` field, and
  the pin test added in the fourth-pass commit
  (zig-extractors.test.ts, `childForFieldName('return_type')` is null,
  `returnType` = `void` / `!*Counter`) proves it.
- "`DIR_LANG` has no zig entry": carried over unchanged for the third time;
  the entry exists since the second-pass commit.

* feat(zig): export fn visibility, opaque containers, named test blocks, member ownership

Ports the parts of upstream PR #305 (closed, unmerged) that our Zig
provider lacked, plus two gaps found while porting.

- `export fn` / `export var` (C-ABI linkage, never `pub`) are exported;
  the pub/export predicate is now shared by the export checker and the
  method/variable extractors' visibility (`hasZigVisibilityKeyword`).
- `const H = opaque { … }` is a Struct-labelled container (it may own
  methods, never fields) in both the structure queries and the scope
  query; ZIG_CONTAINER_TYPES is the single source for the extractor
  configs.
- `test "name" { … }` blocks are Function nodes named by the string
  node WITH quotes, so `test "add"` beside `fn add` cannot merge onto
  Function:<file>:add; `test_declaration` joins FUNCTION_NODE_TYPES and
  the Zig method config names it in the enclosing-function walk, so
  calls inside a test attribute to the test. Anonymous `test {}` and
  decl-tests `test add {}` are scopes without a node (an empty-name hook
  result stops the walk instead of falling through to the identifier of
  the function under test).
- Empty container bodies (`struct {}`, `opaque {}`) no longer mint a
  nameless Property: tree-sitter-zig 1.1.2 recovers them as a
  container_field with a MISSING identifier; #not-eq? guards in both
  queries and the field extractor drop it.
- Owner walk (`findEnclosingClassInfo`): an anonymous container bound
  by the enclosing `variable_declaration` takes the binding identifier,
  same shape as the Go `type_spec` branch. Before this NO Zig member
  had an owner — zero HAS_METHOD / HAS_PROPERTY edges for Zig.

Not ported from #305, deliberately: `builtInNames` (a bare-name call-site
drop filter; `alloc`/`free`/`append`/`print` are the most common user
method names in Zig and `std.*` receivers are already external via the
import binding), `usingnamespace` (removed in Zig 0.15), `@cImport`,
`build.zig` ignore, and the web-app changes.

* test(zig): absent-grammar case owns GITNEXUS_SKIP_OPTIONAL_GRAMMARS

The loader parses the opt-out variable lazily once per module copy, so
under `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=zig` (or `all`) — a supported way
to run — the fresh loader took the opt-out branch and the "not the
opt-out path" assertion failed before the absent-binding path ran.

Clear the variable for the fresh module and restore it in `finally`,
instead of returning early: the branch stays exercised in every
environment. Verified: fails on the previous code under `=zig`, passes
with and without the variable now.

* fix(zig): declare the Union relation pairs — analyze aborted on any union

`gitnexus analyze` exited 1 on every Zig repository that declares a
`union` (including the zig-basic fixture itself): the member-ownership
commit made `union_declaration` a MEMBER_OWNER, so HAS_PROPERTY /
HAS_METHOD edges are emitted FROM a Union node, but `Union` was not in
LINKABLE_LABELS, so the schema's scope-bridge cross product never
generated a `FROM Union` pair and LadybugDB rejected the edge
(`labelPair: "Union|Property"`). Resolver tests stayed green because
they never write to the DB.

- `Union` joins LINKABLE_LABELS (also bridges `Tag{…}` constructor
  references); the three hand-written `→ Union` target pairs move to the
  generated half, per the STRUCTURAL_PAIR_DDL rule.
- structural-pair-coverage gains an optional-grammar corpus with
  zig-basic (`Union|Property`, `Union|Method` sentinels), skipped when
  the grammar is absent.
- Rust `union_item` note updated: the three gates it cited are widened.

Note for reviewers: the DDL fingerprint changes (#2808), so existing
indexes are rebuilt on next analyze.

* feat(zig): resolve path deps through the dep's build.zig and src/root.zig

The bare-name resolver only knew `src/<name>.zig` and `src/main.zig`.
`zig init` has written `src/root.zig` for libraries since 0.12, so the
default library layout never resolved. Now: the root the dep's own
build.zig declares (`b.addModule("<name>", .{ .root_source_file =
b.path("…") })`, name-matched module first), then src/root.zig,
src/<name>.zig, src/main.zig. `normalizeZigDepPath` is shared by the
loader and the resolver.

* feat(zig): Const/Variable defs, member imports, receiver typing, generic type constructors

Coverage gaps found by indexing idiomatic Zig against the branch:

- Const / Variable nodes: ZIG_QUERIES had no @definition.const /
  @definition.variable, so `pub const VERSION`, error sets and type
  aliases were absent and zigVariableConfig never ran. Rules are gated on
  the literal `const` / `var` keyword — tree-sitter-zig 1.1.2 parses
  statement assignments (`x = 5;`, `x += 1;`, `_ = expr;`) as keyword-
  less `variable_declaration`s, and the scope query minted a phantom
  local per assignment and one `_` per discard. Container and @import
  bindings are skipped via `shouldSkipDefinitionCapture`.
- Imports: `const X = @import("x.zig").X` (named / alias), `const X =
  ns.X` where `ns` is an @import binding of the file (promoted to a
  named import), and `pub usingnamespace @import(...)` (wildcard, with
  `expandsWildcardTo`). All three lost the file-level IMPORTS edge.
- Receiver typing: `var x: T = undefined` / decl literals `const x: T =
  .init()` (annotation), `var c = T.init()` / `mod.T.init()` (call
  return), `List(u8){}` (instantiation literal); `normalizeZigTypeName`
  drops the comptime argument list.
- Generic type constructors `fn List(comptime T: type) type { return
  struct {…}; }`: the returned container is a Struct/Union/Enum named
  after the fn, owns its members (HAS_METHOD / HAS_PROPERTY), binds in
  the module scope beside the Function def, and is emitted ahead of it
  so a named import binds the type.
- `export` vs `pub`: `visibility` is now `pub`-only (Zig-module fact);
  `isExported` keeps `pub|export` (visible outside the unit, as C's
  external linkage). `export fn` without `pub` is not reachable from
  other Zig files.
- Extractor configs share `zigContainerName`; ast-helpers' owner walk
  learns the type-constructor shape.

Tests: zig-idioms fixture (10 resolver cases), extractor/interpret unit
cases for each rule.

* fix(zig): address sixth gitnexus-check review pass

- Windows absolute `.path` deps (`C:\x`, `C:/x`) return null from
  `normalizeZigDepPath` like POSIX ones; a `/`-only check let them
  through as repo-relative.
- `parseZigBuildZon` accepts the `.dependencies = .{` header only at
  brace depth 1 (a direct field of the file's `.{`), so a same-named
  field nested in an earlier struct cannot hijack the block.
- `importsExecuteWhereWritten: false` on the provider: `@import` is
  compile-time name lookup (as C `#include`, Rust `use`); a body-level
  `@import` is no longer marked `runsOnlyWhenCalled`.
- Namespace imports record the MODULE as `importedName`
  (`zigModuleNameOf`: last path segment without `.zig`), per the shared
  contract; the local handle stays `localName`.
- Keyword-less `<ident> = @import(…)` (`_ = @import("x.zig")` in a test
  block) is a `side-effect` import: file edge, no binding. Only
  `const`/`var` declarations bind a name or feed alias promotion.
- `extractZigFunctionName` doc: an empty name is falsy, so the enclosing-
  function walk skips the test node and continues to the File; it does
  not "end" there.

Not re-fixed: "DIR_LANG has no zig entry" — carried over for the fourth
pass in a row; `test/helpers/literal-collectors.ts` has had `zig` in
`DIR_LANG` (line 91) and `BASENAME_LANGS` since the second-pass commit.

Regression tests: absolute-path spellings, nested `.dependencies`
decoy, namespace/side-effect interpretation, function-scoped `@import`
not deferred (all four fail on the previous source).

* fix(zig): address seventh gitnexus-check review pass

- The scope query's `@import` binding rules are keyword-gated (`"const"` /
  `"var"`, first-child anchored) like every other binding rule, and the
  keyword-less `<ident> = @import(…)` statement has its own
  `@import.side-effect` rule. Tree-sitter queries cannot express "no
  keyword child", so that rule also matches the keyword shapes and
  `emitZigScopeCaptures` drops those (they are the binding rules'
  matches). Behaviour is unchanged from the sixth-pass fix — the existing
  side-effect test covers it — the query text now carries the guard the
  finding asked for.

Not re-fixed: "DIR_LANG has no zig entry" — fifth pass in a row;
`test/helpers/literal-collectors.ts` has had `zig` in `DIR_LANG` since
the second-pass commit. Left for a human reviewer to close.

* fix(zig): resolve @import of the repo's own build.zig modules (F3)

Bare-name imports were resolved through build.zig.zon path deps only, so
the module a repo's ROOT build.zig declares for itself —
`b.addModule("lightpanda", .{ .root_source_file = b.path("src/lightpanda.zig") })`,
imported by name from 378/567 Lightpanda files — never produced an IMPORTS
edge, and nothing reached through `lp.X` resolved. A repo with a build.zig
but no build.zig.zon got no resolution config at all.

- language-config: `parseZigRootModules` (static scan of the root build.zig:
  `addModule("<name>", …root_source_file = b.path("<p>.zig")…)`, and
  `createModule`/`addModule` bindings named via `addImport("<name>", m)` or
  `.imports = &.{ .{ .name, .module = m } }`; generated / `.url` / computed
  modules are skipped) → `ZigBuildZonConfig.rootModules`.
- `loadZigBuildZon` → `loadZigBuildConfig`: reads the zon AND the root
  build.zig; null only when neither contributes.
- resolver: root modules are consulted before path deps; std/builtin/root
  still never resolve.
- fixtures: zig-idioms gains a Lightpanda-shaped root module (+ decoy
  `addOptions().createModule()`); new zig-rootmodule (build.zig, no zon).

Corpus (Lightpanda): IMPORTS 3014→3389 (378 edges to src/lightpanda.zig,
was 0), CALLS 13885→13989, ns.f() 79.3%→83.3%,
param.m() type=ns-qualified 43→46/417.

* fix(zig): import every @import in expression position; resolve @import("x").f()

Both query sets only saw `@import` as the value of a const/var or under
`usingnamespace`, so an @import in any other position produced no file
edge: Lightpanda's `pub const Interfaces = .{ @import("a.zig"), … }`
registration table (288 modules), call arguments
(`CounterEnum("size", @import("ArenaPool.zig").BucketSize)`), comparison
operands (`JsApi == @import("x.zig").JsApi`) and member-call receivers
(`try @import("dump.zig").root(...)`) — 417 of 3,401 in-repo import pairs
had no IMPORTS edge, and the 80 inline-receiver calls resolved 0 times.

Scope query: a catch-all `@import.inline` rule matches every `@import`
builtin; `emitZigScopeCaptures` drops the ones a binding rule (or the
keyword-less side-effect rule) already claimed (by string-node id) so a
bound import is never doubled, emits the rest as side-effect imports once
per distinct source per file, and binds a member-call receiver as a
namespace import whose local name is the builtin's own text — the
`@reference.receiver` text on that call is identical, so the shared
namespace-receiver lookup (Case 1) resolves the member in the imported
module.

ZIG_QUERIES: the three variable_declaration/usingnamespace-anchored
`@import` rules collapse into the same single builtin rule (the structure
phase only skips import matches; one match per builtin keeps
tree-sitter-languages' exact-capture assertion intact).

Lightpanda corpus (zig-corpus-check, before → after): IMPORTS 3014 → 3426,
in-repo pairs missing 417 → 5 (4 under a default-ignored `cache/` dir, 1 a
commented-out import the census regex counts), `@import(..).f()` 0/80 →
72/80 (the 8 left are `@import("root")` and non-import builtins the census
mislabels), CALLS 13885 → 13959; every other line unchanged.

* feat(zig): model file-structs — a file with top-level fields is a Struct named after the file

In Zig every file is a struct; one that declares top-level fields is an
instantiable type whose name is the file stem (`Page.zig` declares `Page`,
`@typeName` agrees), and its top-level `fn`s taking `self` are its methods.
Lightpanda spells 413 of 567 files this way and, before this, `page.getArena()`
on a `page: *Page` parameter resolved 23 of 993 times (2.3 %) — `impact` on
`Page.getArena` reported 0 callers for 159 call sites, and 2,395 top-level
fields were ownerless Property nodes.

Definition phase: `((source_file (container_field …)) @definition.struct)` +
the class extractor names it from the file path (`zigContainerName(source_file,
filePath)`); top-level fns/fields are owned through the new
`LanguageProvider.resolveFileTypeOwner` hook (consulted by
`findEnclosingClassInfo` when the walk reaches the tree root, and by the
method/field extractors' owner lookup) — ids become `Method:<file>:Page.getArena#0`
/ `Property:<file>:Page.session` with HAS_METHOD / HAS_PROPERTY edges.

Scope phase: `emitZigScopeCaptures` emits a Class scope over the whole file
(same range as the Module scope, nested under it — the pair `canParentScope`
already admits) plus a Struct def anchored on it; member NAME bindings are
hoisted back to the Module scope by `zigBindingScopeFor` so `Page.init()`
(namespace member) keeps working, while ownedDefs stay in the Class scope so
`populateClassOwnedMembers` stamps the owner. The file-level `const Page =
@This();` alias no longer mints a Const (it would shadow the Struct); `@This()`
aliases in type position (`self: *SigHandler` in Sighandler.zig, nested
`Self`) are rewritten to the container name so receivers resolve. A namespace
import of a `.zig` file gets a NAMED twin of the file stem so `x: *Page` in the
importer binds the type as well as the module.

Shared, additive: `resolveFileTypeOwner` hook; `filePath` threaded to
`ClassExtractionConfig.extractName` / `extractOwnerName`; nameless
`definition.struct` passes `getLabelFromCaptures` like `definition.class`
already did (extractor synthesizes the name).

Lightpanda corpus (zig-corpus-check): param receivers typed by a file import
23/993 → 993/993; `self.m()` 99.2 → 100 %; annotated locals 11.6 → 23.2 %;
CALLS 13,885 → 15,689; HAS_METHOD 1,477 → 8,003; ownerless Property 2,395 →
276; Function/Method 8,378/1,518 → 2,562/7,330; no row regressed.
Fixture `zig-filestruct` (Page/Session/Sighandler/util) + unit tests pin the
shape, the stem naming, the alias rewrite, the namespace twin and the
unchanged namespace-file behaviour.

* test(zig): expression-position import case sees the file-struct type twin

* fix(zig): stop reading member calls `x.f(arg)` as direct calls named `f`

tree-sitter-zig spells `field_expression` as `object:`/`member:`; the shared
callable-flow reader only knew `property`/`field`/`method`, so every Zig
member call collapsed to a DIRECT call named after the member and the
solver fanned each argument out to every same-named callable
(4,761 cap warnings on Lightpanda, `Global.deinit -> Global.deinit`
self-loops through `pub const release = deinit;`).

Shared (grammar-neutral, receiver-gated):
- `memberParts` also reads `member` (only C/C++ `offsetof_expression` and
  JS `class_body` expose that field, without a receiver field).
- A member call is a field-stored-callable invoke only when a MEMBER store
  (`o->run = handler`, `self.f = target`) or a declared callable-typed field
  is visible — a same-named plain binding no longer gates it.
- `direct-callee-name` requires a direct designator: `.init(x)`,
  `' '.join(x)`, `string.Join(x)` name no callee to seed by simple name.
Zig: formals are numbered without the leading `self`, so `r.run(target)`
joins `cb` and yields `Runner.run -> target`.

Goldens for python/csharp regenerated: the only drift is the dropped
`direct-callee-name` on `' '.join(...)`, `text.strip().ljust(...)`,
`string.Join(...)`.

Lightpanda: cap-warnings 4761 -> 2, cvf self-loops 10 -> 0,
CALLS 13885 -> 13857 (28 removed, all callable-value-flow: 10 self-loops,
17 same-name fan-out, 1 lost `on -> TypeErased.start`; +1 correct
`Arena.alloc -> allocator`).

* fix(zig): bind container field types so `self.field.m()` resolves (F5)

A container's field types were never bound on its Class scope: the scope
query's `container_field` rule captured only the name, and
`emitZigScopeCaptures` synthesized no `@type-binding.field` group. The
compound resolver reads member types from that scope
(`typeOfMemberOnClass` → `classScope.typeBindings.get(field)`), so
`self.session.name()`, `self.counter.incr()` — Lightpanda's dominant
cross-object call shape — resolved 9 of 2803 times (0.3 %).

- query.ts: capture `type: (_)? @declaration.field-type` on
  `container_field` (enum variants have none).
- captures.ts: per typed field, push a `@type-binding.field` group (name =
  field, type = the type text, `@This()` aliases rewritten to the container
  name like parameter types); anonymous inline containers are skipped. The
  binding lands on the container's Class scope — the file's Class scope for
  a file-struct — since `zigBindingScopeFor` hoists only declaration names.
- query.ts/captures.ts: `const page = self.page;` / `var s = self.session;`
  one-level field aliases become `@type-binding.alias` bindings whose "type"
  is the RHS path; the resolver's member-alias branch re-resolves it as a
  receiver chain. Import aliases (`const Counter = counter.Counter;`) are
  dropped — they are named imports.
- interpret.ts: `@type-binding.field` → 'annotation',
  `@type-binding.alias` → 'assignment-inferred' (an annotation on the same
  binding wins).

Corpus (Lightpanda, zig-corpus-check): self.field.m() 11/2803 (0.4 %) →
1433/2803 (51.1 %); ident.m() bound=local-field-access 29/1902 (1.5 %) →
307/1902 (16.1 %); chain.m() 146/5022 (2.9 %) → 642/5022 (12.8 %);
CALLS 15838 → 18066. self.m() / free f() / ns.f() unchanged.

Tests: unit (zig-extractors) — one @type-binding.field per typed field with
sigils stripped and aliases rewritten; the binding hosted on the container's
Class scope (file-struct: the file's Class scope, not Module) with the
written spelling as declaredSpelling; field aliases bound to the RHS path
and never for import aliases. Integration (zig-idioms `holder.zig`,
zig-filestruct `Page.zig`): `viaField → incr` ×2 into counter.zig,
`viaAlias → get/twice`, `sessionName → name` / `sessionLabel → name` into
Session.zig. All fail without the change. The optional-payload capture
`if (self.opt) |c| c.incr()` is not asserted (F6).

* fix(zig): `pub const X = @import(…)` at file scope republishes X (reexportsName)

Lightpanda's `lightpanda.zig` is one long list of `pub const Arena =
@import("Arena.zig");`, and most files name their types through it (`const
lp = @import("lightpanda"); const Arena = lp.Arena;`, `arena: *lp.Arena`).
The scope side treated those bindings as plain imports of the hub file, so
the hub never published the names it re-exports and a third file's `const
Arena = lp.Arena;` (promoted to a named import of `Arena` from the hub) found
nothing.

`emitZigScopeCaptures` now marks named/alias import groups whose declaration
is a file-level `pub const` — the `@import(...).X` form, the alias promotion
`pub const Bar = ns.Bar`, and the file-struct type twin of `pub const Arena =
@import("Arena.zig")` — and `interpretZigImport` sets the shared contract's
`reexportsName: true` on them (the Python `__init__.py` shape, consumed by
`buildReexportClosures`). Private and fn-local bindings stay unflagged.

Not covered here: a receiver ANNOTATED with the dotted hub path (`arena:
*lp.Arena`) — Case 3 of the receiver-bound pass looks the member up with
`findExportedDef`, which only sees locally declared names; following
re-exports there is a shared change left for a follow-up.

* fix(zig): type receivers through `const X = <type expr>;` aliases (F7)

`const LocalAlias = Local;`, `const T2 = Thing;` (alias of an alias/import)
and `const B = util.List(u8);` (an INSTANTIATED generic type constructor)
were plain `@declaration.variable` bindings, so `LocalAlias.mk()`,
`var l = LocalAlias.mk(); l.go()`, `T2.make()`, `B.init()`, `B{}` and
`var x: B` all typed nothing (review repro r3-flow b1..b9; Lightpanda:
`pub const Proto = HtmlElement;` x68, `const Allocator = std.mem.Allocator`
x104, `pub const KeyIterator = GenericIterator(...)`, fn-local
`const R = ...(...)`).

Model: a `@type-binding.alias` binding of the alias NAME to the value's type
text — Rust's `let x = y` / JS's `const B = Foo`, source
'assignment-inferred' — NOT a TypeAlias def. Reasons: (1) the shared
machinery already chains typeBindings (`followChainedRef` in the extractor,
`followChainPostFinalize` after propagation), so `var l = LocalAlias.mk()`
and `var x: B` reach the target through the alias with no new shared code;
(2) nothing shared follows a `TypeAlias` def to its target — `isShapeLike`
only makes the alias itself a member owner (TS object-type aliases) — so a
relabel would have needed language-named shared code; (3) graph node ids
are UNCHANGED: every alias stays `Const:<file>:X`. `normalizeZigTypeName`
already drops the comptime arguments, so `util.List(u8)` binds `util.List`
and resolves through the namespace import (Case 3). The identifier /
member shapes also take `var` (`var cur = orig; cur.go()` — the cursor
idiom, same binding as Rust's `let x = y`).

Heuristic, stated as such: a CALL value is kept only when the callee's last
identifier is TitleCase (Zig's naming convention for types), because the
grammar cannot tell `util.List(u8)` from `util.makeThing()` and the latter
belongs to the call-return rules; the call-return group is dropped for the
same TitleCase shape so the two never race on match order. Import bindings
(`const Stack = @import("x.zig").Stack`), promoted namespace-member aliases,
enum/decl literals (`.foo`) and the `type:` annotation of
`var b: T = undefined;` are excluded.

Not done: the two-hop `pub const bridge = js.Bridge(T); bridge.accessor()`
chain. `js.Bridge` is a Function that RETURNS `bridge.Builder(T)` (a call,
not a container), so the alias binds `js.Bridge`, Case 3 finds a Function
with no members in js.zig, and Case 3b is skipped for a namespace head.
Following that hop needs a namespace-member return-type route in shared
code (or a Zig `resolveQualifiedReceiverMember` hook that re-implements
member lookup without the model); left for a follow-up.

Corpus (Lightpanda, `harness/zig-corpus-check.mjs`): CALLS 15838 -> 16051
(+213, 0 removed), `ident.m() bound=local-alias` 6/63 -> 36/63,
`local-call` 159 -> 176, `module/unknown` 110 -> 116, `local-other`
271 -> 273; `self.m()`, `free f()`, `ns.f()` unchanged or up.

* fix(zig): one alias rule set — F7's alias rules subsume F5's field-access alias rules

* fix(zig): type locals through try/catch/orelse, return types and payload captures

F6 of the gitnexus-check review. Three gaps in the value flow that types a
local receiver, all measured on Lightpanda:

1. `@type-binding.call-return` needed the `call_expression` as the DIRECT
   value child, so `const p = try Page.init(…)` (2,551 sites), `… catch
   return` (410) and `… orelse return` typed nothing. The rule is now one
   keyword-gated declaration match; `emitZigScopeCaptures` unwraps `try`,
   `catch`, `orelse` and parentheses (`zigUnwrapValue`) and decides what the
   value types (`zigCallReturnTypeOf`): a module-level receiver still names
   the type (`Counter.init()` → Counter, Rust `Foo::new()`); a free call
   binds the callee name (`makeThing`); a member call on a fn-LOCAL receiver
   (parameter / local / payload — Zig forbids shadowing, so "declared in the
   fn" is exact) binds the compound `node.asElement()` the shared resolver
   walks to the method's return type — instead of typing `el` as `Node`. A
   TitleCase callee (`List(u8)`) is a type constructor and binds nothing.

2. No `@type-binding.return` existed. `fn make() !*Thing` now binds
   `make ↦ Thing` in the enclosing scope (Module for free fns, the container's
   Class scope for methods, where the compound resolver reads it). Builtins,
   `type`, `@TypeOf(…)` and comptime type parameters (`?*T`) bind nothing;
   `@This()` / `Self` returns name the container. `normalizeZigTypeName` now
   strips the error union BEFORE the payload's sigils, so
   `Allocator.Error!*Page` → `Page` (it used to leave `*Page`).

3. Payload captures had no binding at all. `populateZigRangeBindings`
   (registered as `populateRangeBindings`) types `for (items) |it| / |*it|`,
   `for (items, 0..) |it, i|`, `if (opt) |v|`, `if (call()) |v|`,
   `while (it.next()) |x|` from the SUBJECT's written type minus one layer
   (`[]T` element, `?T` payload) — declining when the layer is not visible
   (`ArrayList(T)`) — and the same projection for `const t = items[i]` /
   `opt.?` / `ptr.*`. `catch |err|` and `switch` prongs are skipped.

Corpus (Lightpanda, gate before → after): CALLS 15838 → 17853;
`ident.m() bound=local-try/catch/orelse` 11/1298 → 584/1298;
`local-call` 159/1282 → 398/1282; `payload` 82/1085 → 233/1085;
`other-recv:call_expression` 5/991 → 457/991; `local-other` 271 → 364;
`self.m()` 100 %, `free f()` 97.5 %, `ns.f()` 83.4 % unchanged; nothing down.

Not covered: `const t = ns.f()` (a namespace fn's return type across files —
Case 3 has no path from a namespace head to a callable's return binding),
and expression receivers (`items[i].run()`, `o.?.run()`).

* fix(zig): resolve leftover fixture merge markers (Page.zig)

* fix(zig): reconcile F6 value inference with F7 aliases and F5 field bindings

- A fn-local TitleCase receiver (`const R = generic.List(u8); var l =
  R.init();`) is a type alias (F7), not a value local: `R.init()` names the
  type `R` like `Counter.init()` does at module level, so `l` chains
  R → util.List → push. F6's local-receiver rule now excludes TitleCase heads.
- The F6 unit helper only collects the value-inferred / return kinds it
  owns; F5 field and F7 alias bindings for the same names are asserted in
  their own suites.

* fix(zig): give function-local and anonymous containers an identity (F8)

`const R = struct {…}` declared inside a fn (Lightpanda's reflection.zig
has ~20, one per builder) all collapsed onto one `Struct:<file>:R` with one
`R.get`; anonymous containers (`std.sort.pdq(…, struct { fn lessThan … }
.lessThan)`, `const byte_size = struct { fn it … }.it;`, `?struct { min,
max }` field types) had no identity at all, so their fns were OWNERLESS
Methods (`Method:<file>:lessThan#3`) that collided across a file.

`zigContainerName` now yields the graph IDENTITY on both phases:
  - function-local named: `<enclosing callable>$<name>` — `Reflect.string$R`
    (Java local-class `$` chain; `populateClassOwnedMembers` leaves it whole);
  - anonymous: `<host>$<ordinal>` — `build$1`, `Outer$1`, `Page$1`
    (javac's `Outer$1` numbering per host, in source order);
  - a `test` host is keyed `test@L<line>` (its string does not survive the
    class extractor's qualified-name normalization).
`zigContainerBindingName` keeps the spelling code writes (`R`) for scope
bindings and `@This()` alias rewrites (`@declaration.binding-name`).

Structure phase: bare `(struct|enum|union|opaque_declaration)` rules mint the
local/anonymous nodes via the class extractor; `shouldSkipDefinitionCapture`
keeps exactly one rule per container (`zigContainerAnchor`); a new
grammar-neutral `resolveContainerTypeOwner` provider hook lets the shared
owner walk name a container from context, so `Method:<file>:Reflect.string$R
.get#0` and its HAS_METHOD source agree by construction. Scope phase: the
wrapper group splits name/binding-name for locals and anonymous containers
get synthesized `@declaration.<kind>` defs (`is-synthetic`).

Lightpanda: ownerless Methods 14 → 0, fns without a node 55 → 0, ownerless
Properties 276 → 5, HAS_METHOD 8003 → 8074, HAS_PROPERTY 7275 → 7403,
CALLS 15838 → 15857, Struct 1905 → 2199; no resolution bucket dropped.

* fix(zig): re-add the implicit receiver on the call side so both method-call spellings reach the callback formal

`extractFunctionParameters` sliced the leading `self` off the formals, which
lined up `r.run(target)` (target@0 ↔ cb@0) but lost the explicit spelling
`Runner.run(&r, target)` (&r@0, target@1 ↔ cb@0): the callback never joined
its formal and `run → target` was missing (PR #1432 review by koriyoshi2041).

Formals are numbered once per function while the receiver differs per call
shape, so the fix lives in `extractCallArguments`: keep `self` as formal 0
and prepend the receiver as actual 0 when the callee is a member call on a
VALUE receiver — chain head is a fn-local name that is not TitleCase, the
same value-vs-type rule F6 uses. Namespace / type / decl-literal receivers
(`Runner.init(cb)`, `helpers.apply(cb)`, `List(u8).init`, `.init(cb)`) get
no prepend. Known residual gap, documented: a module-level value receiver
(`global_runner.run(cb)`) is not fn-local and still misses.

Tests: the F2 integration case now asserts both spellings plus a namespace
call; the unit contract pins `self@0, cb@1` and the per-call actual index.

* fix(zig): address eighth gitnexus-check review pass

- Named dependency modules: `parseZigBuildModuleRoots` scanned
  `addModule("<name>", .{ … })` with a `[^}]*` regex, so a nested field
  before `.root_source_file` (`.imports = &.{ .{ … } }`) ended the match
  at the inner `}` and demoted the module to an unnamed fallback — the
  first exe/lib root in the file then answered `@import("<name>")`. The
  named lookup now walks the balanced `addModule(…)` argument list
  (same scanner as `parseZigRootModules`, comment-stripped, string-aware);
  the unnamed fallbacks are unchanged. Regression test in
  `zig-import-resolver.test.ts`.
- Type-position `@import`: `var x: @import("m.zig").T = undefined;` was
  read as an import binding of `x` on both sides — the query rules match
  the `type:` child like a value, and `isZigContainerOrImportBinding`
  scanned every named child — so `x` was never declared and became a
  named import of `T`. The helper now skips the `type:` field, and
  `emitZigScopeCaptures` drops binding-rule matches whose `@import` sits
  in the annotation (`isZigTypePositionImport`) without claiming the
  source, so `x` binds as a variable and the file edge survives as a
  side-effect import. Regression tests in `zig-extractors.test.ts`
  (variable extractor + scope captures).
- Union in the class-capture skip guard: the parse-worker's inline
  class-like predicate lacked `Union`, so a `Union` definition bypassed
  `shouldSkipClassCapture` unlike every other `ClassLikeNodeLabel`.
  Added the label; no Zig behavior changes (Zig defines no skip hook), so
  no test.
- File-owned method ids in `findEnclosingFunctionId`: the arity lookup
  used `findEnclosingClassNode` while the owner came from the file-owner
  aware `cachedFindEnclosingClassInfo`, so a Zig file-struct's top-level
  fn produced `Method::Page.get` without the `#<arity>` suffix. It now
  uses `findEnclosingClassNodeOrFileOwner`, the definition-phase lookup.
  Consistency fix only: `ParseWorkerResult.calls` / `.assignments` (the
  sole consumers of this id) are merged but not read since #942 — CALLS
  edges come from the scope pipeline, whose ids were already right — so
  no observable graph change and no test.

Not re-fixed:
- `test/helpers/literal-collectors.ts` `DIR_LANG` has no `zig` entry
  (raised for the seventh time): the entry exists (`zig:
  SupportedLanguages.Zig`, added by the second-pass commit), so
  `languages/zig/**` literals are already validated against the Zig
  grammar alone; documented in the PR body since the fifth pass.

* fix(zig): keyword-gate the constructor type-binding rules

The three `@type-binding.constructor` rules (`const p = T{…}`, `mod.T{…}`,
`List(u8){…}`) matched any `variable_declaration` with an identifier and a
`struct_initializer`, keyword or not — and tree-sitter-zig 1.1.2 parses a
re-assignment `p = T{…};` (and `_ = T{…};`) as the same node type. So an
assignment minted a constructor binding for `p` in its own block, and one
for `_`. Zig's static typing makes the extra binding redundant (`p`
already carries its type from its declaration: annotation, constructor or
inferred value), so it cost little, but it declared nothing and stood out
against every other binding rule (`@declaration.variable`, the import
rules, the call-return rules), which are keyword-gated for exactly this
shape. Split each rule into `"const" .` / `"var" .` variants, like the
call-return rules.

Regression test in `zig-extractors.test.ts`: `p = T{…}`, `q = mod.T{…}`,
`l = List(u8){}` and `_ = T{…}` after their declarations yield only the
three declaration bindings (fails on the previous query). The zig,
callable-value-flow, grammar-literal and tree-sitter-languages suites are
unchanged.

Raised twice by gitnexus-check (passes on 2026-08-18 12:42 and 12:56).

* fix(zig): re-baseline the callable-flow capture fingerprints, keep `await f<T>(x)` a direct callee

The `benchmarks (GITNEXUS_BENCH)` CI job gates two capture fingerprints
that this branch's shared callable-flow change (3e62b99a) moved without
re-baselining: `bench/python-scope/baseline-fingerprint.txt` and four
languages in `bench/scope-capture/baselines.json` (csharp, cpp,
typescript, kotlin). Both `--check` runs pass on origin/main and failed
on this branch; every other language matched its baseline on the same
run.

Drift, verified by dumping the canonical matches on both trees:
- csharp / kotlin / python: exactly the intended change — a MEMBER call
  (`string.Join(x)`, `.forEach { }`, `' '.join(x)`, `.ljust(w)`) no
  longer carries `direct-callee-name`; the argument fact is unchanged.
- cpp: `choice.select(1)` (cpp-deleted-overload) drops an INDIRECT
  invoke + its synthetic `@reference.call.free` that were gated only by
  the same-named free `select` binding; the site is a genuine method
  call already captured as `@reference.call.member`. capture_groups_fp
  4605 -> 4601.
- typescript: `await svc.verify<T>(x)` (member) and `initializer()(cb)`
  (call-of-call) drop the name as intended. But `await verifyToken<T>(x)`
  — a DIRECT call — lost it too, because tree-sitter-typescript parses
  `await f<T>(x)` as `call_expression(function: await_expression(f),
  type_arguments, …)` and the new direct-designator gate saw an
  await_expression, not `f`. `wrappedExpression` now unwraps
  `await_expression` (no named field, so the field-based unwrap missed
  it), restoring parity with main for the direct spelling while the
  member spelling stays nameless. Regression test added; it fails
  without the unwrap on both assertions.

Gates run locally: scope-capture --check (15 languages), python-scope
--check, import-target --check, tsc, eslint, prettier, the callable-flow
/ golden / tripwire / resolver test files (33 files), full suite with
coverage (82.9/71.5/89.1/86.4 vs 26/23/28/27 thresholds).

* test(zig): fail CI when the optional Zig grammar is absent

`@tree-sitter-grammars/tree-sitter-zig` is an optionalDependency, so every
Zig suite gates on `isLanguageAvailable(Zig)` and the ABI load-smoke accepts
a clean load failure for an optional grammar. Both are the right contract for
a platform with no prebuild, and together they leave a hole: if the grammar
never installed on any CI runner, this PR would merge with all eight Zig
resolver suites plus the structure-phase suite reported green-by-skip, having
never executed the native Zig parser once.

Close it with the `GITNEXUS_REQUIRE_FTS` idiom already used for the FTS
suites. `GITNEXUS_REQUIRE_ZIG=1` declares "this runner has a prebuild, the
grammar MUST be here", and a missing grammar becomes a failure instead of a
skip. tree-sitter-zig@1.1.2 publishes prebuilds for {darwin,linux,win32}-
{x64,arm64}, so the flag is set on two required jobs that all run on covered
platforms: the sharded ubuntu `tests` job and the three-OS `abi-assert` job.

- test/helpers/optional-grammar.ts: the registry mapping a language to its
  require-variable, plus `describeGrammarPresence`, a presence assertion that
  FAILS when required-but-absent. Deliberately a separate test rather than
  flipping the suites from skip to fail: a skipped suite reports success, so
  only a failing test can turn "Zig never ran" into a red job.
- parser-loader-abi.test.ts: the optional exemption is revoked for a language
  the environment declares required, so an ABI-broken Zig binding fails the
  smoke instead of passing as a clean absence.
- optional-grammar-gate.test.ts: pins the two ways the gate could silently
  never fire — reading a variable name CI does not set, or accepting a value
  CI does not write.

Nothing changes for a run that leaves the variable unset: local runs and any
future prebuild-less platform still skip. Verified both directions --
`GITNEXUS_REQUIRE_ZIG=1` alone: 149 passed, 0 skipped; with
`GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` forcing the grammar away it fails 2 tests
with an actionable message; with the skip flag but no require flag it is
green-by-skip exactly as before.

Addresses the test-only blocker in the gitnexus-check review of #1432.

* fix(zig): type the optional-grammar gate by grammar key, not language

The gitnexus-check finding on parser-loader-abi.test.ts:155 is right, and
none of the gates caught it: `tsconfig.json` includes only `src/**/*`, so
neither `tsc --noEmit` nor CI typechecks the test tree, and vitest strips
types without checking them. Confirmed with a scoped tsc run over the file:
`error TS2345: Argument of type 'string' is not assignable to parameter of
type 'SupportedLanguages'`.

Not fixed with the proposed `key as SupportedLanguages` cast, which would
assert something false: `listGrammarSources()` yields one row per SOURCES
entry, including variants like `typescript:tsx` that are not enum members.
`isOptionalGrammarRequired` now takes the grammar KEY it is really given,
and the registry keeps a `satisfies Partial<Record<SupportedLanguages,
string>>` so every key we write is still pinned to a real language.

Two new cases cover the failure mode the type error was pointing at — a
registry key that can never match what the ABI smoke passes, leaving the
gate configured-looking and permanently inert: every OPTIONAL_GRAMMAR_ENV
key must be a key `listGrammarSources()` yields and must be marked optional
there, and an unregistered variant (`typescript:tsx`) must not be required
even with the variable set.

Same blind spot, two more latent errors in files this PR adds, both fixed:
`Parser.Language` is not an exported member (use the `setLanguage` parameter
type, as parser-loader-abi.test.ts already does), and the `ParsedImport`
filter did not narrow the union, so `localName` was read through a `!` on an
arm that has no such property — now a type predicate. The one remaining
error under the same probe, in `resolvers/callable-value-flow.test.ts:319`,
predates this branch (authored 2026-07-17, on main) and is left alone.

structural-pair-coverage's optional-grammar case switches from
`it.concurrent.each` to `it.concurrent.for`: only `for` passes the test
context as a second argument (`each`'s callback is `(...args: T[])`), and
that context carries the dynamic `skip()` the per-language gate calls.
Behaviour is unchanged — grammar present: 10 passed; grammar forced away:
9 passed, 1 skipped.

Whole test tree typechecking is a separate, much larger job: the same probe
over `test/**` minus fixtures reports 734 pre-existing errors across the
repo. Out of scope here.

* fix(zig): address tenth gitnexus-check review pass

- `normalizeZigDepPath`: normalize backslashes BEFORE the absolute-path
  check. A UNC dep (`\\server\share\dep`) used to slip past the check and
  normalize to the repo-relative `server/share/dep`; root-relative `\dep`
  had the same hole. Both now return null. Regression case added to the
  absolute-spellings test with the files those misreadings would resolve.
- `bindPayloads`: a pointer capture `for (pages) |*p|` now records `*Page`
  (declaredSpelling) instead of `Page` — the `*` is an anonymous payload
  child before the identifier. Method dispatch is unchanged (`rawName`
  strips the sigil), but a deref projection `const q = p.*;` now sees the
  pointer layer. New fixture fn `viaPtrCaptureDeref` + assertion; fails on
  the previous code (verified by stashing the src fix).
- `optional-grammar-gate.test.ts`: renamed the `typescript:tsx` case — the
  key IS a registry row; what makes it inert is the missing gate entry. Now
  also asserts a key no registry yields.
- `structural-pair-coverage.test.ts`: header updated — ten tables (not
  eleven) are absent from every rule's target side; `Union` left the set
  when Zig made it linkable.
- `language-classification.ts`: doc comment now names zig in the
  experimental set (added after Ring 1).

Not re-fixed (invalid findings):
- "owner-hook contract wired to an undeclared variable": stale-diff read —
  `findEnclosingClassInfo` declares `resolveFileTypeOwner` /
  `resolveContainerTypeOwner` as optional parameters (ast-helpers.ts:905,
  917) and parse-worker threads them at every call site; tsc compiles clean.
- "optional Zig grammar added unconditionally to the parsing fixture
  suite": the cited block only `fs.readFile`s the committed fixture file to
  assert it is non-empty — no parser or grammar load is involved.

* fix(scope-resolution): mark construction-site CALLS edges in reason (opt-in), enable for Zig

PR #1432 human review, item 2: a Zig struct literal `T{ .f = x }` (no
parens) is modelled as a CALLS edge to the type — the Rust `T { .. }` /
Go `T{}` shape — and nothing on the edge told it apart from an invocation
(`get_next_spawn → SpawnRequest` from seven `return SpawnRequest{ … }`).

`ScopeResolver.markConstructionSites` (default off): when set, the edge
emitted for a `callForm === 'constructor'` site gets ` (constructor)`
appended to its reason, in both emit paths — `local-call (constructor)` /
`import-resolved (constructor)` in the free-call fallback and
`scope-resolution: call (constructor)` in the reference bridge. The Zig
resolver opts in. `Reference` gains an optional `callForm`, copied from
the site by `buildReference`, so the bridge can see the form.

Why `reason` and not a property or edge type: relationships carry no
arbitrary properties, a new column changes the relation DDL and moves
SCHEMA_FINGERPRINT, and `reason` is the channel the IMPLEMENTS `-pointer`
receiver form already uses. Why opt-in: the unsuffixed strings are a
pinned contract asserted verbatim by the other language suites
(php/cpp constructor calls expect exactly `import-resolved`); every
non-Zig edge stays byte-identical.

Tests: `references-to-edges-call-form.test.ts` pins both vocabularies
and the default-off behaviour; `zig.test.ts` asserts
`Reflect.string → Accessor` / `Reflect.url → Accessor` carry
`local-call (constructor)` next to a plain invocation, and that every
marked edge targets a Struct.

* feat(zig): track qualified struct literals (`mod.T{…}`) as construction sites

PR #1432 re-test (issue comment on 97571d23): 163 qualified literals
`mod.Type{ … }` in a real project produced no CALLS edge at all, so only
same-file and imported-name literals were tracked as construction sites.

One query rule captures `(struct_initializer (field_expression object
member))` as `@reference.call.constructor` WITH the receiver. Captured as a
free constructor instead, the site resolves by its simple tail and a
workspace-unique `Thing` answers for `c.Thing{}` whichever module the
source named (measured: c.zig defines no `Thing`, the edge went to a.zig's).
With the receiver the site takes the receiver-bound namespace case — the
path `mod.fn()` takes — which resolves inside the module the receiver is
bound to: `a.Thing{}` / `b.Thing{}` bind their own files, `c.Thing{}` binds
nothing, `std.Thread.Mutex{}` binds nothing next to a local `Mutex`.

That case's edge now goes through `constructionSiteReason` too, so the
opt-in marker (`import-resolved (constructor)` / `global (constructor)`)
reaches it; `markConstructionSites` joins `ReceiverBoundProviderSubset`.
Byte-identical for every provider that does not set the flag.

Also answers the twelfth gitnexus-check pass: the `bodyNodeSet.size === 0`
guard on the extractor factories' no-wrapper branch is deliberate (a config
with wrappers whose node lacks one is a bodiless declaration); the two
comments now say so instead of reading as a universal last resort. Go's
method config, the only other empty-`bodyNodeTypes` config, never reaches
the branch (its `extract()` gates on method/function nodes the class-node
caller never passes).

Tests: new `zig-qualified-literal` fixture (same-named `Thing` in two
modules, a module without it, an external `std` qualifier next to a local
and an imported `Mutex`); `zig-basic` pins `pioneer.Pioneer{…}` and the
union `pioneer.Tag{…}` as marked construction sites.

* feat(zig): resolve hub re-exports, enum-variant receivers and type-named receivers (real-project audit)

Audit of three real Zig projects indexed with this branch (tigerbeetle 246
files, mach 132, ghostty 788): method reachability was 63 % / 35 % / 55 %,
and three shapes accounted for most of the misses.

1. Hub modules. Zig projects publish types through a file made only of
   re-exports (`pub const Terminal = @import("Terminal.zig");`, `pub const
   PRNG = @import("prng.zig");`, `pub const Thing = @import("thing.zig")
   .Thing;`). Such a file owns NO local binding, and `findExportedDef` reads
   local bindings only — so `terminal.Terminal.init()`, `t: stdx.Thing`,
   `var p = stdx.PRNG.from_seed()` and `h: stdx.BoundedArrayType(u8, 4)`
   all resolved to nothing. Measured before → after: CALLS into ghostty's
   `src/terminal/` from outside it 46 → 253 (150 `terminal.Terminal.` sites
   alone); into tigerbeetle's `stdx` hub from outside it 837 → 1500 (136
   static calls, 289 annotations). Method reachability: tigerbeetle
   2249 → 2272 of 3544, ghostty 2766 → 2865 of 5016, mach 1047 → 1051 of
   2967 (mach's hub publishes generic instantiations, `pub const Quat =
   q.Quat(f32)`, a shape this commit does not cover).
   `findExportedDefIncludingImportedNames` reads the finalized channel
   (origin import / namespace / reexport, def already resolved to the
   declaring file), refusing a name bound to two distinct defs. Opt-in per
   provider (`namespaceExportsIncludeImportedNames`): a module's imports are
   not its exports in most languages; Zig opts in because a hub member a
   consumer can name is public by construction. Used by receiver-bound Case
   1, Case 3, the compound resolver's namespace branch, and a new Case 2
   route that resolves a namespace-qualified class receiver (`stdx.PRNG`)
   through the same lookup.

2. Enum variants as receivers. `Operation.create_accounts.event_max()`
   (147 sites in tigerbeetle): a variant has no written type, but it has
   one — the enum itself. `emitZigScopeCaptures` now emits a field type
   binding per enum variant, so the field walk that already handles
   `self.session.name()` types `Op.create` as `Op`.

3. Receivers named after their type. `self` is a convention, not a rule:
   tigerbeetle writes `replica: *Replica` (777 of 1127 methods), mach
   `pool: *@This()` (764 of 833). Reading only `self` as the receiver
   labelled all of them `isStatic: true`, counted the receiver in their
   arity (`Counter.incr#1`) and sourced the scope binding as a plain
   parameter. `zigReceiverParameter` is the single rule for both phases:
   the FIRST parameter when named `self`, or typed as the enclosing
   container (`@This()`, its binding name, a `const X = @This();` alias),
   pointers / const / optionals stripped.

Fixtures `zig-hub` and `zig-receivers` pin each shape, including the
refusals: a private hub import does not leak, a foreign-typed first
parameter is not a receiver, a factory stays static.

* fix(zig): address thirteenth gitnexus-check review pass

- File-struct receivers named after the file stem were always static: the
  method builder called `isStatic` / `extractReceiverType` /
  `extractParameters` without the extractor context's `filePath`, so
  `zigReceiverParameter` could not name a file-struct (`fn add(ledger:
  *Ledger)` in `Ledger.zig`, no `Self` alias) and the fn came out static
  with the receiver in its arity (`Ledger.add#2`) — an id the scope side,
  which always has the path, never produces, so its CALLS edges went
  nowhere. `MethodExtractionConfig` now passes `filePath` as an optional
  trailing argument to those three hooks (same shape as
  `extractOwnerName`); the Zig config threads it through, every other
  config ignores it. Regression tests in `zig-extractors.test.ts` (unit)
  and `resolvers/zig.test.ts` (new `Ledger.zig` in the `zig-receivers`
  fixture: ids, `isStatic`, and the three CALLS edges); both fail on the
  previous source.
- `LINKABLE_LABELS` comment: the remaining `CLASS_KINDS` entries include
  `Namespace`.

Not re-fixed:
- "Ownerless-method assertion regex cannot match `.zig` graph IDs": the
  `[^:]+` segment consumes the whole file path (dots included) up to the
  second colon, and `[^.]+#\d+$` then matches only an owner-less name —
  `Method:src/Sorter.zig:lessThan#3` → true, `…:Sorter.sortBoth$1.lessThan#3`
  → false, checked with node.
- "Public namespace imports are never marked as re-exports": the shared
  `ParsedImport` namespace variant has no `reexportsName` field and
  `contributesReexportEdge` excludes namespace drafts on `base.kind` by
  contract; a `pub const X = @import("x.zig")` hub member is exposed
  through `findExportedDefIncludingImportedNames` instead, which is what
  the audit commit added for exactly that shape.
- "Private Zig namespace imports are treated as public hub exports": a
  private import cannot be named through the hub in code that compiles,
  and `findExportedDef` applies the same no-visibility rule to local
  defs; the finalized binding channel carries no `pub` bit to check.
- "Range binding mutates finalized scopes" and "unconditionally adds an
  optional Zig grammar to the fixture suite": refuted in the tenth and
  eleventh pass notes of the PR body, unchanged since.

* fix(zig): close the adversarial review's ten findings (8.2–8.12)

PR #1432 review 5095267917 on 34c53473 retained eight P1 and two P2
findings; each is reproduced on the new `zig-chains` / `zig-buildmodules`
fixtures with the decoy that made the old answer wrong, and pinned by a
test named after its number.

- 8.2 per-build-module import tables (`parseZigBuildModules`): a source
  resolves a bare name through its own module's `addImport` table (root
  file, else deepest root directory), fails closed when same-directory
  modules disagree, and follows `addImport("api", dep.module("core"))`
  through the dep's `addModule`; repo-wide names and zon deps remain the
  fallback.
- 8.3 module-level value receivers (`zigHostValueNames`) prepend the
  implicit `self` like fn-locals, so `global_runner.run(cb)` joins `cb@1`.
- 8.4 deep member aliases (`@import("lib.zig").B.work`, `lib.B.work`)
  keep the written owner: the module is bound as a namespace and the
  alias's use sites are rewritten to `receiver . member`; only one-level
  aliases are promoted to named imports.
- 8.5 container-hosted containers get owner-qualified identities
  (`A.Item`, `B.Item`, `Outer.Inner`), minted by the bare-container rule,
  while the scope keeps the lexical binding.
- 8.6 result-location `.init(…)` / `.{…}` under an annotation, a return
  type or a field type emit the call / construction site with the
  expected type as receiver.
- 8.7 Zig arm in bench/import-target (five dispatchers, config-free
  fingerprint) + baselines row; `--check` passes.
- 8.9 fn-local `@import` bindings and their uses are keyed per callable
  (`m$f_sib_a`), so sibling fns no longer share one namespace bucket.
- 8.10 `ScopeResolver.resolveNamespaceChains` (opt-in, Zig only): Case 1
  / Case 2 / Case 3 and the compound resolver walk a qualified receiver
  segment by segment — republished modules, nested types, enum variants
  through the module — refusing ambiguous hops. Off, every lookup keeps
  its one-hop split; the 70 resolver suites are unchanged.
- 8.11 `@import("a.zig").Thing{}` binds the module as a namespace in type
  position; `List(u8){}` / `lists.List(u8){}` get constructor sites.
- 8.12 a fieldless file whose top-level fn takes the file's own type
  (`self: *@This()`, `self: *Self`) is a file-struct; two over-matching
  ZIG_QUERIES rules are filtered by `shouldSkipDefinitionCapture`.

Also asserts the committed `opmod.Op.lookup.event_max()` call in zig-hub.

* fix(zig): address gitnexus-check findings on 215f70e3

- receiver-bound Case 3 wraps its reason in constructionSiteReason, like
  Case 1 and the nested-type route of Case 2 (one vocabulary per provider)
- resolveZigImportInternal rejects drive-qualified absolute imports
  (C:\foo.zig), the same test normalizeZigDepPath applies; unit case added
- the compound resolver's chain seed also tries the whole receiver as the
  qualified class (opmod.Op), as a bare class-name head already does
- markConstructionSites contract text names the receiver-bound routes

The return_type field claim is refuted: tree-sitter-zig exposes a fn's
return type as the type field (checked on the grammar).

* fix(zig): a build-module alias bound to an unindexed root fails closed

resolveThroughBuildModules returned undefined when the containing module
bound the alias to a file that is not indexed, which let the repo-wide
addModule map answer under the same name (gitnexus-check on 5299c552).
The module's table is the authority for its aliases: bound-but-unindexed
is null, only an unbound name falls through. Unit case with a same-named
repo-wide decoy, plus the outside-module file that still reaches it.

* fix(zig): an unindexed root module fails closed, never a same-named zon dep

The root build.zig's addModule declaration is authoritative for a bare
name when it binds it; a root that is not indexed used to fall through
to a build.zig.zon path dep of the same name — a different declaration
answering under the name (gitnexus-check on fe24b37f). Same rule as the
build-module tables. Unit case added.

* Address PR review feedback (#1432)

Tighten Zig build-module parsing, receiver/merge helpers, and container queries that gitnexus-check flagged on the open threads.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#1432)

Attach the paren-matcher doc comment to findZigParenEnd instead of zigTopLevelStaticRoot.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Simplify Zig review-feedback helpers after #1432.

Reuse ZON brace/string walkers for top-level root_source_file, drop the dead bind flag and one-off staticRoot wrapper, and merge bindings via a first-wins map.

Co-authored-by: Cursor <cursoragent@cursor.com>

* bench(receiver-resolution): rebaseline for the Zig lang-resolution fixtures

The receiver-resolution gate (#2856/#2899) landed on main after this branch
forked and counts call drops over test/fixtures/lang-resolution, which this
branch extends with the zig-* fixture projects. Regenerated with
`measure.mjs --update-baseline`: callDrops 102 -> 113, all 11 new drops in
.zig files, shape `no-chain`.

Every new drop is a call whose callee has no node in the corpus, not a
resolver regression: `std.Build.Module.addImport` in the three build.zig
fixtures (7, classified in-program), `std.sort.pdq` in Sorter.zig (2,
unknown), `std.Thread.Mutex{}` / `std.mem.Allocator{}` literals in
zig-qualified-literal (2, unknown, the fixture asserts std stays external),
and one `.init()` decl literal on a generic instantiation
(`const u: Stack(u16) = .init()`, in-program). Shape arm unchanged.

---------

Co-authored-by: Garrett Griffin-Morales <grgisme@gmail.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-03 13:29:42 +01:00
glier
dea396a13c
feat(ingestion): resolve Spring messaging destinations into Destination nodes (#3132) 2026-09-02 11:46:50 +01:00
Gergő Magyar
e1e2464960
feat(kotlin): bind Spring config consumers on Kotlin sources (#3126)
* feat(kotlin): bind Spring @Value and ConfigurationProperties consumers

Kotlin sources were skipping the Java-only config-binding attach path, so mixed JVM apps under-reported blast radius for Kotlin placeholders. Capture from the live AST, serialize on the existing side channel, and reuse the shared binder. (U1-U4)

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(review): require exact Spring annotation FQNs and skip raw-string escapes

Reject similarly named third-party imports and leave Kotlin triple-quoted bodies undecoded so capture stays fail-closed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kotlin): use only grammar-valid Kotlin string and class-name nodes

The coverage shard failed the #1920 literal gate on invented string node types and a Java-style name field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kotlin): fail closed on non-literal prefixes and persist unresolved config markers

Reject constant and boolean @ConfigurationProperties arguments, scope nested Value shadows to their owner, and rewrite drifted consumer files when a config key is deleted.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(kotlin): add config-consumer capture benchmark and fix JVM feature stamps

The Kotlin capture path had no performance or behavior guard: a file-wide
lexical shadow regression silently dropped two of every three facts. The new
bench arm fingerprints @Value / @ConfigurationProperties facts from an
explicit-import control against a wildcard-import corpus whose files each
declare a sibling nested `Value` type, so parity between the arms is the
regression gate, and CI runs it with scaling + widening budgets.

Broadening spring.config-bindings to .kt also means Kotlin-only JVM repos now
stamp the feature, which the two exact-map orchestration expectations still
denied.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kotlin): let an explicit import shadow the Spring wildcard import

importedAs accepted a star import of the Spring package even when the same
simple name was explicitly bound to another type, so a file importing
com.example.Value alongside the Spring annotation package emitted a false
@Value consumer fact. Kotlin resolves the explicit import first, so the
wildcard branch now only applies when the name is otherwise unbound.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:32:09 +00:00
MyShining
c217a0f257
feat(spring): import optional Actuator runtime data with Kotlin JVM mapping (#3107) 2026-09-01 06:22:38 +01:00
ChunxueLi
72ff2b9453
feat(jvm): fold Java static wildcards and Kotlin star imports for route constants (#3110)
* feat(java): expand wildcard static imports + POSIX-space import resolution

- `import static a.b.C.*` records the class FQN at extract time
  (ModuleConstants.wildcardImports) and expandJavaWildcardStaticImports
  materializes the bare-name bindings once the repo constants map exists,
  mirroring explicit single-member static imports. Wired in both the
  group-side prepareRepo fold and the ingestion-side parse-impl pass so
  the two query surfaces cannot diverge. Without this, wildcard-imported
  route constants (~693 routes on our monorepo) silently failed to fold
  and their provider contracts were dropped.
- resolveJavaImport compares in POSIX space (backslash-normalized repo
  keys) — on Windows the '/'-joined class file never matched a
  backslash-keyed repo (observed: 675 calls, zero hits).
- Wildcard bindings never overwrite single imports (a member shadowing
  its own wildcard is honored); unresolved wildcards degrade to the
  existing skip floor.

Rebased onto current main: the parse-worker gate this originally carried
is superseded by the provider moduleConstantHeuristic architecture; only
the resolver-side wildcard expansion and POSIX tolerance remain.

* chore(cache): claim SCHEMA_BUMP 83 (82 taken upstream by Spring lookup facts)

* style: prettier

* fix(java): make wildcard static imports actually resolve route constants

extractJavaModuleConstants never populated wildcardImports, so
`import static a.b.C.*;` was inert: the asterisk is a sibling of
scoped_identifier in tree-sitter-java, not a path segment. Record the
class FQN there and stop binding the class simple name as a field.

Wildcard-only files also fell through every harvest gate (the java
provider heuristic regex, the parse worker emit check, and the group
prepareRepo filter), so the constants never reached either layer.
Expansion now resolves targets against constant-defining files only,
keeping ingestion and group in parity (#2980 R4).

Adds unit coverage for extraction, expansion/shadowing, the unresolved
skip floor, the harvest heuristic, Windows path keys, and group <->
ingestion parity.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(kotlin): fold package-star imports for route constants

Kotlin `import pkg.*` now records star scopes and resolves unique
top-level names after local and explicit imports, matching the Java
wildcard path without accepting invalid object-star imports.

Reuse a Java constant-file suffix index across expansion so
wildcard materialization stays linear as both constants and
importers scale. Add named-vs-star benches and CI --check gates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(jvm): fold wildcard imports without shadowing locals or skipping same-package names

Keep Java unfoldable declarations from being resurrected by static wildcards, bind only the target type's members, and prefer Kotlin same-package names over package-star imports. Move repo-wide preparation behind a language-provider hook so the shared parse phase stays language-agnostic.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3110)

- Stop treating a Java static wildcard as a type import for qualified refs.
- Harvest only class and static (including on-demand) imports, not import pkg.*.
- Let harvested Kotlin top-level names shadow same-package star imports.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3110)

- Fold Kotlin classifier-star imports (`import Type.*`) the same way package stars already fold.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3110)

- Rebuild the Kotlin constant index when overlay replaces a contributing file.
- Document the wildcard shape in the Java pipeline e2e fixture.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: l.cx <l.cx@winning.com.cn>
Co-authored-by: ChunxueLi <mecoloud@users.noreply.gitee.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 21:58:40 +00:00
ChunxueLi
4aa6bddd0a
feat(jvm): synthesize Lombok and Kotlin JVM accessor methods (#2885)
* feat(java): synthesize Lombok @Data/@Getter/@Setter accessor methods

* fix(lombok): resolve class identity by AST node id, not simple name

Root-cause fix for the bot review's name-ambiguity findings:

1. Cross-file collision: the owner map was rebuilt per file from
   result.symbols, which accumulates across the whole language group —
   a later Java file with the same simple class name resolved to the
   earlier file's class node. The map is now filled INSIDE the capture
   loop (per-file scope) and keyed by the class_declaration AST node
   id (SyntaxNode.id), which is unique by construction.

2. Same-tail nested classes (Outer.A vs Other.A): a name-keyed map
   overwrote one with the other; AST-node-id keys cannot collide.

3. Synthesized method ids now follow the SAME convention real nested
   member ids use (keyed by the class's own simple name, matching
   findEnclosingClassInfo().className), so call resolution can hit
   synthesized accessors exactly like hand-written ones.

4. Lombok semantics: setters are no longer generated for final fields
   (Lombok never emits those) and @Setter(AccessLevel.NONE) now
   suppresses setters, symmetric to the existing getter suppression.

Also tightens two vacuous test loops flagged by the bot (empty-array
for..of passed trivially): counts are asserted before property loops,
and a new regression test pins distinct owners for same-tailed nested
classes plus the real id convention for nested accessors.

* feat(java): synthesize Lombok accessors via provider hook and scope dual-path

Replace the worker language===Java branch with LanguageProvider.synthesizeStructureMembers,
align MethodRegistry ownership through scope captures, and bump parse-cache schema to 83
so warm caches cannot replay pre-synthesis worker output.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(java): cover Lombok synthesis semantics, cache replay, and CI bench

Add unit/integration matrices (including durable cold/warm/historical parse-cache),
a permanent no-Lombok vs Lombok-heavy harness with fingerprint budgets, and a CI
--check step. Document that Kotlin→Java member CALLS remains a pre-existing gap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(lombok): drop dead state and redundant scans from accessor synthesis

Collapse Lombok import provenance into one compilation-unit scan with a cached
wildcard flag, remove unused planned-accessor fields and the duplicate @Data
enable flag, and plan scope captures without wrapping a fake Parser.Tree.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#2885)

Give each Lombok accessor a unique scope range so multi-declarator fields do not share @scope.function IDs, and type the owner map as ReadonlyMap to match the provider hook.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bench): pin the real Lombok synthesis fingerprint (#2885)

The committed baseline held a fingerprint no revision of this branch ever
produced, so the CI guard failed on every push. Re-pin it to the value the
synthesizer deterministically emits and correct the method count the comment
claims (800 x 4 x 2 = 6400, not 12800).

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(kotlin): synthesize JVM accessors using shared beanspec helpers (#2885)

Kotlin val/var properties now emit the same JavaBeans get/set Methods as Lombok, via jvm/beanspec + jvm/synthetic-accessors. SCHEMA_BUMP 84 invalidates warm caches that would omit those callables.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kotlin): match kotlinc JVM accessor ABI (#2885)

Emit custom getters, preserve is-prefix names, and convert synthetic
graph lines to 0-based so same-name accessors resolve to the owner.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#2885)

Restrict Lombok provenance to lombok/experimental FQNs and match Kotlin existing methods by exact JVM name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(jvm): consolidate accessor synthesis (#2885)

Keep language-specific discovery in Java and Kotlin adapters while centralizing owner orchestration, collision policy, graph emission, and captures.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(jvm): align accessor synthesis with compiler ABI (#2885)

Match Lombok and kotlinc provenance, companion owners, and collision
arity so mixed-JVM CALLS bind to the Methods compilers actually emit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#2885)

Mark Kotlin interface accessors abstract, pin the Lombok case-fold collision test, and document the non-lowercase is-prefix rule.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#2885)

Honor explicit Getter/Setter over @Data regardless of order, and let field @Accessors replace class-level fluent/chain.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): pin Kotlin scope-capture fingerprint after interface accessors (#2885)

Invalidate warm parse cache so interface property Methods are not replayed as concrete.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: ChunxueLi <mecoloud@users.noreply.gitee.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 17:55:52 +00:00
Gergő Magyar
72edf40087
perf(store): V8 sidecars plus hardlinked ParsedFile restore (#3099)
* perf(store): add best-effort V8 sidecars beside canonical JSON caches

Warm ParsedFile and parse-cache loads skip JSON.parse when a sidecar is present. JSON remains authoritative: envelope validation plus v8.deserialize decide the hit, and any failure falls back without reparsing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(store): require generation bind or sidecar drop before cache overwrite

A same-length JSON rewrite could accept a leftover V8 sidecar if both
generation rotation and unlink failed. Refuse the new generation unless
at least one of those invalidations succeeds; skip publishing a sidecar
when only the drop succeeded.

detect_changes --scope all: 7 files, risk low, no affected processes.
tsc --noEmit clean; 115/115 relevant unit tests; cache-related
integration tests pass. parse-impl-env-reads worker-ready timeout is
pre-existing (same 5 failures with this change set stashed). ESLint 0
errors; remaining warnings are pre-existing and not on changed lines.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* refactor(store): share V8 overwrite invalidation across persist paths

The bind-or-drop gate lived in five writers. One helper keeps the
protocol in a single place and lets bind/drop run together on the
async path.

detect_changes --scope all: 3 files, risk low, no affected processes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(store): hardlink durable ParsedFile shards into the run store

Warm restore of parsedfile-cache into parsedfile-store now publishes all four shard files via fs.link, falling back to copy-into-tmp + rename so a leftover dest hardlink can never be written through. JSON remains the canonical cache; V8 sidecars ride the same path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(store): load immutable V8 shards in place, drop JSON fallback

Warm analyze was still paying JSON.parse plus a restore copy. One .v8 envelope per shard and SCHEMA_BUMP 81 make a miss re-extract instead of serving a stale JSON twin.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(store): validate durable V8 warm-cache restores

Reject incomplete or corrupt durable generations and snapshot valid shards before skipping parse workers, preserving ParsedFiles when persistence fails.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(store): drop unused durable load path

Load ParsedFiles only from the run-store snapshot and share one checksummed payload reader so inspect and deserialize stay consistent.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-30 20:50:20 +00:00
DuduPhudu
0f793558ad
fix(group)!: stop group sync claiming matching it never did (#3020)
* fix(group)!: remove the matching cascade that was advertised but never built

`gitnexus group create` wrote `matching.bm25_threshold` and
`matching.embedding_threshold` into every generated group.yaml, and no matcher
ever read either one. That was not the whole of it — an entire feature surface
described a BM25/embedding cascade that does not exist:

- `matching.bm25_threshold` / `matching.embedding_threshold` — parsed, persisted,
  unread
- `detect.embedding_fallback` — defaulted and templated, unread
- `MatchType` declared `'bm25' | 'embedding'`; both variants unreachable
- `SyncOptions.skipEmbeddings` — declared in sync.ts and never read
- `gitnexus group sync --skip-embeddings` — accepted, threaded through
  GroupService, ignored
- CLI help in en and zh-CN promised "Exact + BM25 only (no embedding fallback)"
- the MCP `group_sync` schema exposed `skipEmbeddings`, described as
  "Exact + BM25 only (Demo PR: same as default exact path)"

`sync.ts` imports exactly `buildProviderIndex`, `runExactMatch` and
`runWildcardMatch`, and the printed cascade has one stage. An operator whose
links do not match reaches for those thresholds first, and turning either knob
changes nothing — config that silently does nothing is how people conclude a
feature is broken.

Evidence that the cascade should be deleted rather than implemented, from a real
backend/frontend pair: of 165 consumer contracts, 149 link exactly and 16 do not.
Nine of the sixteen are third-party APIs (Google OAuth, Apple public keys,
PostHog, image annotation) with no in-group provider by construction — similarity
matching cannot recover them, it can only invent false links. Two are verb
mismatches: the frontend calls `POST /links` and `GET /links/check-exists` while
the backend declares `GET /links` and eleven other `/links/*` routes but neither
of those, so a fuzzy path match would link a POST consumer to a GET provider. The
rest are path-extraction artifacts. Roughly none of the sixteen would be
correctly recovered, and several would be actively mis-linked.

BREAKING CHANGE: `gitnexus group sync --skip-embeddings` and the MCP `group_sync`
`skipEmbeddings` parameter are removed. Both were accepted and ignored, so no
behavior changes — but a script passing the flag now fails with `unknown option`
instead of being silently misled. Existing group.yaml files keep loading: the
removed keys are simply no longer part of the schema, and a regression test pins
that a legacy config carrying all three still parses.

Closes #3006

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(group)!: honour --exact-only, drop inert --allow-stale, report every matching stage

Addresses the review findings on #3020, all of which are the same defect the PR
itself is about: group-sync surface that describes behaviour the pipeline does
not have.

`exactOnly` was inert in exactly the way `skipEmbeddings` was — declared on
`SyncOptions`, threaded through the CLI and the MCP tool, and read by nothing —
and strictly worse, because the stage it promised to suppress DOES run and DOES
write `matchType:'wildcard'` links into contracts.json and the bridge, which
`group impact` and cross-repo `trace` then traverse. It is now honoured rather
than deleted: unlike the never-built BM25/embedding stages, the stage it names
exists, so the flag describes a real choice. The substituted result is
`{ matched: [], remaining: unmatched }`, not an empty result — `wildcard.remaining`
IS `SyncResult.unmatched`, so skipping the stage has to leave its input unmatched
rather than dropping it from the count an operator reads.

`allowStale` had no such stage to gate: `syncGroup` emits no stale warning at any
point (the `checkStaleness` call lives in `groupStatus`, a different path), so it
is removed under the same rationale as `skipEmbeddings`.

`group sync` now prints every matching stage instead of `exact` alone. The old
block printed a `Matching cascade:` header and counted only exact links while the
next line reported `result.crossLinks.length` — which also includes `manifest` and
`wildcard` — so for any group with those the two numbers disagreed with nothing on
screen explaining why. Counting is an exhaustive `Record<MatchType, number>`, so a
new MatchType fails the build here instead of going silently uncounted, and reads
through `?? 0` so a legacy registry carrying a removed matchType prints an honest
count rather than `NaN`.

Also: the MCP `group_sync` description no longer omits the wildcard stage that
always runs, `exactOnly`'s description no longer refers to a "cascade", and
bench/cross-repo-trace/verify.mjs no longer generates the removed threshold keys
into a fresh group.yaml.

Tests: `sync-exact-only.test.ts` pins both directions of the gate (mutation-verified:
removing the gate, or returning `remaining: []`, both go red). `group-tools.test.ts`
pins that the MCP schema dropped `skipEmbeddings` and kept `exactOnly`.
`group-cli.test.ts` pins that both removed flags are rejected, with `--exact-only`
as an accepted-flag control. `config-parser.test.ts` now pins that legacy keys are
PRESERVED (measured, not assumed) rather than only that parsing does not throw.

The type narrowing's fallout in test files is cleared: `tsc -p tsconfig.test.json`
is 987 errors at head against 987 measured on origin/main, with the two error sets
identical — zero net, zero new, zero masked.

Verification: `tsc --noEmit` exit 0; prettier clean; eslint 0 errors (2 warnings,
both pre-existing on base); 69 test files / 1169 tests green across
test/unit/group, test/integration/group, tools, cli-i18n and cli-index-help.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(group): reject malformed and retired group_sync parameters (U1)

`GroupService.groupSync` read `exactOnly` off an untyped MCP payload with
`Boolean(params.exactOnly)`. While the flag was inert that coercion was
harmless; now that it gates the wildcard matching stage, the string "false" --
a routine shape for an LLM caller emitting JSON -- is truthy, so a caller that
asked to KEEP wildcard matching got it suppressed and a registry with fewer
cross-links persisted to disk. The opposite of the request, written down.

Validate instead of coercing, at the service boundary: the MCP SDK does not
enforce a tool's advertised inputSchema and `callTool` is reachable directly,
so this method is the real gate. The validator mirrors `validateImpactMode`'s
`{ value } | { error }` shape -- the established idiom for this boundary, and
the one groupSync's other guards already return through.

Also refuse `skipEmbeddings` and `allowStale` by name. The CLI rejects them
outright because commander errors on an unknown option; the MCP path accepted
and silently dropped them, so an agent working from a cached tool schema was
never told. Removing them took away discoverability, not acceptance.

Both guards run before the group is read off disk, so a rejected call performs
no work. Every test asserts the sync did NOT run -- an error string alone
cannot distinguish "refused" from "refused but synced anyway".

The tool description gains the validation note AFTER the registryOutcome
paragraph: `tools.test.ts` slices that description by ordinal position of the
'preserved' / 'superseded' / 'no-prior-registry' literals, so appending past
all three leaves those slices intact (verified, 44/44).

tsc clean; 1039/1039 group unit tests pass.

* fix(group): record the matching stages a sync was told to skip (U2)

An `--exact-only` sync wrote a contracts.json and bridge with fewer cross-links
and nothing recorded that the wildcard stage had been suppressed by request.
`group_impact` and cross-repo `trace` read that registry as authoritative, so a
narrowed graph was indistinguishable from a complete one -- and because
`group_sync` is MCP-exposed, one agent call durably narrowed the shared answer
for every later reader with no signal at all.

Add `suppressedMatchStages` to ContractRegistry and SyncResult, following the
`unreadableRepos` tri-state end to end: absent means a registry written before
the field existed, `[]` is the measurement "this run suppressed nothing", and a
populated list names the stages. The writer always emits it, because omitting
the empty case is what made "measured, none" unreachable for `unreadableRepos`.

Two properties that are easy to get backwards, and are why the split matters:

- SyncResult carries the marker on EVERY outcome. The sync genuinely did skip
  the stage whatever happened to the file afterwards, and the CLI summary (U3)
  renders from this rather than re-deriving it from the caller's options.
- The PERSISTED registry stamps it only on the `written` outcome. The preserve
  path re-writes `{ ...prior }`, so a carried-forward registry keeps the marker
  of the sync that actually produced its contracts instead of being relabelled
  with this run's request. That holds by construction: the registry literal
  carrying the field is only reachable on the written path.

`loadContractRegistryResilient` gains an explicit line, because it rebuilds the
envelope field by field with no spread of the parsed root -- a new on-disk field
is silently dropped unless named there.

Its reader is `recordedMatchStages`, not the existing `recordedRepoList`: that
one validates `string[]`, which is right for repo names and one notch too weak
here. This repo has already retired MatchType members ('bm25', 'embedding'), so
a stale value on disk is a real shape, and dropping non-members keeps an unknown
stage name from reaching a caller typed as a live one.

Surfaced on `group_contracts` and on `group_sync`'s own return -- deliberately
kept separate from the truncated/truncationReason/riskEpistemic triple. That
triple reports limits a run hit by accident, whose remedy is to fix the repo; a
suppressed stage was asked for, and its remedy is to re-sync without the flag.
Conflating them would tell an agent to retry something that returns identically.

tsc clean; 1043/1043 group unit tests pass.

* fix(group): name a skipped matching stage as skipped, and pin it (U3, U4)

Two facts were printing as the same line. `wildcard: 0 cross-links` meant both
"the stage ran and matched nothing" and "the stage never ran because you passed
--exact-only" -- the same conflation this summary block was introduced to remove
one line up, reintroduced by the flag that made the block necessary.

Render a suppressed stage as `skipped (--exact-only)`, driven by the sync's own
`suppressedMatchStages` rather than by `opts.exactOnly`. The renderer reports
what the sync did, not what the caller asked for, so it stays correct on the
outcomes where the run ended without writing a registry -- which is where the
summary is least legible and a re-derivation from the options would have been
wrong.

Also drops the `?? 0` fallback and the comment justifying it. The comment
claimed a legacy registry could carry a retired matchType into this loop. It
cannot: `syncGroup` returns a freshly computed `crossLinks` array on every
outcome, and even on the preserve path the prior links go to disk while the
fresh array is returned. The code was harmless; the stated reason was false, and
a comment that explains an unreachable path is worse than no comment.

U4 pins both halves through the CLI. A manifest fixture is sufficient: the stage
counts must sum to the total on the `Wrote contracts.json (…)` line, and the
skipped rendering does not need a stage to have matched anything, because
--exact-only records the suppression whatever the fixture holds. That is why
this coverage did not need indexed gRPC/Thrift fixture repos.

Verified by mutation, not assertion: removing the skipped-rendering branch turns
`names a stage it was told to skip as skipped` red and leaves the other 22
green. A control case pins the opposite direction -- the same group without the
flag still reports the stage as zero -- so `skipped` cannot be printed
unconditionally and pass.

U3 and U4 land together: the test has no value without the renderer, so one
commit keeps a revert clean. Both depend on U2, which introduced the field they
read.

tsc clean; 1066/1066 across the group unit and CLI integration suites.

* fix(group): make every description of --exact-only match what it does (U5)

Two descriptions this branch wrote or touched still misstated behavior.

The MCP `exactOnly` description carries "Manifest links still apply." The CLI
help and both locale strings, rewritten in the same commit, omit it -- so the
surface most operators read understated what still runs. Manifest cross-links
are computed before the gate and are genuinely unaffected by the flag, so the
caveat is the accurate half and the CLI now says it too.

The `group_sync` tool description opened with "extract HTTP contracts". That
clause was carried forward byte-identical while only the trailing cross-linking
half was rewritten, and it is wrong: the detect config has six non-HTTP
extraction toggles, and this branch's own new test fixture is Thrift.

`help-i18n.ts` is deliberately untouched. It maps an option to its translation
key and that key already exists; only the commander string and the two locale
values carry text, so a text-only change does not reach it.

The tool-description edit sits ahead of the registryOutcome paragraph, leaving
the relative order of the 'preserved' / 'superseded' / 'no-prior-registry'
literals intact -- `tools.test.ts` slices that description by their positions.

tsc clean; 64/64 across the locale-parity, help-registration, tool-schema and
group-tool suites.

* fix(group)!: remove max_candidates_per_step and shared_libs (U6)

Both keys were declared, defaulted, written into every generated group.yaml,
and read by nothing -- the same three-station dead surface this PR removed for
bm25_threshold, embedding_threshold and detect.embedding_fallback. Every other
DetectConfig field gates a real extractor in sync.ts; shared_libs gates nothing,
because 'lib' contracts come only from the operator-declared manifest extractor.
MatchingConfig reaches matching.ts solely through buildNoisyContractFilter,
which reads exclude_links_paths and exclude_links_param_only_paths and nothing
else.

Existing group.yaml files keep loading and keep their keys. parseGroupConfig
spreads the raw block over its defaults, so a key the schema no longer knows
about survives into the returned config -- which matters because `group add` and
`group remove` round-trip the operator's file through loadGroupConfig ->
yaml.dump -> write, so anything the parser dropped would be deleted from their
checked-in file. The legacy-config test now pins both keys in the same cast form
as its three siblings, and the fixture carries shared_libs so that assertion is
not vacuous.

Two stations that are easy to miss and are swept here:

- gitnexus/bench/cross-repo-trace/verify.mjs GENERATES a fresh group.yaml. It is
  not a preserve-path fixture, so "leave YAML fixtures alone" does not cover it;
  the repo has two generators and both are updated. It is a .mjs file outside
  tsconfig's include, so no type gate would have caught it.
- config-parser.test.ts asserted the removed default at runtime, which vitest
  DOES run. That assertion is gone from the defaults case (the key no longer has
  a default) and re-formed as a preserve assertion in the legacy case.

Verification gate, corrected: "zero net new errors against origin/main" would
have measured the whole branch delta and been red through no fault of this
commit. Measured instead against the branch tip immediately before it --
tsc -p tsconfig.test.json --noEmit reports 989 before and 989 after. Twenty-four
typed-literal sites across ten test files, none of them CI-gated, plus the two
runtime sites above which are.

Note the deliberate side effect: removing a key from the defaults also stops the
group add round-trip from re-adding it to a file that never carried it. Nothing
in src reads either key, so no behavior changes.

BREAKING CHANGE: `matching.max_candidates_per_step` and `detect.shared_libs` are
no longer part of the group.yaml schema and are no longer written into generated
templates. Existing files carrying them continue to parse and retain them.

src tsc clean; 1189/1189 across the group unit, group integration, locale-parity,
help-registration and tool-schema suites.

* docs(group): map PR #3020 review findings to the commits that close them

Retitles the ledger to hold one section per reviewed PR and adds #3020's ten
findings. Two things are stated rather than claimed away: `abda0d041` closes
three findings because they are one code block plus the test that pins it, and
the suppressed-stage marker is a coupled set because the renderer consumes the
field the earlier commit introduces.

Also records what is NOT closed here -- the PR description's false claim about
`max_candidates_per_step` lives outside this branch.

* refactor(group): apply simplify-pass findings

Four cleanup agents (reuse, simplification, efficiency, altitude) over this
run's diff. Efficiency was clean. The rest found five things worth fixing, two
of which were real gaps rather than style.

`recordedMatchStages` filtered unknown values instead of rejecting the list.
That inverted the tri-state on the one field built to prevent exactly this
conflation: a stale `['bm25']` -- the scenario its own comment cites as the
motivation -- survived as `[]`, which on this field MEANS "measured, nothing was
suppressed". A confident clean answer manufactured from a value we could not
read. Now all-or-nothing, matching `recordedRepoList`.

`gitnexus group contracts` showed nothing after an exact-only sync. The human
renderer destructures a fixed field list and gates its incompleteness warning on
`truncated`, so the marker reached the MCP payload and the JSON output but not
the listing an operator actually reads. It now warns, separately from the
`truncated` warning, because the remedies differ: one says fix the repo, this
one says re-run without the flag.

`verbose` was still coerced with `Boolean()` in the same call whose tool
description this branch changed to promise "PARAMETERS ARE VALIDATED". Validated
now, and added to the tool schema -- it was read by the backend and advertised
nowhere.

Reuse: the thrift wildcard-matchable pair existed twice, near-verbatim, in
`sync-exact-only` and `registry-suppressed-stages`. Both now call a shared
`makeWildcardPair` fixture, so the shape `runWildcardMatch` fires on is defined
once.

Simplification: dropped a `Set` built per sync over a list that only ever holds
zero or one entries; iterating `Object.keys(STAGE_COUNTS) as MatchType[]` also
keeps the exhaustiveness the `Record` was built for, which `Object.entries` had
discarded.

Deliberately not done, with reasons: a schema-driven unknown-parameter layer at
the MCP chokepoint (five parameters are read by backends and declared in no
schema, so a strict layer rejects working calls today, and it cannot produce the
"was removed" message finding 3 is about); folding the marker into
`GROUP_IMPACT_TRUNCATION_REASONS` (reverses a recorded plan decision and the
bridge scope is an open question for the maintainer); a per-stage suppression
cause `Record` (no second suppressor exists -- speculative); collapsing the six
`detect` extractor branches into a table (a real generalization, but a refactor
outside this diff); and converging an untouched pre-existing CLI test onto the
new manifest helper (it captures a value the helper does not return, so the
change risks more than the duplication costs).

tsc clean; eslint 0 errors (1 pre-existing warning); 1085/1085.

* docs(group): remove REVIEW-FINDINGS-MAP.md

Removes the findings-to-commits ledger from the source tree.

Note for anyone reading this in history: the file was introduced on main by
#3012 and carried that PR's findings map; this branch had appended a #3020
section. Deleting it drops both. #3012's content is recoverable with
`git show 2c0fb7753:gitnexus/src/core/group/REVIEW-FINDINGS-MAP.md`.

* fix(group): stop cross-repo impact and trace claiming a narrowed graph is complete

Closes the half of the suppressed-stage finding that was deferred. The reviewers
were right that deferring it was the weak point: the motivating harm was named
as `group_impact` and cross-repo `trace` traversing a graph missing real edges,
and those were exactly the surfaces left uncovered.

The deferral rested on an assumption that does not hold. "It is already blind to
this, so we do not make it worse" is false: `--exact-only` was inert before this
PR, so the number of narrowed registries in the world goes from zero to nonzero
exactly when this lands. The blindness was harmless only while narrowing was
impossible. And silence there is not neutral -- `cross-impact.ts` documents
`truncated: false` as an affirmative completeness claim, so those tools were
about to start asserting a complete answer over a knowingly short graph.

`suppressedMatchStages` now rides the bridge the same way `unreadableRepos`
does: persisted in meta.json (no BRIDGE_SCHEMA_VERSION bump -- meta fields have
this precedent), read back all-or-nothing, and carried across the preserve path
through `refreshPreservedBridgeMeta`'s diagnostics so a preserved bridge keeps
the marker of the sync that actually built it.

`crossRepoCompleteness` folds it in, which is what makes this one change reach
all three surfaces -- that function is by design the ONE computation behind the
truncation triple. Precedence is explicit: an unreadable or unaccounted repo
outranks a suppressed stage, because it is the more serious structural gap and
its remedy has to be the one reported.

`'suppressed-stage'` is a new member of the truncation-reason union rather than
a reuse of `'incomplete-sync'`. The earlier decision not to touch that union was
about not conflating remedies -- telling an agent to repair a repo that read
fine, for a narrowing it requested. A distinct member preserves that reasoning
while letting the answer stop claiming completeness, which is what reusing the
existing member would have destroyed.

The union's guard test did its job: adding a member failed the check that every
reason is explained on the agent-facing surface, so the impact tool description
now names this one and its distinct remedy (re-run WITHOUT the flag; nothing
failed to read).

`group status` and its CLI renderer surface it too, on the populated case only
-- absent is a registry predating the field and empty is the ordinary clean
sync; neither earns a line.

Deliberately still not done, and why: a repo-wide unknown-parameter layer for
every MCP tool. Five parameters are read by backends and declared in no schema
(`subgroupExact`, `unmatchedOnly`, `showClusters`, `showProcesses`, and
`verbose` until this branch declared it), and three tools dispatch with no
schema entry at all, so a strict layer rejects working calls until each is
reconciled. That reconciliation is the work; the layer is the cheap part. It
also cannot produce the "was removed and is no longer accepted" message the
retired-parameter guard exists to give.

tsc clean; eslint 0 errors (2 pre-existing warnings); 1159/1159 across the group
unit, group integration and tool-schema suites.

* fix(group): make the suppressed-stage signal actually reach its readers

Applies the mechanical findings from the code review of the previous commit.
That commit claimed cross-repo impact and trace stop reporting a narrowed graph
as complete. Trace did; impact did not, and two operator-facing messages said
something false. Four reviewers plus the cross-model pass converged on the same
two defects, and the untested seams were exactly where they were.

`runGroupImpact` recomputed the truncation reason and hardcoded its fallback, so
it could never emit 'suppressed-stage' -- the value the previous commit added to
the union and documented in the tool description. Every narrowed-but-readable
bridge was reported as 'incomplete-sync', telling the caller to repair a repo
that read fine. It now propagates the bridge's own reason, as cross-trace.ts
already did.

The preserve path stamped this run's request onto an older bridge. When no repo
can be read the database and registry are kept from an earlier sync, so
meta.json has to keep describing that sync; instead `{ ...existing,
...diagnostics }` overwrote its marker, leaving contracts.json, meta.json and
bridge.lbug describing three different runs. Currently masked by unreadable-repo
precedence, one loosened condition from a live wrong verdict.

`group contracts` printed "the last sync did not record which repos it could
read" after any exact-only sync: truncated was set with both repo lists empty,
so the message fell through to the wrong branch. It is now gated on the reason,
not the flag. `group impact` likewise blamed the local walk for a floor the flag
caused.

The tri-state reader is now defined once, in the leaf module whose own comment
says it exists so this exact duplication cannot recur -- it had been copied into
bridge-db.ts within one commit of that comment being true.

Both agent-facing descriptions now name the field. The previous commit added it
to three payloads and documented it on none.

Tests cover what shipped green: the preserve path for both artifacts (verified
by mutation -- reintroducing the stamp turns exactly one test red), and the
reason's REACHABILITY. The existing guard only asserted each reason is
described, which is why a documented-but-unemittable value passed it.

Also corrects a comment that said the marker is deliberately not folded into the
truncation triple. True when written; false one commit later.

tsc clean; eslint 0 errors; 1163/1163 across the group unit, group integration
and tool-schema suites.

* fix(group): drop verbose from the MCP surface, fail a superseded bridge closed

Two maintainer-directed findings from the review.

verbose is removed from the group_sync MCP schema and from GroupService, and
kept on the CLI. The parameter never did what either description claimed: the
gates emit workspace-dependency discovery stats and one aggregate manifest line,
not "each cross-link". Worse, they emit them through the server's logger, which
an MCP caller cannot read at all -- so advertising it introduced precisely the
kind of knob this PR exists to delete, in the PR that deletes them. SyncOptions
keeps the field and the CLI keeps --verbose, because a CLI user really can see
that output; its help now says "Show additional sync diagnostics", which is
what it shows. It was added to the MCP schema earlier in this same PR, so there
is no published compatibility burden in taking it back out. A caller that still
sends it is ignored rather than refused: it was never a documented parameter,
and the retired-name guard is reserved for ones this tool actually withdrew.

The second fixes a split-brain the completeness work made materially worse. When
contracts.json commits and the bridge write then fails, the previous database
stays in place describing an EARLIER sync. Until now it kept vouching for
itself, so group_impact could traverse the superseded graph and call its answer
complete while group_contracts reported the advanced registry -- two public
surfaces making contradictory epistemic claims out of one sync. That was
tolerable when the disagreement was about counts. It is not, now that
suppressed-stage makes completeness a correctness property.

markBridgeProvenanceUnknown withdraws the claim without touching the database:
bridgeMetaMatchesFile already gives provenanceUnknown highest precedence and
refuses to vouch for the pair, so cross-repo answers downgrade to a floor until
a sync succeeds. Deliberately not a re-stamp -- the metadata still describes the
database it was written for, and saying otherwise recreates the mis-pairing the
preserve path avoids. Deliberately not a delete -- the old graph is still worth
having as a floor, it just stops being called complete. Best-effort, because it
runs inside a failure handler and must not replace a reported bridge failure
with an unrelated one; the warning now states which of the two happened.

Shared registry+bridge generation identity is the architectural fix and is
deliberately NOT attempted here. This is the PR-sized containment.

Verified by mutation, both directions: neutering the withdrawal turns the new
test red, and a control pins that a healthy sync does not withdraw provenance --
otherwise every successful run would report its own answers as a floor.

tsc clean; eslint 0 errors; 1185/1185.

* refactor(group): apply simplify-pass findings

Four cleanup agents over the last five commits. Efficiency was clean and traced
why: the containment helper is failure-path only, the reason ternary sits after
the fan-out loop, and the tri-state readers run once per artifact read.

The strongest finding was one the diff itself proved. `refreshPreservedBridgeMeta`
enforced the never-persisted rule for `repoListsUnreadable` and
`pairedWithDatabase` with two deletes in its own body, under a comment noting it
was the only code that read metadata and wrote it back. That held exactly as
long as there was one such caller. `markBridgeProvenanceUnknown` made it two,
and inherited nothing. The strip now lives in `writeBridgeMeta`, so every writer
gets it and no future one can forget; `pairedWithDatabase` is the dangerous one,
because persisted it tells every later reader the pair was verified when nothing
verified it.

`group impact` still printed "fan-out stopped early" whenever `truncatedRepos`
was non-empty — but the bridge's incomplete repos are unioned into that list
even when zero crossings were attempted, so a structural gap was reported as a
runtime one, with the only working remedy omitted. That is the same false-cause
shape the contract listing was re-gated for one commit ago, left live one
command over because the new reason was bolted in front of the old branch rather
than replacing the thing it branched on. Now keyed on the reason.

The `?? 'incomplete-sync'` arm in cross-impact was unreachable: reaching it
needs `truncated` true with all three of its inputs false, which
`truncated = runtimeTruncated || bridge.truncated` forbids. Flattened.

Also: a `recordedMatchStages` insert had split `crossRepoCompleteness` from its
own JSDoc; one new test was a strict subset of another; and the bridge-failure
warning interleaved concatenation with a mid-chain ternary.

The new invariant assertion was caught being VACUOUS by mutation before it
shipped — seeded with a valid repo list, `readBridgeMeta` never sets the
reader-only field, so it passed with or without the strip. The fixture now seeds
an unreadable list, and both it and the pre-existing assertion go red when the
strip is removed.

Deliberately skipped, with reasons: a shared `firstTruncated` fold over
`TruncationFields` (the right altitude, but it changes cross-trace's return
assembly and that surface separately documents a 'timeout' rung it cannot emit —
a behavior change, not a cleanup); a reason-keyed `explainFloor` helper across
all four CLI renderers (real, but a four-site refactor); narrowing the persisted
stage vocabulary to a `SuppressibleStage` alias (would be undone by the very
extension the field was modelled as a list to allow); moving `verbose` to
`logger.debug` and deleting `SyncOptions.verbose` (the maintainer explicitly
directed keeping both); and merging the two tri-state readers behind a predicate
(they are adjacent in one file now, so a tightening applies to both by
inspection — the duplication the comment warned about was cross-FILE).

tsc clean; eslint 0 errors; 1164/1164.

* fix(group): address gitnexus-check findings

Seven bot comments across two review rounds; five distinct after dedup. Four
were valid and are fixed, two were already resolved by later commits the bot
had not seen.

The validator could throw from its own error path. `JSON.stringify` is the right
renderer there — it is what distinguishes the string "false" from the boolean,
which is the entire point of the message — but it throws on a BigInt and on a
cyclic object. So a validator promising a structured `{ error }` instead
rejected, and `callTool` is reachable directly, so neither input is
hypothetical. Guarded, keeping the distinction and falling back for the shapes
that cannot serialize.

An unreadable suppression record read as "nothing was suppressed".
`recordedMatchStages` is all-or-nothing by design, so garbage collapses to
`undefined` — and the consumer treated `undefined` as an empty measurement,
throwing that safety away and reporting a registry it could not parse as
complete. Present-but-unreadable now forces the floor, while absent stays
legitimate: a registry written before the field existed has no opinion and
should not be dragged to a floor for it.

Two test-side findings, both real and both invisible to CI because
`tsconfig.json` is src-only. Three `mock.calls[0][1]` accesses did not
type-check against a zero-arg mock, and four assertions read `truncationReason`
/ `riskEpistemic` straight off `CrossRepoCompleteness`, which is a discriminated
union carrying them on one arm. Also removed a `StoredContract` import that went
dead when those fixtures moved to `makeWildcardPair`.

Worth recording: U6 set a test-config gate at 989 errors and later commits
walked it to 994 without anyone re-measuring — the bot caught three of the five.
Now 987, below the original baseline.

Already fixed, not by this commit: the preserve-path stamp the bot flagged
against 6ceac8b1f (fixed in 1fbe0dc6b) and the displaced completeness JSDoc
(fixed in 2d2ef8c47).

Both behavior fixes are mutation-verified: restoring the unguarded stringify
turns the new unserializable-value test red, and a control pins that an absent
record still reads as complete so the fails-closed change cannot pass by forcing
every registry to a floor.

src tsc clean; eslint clean; 1167/1167.

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-27 18:27:32 +01:00
azizur100389
9d4f029001
fix(impact): mark Convex caller results incomplete (#3044)
* fix(impact): mark Convex caller results incomplete

* fix(storage): align Convex Const persistence
2026-08-26 12:54:00 +01:00
azizur100389
e87b1c3ffd
fix(php): gate imports by Composer autoload map (#2987)
* fix(php): gate imports by Composer autoload map

* fix(php): handle Composer catch-all mappings

* test(php): clarify Composer fallback coverage

* bench(php): fold Composer into canonical arm

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-25 07:56:41 +01:00
azizur100389
b77d6f662b
fix(kotlin): resolve imports from declared packages (#2990)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / Classify release event (push) Waiting to run
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Skill copy sync / shipped skills drift guard (push) Has been cancelled
2026-08-18 20:47:30 -07:00
azizur100389
87dc6c4d00
fix(go): gate imports by module path (#2984) 2026-08-18 14:31:40 +01:00
azizur100389
dac33d8056
fix(java): resolve imports from declared packages (#2955)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Resolve Java imports against parsed package declarations, expand package wildcards deterministically, and keep external imports unresolved when no in-repo package declares them.

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-15 11:53:28 +01:00
Gergő Magyar
28187bb3a7
fix(typescript): resolve imports against declared config, not path suffixes (#2953) (#2956)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(typescript): resolve imports against declared config, not path suffixes (#2953)

TypeScript/JavaScript/Vue import resolution ended in `suffixResolve`, which
answers "does any file in this repo have a path ending in this specifier?" and
answers it by dropping leading segments until something matches. That is not
module resolution, and it failed in both directions at once:

  - `@acme/telemetry/nest`, a registry dependency with no in-repo file, landed
    on the repo's only path ending in `nest/index.ts` — a false IMPORTS edge at
    confidence 1.0, indistinguishable downstream from a real one. The reporter
    measured 44 of 74 `apps/ -> packages/` edges landing on two such files.
  - `@repo/utils`, a first-party workspace package, resolved to nothing: its
    name lives in `packages/utils/package.json` and appears in no file path, so
    a path matcher cannot find it. Zero CALLS from 75 import statements.

Both come from the same missing input — nothing read the config that says what
exists — so both are fixed by reading it.

Replaces the suffix matcher on this path with the algorithm tsc and Node
actually run, in their order: relative/absolute, `#imports`, tsconfig `paths`
(longest literal prefix wins, every target tried), tsconfig `baseUrl`, then the
workspace package's own `exports`/`main`. A specifier none of those declare is
external, and resolves to nothing. There is deliberately no fallback.

New:
  - `typescript/tsconfig.ts` — every tsconfig/jsconfig in the repo with
    `extends` chains resolved, nearest-config-wins per file. The old loader read
    three filenames at the repo root, required `paths` to exist, and kept only
    `targets[0]` — none of which describes a monorepo, where `apps/web/
    tsconfig.json` is what governs `apps/web/src/main.ts`.
  - `typescript/module-resolution.ts` — the algorithm.
  - `typescript/file-candidates.ts` — 11 TS-family extensions, replacing a
    shared 39-entry list spanning every indexed language, so a TypeScript
    import can no longer resolve to a `.py` file.
  - `import-resolvers/node-workspace-packages.ts` — in-repo manifests, with
    `exports` subpath maps, patterns, condition nesting, and the restriction
    that a package declaring `exports` exposes only what it lists.

The per-pass `SuffixIndex` is gone from these three adapters: real resolution
derives nothing from the file list — every candidate comes from a declared
source and is checked with one `Set.has` — so there is nothing left to cache.
Their `*-import-index-reuse` guards and the JS index-vs-scan differential are
deleted with the mechanism they measured; the cross-language contract test
moves the three languages to its existing `KNOWN_UNINDEXED` channel, and pins
the exemption as a list so a fourth arrival is deliberate.

Python, Ruby, Java, Go and the rest still route through `suffixResolve` and are
untouched here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* test(scope-resolution): assert every resolver refuses external imports (#2953)

One property, for all 16 registered resolvers: a specifier naming something
outside the repository must not resolve to a file inside it.

That is the property #2953 was filed against, and its violation is not a
missing edge but a fabricated one — an IMPORTS edge at full confidence between
two files with no relationship, which `impact` then reports as blast radius.
The mechanism is shared (`suffixResolve`), so the guard is too.

Every case pairs an external specifier with a DECOY: an unrelated in-repo file
whose path ends the way the specifier does. Without one a resolver that merely
found nothing would pass while holding no property at all, so each case also
asserts the decoy is reachable by the spelling that SHOULD find it — a typo in
a fixture cannot manufacture a pass.

Two fixtures had to be corrected before the results meant anything, and both
would have recorded a false gap:

  - C# reads its #1881 gate from scanned namespace evidence and fails OPEN
    without any, so passing `undefined` measured nothing. Armed, C# holds.
  - C++ was posting a pass on an extension mismatch (`vector` could never match
    `src/vector.hpp` whatever the resolver did). Given the header spelling, it
    does not hold.

Result: six hold it — TypeScript, JavaScript and Vue because they resolve
against declared config only (#2953); Python (#898) and C# (#1881) because they
gate the fallback on in-repo evidence; Rust because `::` never decomposes into
a path suffix, which the decoy-reachability arm confirms is a real pass rather
than a vacuous one.

Ten do not, and are recorded in KNOWN_GAPS with what each currently answers:
Java, Kotlin, Go, Ruby, PHP, Dart, Swift, C, C++, COBOL. The map is a work
list, not an allowance — the entries are ASSERTED, so a language that starts
holding the property fails here and its line gets deleted deliberately rather
than rotting into a lie.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* fix(typescript): admit only declared workspace packages, and fix four resolver defects (#2953)

Review of #2956 found one boundary bug and four correctness defects. The
boundary one is the same defect class this PR exists to fix, arriving from a
different direction.

## The workspace boundary (review)

`loadNodeWorkspacePackages` registered every `package.json` the repo-wide scan
found, and never read `pnpm-workspace.yaml` or a root `workspaces`
declaration. Finding a manifest is not the same as the workspace admitting one:
an app importing registry package `foo` would bind to an excluded fixture or
example that happens to declare `name: "foo"` — the false-positive half of
#2953, from a new source of evidence. This repository is the example, since
`test/fixtures/**` declares `@repo/utils` among others.

The admitted set now comes from the declaration — `workspaces` (array and yarn
object form), `pnpm-workspace.yaml`, `lerna.json`, with `!` exclusions and
`*`/`**` — plus the root package itself. A repo that declares no workspace has
exactly one package: the root. A negative fixture pins it, with a named package
outside the declared globs that must not resolve.

## Four defects

  - tsconfig `paths` targets were resolved against the config's own directory
    when it declared `paths` but inherited `baseUrl`. tsc resolves them against
    the EFFECTIVE base, so an extending config loaded the right alias pattern
    and pointed every target at the wrong directory.
  - two configs in one directory were ranked by directory-listing order, so
    `tsconfig.base.json` could govern instead of `tsconfig.json` and a config's
    own `paths` went invisible. Found by the test written for the fix above.
  - an unexported package subpath also tried `<dir>/src/<subpath>`. Nothing
    declares that mapping; it is the same kind of guess this PR removes, and
    the import it "resolved" is broken in the real project too.
  - `imports` pattern keys (`"#internal/*"`) were looked up exactly, so a valid
    `#internal/foo` never matched. `exports` and `imports` now share one
    matcher, which is where they should never have diverged.
  - a relative specifier climbing past the repo root was silently clamped, so
    `../../../secret` from `src/main.ts` became `secret` and could resolve a
    root file it never named.

## Test rigor

The conformance suite asserted less than it claimed. The decoy-reachability arm
only checked non-empty, so five cases paired `reachesDecoy` with a different
file than `decoy` and passed while establishing nothing; the KNOWN_GAPS arm
likewise accepted any in-repo answer instead of the recorded one. Both now
assert the exact file. The reachability arm runs only for languages that HOLD
the property — for a gap language the recorded-answer assertion IS that proof,
and for Swift and COBOL no other spelling exists, since `Foundation` and
`EXTERNAL` name the in-repo directory and copybook as well as the external
module, which is precisely why those resolvers cannot tell them apart.

## Benchmarks

Both `--check` guards were red, and both were reporting something true.

`import-target`: the ts-family arms resolved 0 of 3200 imports. Their corpus is
bare specifiers with no config, which the deleted `suffixResolve` answered
without one — so the arms measured an empty branch while printing a clean
scaling ratio. Each now carries the config its corpus is spelled for, and the
`deep` arm's uniform prefix reaches it. THE FINGERPRINTS THEN MATCHED THE
RECORDED BASELINES EXACTLY: same corpus, same targets, once the config it
always implied is passed explicitly. Retained per-pass index went from
26 745 296 B (js, ts) and 28 884 016 B (vue) at 32 000 files to 0-16 B, because
these resolvers no longer build one; they move to the `HEAP_BOUNDED` tier rust
already occupies for the same reason. Depth ratio moved 2.0 -> ~2.2 and the
budget goes to 2.6: candidates now carry the 16-segment baseUrl prefix, so each
`Set.has` hashes a longer string — linear in path LENGTH, independent of file
COUNT.

`scope-capture`: TypeScript capture fingerprint drift, caused by this PR's 12
new `.ts` fixtures entering the corpus. Attribution is exact rather than
inferred — moving that one fixture directory aside returns the fingerprint to
`f719163e…` byte-for-byte with `fixture_count` back at 155 and all 15 languages
passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* fix(typescript): honour exports fallback arrays, paths precedence and package extends (#2953)

Second review round. Four findings, judged against what this tool is: a static
analyser building a code graph, not a compiler. The bar is resolving what the
project DECLARES, on a checkout that may never have been built or installed,
and never inventing an edge.

  - `exports` and `imports` ARRAYS were skipped. An array is Node's ordered
    fallback list, and `{"./feature": ["./dist/feature.js", "./src/feature.ts"]}`
    is exactly what a workspace package publishes to mean "built output, or
    source". Skipping it dropped the declaration entirely and left the package
    looking as though it exported no subpaths. The source arm is the one that
    matters here, because `dist/` is build output and is not indexed — and for
    a static analyser the build need not have run at all.
  - an exact `paths` pattern did not reliably outrank a wildcard. `a` and `a*`
    both match `a` with the same literal prefix length, so sorting on length
    alone left tsc's exact-wins rule to declaration order.
  - package-form `extends` (`"@acme/tsconfig"`) was refused outright. Not
    indexing `node_modules` is different from not READING it, and a shared
    internal base is where a monorepo puts the `paths` its packages import
    through. It is now read from disk, walking `node_modules` up from the
    extending config the way Node does, and absent on an un-installed checkout
    it degrades to whatever that config declared itself.

    The test pins what tsc actually does with such a base rather than what one
    might hope: `extends` never rebases `baseUrl`, so a package base's paths
    point at the package's own directory. That is why a published base rarely
    contributes aliases a repo's files resolve through, and why the
    `@tsconfig/*` family — which sets `target` and `lib`, never `paths` — is a
    no-op here either way.
  - CodeQL flagged `String.replace('*', …)` in two places as replacing only the
    first occurrence. Node subpath patterns and tsconfig `paths` both allow AT
    MOST one `*`, so that IS the specified behaviour — but the spelling states
    it by accident and reads as the replace-all footgun. `substituteStar`
    slices at the known index and says the rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* fix(typescript): treat `exports` as the whole interface, and keep empty tsconfig scopes (#2953)

Third review round. Two findings, both valid, both cases of this resolver being
laxer than the thing it models — which is the direction that fabricates edges.

  - `exports`, when a manifest declares it, is the package's ENTIRE public
    interface: Node ignores `main` outright and refuses any subpath the map
    does not list. This resolver already honoured that restriction for
    SUBPATHS and not for the package ROOT, which is the same rule. A manifest
    exporting only `"./feature"` therefore still answered a bare `@repo/pkg`
    with `main` or `src/index` — an edge for an import that does not resolve in
    the real project. Legacy and conventional root candidates are now offered
    only when there is no `exports` field at all.

  - a tsconfig declaring neither `baseUrl` nor `paths` was dropped rather than
    kept as an empty scope, so `tsconfigFor` fell through to an enclosing
    config. A package whose own tsconfig declares no `baseUrl` — meaning its
    non-relative specifiers are package lookups — silently inherited the repo
    root's aliases instead. An empty scope is the accurate answer for such a
    file, and only a scope can express it.

Both are pinned at the level they broke: the manifest arms assert what
`readManifest` produces, not a hand-built package, since the resolver honouring
empty entries and the loader producing them are different claims.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 10:13:04 +01:00
Gergő Magyar
77360e1043
fix(scope-resolution): make interface dispatch generic-instantiation aware (#2912) (#2939)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(scope-resolution): make interface dispatch generic-instantiation aware (#2912)

Interface-dispatch fan-out walked the subtype closure with generic arguments
erased, so `IValidator<string>` and `IValidator<int>` — one declaration, one
subtype list — were indistinguishable and a call through the first reached
`IntValidator.Check(int)`, a target no runtime dispatch can produce.

The arguments were already in the capture, unread: every language anchors
`@reference.inherits` on the whole base node while `@reference.name` keeps the
erased base. `ReferenceSite.typeArguments` is therefore derived generically in
`scope-extractor.ts` from the anchor's own spelling — no per-language query
changed — covering C#, Java, TypeScript, Kotlin, Go (`Base[int]` embedding),
Python (`Base[User]`) and Swift; Rust and Dart anchor on the bare name and get
nothing, which reads as "unknown". `preEmitInheritanceEdges` is the only code
that pairs a heritage site with a resolved (subtype, supertype), so it records
the instantiation there and hands it to the dispatch pass.

The closure is then walked carrying a substitution, as a type checker would:
`Wrapper<T> : IValidator<T>` binds T to the receiver's argument and stays
reachable from every instantiation, while its own subtypes are matched against
that binding. An incompatible hop is skipped without being marked seen, so a
type reachable by a second, compatible path still gets its edge, and without
descending, since its subtypes inherit the mismatch.

Pruning happens only on positive evidence that two instantiations differ.
Unknown arguments on either side, an arity that does not line up, an unresolved
qualified spelling of the same simple name, or an argument that might be a type
variable the language never captured all keep the target. Telling an uncaptured
type VARIABLE from a concrete type is the crux: `typeParameters` is absent both
for a non-generic declaration and for every declaration in a language whose
query omits `@declaration.type-parameters`, so the pass reads the evidence in
front of it — one run resolves one language, so a single generic declaration
anywhere in it proves the captures record parameters. A language recording
neither arguments nor parameters keeps exactly its pre-#2912 fan-out.

Type arguments are compared as resolved declarations rather than spellings, so
`Models.User` and an imported `User` are one type; the new optional
`ScopeResolver.normalizeTypeArgument` hook canonicalizes a language's predefined
aliases, implemented for C# (`string` ≡ `String`) where mixing the spellings
would otherwise delete a real implementor.

Fan-out cap, skipped-target reporting, overload selection and non-generic
closure behaviour are unchanged. SCHEMA_BUMP 60 -> 64: the heritage arguments
are a parse-time capture, so a warm cache would replay pre-fix sites and leave
the filter silently inert on unchanged files (61/62/63 are claimed by open PRs).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StNKYi7Qxv5DnSURZuFBef

* fix(scope-resolution): close the two generic-dispatch gaps (#2912)

The first commit left two shapes on the pre-#2912 fan-out. Both are now
covered, and the second one turned out to need a route the pipeline did not
have at all.

**Folded receivers (Cases 0 and 3b).** `this._validator.Check(x)` is typed by
the compound fold, and the fold answers with a CLASS — which is exactly what
loses the instantiation, since `IValidator<string>` and `IValidator<int>` fold
to one declaration. The fold now reports the SPELLING it typed each receiver
position from, through a pure side channel (`recordReceiverType`) added to the
one helper every declared-type route already shares plus the two return-type
routes; resolution is unchanged whether or not a caller passes it. The reader
keeps the last report and uses it only when it names the class the fold
returned, so an intermediate position cannot lend its arguments to another
class. This covers the dependency-injection shape the issue is really about —
a field-held generic interface — and multi-hop chains, where it is the last
hop's spelling that types the receiver.

**Rust and Dart heritage.** Neither recorded arguments, for two different
reasons, so both routes exist now:

  - Rust's `@reference.inherits` anchor is the trait identifier INSIDE a
    `generic_type`. Widening the anchor would move the site's range, and that
    range is part of every inheritance edge's id, so the arguments arrive
    through a new `@reference.type-arguments` sub-tag instead.
  - Dart's `implements` / `with` never become reference sites at all: they
    travel as heritage MARKERS and their edges are emitted by the language
    hook. The arguments ride the marker payload as an optional fourth field
    (dropped, not encoded, when the spelling contains the marker delimiter),
    and `ScopeResolver.emitHeritageEdges` now receives the same sink
    `preEmitInheritanceEdges` writes to, so whichever pass emits an edge
    records that edge's instantiation. Dart also gained the
    `@declaration.type-parameters` capture, without which its own type
    VARIABLES are indistinguishable from concrete arguments and
    `class Box<T> implements Validator<T>` would be pruned from every
    instantiation.

Note this makes Rust and Dart record their instantiations; it does not make
them fan out. Interface dispatch still fires only for a receiver whose folded
type is an `Interface` symbol, so a Rust `Trait` or a Dart abstract `Class`
receiver has no secondary targets to filter. Widening that gate emits new
edges for several languages and belongs to its own issue.

**Two matcher rules the wider coverage exposed.** A WILDCARD names a set of
types rather than one — `Repo<? extends User>` holds a `Repo<User>`, and
Kotlin's `Repo<*>` / `Repo<out User>` say the same — so a position with one on
either side is unknown; nullable spellings trip the same test, which costs a
little precision in the safe direction. And insignificant whitespace inside a
nested spelling (`Map<string, User>` vs `Map<string,User>`) is no longer a
difference.

One expectation changed in the #2833 field-receiver matrix: a
`Repo<Repo<User>>` receiver no longer reaches `UserRepo implements Repo<User>`.
That edge is precisely the false positive this issue is about, and the primary
edge to the interface's own declaration — which is what the matrix row exists
to prove — is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StNKYi7Qxv5DnSURZuFBef

* refactor(scope-resolution): apply the quality pass to the #2912 change

Three cleanups, no behaviour change.

**One balanced-list scanner, not two.** `erasedTypeApplication` and
`typeApplicationArguments` each carried a copy of the same fiddly scan — one
bracket list, balanced, closing on the last character, non-empty — differing
only in what they did with the result. Both now call `balancedTailList`; the
rule that rejects `User[][]` and `Repo<User>?` lives in one place instead of
being free to drift between two.

**The receiver's arguments are parsed after the gates, not before them.**
`emitInterfaceDispatchFor` takes the receiver's declared SPELLING and parses it
itself, once the owner is known to be an Interface with subtypes. Every one of
the five cases calls it unconditionally and the overwhelming majority of
receivers are concrete classes that return at the first line, so the parse was
running per resolved receiver site to be discarded immediately. Case 4 and Case
6 now hand over the string they already hold, and the folded-receiver helper
returns the recorded spelling rather than parsing it.

**One question gates the whole instantiation apparatus.** Inside the closure
walk, the graph-id lookups now hang off "is the supertype's instantiation
known?" — false for every non-generic receiver and for every language that
captures no heritage arguments, which is what makes those walks cost exactly
what they cost before #2912.

Also lifted the argument-route choice in `pass5CollectReferences` out of a
nested ternary into a named `heritageTypeArguments`, where the reason the
explicit sub-tag wins over the anchor text can be stated once.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StNKYi7Qxv5DnSURZuFBef

* test(scope-resolution): cover generic interface dispatch in Kotlin and Go (#2912)

Extends the #2912 dispatch coverage past C#/Java/TypeScript. No production
code changes — the derivation is language-agnostic by construction
(`heritageTypeArguments` reads the heritage anchor's own spelling), so the
question was only which languages actually reach the filter.

Kotlin rides the shared heritage pre-pass; Go reaches the same filter from
the other side, matching implementors structurally while the receiver's
`Validator[string]` spelling carries the instantiation. Both are confirmed
to prune the mismatched implementor.

Each language gets a NON-GENERIC control asserting the fan-out still reaches
every implementor. Without it the `not.toContain` assertion passes just as
well when a language emits no dispatch edge at all — which is what Dart,
Python and Rust were measured doing for this receiver shape, generic or not.
They are deliberately not asserted on here: a "filtered correctly" test over
a path that never fans out measures nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* test(bench): re-baseline the Rust and Dart capture fingerprints for #2912

The Rust trait-impl and Dart heritage capture changes this branch makes are
additive TEXT on existing matches — each carries the instantiation the clause
was written with — so they drift the scope-capture digest without adding or
removing a match. The baselines were never re-measured when those captures
landed, which left `measure.mjs --check` red on this branch independently of
the merge.

Re-measured rather than hand-edited. Rust's capture_groups_fp (3556) and
fixture_count (202) are unchanged across the move, which is the evidence that
this is digest drift and not a capture-set regression. The other 13 languages
are byte-identical; 15/15 pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* refactor(scope-resolution): quality pass over the #2912 change

Cleanup only — no behavior change. Findings from a four-angle review (reuse,
simplification, efficiency, altitude), applied where they were verified.

Reuse / duplication:
* `stripTrailingCallSuffix` was a second copy of `matchingOpenParen`'s backward
  balanced-paren scan. Both now live in `template-arguments.ts` beside
  `balancedTailList`, for the reason that helper was shared in the first place:
  two copies of a scan this fiddly are free to disagree.
* The two call-return arms of the compound fold repeated the same four-part
  expression character for character; they share `classOfReturnType` now, the
  return-type twin of `classOfDeclaredType`, which keeps the "look up by
  rawName, report the erased application" pairing in one place.
* `pipeline/run.ts` implemented first-writer-wins twice — once in the pre-pass
  and once in the provider sink. One store, one sink, one rule; the pass keeps
  its `Set<string>` return and the callable-flow-only arm stops building an
  empty map to satisfy a widened return shape.

Simplification:
* `subtypeParametersComplete` dropped a disjunct that could never decide: every
  `subDef` reaching it comes out of the same loop that sets
  `languageCapturesTypeParameters`, from exactly those defs.
* The heritage-argument lookup asked "is the supertype's instantiation known?"
  three times; `superGraphId` now gates the block once.
* `TypeArgumentResolver` and `HeritageInstantiationResult` un-exported — no
  consumer outside their module.

Efficiency (all on the per-site dispatch walk):
* `resolveSupertypeArgument` captures only the site, so it is built once per
  site instead of once per subtype visited; the subtype's scope id is looked up
  once per subtype instead of once per argument position.
* `erasedTypeApplication` no longer runs on every fold hop through a call — the
  spelling is built only once the lookup has found a class, since it is
  discarded otherwise.
* `normalize`+`compact` computed once per side rather than twice.
* Regex literals and the identity `normalize` fallback hoisted to module scope.
* C# `System.` prefix stripped with `startsWith`/`slice` instead of a regex.

Verified: tsc clean, build clean, 1994 scope-resolution unit tests, 171
generic-dispatch + generic-field-receiver integration tests, 15/15 capture
bench fingerprints unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* style: apply Prettier to the two files the quality pass reformatted

Whitespace only — `quality / format` (npx prettier --check .) was red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* fix(scope-resolution): close the generic-dispatch review findings (#2912)

Addresses the gitnexus-check review on #2939.

A repeated type variable was rebound rather than unified: `class C<T> :
Pair<T, T>` accepted a `Pair<string, int>` receiver, with `T = int` silently
replacing `T = string` and the bogus substitution carried to the next hop.
It now unifies, and prunes only on the same positive evidence the concrete
path demands — an undecidable repeat keeps the target with no binding.

A type PARAMETER of the declaration enclosing either side is now recognised
and never compared. `subtypeParametersComplete` is evidence about the
SUBTYPE's parameter list and says nothing about a `T` written at the call
site, so `void Run<T>(IValidator<T> v) { v.Check(x); }` pruned every
implementor: unbounded, `T` grounds to nothing; bounded, it grounds to its
BOUND. Both read as a difference of type. That is the missing-edge failure
this filter is built to avoid, and it is the common dependency-injection
shape in C#, Java and Kotlin.

Making that recognition reliable is why generic METHODS now capture
`@declaration.type-parameters` in C#, Java and Kotlin — TypeScript already
did, which is why its generic functions never had the defect. The capture
feeds the existing `bindsTypeParameter` guard, so a method-level `T` also
stops resolving to a same-named class in every other lookup.

C# alias normalization additionally strips the `global::` qualifier, which
`import-decomposer` already unwraps elsewhere: `global::System.String` read
as unequal to `string` and pruned a live implementor.

The C# captures golden fixture is regenerated for the new capture; the
extractor reads `@declaration.type-parameters` generically, so no reader
changed. SCHEMA_BUMP 64 already covers these capture changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* fix(scope-resolution): close the two remaining gitnexus-check findings (#2912)

`balancedTailList` counted ONE bracket family, so a crossed pair slipped
through: scanning `Foo<Bar]>` it never sees the `]`, reaches the final `>` at
depth zero, and reports `Bar]` as a balanced argument list — which
`typeApplicationArguments` then splits and `erasedTypeApplication` rebuilds a
spelling from. It now tracks a stack of expected closers, so every closer must
match the opener it actually closes and a crossed pair declines to `undefined`,
the "unknown" both callers already fail open on. Well-formed mixed nesting
(`List<Dict[a, b]>`) is unaffected.

C# `normalizeTypeArgument` stripped `System.` from every qualified spelling, so
`System.Custom` answered `Custom` and compared equal to an unrelated `Custom`
elsewhere in the workspace. The strip is now earned: a keyword answers from the
alias table first, and the qualifier is dropped only when what remains IS a
predefined type. `System.Custom` is returned as written and goes to the identity
comparison instead — the step that can actually tell two declarations apart.
`global::System.String` still meets `string`.

Both are pinned by unit tests, including the well-formed mixed nesting and the
`global::`-qualified ordinary type that must keep its qualifier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* docs(csharp): record why a shadowed `String` keeps its implementor (#2912)

Answers a review finding rather than changing behavior.

A workspace may declare its own type named `String`, shadowing the BCL simple
name, and the alias table then reads `IValidator<String>` as the `string`
instantiation and keeps that implementor. That is the SAFE direction, not an
oversight: pruning instead would rest on the belief that two spellings differ,
which is the missing-edge failure `generic-instantiation.ts` exists to avoid.

Resolving rather than normalizing cannot settle it either — the identity
comparison needs a `definitionId` from both sides, and a built-in name carries
none, so "built-in versus workspace-declared implies different" would be a new
prune with no positive evidence behind it. The cost is one surplus edge for that
pair, which is exactly the pre-#2912 fan-out and no worse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* refactor(scope-resolution): pair the receiver spelling with the class structurally (#2912)

The fan-out needs the spelling a receiver position was typed from, because the
class the fold returns has lost the generic arguments. That was carried by a
PASS-LEVEL mutable holder, written by every declared-type lookup anywhere in the
fold and read back through a def-id coincidence check, with the holder cleared
by hand before each call site. Three things were load-bearing and none were
enforced:

* the reset had to be remembered at every call site. It was not: the Case 3b
  retry (`rawName` then `rawName + '()'`) reset once, BEFORE the first attempt,
  so a spelling reported by the attempt that failed could be attributed to the
  one that succeeded.
* the holder outlived every resolution, so a site that resolved through a route
  reporting nothing could read the previous site's spelling if the def ids
  happened to line up.
* the pairing itself was inferred from "whichever lookup reported last", not
  from the fold's own bookkeeping — losing branches (an MRO walk that moved on,
  a step later folded past) report too.

`foldReceiverChain` already had the answer and threw it away: its final
`FoldState` holds `def` and `declaredType` produced by the SAME step. It now
reports that pairing last, so the structural route is the one that stands.

`resolveCompoundReceiverTyped` returns `{def, declaredSpelling}` and owns a sink
created and read within the single call, which is what removes the reset
discipline — a local cannot be forgotten, and each of the two retry attempts
carries its own. The def-id guard stays as the check that a report names the
class actually returned.

Behavior is unchanged: 1975 scope-resolution unit tests, 177 generic-dispatch
and generic-field-receiver integration tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 14:50:53 +01:00
azizur100389
3d4a95360d
fix(java): materialize record component accessors (#2936)
* fix(java): materialize record component accessors

* fix(java): ignore receiver params in record accessor arity

* fix(java): address record-accessor review findings (#2917)

Five findings from the tri-review of #2936.

P1 — a synthesized callable evicted a source-written one from the method map.
`getMethodInfo` keyed its per-class map by `name:line`, but a callable that is
SYNTHESIZED at a position that is not its own declaration shares its owner's
line: a record's implicit accessor is minted at the component, and a C# 12
primary constructor at the owner's `parameter_list`. Both are appended last by
their extractor, so on a single line the synthesized entry overwrote the
explicit method's MethodInfo and both definitions collapsed onto one id —
`record P(int x, int y) { int x(int s) {...} }` lost `P.x#1` and rebound the
arity-1 call to the zero-argument accessor. Adds a required `MethodInfo.column`
and keys the map by `name:line:column` through a single `methodInfoKey` helper.
Required, not optional: an absent column would key an entry no lookup could
reach — a silent, whole-language loss of enrichment instead of a compile error.
All three lookup sites move together; the file's own lockstep docblock warns
that a half-applied change loses caller edges silently rather than dangling.
This also fixes the same collision in C#, which never touched record code.

Degenerate component names no longer mint a node. tree-sitter's zero-width
MISSING recovery token satisfies `name: (identifier)`, so `record M(int x, y) {}`
minted an empty-named Method whose returnType was the neighbouring `y`; and the
grammar admits `underscore_pattern` in the same field, which the query rejected
but the scope path accepted, so `record R(int _) {}` left a scope declaration
with no node behind it. One `isRecordComponentName` predicate now gates all
three emitters — query suppression, scope synthesis, and the method extractor —
so they cannot drift apart again.

Component annotations reach the implicit accessor (JLS 8.10.3 / 9.7.4) by
reusing the shared `extractAnnotations` helper. Deliberately over-approximate
and commented as such: `@Target` lives in another file and parsing is per-file.

`explicitZeroArgAccessorNames` is memoised per record node. It was rebuilt on
every component capture — O(components x body members) for one record, measured
at ~4x per 2x input — while the scope path already hoisted the identical call.

Docs: the `java-local-types` baseline now stores the `capture_groups_fp` its own
note cites, the SCHEMA_BUMP ledger no longer claims a v65 that nothing holds,
and `shouldSkipDefinitionCapture` documents that `defaultLabel` may be ignored.

Scope-capture fingerprints are unchanged (`measure.mjs --check` PASS, 15
languages): the bench corpus contains no degenerate components, so the new
predicate is inert on it. SCHEMA_BUMP stays 67 — this branch's existing claim
already covers the changed worker output; re-check it against origin/main before
merging.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

* docs(ingestion): reunite the overload-suffix JSDoc with typeTagForId

The block describing the `~type1,type2` same-arity discriminator was stranded
above `buildCollisionGroups` when that function was inserted between it and the
`typeTagForId` it documents (#658). Adding `methodInfoKey` in this branch parked
it directly above yet another unrelated function, which gitnexus-check flagged.

Moves the comment down to the function it describes. No behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Px638Zyqa9CJMUU7DsJoB

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 08:43:11 +00:00
azizur100389
cdc98a9cf8
fix(java): capture enum interface heritage (#2935)
* fix(java): capture enum interface heritage

* fix(java): harden enum heritage dispatch

* test(java): refresh synthetic capture baselines

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-13 07:04:00 +01:00
Gergő Magyar
d540b00184
fix(check): stop reporting erased and deferred imports as initialization cycles (#2934)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Skill copy sync / shipped skills drift guard (push) Has been cancelled
2026-08-12 17:09:32 +00:00
Gergő Magyar
054641cafa
fix(scope-resolution): resolve a package whose directory name repeats higher in the path (#2881) (#2929)
* fix(kotlin): resolve a root-level package whose name repeats higher in the path

`getKotlinFileIndex` built its `dirChildren` buckets under two guards
inherited from the pre-index per-import scan rather than from anything
Kotlin requires: a `startsWith` test that skipped the bucket when the
path began with the package name, and an `indexOf` equality that
demanded the parent be the FIRST occurrence of `/<name>/` in the path.

`s` is taken as `dir.slice(i + 1)` at each `/`, so `dir` ends with `/s`
by construction and the file always IS a direct child of a directory
named `s`. The guards therefore dropped legitimate buckets:

  data/src/main/kotlin/com/example/data/Repo.kt   (leading, startsWith)
  top/data/mid/data/Repo.kt                       (mid-path, indexOf)

`import data.helper` resolved to null against both. Only the fan-out
tier was affected — `data.Repo` answers from `suffixByStem`, which
carries no such guard — which is why the shape looked narrow enough for
#2872 to preserve rather than change inside a performance PR.

Both guards are removed. The rule stays "the parent directory is named
`s`" — a name that appears in the path without being the parent
(`top/data/mid/Repo.kt` for `data.something`) is still not a child, and
a new case pins that.

Widening is filtered downstream for the fan-out tier, which hands the
finalize pass a candidate list (#1759), but NOT for the tier-1 fallback,
which commits to `children[0]` unfiltered — and that is where most of
the change lands: 149 of the 235 moved corpus records are a different
first child against 32 wider arrays. Both are deliberate. A narrower
bucket for the first-child tier alone would keep its answers identical
and would also leave `import data.*` — a wildcard, which strips to
`data` and lands on exactly that tier — resolving to null on the very
shape this fixes.

Both Kotlin benches are re-baselined deliberately, with the drift
measured rather than accepted:

  - bench/kotlin-import-target: 235 of 19968 distinct records moved.
    54 null -> resolved (the fix, and exactly the +54 in non_null),
    181 answers that changed within a now-larger bucket. Zero buckets
    lost a member, zero results were dropped, and every reselected
    answer's parent directory is the queried package segment. The
    corpus is untouched, so `cases` is unchanged and the fingerprint
    covers the same surface as the value it replaces.

  - bench/import-target: the collide arm needed a corpus edit beside
    the new numbers. Its `d % 7` slice imported `com.example.vendor{d}`,
    a package that exists nowhere, purely to mirror the unique arm's
    nested-slice MISS; with that slice now resolving, leaving it would
    have left collide at 1100 against small's 1153 and broken the
    same-workload invariant the arm is built on. That assertion is what
    caught it.

The gate controls were re-run against the new baseline, including one
the fix makes newly plausible: a HALF fix that drops only `startsWith`
and keeps the `indexOf` check still fails the fingerprint, so a partial
fix cannot land quietly.

Two gates moved with the code rather than being left behind:

  - kotlin `heap_reading_bytes` and `heap_ceiling_bytes` are re-recorded
    together as `_heap_reading_note` requires (48073096 -> 48200224,
    +0.264%, ceiling still 1.5x). The note says why that is small: the
    heap corpus is built with HEAP_PAD 8, so no path can begin with a
    suffix of its own directory and the leading-segment half of the old
    rule is invisible to that arm.

  - `depth_budget` 2.4 -> 2.2. Deleting two string comparisons per
    directory component is per-depth work, so the depth band fell from
    1.44-1.51 to 1.27-1.40; left at 2.4 the gate's headroom would have
    drifted from ~1.6x to ~1.8x without anyone deciding to loosen it.

`package-dir-index.ts` documents the same first-occurrence rule as
universal, and it is not any more: Go, Java and C# still carry it and
still have the shape. Fixing them means re-baselining three languages
and editing the verbatim pre-change scans that
import-target-index-parity.test.ts keeps as the specification, so it is
a separate change — the comment now says so instead of describing a rule
one of its readers no longer follows.

Fixes #2881.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e

* fix(scope-resolution): drop the first-occurrence directory rule for Java, Go and C# too

#2881 was reported against Kotlin, but the rule it removed was never
Kotlin's. It is what the pre-index per-import scan happened to compute —
`indexOf` for the package directory, then "nothing after the match holds
a slash" — and every resolver built to reproduce that scan inherited it.
Three still had it, and all three reproduced the reported defect:

  java   data/src/main/java/com/example/data/Repo.java  `import data.*`  -> null
  java   top/data/mid/data/Repo.java                    `import data.*`  -> null
  csharp Models/src/App/Models/User.cs                  `using Models;`  -> null
  csharp a/Models/b/Models/User.cs                      `using Models;`  -> null
  go     a/internal/auth/b/internal/auth/svc.go   import "internal/auth" -> null

Controls (`top/data/Repo.java`, `a/b/internal/auth/svc.go`) resolve, so
these are the rule firing rather than an unrelated miss.

Four sites, all reduced to "the file's parent directory ends with the
queried path":

  - `package-dir-index.ts` `matchingDirs` (Go, Java, C# without csproj):
    the `indexOf` equality becomes `endsWith`, which also subsumes the
    length guard it needed — a shorter haystack is false instead of
    comparing -1 to -1.
  - `csharp.ts` `matchingDirPositions` (csproj step 3): same, and still
    deliberately UNANCHORED, so `src/SubModels` keeps answering `Models`.
  - `csharp.ts` csproj step 2: `indexOf` -> `lastIndexOf`, EXCEPT for an
    empty `dirPrefix`, which must keep `indexOf`. Its needle is a bare
    '/', and step 3 answers that query from `singleSegmentDirs` ("exactly
    one directory deep"), which only the first occurrence expresses; with
    `lastIndexOf` there, step 2 accepts every `.cs` in any directory and
    diverges from step 3. The csproj parity test catches it.
  - `go.ts` `resolveGoPackage`: `indexOf` -> `lastIndexOf`. No production
    caller, but the parity harness copies it verbatim as its spec.

The two C# csproj sites must move together. Fixing only step 3 makes
`Lib.Models` return step 3's superset instead of step 2's segment-aligned
answer.

Risk is not symmetric across the three. Go's consumer is a fan-out list
and the finalize pass materializes one IMPORTS edge per element, so
widening only ADDS edges. Java and C#-without-csproj commit to a single
file through `firstFileDirectlyInPkgDir` with no downstream filter, so a
widened bucket can also change which file an already-resolving import
binds to — java's collide fingerprints moved while its resolved count
did not, which is exactly that. C#'s leg is additionally gated by
`csharpSuffixFallbackAllowed` (#1881) before resolution runs.

Gates:

  - Twenty fingerprints re-baselined across go, csharp and java (five
    arms plus the top-level alias each). resolved 979 -> 1153 small,
    4064 -> 4681 large for go and csharp; 1100 -> 1153 / 4456 -> 4681 for
    java. No `distinct_outcomes` moved.

  - csharp and java hit the same collide-arm trap Kotlin did: both sent
    their `d % 7` slice to a namespace that exists nowhere purely to
    mirror the unique arm's nested-slice MISS, so once that became a hit
    the arms resolved fewer imports than `small` and the same-workload
    assertion failed. Both now use their arm's ordinary spelling.

  - GO WAS NOT GATED AT ALL and the corpus had to change to make it so.
    Its nested slice repeated only the last segment (`src/pkg{d}/internal/
    pkg{d}`) while a Go query addresses the whole package path, so the
    directory never ended with the query and the rule was never reached —
    every go arm sat unchanged through the resolver fix. `uniqueDir` and
    `collideDir` now repeat the shape at the granularity Go queries.
    `languages.go.heap.path_segments` 13 -> 14 follows from that.

  - `csharp_csproj`'s heap reading moved -0.79% (stable across runs) and
    is re-recorded with its ceiling: the step-2 filter decides which lazy
    `getFilesInDir` maps the probe forces. Everything else stayed within
    +/-0.03%, which is this box's jitter — `_heap_reading_note`'s claim
    that the readings reproduce to the byte across processes did not hold
    here, and the note now says so.

The three parity harnesses keep VERBATIM copies of the pre-change scans
as their specification, so each copy was updated with the resolver and
the cases that pinned the rule now pin its removal. Two of them left the
`mustBeNull` set in the shared harness — they resolve now, which holds
them to the stronger "pin a winner" bar the rest of that arm uses.

Refs #2881.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e

* perf(kotlin): intern dirChildren keys per directory and compact the buckets

Two optimizations to `getKotlinFileIndex`, both output-identical, kept
because they were measured and a third was dropped because it was not.

1. PER-DIRECTORY KEY MEMO. The component walk over `dir` cut one `slice`
   per component per FILE, and every slice after the first file of a
   directory is a freshly allocated string that hashes to a key the map
   already holds and is then dropped. The key list is a pure function of
   `dir`, so it is interned once per DIRECTORY. Measured -18.4% to -21.7%
   of the build at 32 000 files; zero retained cost, the memo dies with
   the frame.

2. BUCKET COMPACTION at the freeze loop. `addChild` mints `[raw]` and
   pushes, and V8 grows a backing store by `old + old/2 + 16`, so the
   SECOND child takes a 1-slot store to 17 and every bucket then retains
   its overshoot. 61 144 buckets at 32 000 files, 52.9% of their slots
   empty, 88 B each. `slice()` on freeze: -5 397 768 B, -11.20%, and the
   predicted 5 382 507 B lands within 0.03% of it. Same fix and the same
   accounting as the python `byBasename` note this repo already carries.

   `length === 1` is skipped deliberately. A bucket that never grew is
   already exact, so slicing it allocates a second array to save nothing
   — on a corpus of single-file packages the unguarded form costs 31% of
   the build for zero bytes.

DROPPED: merging the `dirChildren` walk into the `suffixByStem` walk.
It measures -0.10% at 32k, +0.23% at 100k and +0.40% at one file per
directory, all inside a base-vs-identical-copy noise floor of -2.3% to
+3.1%, and it does not compose usefully with the memo — the second scan
it deletes is exactly the scan the memo makes rare. Only its provably
free half is kept: `stem.lastIndexOf('/')` in place of `norm.lastIndexOf`,
one backwards scan instead of two, exact because an extension carries no
'/'.

Neither optimization is visible to the correctness fingerprint, which is
the point and also the risk: it observes the index only through the four
resolver tiers, so a key-order move no corpus query reaches would survive
it. Correctness therefore rests on a structural comparison of all three
maps — key insertion order, values, bucket contents in order, frozen-ness
— over 1234 corpora in both iteration orders, 14 808 comparisons, zero
failures. The fingerprint, `cases` and `non_null` are unchanged and MUST
NOT be re-baselined by this commit.

Gates that did move, both because a reading and its budget move with the
code rather than when CI goes red:

  - `heap_reading_bytes.kotlin` 48 200 224 -> 42 802 456 with its ceiling
    at 1.5x. A memory WIN passes every arm, so nothing forced this.
  - `depth_budget` 2.2 -> 2.0. The memo turns a per-file component walk
    into a per-directory one, which is precisely the per-depth work this
    arm exists to see: the band went 1.27-1.40 -> 1.20-1.26, and 2.2 held
    over it would have drifted from ~1.6x headroom to ~1.9x.

The gate controls were re-run against the optimized builder, including
one this change makes newly plausible: keying the memo on the directory's
LAST SEGMENT instead of its full path drifts the fingerprint
(36a4e9dad313, non_null 13310 -> 13305). That is the memo's whole
safety argument stated as a test — its key decides which key set a
directory contributes — and it is the one way this optimization could
move an answer. The bucket-cap control was re-run too, since compaction
now rewrites the same buckets.

Also recorded, from measuring a reuse this repo had been invited to make:
replacing `dirChildren` with the shared `package-dir-index` is
output-identical (0 divergences over 107 948 answers) and passes every
arm of the kotlin bench at 1.37x-1.50x — while costing 409x per fan-out
and 8114x on `import data.*` at 200 matching directories on a corpus this
bench does not carry. `_blind_spot` in the kotlin baselines now says so,
with the memory the trade would have bought (26.2%, 12.18 MiB) and the
corpus arm that would have to exist first.

Refs #2881.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e

* fix(scope-resolution): close the gaps a four-lens review found in the #2881 change

Correctness review found no defect in the shipped resolvers — the `endsWith`
rewrites, the C# empty-prefix guard, the memo's purity, the `stem` vs `norm`
derivation, the Map re-`set` during iteration and Go's `substring` arithmetic
were each attacked with running code and each held. Everything below is a gap
in what the change ASSERTS, measures or claims.

UNGATED BEHAVIOUR, now covered:

  - `import-resolvers/go.ts` had no test at all. Its rule changed, the shared
    bench drives the indexed leg rather than this one, and a revert was caught
    by nothing. `go-package-resolve.test.ts` pins the membership rule and, more
    usefully, pins that Go's two independent legs agree on it — they disagreed
    before #2881 and a divergence here means the LanguageProvider hook and the
    ScopeResolver hook hold different views of a package.
  - The memo and the compaction are output-identical, so no fingerprint sees
    them and reverting either leaves both benches green. `kotlin-index-internals
    .test.ts` asserts them directly: the memo hit path against the miss path,
    two directories sharing a component-suffix keeping separate buckets (the one
    way a coarser memo key could move an answer), and that the bucket handed out
    is the cached, frozen, compacted array on both the sliced and the skipped
    path. The comment claiming this was "asserted structurally" previously
    pointed at nothing in the repo.

GATES:

  - Six ratio budgets in `bench/import-target` were slack: the measurements they
    bound got faster and the numbers were left alone. kotlin depth 3.4 -> 2.8,
    go 1.6 -> 1.4, csharp 2.2 -> 2.0, java 2.2 -> 2.1, kotlin collide_scaling
    1.8 -> 1.65, go 5.5 -> 5.1, each holding the headroom the old value
    expressed. The absolute ms ceilings are deliberately untouched: they carry
    runner-contention headroom, and a ratio is runner-speed-invariant where a
    millisecond is not. This is the failure the branch already fixed one
    directory over and missed here.
  - The `csharp_csproj` heap re-baseline is REVERTED. Base and branch both
    measure ~73.10e6 three runs each; the recorded 73703384 was simply not
    reproducible, and re-recording it would have dropped that language's derived
    floor 0.8% for no reason belonging to this change.
  - kotlin's collide arm was blind to the rule it was re-baselined for — a full
    revert of the Kotlin guards left both its fingerprints unmoved, because
    `com/example/models` is not a suffix of `…/models/inner/models`. Deepened to
    repeat the whole queried path; those two fingerprints are the only ones that
    moved for it. The same deepening on the java and kotlin UNIQUE arms was
    measured and REVERTED: ten more fingerprints, java's heap reading up 43%,
    and no coverage gained, because progressive stripping lands those queries on
    the same file either way.

SIMPLIFICATION:

  - `go.ts` now states the predicate as ends-with like its three siblings,
    instead of keeping the `indexOf` shape with `lastIndexOf` swapped in.
  - C# csproj step 2's direct-child filter is dead for a non-empty prefix —
    `getFilesInDir`'s keys ARE segment-aligned directory suffixes, so it cannot
    reject, and measurement agrees over 12 008 pairs. Only the empty-prefix case
    does work, and only that case remains.
  - `addChild` had one call site left; inlined. The memo's double read of its
    own lookup is gone. The V8 byte accounting duplicated verbatim between the
    resolver comment and the baselines note now lives only in the note.
  - Four copies of the same ternary in the csproj parity harness collapse onto
    one hoisted `dirTrail`; two locals in the java harness were named for the
    branch that was deleted.

CLAIMS THAT WERE WRONG:

  - `package-dir-index.ts` said "the four resolvers agree again". It is six, and
    the sixth is the evidence: `import-resolvers/jvm.ts` has answered the same
    question with `lastIndexOf` since #488, so before #2881 Java's and Kotlin's
    LanguageProvider hook and their ScopeResolver hook disagreed about which
    files a package holds.
  - The `uniqueDir` docblock claimed the last segment IS the query granularity
    for csharp/java/kotlin. They query the whole dotted path first and reach the
    tail only through stripping — which is why the partial-revert control fires
    on the go arm alone, now stated instead of implied.
  - Three parity harnesses described themselves as verbatim copies of the
    pre-change implementations; they were edited by this branch, so they are
    re-derivations of the current spec, a weaker claim their headers now make.
  - The shared harness header still listed the removed rule as current, the
    `DIRS` docblock still justified shapes by a divergence that no longer
    exists, and `measure.mjs`'s tier-two docblock plus `_heap_bound_note` still
    counted nine bounded languages when `HEAP_BOUNDED` derives to three — this
    branch had dutifully updated a kotlin bound in a list no gate reads.
  - `_blind_spot` told the next reader to build a repeated-leaf arm that already
    exists in the sibling bench, with a budget that already fails the swap.

Both baselines are also re-serialized to preserve each note's original escaping,
undoing ~20 KB of no-op churn an earlier revision introduced by round-tripping
the JSON.

Refs #2881.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e

* perf(scope-resolution): drop the string each membership test built per candidate

The three `endsWith` membership tests each minted a decorated copy of the
directory once per candidate, per import. The decoration cancels:

  ('/' + D + '/').endsWith('/' + P + '/')  <=>  D === P || D.endsWith('/' + P)
  (D + '/').endsWith(P + '/')              <=>  D.endsWith(P)

Verified exhaustively rather than argued — every pair of strings up to length 5
over `{a, b, /}` including the empty string, 132496 pairs, 0 divergences, with
the match count reported beside it because two predicates that agree on `false`
everywhere also show 0 divergences. `matchingDirs` 32.58 -> 8.22 ns/candidate
(3.96x), `matchingDirPositions` 64.9 -> 18.4 ns. C#'s deliberate unanchoredness
survives verbatim: `src/SubModels` still answers `Models`.

`resolveGoPackage` was the opposite of a win — the rewrite in this branch left
the `'/' + path` cons the old `includes` guard used to short-circuit, and the
first `endsWith` forces V8 to flatten it once per file. Working on the raw path
with an explicit start index is 4.8x faster than that and 1.78x faster than the
code before this branch. It also now reuses `resolveGoPackageDir` instead of
re-deriving six of its lines.

Three claims these files make are corrected while they are open:

- `package-dir-index.ts` argued the rule was accidental because a sixth
  implementation never had it, "wired as `importResolver` by
  `languages/{java,kotlin}.ts`" and therefore live. It is wired and not read:
  `provider.importResolver` is consumed only at `import-target-adapter.ts:74-75`,
  and that module's exports have no importer outside their own unit test, while
  its docblock claims it is threaded through `finalizeScopeModel`. The argument
  survives on the pre-index-scan derivation; `jvm.ts` is evidence about how the
  predicate was written, not about live behaviour. Whether those resolvers
  should be deleted or wired is left as an open question.
- `csharp.ts` derived the empty `dirPrefix` case from "any path whose first
  slash is its last", which is wrong in both directions: `src/X.cs` satisfies it
  and emits no empty key, `a//X.cs` violates it and does. The conclusion stands
  and the filter stays — it is what rejects `a//X.cs`.
- Step 2 returns on its first push, so widening it also suppresses step 3's
  unanchored leg. The narrower answer is the more precise one, but it was an
  unstated output change.

`SuffixIndex.getFilesInDir` now states the segment-alignment its callers rely on,
bounded as a guarantee about what may be RETURNED — php's root-anchored index
answers only the equality arm.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus

* fix(scope-resolution): say what the widened bucket actually does downstream

The comment justifying the widening claimed a bucket that is too wide is
"filtered downstream" by the finalize pass. It is not, for the edge that
matters. `finalize-algorithm.ts` mints one draft per candidate, each keeping its
own `targetFile`, and the File->File emitter in
`graph-bridge/imports-to-edges.ts` tests only `targetFile === null` and
`targetFile === sourceFile` before adding an `IMPORTS` relationship at
confidence 1.0 — it never reads `linkStatus`. The `localDefs` filter from #1759
constrains `targetDefId` and the `BindingRef`; every extra bucket member is an
unconditional file-level edge regardless. Measured on an Android-shaped layout,
one `import data.load` goes from 5 to 6 edges, all six unresolved.

No filtering is added here. Whether an unresolved candidate should produce that
edge at all is a design question about the graph bridge, not about this bucket.

The published drift census — 149 first-child reselections, 32 wider arrays, 54
null -> resolved — has no bucket for a fourth class this change introduces.
Tier 3 precedes tier 4, so a bucket the guards used to leave empty returned null
and let the progressive strip run; a populated bucket stops tier 4 entirely,
turning a bound answer into a candidate list that need not carry the symbol.
Re-running the census with a shape classifier finds that class ZERO times over
the corpus, and the zero is the finding: the shape reproduces by hand, and this
bench's own generator at 4000 repositories hits it 4-12 times per seed. The
fingerprint cannot gate what the corpus cannot express — the same blindness the
go arm carried until #2881 widened it.

Two further claims are brought back in line with what shipped. The memo's
docblock said `kotlin-index-internals.test.ts` asserts the key set, key
insertion order and bucket order "over the built maps"; that file says it works
through the resolver's observable surface and omits key order deliberately. The
mutation matrix bounds it honestly: a mis-keyed memo is caught, a deleted one is
not, and the compaction's only instrument is the bench heap ceiling.
`findKotlinDirectoryChild` no longer claims to return "the same file the scan
used to return" — that is precisely what moved.

Structural, no behaviour: `let keys` sits with its consumer instead of 33 lines
above it, the archaeology moves to the docblock, `tight` -> `compacted`,
`dirEnd` -> `lastSlash` (the name three sibling builders use), and the one-use
`MutableDirChildren` alias goes with the `addChild` it existed for.

`finalize-algorithm.ts` annotates `targetFiles` as `readonly string[]` so
`Array.isArray`'s `any[]` predicate can no longer widen a frozen cached bucket
into something `.sort()` compiles against. The runtime freeze stays; it is the
backstop for every other call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus

* test(scope-resolution): gate the edges #2881 moved but nothing watched

Every widened-shape test in the branch used a one-file corpus, so not one of the
149 first-child reselections was pinned — the tier that commits to `children[0]`
unfiltered had no test that could see which file it commits to. Kotlin and Java
now pin that choice absolutely, in both insertion orders, for the member path
(tier 3, both members) and the wildcard path (tier 1, one file) separately,
saying plainly that both candidates are valid members and the only tie-break is
file-set iteration order.

The tier-3-preempts-tier-4 class gets its first gate, with a control that makes
it a transition rather than a fact. The bench corpus holds zero instances, so
this case is the only thing standing between that behaviour and a silent
revert.

C# gains three absolute arms, because its differential harness cannot see any of
them — the legacy copy was edited in lockstep with production, which the file's
own header admits. One pins the empty-`dirPrefix` filter the branch calls
load-bearing and which nothing defended: deleting the guard leaves the whole
suite green but changes the answer, so the arm was verified to fail with the
guard removed and pass with it restored. Java gains the negative control Kotlin
already had.

`kotlin-index-internals.test.ts` stops implying coverage it does not have. The
mutation matrix is recorded in its header: deleting the memo passes every arm
(it is output-identical by construction), deleting the compaction's `slice()`
passes every arm (a JS array's capacity has no reflective surface), while
mis-keying the memo fails three and compacting-but-never-storing fails two. Four
arms were added that do fail under those mutations. V8's growth steps were
re-measured — 1, 19, 46, 86 with growth at lengths 2, 20, 47, 87 — so the old
1/17/41 model, which under-counted the slack at 40 files by 6x, is gone.

`go-package-resolve.test.ts` drops four `as never` casts that were hiding
nothing (`GoModuleConfig` is structurally satisfied), and pins vendor/, testdata/
and nested-go.mod directories, which merge into the importing package — a
pre-existing unmodelled gap, verified present before #2881 and documented as
such rather than blamed on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus

* test(bench): gate the bucket compaction, and publish the whole drift taxonomy

The compaction shipped with no gate anywhere. Deleting `bucket.slice()` while
keeping the freeze moves no fingerprint, no count and no test — only retained
heap, 42805256 -> 48184784 B (+12.57%), byte-identical across three runs. Note
the direction: compaction reclaims, so losing it makes the reading GROW, which
no floor can see. `heap_ceiling_bytes.kotlin` tightens 64203684 -> 46000000
(1.5x -> 1.0747x of the reading), leaving the regression 4.8% clear above the
ceiling and the reading 7.5% below it. The band is derived from first principles
in `_heap_compaction_gate` (~61000 buckets x 11 spare slots at Node 22's 1->19
step) so it can be re-checked rather than trusted, and the note carries the
triage rule: heapUsed accounting drift moves every arm, so kotlin alone over its
ceiling is a lost compaction.

`_gate_controls` claimed the two optimizations rest on a structural comparison
over 1234 corpora in both iteration orders. No such probe exists in the tree. It
now names the test that does exist and lists what it actually pins, and says
key insertion order is unasserted by design.

`_provenance` gains the full shape classification behind the 235 moved records:
149 string -> string, 38 null -> string, 16 null -> array, 32 array grew, and
zero of every other transition — including `string -> array`, the
resolved-becomes-unresolved class the old taxonomy had no bucket for. The
harness was validated byte-exactly first: driven over this corpus the base
resolver reproduces ebf1790bf1 / 13256 and head reproduces d91110bee3 / 13310.

`measure.mjs` loses a paragraph asserting the C# unique slice repeats the whole
queried path, directly above the paragraph explaining it is leaf-only
deliberately and the code that makes it so. Acting on the deleted half resolves
the csproj arm to zero. While measuring: the csharp collide arm is NOT blind —
its fingerprint already moves across #2881 — but both csharp_csproj arms are,
because `getFilesInDir` keys on segment-aligned suffixes and neither nested slice
is one. Closing that needs a corpus redesign and four re-baselines; recorded, not
attempted.

One number changes in either baselines file, and it tightens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 10:26:39 +01:00
Gergő Magyar
135bcae03d
fix(go): resolve out-of-repo package qualifiers, and stop reporting an undecided interface check as a decided negative (#2873) (#2921) 2026-08-11 10:36:59 +01:00
Gergő Magyar
18bc51dfd2
perf(import-resolvers): index every scanning resolver, consolidate the memo, gate every registered language (#2911)
* perf(import-resolvers): build buildSuffixIndex's dirMap lazily (#2903)

`buildSuffixIndex` eagerly built three maps. `dirMap` is the array-valued one —
one entry per directory suffix per file, so O(files x depth) in entries and
array churn — and only four call sites ever read it, all via `getFilesInDir`:
`import-resolvers/{php,csharp,jvm}.ts` and `import-resolvers/configs/python.ts`.

Ruby (through workspace-file-index), the TypeScript scope resolver, Vue's
import-target and the include-extractor never ask a directory question, and
built it anyway. Since #2880 these indexes are retained for a whole resolution
pass rather than rebuilt per import, so that waste is now resident memory.

Deferring it to the first `getFilesInDir` call is behaviour-identical — same
key, same descending-suffix order, same per-bucket push order, same
`substring(lastIndexOf('.'))` extension clamp. The builder assigns the MAP on
completion, so a repeated miss cannot rebuild it.

Measured on `buildSuffixIndex` alone, 32k paths, index built and
`getFilesInDir` never called:

  C# layout, 13 segments   79,018,680 -> 66,580,488 B   -15.74%
  Ruby layout, 11 segments 60,752,792 -> 48,656,856 B   -19.91%

and on the whole retained WorkspaceFileIndex the bench measures:

  csharp 32k  73.62 -> 61.76 MiB   ruby 32k  55.26 -> 43.69 MiB

When `getFilesInDir` IS called the footprint is unchanged, so the deferral is
never a loss. No new retention: all five construction sites already hold both
input arrays alive beside the index.

The laziness is pinned structurally rather than by timing. The test's corpus is
a `string[]` whose elements are accessor properties, so an indexed read is
observable and the read count IS the pass count: 14 after construction, still
14 after any number of get/getInsensitive, 28 after the first `getFilesInDir`,
28 after five more. Memoizing the decision instead of the map would read 42.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(php): resolve imports from a per-run index, not a scan per import (#2901)

PHP was the last language whose import resolution scanned the workspace per
import. Both `resolvePhpImportTarget` and `resolvePhpImportTargetInternal`
materialized two full arrays from the Set on every call, then passed
`undefined` as the `index` argument — so `resolvePhpImportInternal` fell
through to `suffixResolve`'s linear `findIndex`, once per extension per path
part. Measured at 20,000 files: 96.40 ms per import.

**Handing it the shared SuffixIndex would have moved IMPORTS edges.** All three
index-fed sites answer a different question than the scan they short-circuit,
each found by differential with a concrete witness:

  1. `getInsensitive` — the scan leg is `allFiles.has(path)`, exact whole-path
     with no case-insensitive counterpart; the shared index answers a ci SUFFIX
     probe.
  2. `getFilesInDir` — the scan is root-anchored `startsWith(nsDir + '/')`;
     `dirMap` is keyed on every directory SUFFIX, so a vendor copy can win.
  3. `suffixResolve` — the scan's `endsWith('/' + S)` matches only a PROPER
     suffix; `buildSuffixIndex` indexes j=0, so a root-level `Foo.php` starts
     resolving `use Foo` where it returned null.
  3b. the scan's `endsWith(p) || lower.endsWith(lower(p))` has a second
     disjunct that subsumes the first, so it is purely first-in-Set-order and
     case-insensitive; `get(S) || getInsensitive(S)` lets a case-exact hit
     anywhere beat an earlier ci hit.

So this is not Ruby's #2880 shape. Both sites take `getWorkspaceFileIndex` for
the memoized arrays and hand the internal resolver a PARITY `SuffixIndex`
memoized on the same Set identity: `getInsensitive` disabled, `get`
implementing the scan's real rule via the shared ci lookup plus one O(files)
whole-path correction map, `getFilesInDir` root-anchored in Set order.

  no composer.json    96.40 -> 0.036 ms/import steady state
  with composer.json 100.19 -> 0.068 ms/import steady state

Also closes PHP's last per-import traversal, in `import-resolvers/php.ts`: its
namespace-directory scan ran whenever `getFilesInDir` came back EMPTY, not
merely when no index was supplied — despite the comment above it claiming
"only when SuffixIndex unavailable". An empty bucket is already the answer, so
the scan could only confirm it, at one full pass per import whose namespace
matches a PSR-4 prefix but whose directory has no direct `.php` child
(measured 11 traversals for 10 imports; now 1). Moving it into the `else` is
safe because the bucket is a SUPERSET of what the scan finds — a root-anchored
direct child `nsDir/<x>.php` has its directory exactly equal to `nsDir`, and a
directory is always one of its own suffixes, so both index shapes contain it.

Nine mutations of the new code are caught, including M1 "pass the raw shared
index" (the naive fix) at 23 arms. The adapter guard reads 600 instead of 1
under a defensive `new Set(allFilePaths)` — the #1918 P1 hazard the unit
differential is structurally blind to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(java): index import resolution instead of scanning per import (#2908)

Java scanned the whole workspace twice per import: once for the three-tier
direct match, and again INSIDE the progressive prefix-stripping loop — so a
single unresolvable import cost one full pass per stripped segment. No WeakMap,
no index, and it is registered in `SCOPE_RESOLVERS`, so it ran in production.

This is byte-for-byte the C# shape #2878 fixed, so Java now reads the same
machinery: `getWorkspaceFileIndex` for `normToRaw` + the segment-suffix index,
and a Java-owned `PackageDirIndex` WeakMap over `buildPackageDirIndex(_, n =>
n.endsWith('.java'))` read through `firstFileDirectlyInPkgDir`. Structure
mirrors C#'s `narrowContext` / `resolveDirectMatch` /
`resolveByProgressiveStripping`.

  20k files, 256 imports, 7-in-8 unresolvable:  8.05 -> 0.62 ms/import
  steady state once the index is built:         0.0036 ms/import

Tie-breaks preserved, and Java's are NOT identical to C#'s:

  - tier 1 `break`s on the exact match, so an exact whole-path hit wins even
    when a suffix or directory-child hit came earlier in iteration order —
    hence `normToRaw.get` before `index.get`, which conflates them;
  - the stripping loop instead returns at the FIRST hit of `f === tailFile ||
    f.endsWith('/' + tailFile)` and only yields its directory child after the
    scan completes, so the conflated `index.get` is the correct lookup THERE.
    Applying tier 1's exact-wins rule inside the loop is a real behaviour
    change (mutation M6);
  - `.*` wildcard stripping stays ahead of everything;
  - `firstFileDirectlyInPkgDir` reproduces Java's at-root/at-nested predicate
    exactly, including the first-`indexOf` rule — proved algebraically rather
    than assumed: the `atRoot` branch matches iff `dir === pathLike`, which is
    `D.indexOf(P) === 0 === D.length - P.length`, and the `atNested` branch's
    first occurrence in `f` is the first occurrence in `D` shifted by one.

Six mutations are caught; a seventh (swapping the two index builds) is a true
equivalence and is recorded as such. Hand-derivation also corrected four cases
where the legacy code resolves and I had predicted null — including
`java.util.List` reaching a local `util/List.java`, because Java has no
in-repo-namespace gate like C#'s #1881. That is preserved here and filed
separately as #2910; the parity test pins it so the fix is visible.

The adapter guard reads 800 instead of 2 under a defensive
`new Set(allFilePaths)`. Two traversals is correct: the workspace index and the
package-dir index are separate WeakMaps and each iterates the Set once, the
same accounting as C#.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(cobol): index COPY resolution instead of two scans per statement (#2908)

`cobolScopeResolver.resolveImportTarget` ran two full workspace scans per
`COPY`, each calling `path.extname` + `path.basename` + `.toUpperCase()` on
every entry: tier 1 over `.cpy`/`.copybook`, tier 2 over `.cbl`/`.cob`/
`.cobol`. No WeakMap, no index, and registered in `SCOPE_RESOLVERS`.

Two uppercased-basename maps, one per tier, filled in a SINGLE pass over the
Set and memoized on Set identity. Lookup is
`copybooks.get(upper) ?? sources.get(upper) ?? null`.

  20k files, 500 COPY operands:  3879-4082 -> 10.5-11.7 us/import  (~350-369x)
  steady state once built:       0.253 us/import

Tie-breaks preserved:

  - TIER ORDER. A `.cpy` match beats a `.cbl` match even when the source file
    appears EARLIER in Set-iteration order. This is the one a naive
    single-map rewrite silently breaks, so it gets its own fixture.
  - Within a tier, first in Set-iteration order wins (`if (!tier.has(...))`,
    mirroring the scans' first-match return).
  - The key is built with the identical call sequence,
    `basename(fp, extname(fp).toLowerCase()).toUpperCase()`, so `Foo.CPY` still
    keys under `FOO.CPY` rather than `FOO`.
  - `path` stays in the loop rather than hand-rolled `/`-slicing, so backslash
    handling is unchanged on every platform — pinned by a `dir\sub\BOOK.cpy`
    case.

All six mutations are caught: collapsing the tiers, within-tier last-wins,
dropping the target uppercase, dropping the extension lowercase, hand-rolled
slicing, and the adapter's defensive copy. The first five are caught by the
differential and are invisible to the adapter guard; the sixth is the reverse,
which is the layering working as intended — the guard reads 600 instead of 1.

`COBOL_SOURCE_EXTENSIONS` was being re-allocated on every call; hoisted to
module scope beside `COPYBOOK_EXTENSIONS`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(csharp): index the csproj leg's namespace-directory scan (#2902)

#2878 moved C#'s no-csproj leg onto memoized indexes; the csproj leg kept a
per-import full scan in `resolveCSharpImportInternal` step 3, measured at
~1.10 ms per import at 50,000 `.cs` files.

**The fix the issue proposed would have moved edges.** It suggested skipping
the fallback when an exhaustive index is available, on the assumption that
step 2's `getFilesInDir` answers the same question. It does not: step 2's
`dirMap` is keyed on segment-aligned directory suffixes, while step 3's
`normalized.indexOf(dirPrefix + '/')` is an UNANCHORED substring match, so
step 3 finds a strict superset — and it runs only when step 2 came back empty,
so those extra hits are observable, not shadowed:

  dirPrefix 'ubModels'  step 2 []  step 3 ['src/SubModels/Widget.cs']
  dirPrefix 'rc/Models' step 2 []  step 3 src/Models/* AND vendor/mysrc/Models/*

So the predicate is kept byte-for-byte and made fast instead. It depends only
on the file's directory (the needle ends with `/`, so every occurrence lies
wholly inside `D + '/'`), which reduces to the `package-dir-index` formula
minus the anchoring leading slash. `PackageDirIndex` itself cannot be reused
for the same reason — its matcher is anchored.

The index is memoized on the `normalizedFileList` array identity and built
lazily at the point step 3 is first reached, so BCL usings — which `continue`
out at the root-namespace gate — never pay for it. Candidates come from an
exact last-segment bucket when `dirPrefix` contains a slash, a last-segment
key sweep when it does not, and `singleSegmentDirs` when it is empty.
Positions rather than paths, merged and sorted when several directories match,
so file-list order survives.

  App.Missing @ {App, src}  1103.0 -> 7.6 us   (145x, and flat in file count:
                                                7.3 @10k, 7.6 @50k, 8.4 @200k)
  App.Missing @ {App, ''}    626.7 -> 108.5 us
  App @ {App, ''}           1077.9 -> 2.0 us   (539x)
  App.Ns8 @ {App, src}         0.6 -> 0.6 us   (step-2 hit, untouched)

`relative === ''` is preserved exactly, including the no-`projectDir` case
where the needle is a bare `/` and the answer is "every `.cs` whose directory
has no slash of its own" — `getFilesInDir('', '.cs')` cannot answer that over
repo-relative paths, so it has its own arm.

13 of 14 mutations are caught, including M1, the naive skip-when-indexed
cleanup, at 9 arms. The survivor drops the empty-prefix fast path and is a
true equivalence. M9 initially survived and exposed a real corpus gap — no
non-`.cs` file lived inside a directory — now covered.

The remaining non-constant term is the slash-free sweep, O(distinct last
segments): 456 us at 200k files on a unique-name layout, but 7.9 us on a
`SrcN/Models` layout, which is how C# repos are actually laid out. Closing the
unique-name case needs a character-suffix map over segments — the
O(files x depth) memory shape `package-dir-index.ts` cites #2649 to avoid — so
it is documented in the code as a design change rather than tuned here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* test(scope-resolution): assert index reuse for every registered language (#2909)

Index reuse was asserted by nine hand-written per-language files, so the
guarantee existed exactly for the languages someone remembered — and #2908 is
the proof that is not good enough: Java and COBOL were registered, quadratic
and unguarded until this branch. `resolveImportTarget` is a required member of
`ScopeResolver` with one signature and 16 registrations, so "calling it N times
against a stable `allFilePaths` must not traverse the set N times" is a
property of the CONTRACT.

`import-target-index-reuse.contract.test.ts` drives every entry of
`SCOPE_RESOLVERS`, modelled on `construction-syntax-wiring.test.ts` — the
established shape here for a property plus a justified inventory. Measured
counts, all memoized:

  c 1  cobol 1  cpp 1  csharp 2  dart 1  go 1  java 2  javascript 2
  kotlin 1  php 1  python 1  ruby 1  rust 0  swift 1  typescript 2  vue 2

**`KNOWN_UNINDEXED` is empty.** The audit that produced it also cleared C, C++,
Rust, Swift, TypeScript, Vue and JavaScript by hand — Rust's memo lives in
`qualified-call.ts::moduleIndexFor`, C's and Swift's loops are inside their
WeakMap builders. The empty map stays as a mechanism: a 17th language cannot
opt out silently, and the inventory arm fails when a registered resolver has no
fixture.

Two things the assertion had to get right:
  - it is `scans(200) === scans(2)`, not `scans === 1`. Per-language counts
    legitimately differ (C# and Java build two indexes), and comparing two
    counts needs no per-language expected value.
  - Rust legitimately scans ZERO times — it answers every leg with
    `allFilePaths.has(candidate)` probes — so the floor is a per-language
    `minimumScans`, 1 for fifteen languages and 0 for Rust with the reason on
    the interface. Paired with a `hitTarget` that must resolve non-null, so the
    property cannot pass vacuously on a resolver that stopped answering.
Miss targets are distinct per import, which defeats the TS/JS/Vue per-target
`resolveCache`.

Also unifies the instrument. Kotlin and Python counted index BUILDS from
production; the other seven count traversals of a `CountingSet`. The build
counter is strictly weaker — a scan added BESIDE a reused index moves no build
count, which is exactly the mutation `baselines.json` `_blind_spot` records as
invisible to every timing arm — and it costs two production modules that ship
in the bundle purely for tests, holding module-global state every test must
`reset()`. Both guards migrate to `CountingSet`, and
`languages/{kotlin,python}/index-stats.ts` plus both call sites are gone, for
-59 lines of shipped source.

(Mechanical note: the two `index-stats.ts` file deletions appear in the #2901
commit rather than this one. They were staged with `git rm` while a concurrent
commit swept the index. The final tree is correct; only that attribution is
off, and rewriting a sibling commit to move them was not worth the risk.)

Coverage went up in the swap: Kotlin's old "rebuilds when the file set is a
different object" arm (3 sets, 3 builds) would have PASSED under a defensive
adapter copy. Its replacement fails, as do all six arms across the two files.

Verified by mutation: `new Set(allFilePaths)` inserted into the kotlin, python
and go adapters fails exactly those three and no others —
`python: 200 imports cost 201 traversals, 2 cost 3`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* test(import-target): gate the four newly-indexed resolvers, retighten heap

The bench covered go/csharp/dart/ruby/kotlin. The four resolvers indexed on
this branch shipped unmeasured, and #2903's memory win was not locked in.

**php, java and cobol join the shared corpus**, each with the two load-bearing
properties the header requires: imports scale with file count, and most imports
MISS so the full cascade runs (resolve rates php 36.0%, java 34.4%,
cobol 36.0%). Java's miss families were measured rather than assumed, since it
has no in-repo-namespace gate (#2910): `java.*` 1041 imports and
`com.google.*` 1006, both resolving 0. COBOL's collide layout repeats a
bookname across BOTH extension tiers, so it reaches the copybook-over-source
tie-break rather than only the basename map.

**`csharp_csproj` is a sixth LANGS entry**, not a new arm dimension — an entry
needs five small additions and inherits all five arms and all seven gates,
where a context axis would have to be threaded through `buildRepo`,
`resolveAll`, `identityPass`, the report shape and every gate. `buildFiles`
aliases it to `csharp`, so the two share one corpus by construction and cannot
drift. Two configs (`{App, 'src'}`, `{Lib, ''}`) produce all three `dirPrefix`
shapes — slashed, slash-free and empty — in five arms instead of ten:

  App.Ns{d}      30.6%  src/Ns{d}        step 2 hit
  App.Missing{n} 25.5%  src/Missing{n}   step 3, last-segment bucket
  Lib            14.0%  (empty)          step 3, singleSegmentDirs
  Lib.Missing{n} 12.0%  Missing{n}       step 3, KEY SWEEP — the one
                                         non-constant path
  BCL / Ghost    12.4%  —                root-namespace-gate control

**2221 of 3200 imports reach the indexed leg**, only 12.4% `continue` out. What
that arm pins is stated plainly rather than overclaimed: step 3 answers null
for all 2221 here (the hits land at step 2), so it gates that leg's COST and
its null answers; its positive tie-breaks stay pinned by the unit parity test.

**Heap ceilings retightened.** #2903 dropped the measured figures, leaving the
1.5x ceilings at ~1.9x — a straight revert to the old size would have passed:

  csharp 116,000,000 -> 98,000,000 B   (measured 61.76 MiB)
  ruby    87,000,000 -> 69,000,000 B   (measured 43.69 MiB)
  php    new 106,000,000 B             (measured 67.29 MiB)
  java   new 154,000,000 B             (measured 97.32 MiB, the largest in the
                                        file — Maven layout is 18 segments)

php and java are gated because both retained NOTHING across imports at BASE and
now retain the O(files x depth) suffix index — the same argument that gates C#.
cobol is not: two `Map<basename, path>`, O(files) with no depth term, and its
retained delta does not clear measurement noise, so a ceiling would gate
nothing. `csharp_csproj` is not: same corpus, same index, a duplicate number —
its one distinguishing footprint, the lazily-built `dirMap` its `getFilesInDir`
forces back, is measured at +20.8% and recorded as a residual instead, because
gating it would licence eager-dirMap everywhere.

csharp's `depth_ratio` also fell 3.318 -> 2.31 (the no-csproj leg never asks a
directory question, so the deep arm stopped paying an eager dirMap build).
Budget 5 -> 3.5, restoring the file's 1.5x convention — and `_arms_note` says
plainly that 3.5 does NOT lock that win in, because locking it needs ~2.9,
which is 1.25x over a 1.05x spread and the kind of tightening `_triage` warns
buys flake rather than signal.

All five pre-existing languages are byte-identical: 25 cells x 5 fields = 125
values, 0 mismatches. The new arms were proven live by a doctored baseline
(cobol ceiling 0.01, php heap 1000 B, java resolved 999) producing three
correctly-worded failures and exit 1.

Wall-clock 10.9 -> 26.1 s, php and csharp_csproj ~11 s of it — both cascades
end in `suffixResolve`'s ~50-extension probe, and both gate the two largest
wins on this branch, so neither is a candidate to drop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(javascript): build the suffix index JS resolution never had

JavaScript's `PassCache` was TypeScript's minus one field: `index`. So JS
called the shared `resolveTsTarget` with `ctx.index === undefined`, and
`import-resolvers/standard.ts` fell through to `suffixResolve`'s linear
`findIndex` — scanning the materialized path list once per extension (~39)
per path part, per import.

  2000 files   6448.9 -> 28.5 us/import   (TypeScript: 25.0)
  8000 files  25972.6 -> 27.4 us/import   (TypeScript: 27.0)

Per-import scaling over 4x the files: 4.12x -> 1.09x.

**Every instrument on this branch was blind to it.** `CountingSet` counts
traversals of the Set; this walked the array the adapter had already
materialized — the blind spot `counting-file-set.ts` documents in its own
header and `baselines.json` records under `_blind_spot`. Under mutation M1,
which drops `index` and reproduces the shipped defect exactly, the sixteen-
language contract test stays GREEN for javascript, because the pass cache is
still reused and `files.scans` reads 2 either way. Two new arms do catch it: a
`suffixResolve` linear-branch counter that runs the legacy adapter first as its
control (135 entries legacy, 0 now), and a mock-free behavioural assertion that
a repo-root module resolves by bare specifier.

Adding an index moves output, exactly as it did for PHP in #2901, so it was
characterized rather than assumed — 211,200 pairs (400 corpora x 3 importers x
176 targets) plus 184 hand cases. **Two classes move and there is no third:**

  A  null -> repo-root file (108)   `require('config')` with root `config.js`.
     The scan tests `endsWith('/' + suffix)`, so a path with no slash has no
     proper suffix and was unreachable through that leg — while `./config`
     from the root already resolved via the exact `Set.has` branch. JS was
     internally inconsistent.
  B  file -> different file (5679)  `import 'app/main'` was resolving to
     `node_modules/dep0/lib/main.js`; the scan skipped the whole-path candidate
     at the 2-segment suffix and fell through to the 1-segment `/main.js`,
     taking the first such file in Set order.
  C  hit -> null                     ZERO, and impossible: proper-suffix keys
     are a subset of the index's keys.

Both moved classes are JS being wrong. **JS-new agrees with TypeScript on all
211,200 pairs and every corpus case, 0 disagreements** — which is the intended
design, since JS delegates to the TS resolver and differed only by this field.

Also swaps the single-slot `let cached: PassCache | null` in JS, TS and Vue for
a module-level `WeakMap`, matching every other language. Two alternating file
sets rebuilt everything on every call: 12.0 -> 1438.2 ms at 4000 files x 400
imports (120x); after, 11.0 -> 15.7 ms. This is LATENT, not live —
`pipeline/run.ts:673` builds one Set per provider pass and the three are
separate providers — but it is why these were the only languages that could not
carry the standard distinct-set guard. They can now: the arm fails on HEAD for
all three (`expected 42 to be 2`) and passes after.

Six mutations caught, including a global `resolveCache` (M5), which needed a
new arm — `expectDistinctFileSetsGetOwnIndex` builds two IDENTICAL corpora, so
a stale answer carried between them is also the right answer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* refactor(ingestion): one per-file-set memo primitive, twenty-one call sites

Every language that indexes its import resolution hand-rolled the same memo:
declare a module-level `WeakMap` keyed on the file-set object, `get`,
`if undefined` build and `set`, return. One concept, written twenty-one times,
and this branch had just added five more.

`import-resolvers/per-file-set.ts` exports it once:

    perFileSet<K extends object, T extends object>(build: (key: K) => T): (key: K) => T

Two decisions, both recorded in the file. `T extends object` rather than
`has`-then-`get`: `WeakMap.get` returning `undefined` cannot distinguish "not
built" from "built as undefined", and the `has` form needs a cast or a non-null
assertion, both banned here — the constraint makes the ambiguous case
unrepresentable instead, and a future caller wanting `string | null` gets a
compile error pointing at the decision. A throwing build stores nothing and
runs again next call, so failures are not memoized and a half-filled index is
never published — inert for these pure builders, and the safer direction.

`K extends object` rather than `ReadonlySet<string>` is what lets C#'s
`readonly string[]`-keyed cache share the helper.

Twenty-one sites migrated across `import-resolvers/` and fifteen languages.
Every existing doc comment was re-homed onto the new call rather than deleted —
several record real invariants (the Set-identity contract, the #1918
pass-through rule, why Rust's memo lives on a different hook).

TypeScript, JavaScript and Vue additionally had byte-identical `PassCache`
interfaces and builders. `import-resolvers/pass-cache.ts` now holds the one
builder, taking a single argument — every difference the three have lives in
the CONSUMER (`tsconfigPaths`, the extension list), not the builder. The
builder is shared, the memo deliberately is not: each adapter keeps its own
`perFileSet`, hence its own index and its own `resolveCache`, because the three
disagree about what a specifier resolves to and one shared cache would hand a
language another language's answers. It buys no runtime reuse and the module
says so — each provider pass builds its own `allFilePaths` Set, so the three
are always different keys.

C and C++'s `augmentedFilePaths` was a two-LEVEL memo, and needed no new
abstraction: the outer memo's value is a function and a function is an object,
so `perFileSet(perFileSet(...))` composes. The two instances stay one per file,
and the reason is now in BOTH doc comments rather than only C++'s — cpp
delegates to `resolveCImportTarget`, whose `suffixIndex` is keyed on the
augmented set, so a shared memo would cross the two languages' indexes.

Two sites are deliberately NOT migrated, each with the reason written at the
declaration so the next sweep does not re-litigate them:
  - `configs/swift.ts` is a two-input memo keyed on one. `targets` is not
    derivable from the key; re-keying on `ctx` would force a banned non-null
    assertion or an unreachable fallback inside a memo builder.
  - `rust/qualified-call.ts` `MODULE_SCOPE_CACHE` is three inputs keyed on one,
    and sits ten lines below a `perFileSet` in the same file — the likeliest
    thing to be "fixed" by mistake.

The other ten remaining `WeakMap`s are different concerns and stay: AST-node
caches, worker-pool runtime state, graph metadata, mutable lazily-filled
accumulators, and the C++ ADL / inline-namespace indexes, which are reassigned
by explicit clear functions and epoch-stamped on read — validity rules beyond
key identity that a closure over a private cache cannot express.

Net −20 lines of code, +22 of the two "why not" notes. The primitive's own doc
is where the cost sits: the Set-identity contract and the two design decisions
are written once instead of being twenty-one implicit facts.

Pure refactor: 1764 unit tests, 42 guard tests, all sixteen contract-test
traversal counts unchanged (c 1, cobol 1, cpp 1, csharp 2, dart 1, go 1,
java 2, javascript 2, kotlin 1, php 1, python 1, ruby 1, rust 0, swift 1,
typescript 2, vue 2), 647 C/C++ tests, and every bench fingerprint unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* test(import-target): gate every registered language, not nine of sixteen

The bench pinned output fingerprints and scaling for 9 of the 16 languages in
`SCOPE_RESOLVERS`. The other seven — c, cpp, javascript, python, rust, swift,
typescript, vue — resolve imports in production with nothing pinning their
output or their cost. JavaScript was the sharpest case: the 25,972 us/import
defect fixed earlier on this branch was gated by unit tests alone.

All 16 are now gated, plus the `csharp_csproj` variant: 17 entries.

**The nine existing languages are byte-identical** — 234 committed values
(9 x 5 arms x 5 fields, plus 9 top-level fingerprints), 0 changed, and no
pre-existing budget touched. Measured both before and after the memo
consolidation in e6f15274e, so it doubles as an independent check that the
refactor preserved behaviour.

Corpora keep both load-bearing rules — most imports MISS, and import count
scales with file count — at resolve rates of 26-36%. C and C++ follow the
`csharp_csproj` precedent: a `LANGS` entry carrying its own context (header
paths through `resolutionConfig`) over an aliased corpus, since cpp delegates
into C's `resolveCImportTarget`. Vue threads `tsconfigPaths` so its alias
branch actually runs; ts/js use bare specifiers only, because relative ones
never reach `suffixResolve`.

Two corrections to my own profiling, both verified rather than assumed:
Swift's `byModule` IS depth-scaled (one bucket entry per interior segment, not
O(files)), and Python's index is depth-free while its RESOLVER is quadratic in
depth — `hasRepoCandidate` and `resolveAbsoluteFromFiles` each rebuild one
ancestor prefix per importer directory component, per import. That is why
python's `depth_budget` is 11 against a 3.5 next-highest; the arm is pinning a
real defect rather than a comfortable number, and it is filed separately.

Rust's collide arm was redesigned rather than budgeted away: it is flat on file
count by construction, so a shared-leaf arm would have asserted nothing. Its
collide corpus varies `::` segment count — the axis its cost actually has — and
the linear 1.8 budget asserts the file-count flatness.

Heap: all 8 measured, 3 gated. javascript (44.07 MiB, retained nothing before
its fix), python (7.27 MiB), c (9.55 MiB). Five skipped with their numbers in
`_arms_note` rather than silently: rust 16 B (no index on this hook), swift
reads 3x SMALLER on a 4x corpus so it is below its own noise floor, typescript
288 B on 46 MB, vue +5.4%, cpp 0.04% from c.

Every gate type was proven able to fail: one run with 10 doctored values fired
10 correctly-worded failures across all 8 new languages, covering per-scale
fingerprint, shape/resolved, shape/distinct_outcomes on a non-small arm, depth,
collide scaling, absolute small ms, absolute collide ms, top-level fingerprint
and heap bytes. That proof found two wrong messages, now fixed: the heap
failure claimed a `buildSuffixIndex` cause that is false for python and c, and
the fingerprint failure pointed at a parity harness covering none of the eight.

Wall clock 26 -> 46 s. The ts/js/vue family is 14.6 s of the 18.8 s added,
because `suffixResolve` probes ~39 extensions per path part on a miss — the
real resolver, not something the bench can tune. Per language the bench got
cheaper (2.7 s vs 3.0 s). If it must shrink, `_arms_note` and the CI comment
record the one cut that removes duplicate work rather than coverage — drop
collide for typescript and vue only, -3.9 s, since all three share
`resolveTsTarget` and javascript keeps the arm covering their common axis.
Explicitly NOT `REPS`: it is 15 because `depth_ratio` flaked 1-in-20 at 5, and
lowering it would re-open that for all 17 languages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(import-resolvers): stop building half of every suffix index

Applies the findings of a four-lane quality review over this branch.

**Half of `buildSuffixIndex` was dead weight for most of its consumers.**
Commit b6ee577e0 on this branch made the THIRD map (`dirMap`) lazy for exactly
this reason and left the two larger ones eager. Tracing every reader: Java and
no-csproj C# call `get` and never `getInsensitive`; PHP calls `getInsensitive`
and never `get`. Measured dead weight at 32k paths: Java 49.98 MiB of a 100.82
MiB index, PHP 34.49 of 69.85.

All three maps are now built on first use, and `lowerMap` is DERIVED from
`exactMap`'s insertion order rather than re-traversed — measured 330 ms against
389 ms today, so it is cheaper even for the two-map consumers. `pass-cache.ts`
hands the builder an already-lowercased list, so for TypeScript, JavaScript and
Vue the derivation is the identity and `getInsensitive` aliases the one map.

  java            80.26 -> 25.61 MiB retained   (-68%)
  csharp no-csproj 57.15 -> 21.52               (-62%)
  javascript       44.07 -> 22.65               (-49%)
  php              60.86 -> 32.09               (-47%)
  build @32k      562.1 -> 119.6 ms  (get-only), 329.5 ms (both)

The derivation is proven, not asserted: keys, values AND insertion order
byte-equal over 968,418 entries across case-colliding, Unicode-adversarial and
pathological corpora, plus 400 seeded-fuzz rounds. Order matters because it is
what makes `getInsensitive` return the first match in file order.

PHP additionally defers `filesByRawDirectory` (statically unreachable unless a
composer.json parses) and `firstProperSuffixMatch` (0 entries and 35.6 ms on
the bench corpus) to the branches that read them.

One suggested micro-optimisation was REJECTED with a counterexample rather than
taken: hoisting `suffixResolve`'s lowercase out of the extension loop assumes
`(s + ext).toLowerCase() === s.toLowerCase() + ext`, which is false for a
segment ending in Greek capital sigma — `("ΑΣ" + ".ts").toLowerCase()` is
`"ασ.ts"`, not `"ας.ts"`, because Final_Sigma is context-sensitive and `.` is
case-ignorable. A file named `ΑΣ.ts` would have stopped resolving. 16
mismatches in 2,171,190 checks, for 8.7%.

**The heap arms had become ceilings over nothing.** `retainedIndexBytes` read
only `index.all.length`, so once the maps went lazy it built none of them and
reported ~0 B — passing every ceiling. All heap arms now route through
`retainedPassBytes`, resolving a real missing import through the real resolver,
so the maps measured are the maps production forces. Two further measurement
defects surfaced while fixing it: PHP reaches the index through a second memo,
so the ephemeron chain needs four GC cycles and was reporting 249,208 B for a
9.3 MB index; and `bytes_large` carried an ~11% rope-flattening bias that made
every ratio read 0.85-0.96 for structures that are linear (now 0.998-1.017).

A `heap_floor_fraction` arm was added — a ceiling can only say "not too big" —
and proven by simulating the exact regression: `16 B at 32000 files < floor
17325000 B — this arm has almost certainly stopped MEASURING`.
`csharp_csproj` is now gated too: its old exclusion as "a duplicate of csharp"
held at +20.8% and is false at 2.47x.

**Three silent-coverage holes in the bench.** `LANGS` was a hand-written
literal claiming to mirror `SCOPE_RESOLVERS` while never importing it — the
seam that let JavaScript ship ungated; it is now derived, with an inventory arm
reconciling both directions. Four per-language budget lookups compared against
a possibly-`undefined` value, so deleting a key deleted the gate. Five
dispatchers ended in bare fallthroughs meaning "ruby" and "csharp", so a
mistyped language would have been benchmarked as Ruby's corpus under C#'s
resolver, forever green.

REPS is now chosen per language (15 below 5 ms, else `clamp(ceil(150/ms),7,15)`)
rather than globally by the noisiest cell: timing phase 39.8 -> 28.7 s, with the
six reduced-N languages showing peak-to-peak 1.008-1.071, no worse than the
eleven that kept 15. Worst headroom across all 85 cells is 0.71 of budget.

`depth_budget` for csharp 3.5 -> 2.2 and java 3.4 -> 2.2: their ratios fell to
1.438/1.402 because the lazy maps stop the deep arm paying for a map it never
reads. The file's own note said 3.5 did not lock that win in; 2.2 does.

Also fixes a raw NUL byte that made `suffix-index-lazy-dir-map.test.ts` BINARY
to git — all 395 lines were invisible to diff, blame and grep. The repo
documents this exact hazard in `route-extractors/dispatch-guard.ts`. That file
now also carries the guard the refactor lacked: eight arms pinning one-map-per
consumer and zero-extra-pass derivation, each proven against four mutations,
including a fused-eager rebuild that moves no total and is caught solely by the
at-construction count.

All 17 bench fingerprints and all 85 per-scale tuples unchanged. 1772 unit
tests, 12 adapter guards, tsc clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

* perf(python): memoize the importer's ancestor chain per directory (#2913)

Python's file index was always depth-free; the resolver was not.
`hasRepoCandidate` and `resolveAbsoluteFromFiles` each rebuilt one ancestor
prefix per directory component of the importer on EVERY import, and the
index's own `dirPrefixes` build inserted one entry per component per file.
So an import from `a/b/c/d/e/f/mod.py` did ~6x the prefix work of one from
`a/mod.py` regardless of corpus size — `depth_ratio` 7.239 where the next
worst language sat at 3.446.

The prefixes are a pure function of the importer's DIRECTORY, so they are
memoized per directory inside `getPythonFileIndex` (`ancestorsByDir`), which
is itself already per-file-set. Three smaller cuts came out of profiling the
same delta: the leading segment is rejected up front against a set of nested
directory names, the module and package buckets are consulted before the
walk instead of inside it, and the `dirPrefixes` build stops at the first
ancestor already stored.

Measured over 6 serial runs: depth_ratio 1.748-1.872 against 7.239, and at a
fixed 400 files the per-import cost at 18 directory components drops 6.761 ->
1.065 us. All five python fingerprints are byte-identical, so this is a
hoist; the budget retightening lands in the following commit, because
`_arms_note` is a single JSON line that also carries the heap-gate rewrite.

Also memoizes `pythonFileExportsName`'s `parsedFiles.find`, which was
O(files) for every import whose package probe resolved — the same shape
#2901 removed, keyed on `parsedFiles` rather than on `allFilePaths`.

The new gate is a count, not a timing: `ancestorsByDir.size` after N imports
from D directories must equal D, paired with a reference-identity assertion
so a memo that rebuilds AND re-stores still fails. `CountingSet` cannot see
this defect — the chain derives from the `fromFile` string and a rebuilt
prefix traverses the file set zero extra times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NtGZN6YSNU8chALSnKYHBp

* fix(import-target): close the eleven findings from the #2911 review

Seven P2s and four P3s. Every one is a gate that could not fail or a
comment that had become false; no shipped behaviour defect was found, and
all 85 per-language fingerprints are unchanged.

GATES THAT COULD NOT FAIL

- The C# namespace-dir memo was keyed on a materialized array, so a
  one-character `[...normalized]` copy at the adapter boundary minted a
  fresh WeakMap key per import while traversing the file set zero extra
  times: 67 tests stayed green and only a timing ratio caught it.
  `resolveCSharpImportInternal` now takes the Set and derives both arrays
  from `getWorkspaceFileIndex`, so there is ONE key shape and ONE
  instrument. Copying the Set now turns three arms red. Established first
  that `configs/csharp.ts` is test-only (`buildImportTargetWorkspace` has
  no production caller) and that both derivations are byte-identical —
  otherwise the rekey would have been a behaviour change, not a hoist.

- The contract test called `resolveImportTarget` with four arguments where
  `pipeline/run.ts:682` passes five, so everything behind `context` was
  ungated for all 16 languages: defeating PHP's `filesByDirectory` memo
  cost 197.0 -> 9,976.2 us/import (50.6x) with 248/248 tests green.
  `CountingSet` provably cannot see it — the builder iterates the
  `parsedFiles` array and touches the Set zero times — so the new gate
  counts own-index reads on `parsedFiles` through a Proxy. Only PHP and
  Python have a context leg; the other fourteen carry the floor anyway.

- Three heap budgets were read with no presence check. `ceiling * undefined`
  is NaN and `bytes < NaN` is false, so deleting `heap_floor_fraction`
  disabled the floor for all eight arms; deleting `heap_ratio_budget` did
  the same; and iterating the baseline's keys dropped a language whose
  ceiling key was deleted out of the gate entirely. All three now fail
  closed with a message naming the broken comparison.

- `HEAP_PROBE_TARGET` decided what each heap arm measured and was compared
  to nothing: repointing csharp_csproj at a non-matching namespace dropped
  it 73.70 -> 59.92 MB with `--check` still exiting 0. The four corpus
  fields are now asserted through the loop the timing scales already use,
  and the floor derives from a recorded reading rather than from a ceiling
  that is itself 1.5x the measurement.

- About 35 of the 86 PHP parity arms were structurally unable to fail:
  both sides called the same production helper, so deleting the `..` guard
  left them green. Every hand case now pins an absolute literal as well as
  the differential. Eight of those literals pin a bug or a documented
  limitation and say so rather than blessing the value.

- The registry inventory arm was weighed and KEPT, against the review's
  suggestion, on a structural number rather than a timing: the benchmarks
  job runs 9m23s against a 12m58s critical path, so its seconds buy no
  merge latency, and moving the arm to vitest would put the registry load
  ON that path while weakening what it reconciles. The "7.3 s" and
  "~46 -> ~42 s" figures it was justified with are corrected, including
  stating that only report mode got faster.

- python's `depth_budget` drops 11 -> 2.6 now that #2913 is in. 1.39x the
  measured maximum rather than the file's usual 1.5x, deliberately: at 2.8
  a revert of the nested-name rejection (2.734) would pass. The two parts
  of that fix this arm cannot gate are named, with the count-based arms
  that do gate them.

COMMENTS THAT HAD BECOME FALSE

- `pass-cache.ts` said it deduplicated "three byte-identical copies".
  JavaScript's had five fields and never called `buildSuffixIndex` — that
  missing field IS this PR's headline defect.
- The per-language census said nine where it is twelve, three of them
  added by this PR. Replaced in seven places with the mechanism that
  enforces it, which cannot go stale.
- `getFilesInDir` handed out the index's live bucket. Now `readonly
  string[]`, so mutation is a compile error; `.slice()` was rejected
  because `configs/python.ts` reads only `.length` and a per-import copy
  would reintroduce the term this PR removes.
- #2910 is the Java in-repo-namespace gap, not the JavaScript index defect.
  13 references corrected, the one correct Java use left in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NtGZN6YSNU8chALSnKYHBp

* perf(python,bench): flatten the bare-import walk, measure the context leg

Two follow-ups the #2911 review surfaced but left open.

BARE IMPORTS (`import os`) still walked every ancestor of the importer.
#2913 fixed the dotted tier; this tier lives in `import-resolvers/python.ts`
and no bench arm can reach it, because every python arm here spells its
imports with a dot and returns at the `pathLike.includes('/')` guard.

It also ran TWICE per `from x import y`: `resolvePythonImportTarget` probed
the package with `targetIncludesImportedName: true`, and on null — the
expensive case, having already walked to the workspace root — fell through
to a byte-identical call. Established that the two cannot differ before
collapsing them: the flag's only effect is to skip
`pythonImportedSubmoduleTarget`, so the recursion re-runs the outer frame's
entire tail on the same three references, and reaching the fallthrough means
that tail already returned null.

The walk itself is now a memoized chain plus an O(1) proof of absence
against the index's basename buckets. Its chain is NOT the one #2913
memoized and the difference is semantic, not accidental — no
`filter(Boolean)`, self excluded, workspace root included — so under an
absolute-path workspace the unfiltered chain probes `/abs/a/` where a
filtered one would probe `abs/a/`, a prefix of nothing. Two negative arms
pin that in both directions. The shared index moved to
`import-resolvers/python-file-index.ts` rather than being reached across a
cycle, which also collapsed a standalone memo into the one per-file-set.

12 / 24 / 72 Set probes at depth 1 / 4 / 16 become a flat 2. At 18 path
components, 11.615 -> 0.740 us/import (15.7x) and the depth curve is gone:
7.843 -> 0.925. Gated by probe COUNT, not timing.

THE BENCH CALLED `resolveImportTarget` WITH THREE ARGUMENTS where
`pipeline/run.ts:682` passes five, so no timing arm entered the `context`
leg for any language. Arity checked against the registry rather than the
comment: php and python declare five, every other hook three or four.
`parsedFiles` is built first and `allFilePaths` derived from it, matching
`run.ts`; fresh per pass, because the memos behind that leg key on the
array identity and `fastest()` takes a min.

Python's `parsedFiles` was structurally unreadable, not merely unread: the
arm passed a `namespace` spelling, which makes `pythonImportedSubmoduleTarget`
return null before the context is consulted. The import KIND had to change
too.

No fingerprint moved anywhere — on this corpus PHP's leg returns the same
file the cascade already did — which is exactly why the new `context` arm
asserts with-context against without-context instead. Defeating PHP's
`filesByDirectory` memo now costs 1003.7 ms against a 148 ms budget; before
this the bench could not see it at all.

Re-recorded on a quiet box, maxima over 5 serial runs: php small 27.762 ->
34.023 and heap 37.6 -> 49.6 MB (`filesByDirectory` is now retained for the
pass), python small 1.76 -> 4.358. `depth_budget.python` moves 2.6 -> 2.2,
because the added work is depth-FLAT: absolute cost doubled while the ratio
FELL to 1.563, so the old budget had gone slack. Both lock-in figures were
re-measured under the new call shape rather than carried over — reverting
the ancestor memo scores 2.524, reverting the nested-name rejection 2.553,
so each fails at 2.2 with 13% to spare.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NtGZN6YSNU8chALSnKYHBp

* refactor(import-target): make the key-shape rule a type, drop three censuses

Cleanup pass over the #2911 review-fix commits. No behaviour change: all 85
per-language fingerprints, every `resolved` and every `distinct_outcomes` are
byte-identical, and the targeted suite is 1851/1851.

MEASURED — `byBasename` was 71% empty array slots

`byBasename` holds roughly one bucket per file, and building each with `[]`
followed by `push` makes V8 grow the backing store to its 16-slot minimum, so
every single-file bucket retained 15 empty pointer slots. Constructing the
one-element bucket directly is byte-identical in contents and 5.50 -> 1.60 MiB
at 32000 `.py` paths. The bench arm reads 10543848 -> 6360936 B (-39.7%);
`heap_reading_bytes.python` and its ceiling are re-recorded. The same edit
shares one `{ raw, norm }` between both maps instead of allocating a second
literal for every `__init__.py`.

THE RULE THAT COST A TIMING RATIO TO FIND IS NOW A COMPILE ERROR

`perFileSet`'s key is narrowed from `object` to
`ReadonlySet<string> | readonly ParsedFile[]`. Reintroducing the #2911 defect
shape — a memo keyed on an array materialized from the file set — now fails
with TS2345 instead of silently minting a fresh `WeakMap` key per import while
traversing the Set zero extra times, which every scan-counting guard reads as
green at its correct value.

That also retires the header's hand-maintained roster of `ParsedFile[]`-keyed
call sites, which listed three — this PR added a fourth in `395c707d4` and did
not update it. A census inside a comment warning that censuses go stale, stale
inside one commit. The header now names shapes; the compiler names sites.

Two more claims that had drifted from their code:

- `per-file-set.ts` asserted "No index derived from the file set is keyed on an
  ARRAY materialized from it". `configs/swift.ts` is, deliberately, with its
  reasons written down. Two files in one directory disagreeing is worse than
  either; the rule now states what the type rejects and names the exception.
- `SuffixIndex.getFilesInDir`'s doc explained that it returns the index's own
  bucket by reference. True of `buildSuffixIndex`; the other implementation of
  that interface, in `languages/php/import-target.ts`, returns a filtered copy.
  The interface now carries only the caller-facing contract (`readonly`, do not
  mutate) and the sharing rationale moved onto the implementation it describes.
- The contract test still described Python as having "NO memo on this key".
  `parsedFileByPath` landed in `395c707d4`; the floor of 1 is now its single
  build rather than a per-import scan.

DEDUP

`importerDirOf` replaces four copies of `replace / lastIndexOf / slice` — two in
production, where one was a memo KEY and the other a memo's query argument, so
the two per-directory memos in one index agreed only by inspection. The tests
keep their own verbatim derivation on purpose: importing production's would
make the key lookup agree by construction and hide a regression.

`buildParsedFiles` maps through `probeFile` instead of repeating its 7-field
literal 900 lines away; `requireNumericBudget` and `expectNoOrphanKeys` replace
three and three copies, with every per-arm `why` kept per-arm. The two Python
memo guards collapse onto shared arms in `test/helpers/counting-file-set.ts` —
1847 tests before and after, and both still go red under mutation.

SKIPPED, with reasons: dropping `normSet` for bucket scans (trades O(1) probes
on the hot path for ~1.6 MB against a 6.4 MB reading); measuring heap for all
17 languages (+9 s and a design decision, not a cleanup); `readonly` on the
five sibling resolvers' array parameters and the `getDirMap` slice/join rewrite
(both correct, both outside this diff).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NtGZN6YSNU8chALSnKYHBp

* perf(import-target): rewrite the dirMap build, gate heap for every language

The three items the /simplify pass deferred, plus what measuring them found.

`getDirMap` BUILD — 226.9 ms -> 173.1 ms at 32 000 paths

It built every key with `dirParts.slice(j).join('/')`: one parts array, one
slice array and one joined string per file per directory component, in the map
its own doc calls "by far the most expensive" of the three. Now a
`lastIndexOf` walk slicing substrings out of the original string — the same
rewrite `getExactMap` already records at 357.4 -> 264.5 ms.

The key set is identical, not merely equivalent: 272 956 keys over a 32 000
path corpus carrying absolute paths, leading/interior/trailing doubled
separators, Windows separators, extensionless files, dotfiles, dotted
directories and colons, run both slash-normalized and raw. Zero differences in
keys, in key INSERTION ORDER, in bucket contents, in bucket ORDER, or across
767 732 probes through the real index. Bucket order matters because `php.ts`
reads `[0]`.

READONLY on the per-pass shared arrays

`WorkspaceFileIndex.normalized`/`.all` and the `normalizedFileList`/
`allFileList` parameters of jvm, php, ruby, go and standard are now
`readonly string[]`. This PR already made that argument for one bucket
accessor; these are the two biggest arrays held for a whole pass, and the
blast radius of an in-place sort is larger. Types only — no cast, no copy —
and it let two pre-existing `as string[]` casts in
`languages/typescript/import-target.ts` be deleted rather than added to.

HEAP IS NOW MEASURED FOR ALL SEVENTEEN LANGUAGES, AND THE PROSE WAS WRONG

Nine were excluded on measurements taken once and never re-checked, with the
re-entry condition stated in a comment and watched by nothing. Measuring them:

- go, dart and kotlin had NO stated reason at all — the header said "six of
  seventeen" against a list of eight. kotlin retains 45.85 MiB, the
  second-largest reading in this file, larger than ruby's and java's;
- swift and cobol were recorded as below-noise (0.29 MB, 0 B). They read
  3.29 MB and 2.21 MB and grow the right way. The arm changed under them —
  #2903's real-import probe, then corpus flattening — and nobody re-took it;
- the header quoted javascript at two different values four paragraphs apart.

Only rust's exclusion survived: 16 B at both scales, identical over five runs.

Six of the nine are now FULLY budgeted rather than merely bounded — ceiling,
floor and ratio — because each grows linearly (0.996-1.004 against a 1.25
budget). cobol, swift and rust keep an upper bound and no floor, deliberately:
a floor over a reading at or below its own noise gates the noise. Proven live:
restating kotlin's reading so its floor clears the real measurement fails with
"this arm has almost certainly stopped MEASURING rather than started saving" —
the failure that once left four arms at 0 B under passing ceilings.

Cost: +1.37 s in the heap phase, measured per language rather than asserted.

`normSet` was NOT removed, and the reason is now in the code. It is derivable
from the two buckets, but `byBasename` is keyed on BASENAME: on a 9 000-file
service tree `utils.py` and `models.py` hold 1 000 entries each, so `import
utils` would scan every `utils.py` in the workspace per import — the exact
defect class #2901/#2902/#2908 removed. ~1.6 MB against a 6.4 MB reading buys
both probes staying O(1).

All 85 per-language fingerprints unchanged; 1854 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NtGZN6YSNU8chALSnKYHBp

* test(php): drop the impossible undefined comparison from the parity copy

CodeQL (js/comparison-between-incompatible-types, alert 945) flags the
`ctx === undefined` arm of the legacy adapter copy: `WorkspaceIndex` is an
object type at that position, so the comparison can never be true.

Optional chaining expresses the same guard without the type-level clash —
an undefined index still fails the `typeof` test and returns null — so the
copy remains behaviourally verbatim against the shipped adapter, which is
the only property this harness relies on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017oEY2i74d1HLa5FuGVuPiT

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 17:22:51 +01:00
azizur100389
4576adfc46
fix(java): emit Record interface heritage (#2916)
* fix(java): emit Record interface heritage

Synthesize inheritance references for Java record implements clauses so scope resolution emits canonical heritage and interface-dispatch edges.

* test(java): cover Record heritage review gaps

Document deferred enum and implicit-accessor behavior, make assertions order-independent, and add Record heritage to the capture benchmark.
2026-08-10 13:35:16 +01:00
Carter LaSalle
81100e2c74
fix(python): resolve calls through __init__.py re-exports (#2864)
* fix(python): resolve calls through `__init__.py` re-exports

A call to a name imported from a package never resolved when the package's
`__init__.py` re-exported it rather than defining it:

    pkg/impl.py       def target_fn(x): ...
    pkg/__init__.py   from pkg.impl import target_fn
    caller.py         from pkg import target_fn
                      def calls_it(): return target_fn(21)   # no CALLS edge

`caller.py` gets no CALLS edge. Both IMPORTS hops are recorded, and all four
functions are extracted as nodes — only the call binding is missing. Because
`__init__.py` re-exports are how Python packages declare a public surface, this
misses a large fraction of real call edges, and the failure is silent: the
defining file looks like dead code with zero callers.

The re-export closure that should carry this already exists and is fully general
(`buildReexportClosures` — SCC over the re-export subgraph, bounded fixpoint for
cycles, transitive `via` chains). Python just never fed it: the subgraph admits
only `kind: 'reexport'` and `kind: 'wildcard'`, and Python emits neither for
`from m import x`.

Python has no dedicated re-export form. A module-level `from pkg.impl import X`
binds X locally AND publishes it as `pkg.X`, so it is both a named import and a
re-export. Emitting `kind: 'reexport'` would be wrong — that form drops the local
binding, which Python's does create. Instead add an optional `reexportsName` flag
to the `named`/`alias` variants, alongside the existing provider-specific
`importedSymbolKind` / `targetIncludesImportedName` flags, and admit flagged
imports into the closure subgraph. Languages with an explicit form keep emitting
`kind: 'reexport'` and leave the flag unset, so nothing changes for them — a
negative-control test asserts a plain named import still does not resolve.

Verified on a fixture covering the three shapes (direct, top-level-via-re-export,
function-local-via-re-export): 1 of 3 CALLS edges resolved before, 3 of 3 after.

On a 12.4k-file Python/Go/TypeScript repository: edges 294,416 -> 301,443
(+7,027) and execution flows 300 -> 813. A previously "100% orphaned" module
(`shared/db/event_writer.py`) now correctly reports its caller.

5 new finalize tests (single hop, 3-hop chain, alias keying, cycle termination,
and the negative control) plus 6 updated Python fixture shapes.
`npx tsc --noEmit` clean in both packages; full unit suite shows no regression
against baseline (remaining failures are pre-existing load-sensitive flakes in
analyzer-identity / evidence-provenance-helper / skip-git-cli / hooks, each
verified passing in isolation).

* fix(python): set reexportsName only for module-level imports

`interpretPythonImport` flagged every `from m import x` as republishing the
name, but only a module-level statement does. A `from m import X` inside a
`def` or `class` body binds locally and puts nothing in the module namespace,
so flagging it fabricates a re-export of a name no importer can reach:

    # pkg/__init__.py
    def loader():
        from pkg.impl import InternalHelper
    # caller.py
    from pkg import InternalHelper      # CPython: ImportError

resolved to `def:pkg.impl.InternalHelper`. Worse, with declaration-order
first-wins in the closure, a scope-blind entry could claim a name ahead of the
real module-level import and give a WRONG def for legal, running code.

`interpretImport` receives a `CaptureMatch`, which is `{name, range, text}`
with no syntax node, so the scope is not recoverable there — and it is not
recoverable downstream either: `pass3CollectImports` applies no scope filter
and `ImportEdgeDraft.fromScope` is hardcoded to the module scope. The decision
therefore moves up to `import-decomposer.ts`, which still holds the live
`import_from_statement` node, and rides down as an `@import.publishes` marker.
Computed once per statement, not once per imported name, with the existing
`findAncestorBeforeBoundary` helper.

Only `function_definition` and `class_definition` suppress publication.
`if` / `try` / `for` / `with` do NOT — Python has no block scope — so the
predicate is an ancestor walk for those two node types and nothing else.
Verified against CPython 3.11 in both directions; both are now pinned by
tests, including the counterpart control that a branch-nested import still
republishes.

Also corrects the docblock in `scope-extractor.ts` that sent this change the
wrong way. It claims pass 3 attaches imports "not to any `Scope` — finalize
reconstructs the owning scope via `provider.importOwningScope` during Phase
2". Finalize does no such thing: `importOwningScope` is declared on
`LanguageProvider` and implemented by a dozen providers, and
`grep -rnE "\.importOwningScope\b" gitnexus/src/` returns exactly one hit —
that doc comment. Nothing invokes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rzsb6mdGtbu66BG1EaF6Zz

* fix(shared): stop guessing ambiguous and namespace re-exports; bound the via chain

Four changes to the re-export closure, all reachable only now that Python
feeds it.

1. AMBIGUOUS NAMES ARE DROPPED, NOT GUESSED. `populateFileClosure` documented
   "declaration order first-wins for duplicates of the same exported name",
   which is sound only where a duplicate export is illegal — two
   `export { X } from …` is a TypeScript compile error, so the rule never
   fires. Python has no such guarantee:

       from .v1 import Client   # legacy, left behind
       from .v2 import Client   # the actual public Client

   CPython binds v2 (verified on 3.11); first-wins attributed every
   `from pkg import Client` in the repo to the DEAD implementation, and
   `impact("Client")` pointed at the wrong file. Last-wins is not the fix
   either: for the equally common `try:`/`except ImportError:` and
   `if sys.version_info` pairs exactly one branch runs, and which one is not
   decidable here. Both directions are wrong on real code, so the entry is
   dropped — the importer stays unresolved, which is exactly the pre-#2864
   answer, and the file-level IMPORTS edge is untouched.

   `collectAmbiguousReexports` runs as a PRE-PASS over data phase 0 froze,
   so the poisoned set is constant across the fixpoint. That matters: a set
   that grew mid-fixpoint would need retraction to propagate to files that
   already inherited the name, would make `myClosure.size > before` an
   unsound progress signal, and would invalidate the `|SCC| + 1` cap. As a
   pre-pass the closure map stays monotone and every existing termination
   argument survives unchanged. Only two flagged drafts resolving to two
   DIFFERENT in-workspace files count; duplicates of one target are
   harmless, and unresolvable targets never entered the closure.

   Checked in both loops. Named re-exports take precedence over wildcards,
   so suppressing only the named loop would hand the name to a later
   `import *` and reinstate an arbitrary winner through the back door.

2. NAMESPACE-RECLASSIFIED DRAFTS ARE EXCLUDED. The admission guards tested
   `draft.source.kind` while `tryFinalize` tests the post-reclassification
   `draft.base.kind`. Python's `from . import logger` is emitted as `named`,
   reclassified to `namespace` by `isNamespaceImport`, and was still
   admitted — republishing whatever def shared the module's simple name. For
   a `logger.py` holding a module-level `logger = logging.getLogger(...)`,
   importers of `from pkg import logger` bound to that Variable instead of
   the module. Reproduced end to end. Both predicates now take the draft and
   test `base.kind`; this is a no-op for TS/Rust, whose only
   `isNamespaceImport` implementation is Python's.

3. `transitiveVia` IS CAPPED AT 32. Each hop copies the inherited path, so
   an unbounded chain is Theta(depth^2) in time AND retained memory, and
   Theta(|SCC|^2) for a cycle whose chain tracks it. `MAX_REEXPORT_DEPTH =
   100` covered this until fc919ad6 removed it — correct for the shallow
   TypeScript barrels that were then the only input, and invisible until the
   input class changed. Measured at depth 400: 67 ms / 145 MB uncapped vs
   25 ms / 40 MB capped. 32 against a real-world worst case of ~6 for
   `__init__.py` chains. Safe because `ImportEdge.transitiveVia` has no
   production reader — it is diagnostic provenance, emitted and typed but
   dropped by graph emission.

4. `localDefs` ARE INDEXED BY SIMPLE NAME. `findExportByName` linearly
   scanned a target's defs on every call, and the phase-3 fixpoint rescans
   the same target once per iteration. Memoized on the array identity, which
   `FinalizeFile` documents as static input. Worth 12-14% where lookups
   repeat and neutral elsewhere.

The 46-line algorithm docblock was also ORPHANED by the helpers inserted
between it and `buildReexportClosures` — AST-verified, that function had zero
jsdoc blocks, so the cross-reference elsewhere in the file landed on an
undocumented function. Helpers move below it (declarations hoist), and its
step 1, precedence and complexity sections are rewritten: they still claimed
regular imports do not contribute to the export surface, and justified the
via-copy cost by TypeScript barrels being shallow.

The `reexportsName` contract consolidates onto `ParsedImport`, where its
"`kind: 'reexport'` would drop the local binding" rationale is corrected —
`materializeBindings` creates a module-scope binding for every linked edge,
re-export included. The real reasons are that `origin` flips, changing
evidence weight and priority, and that it misreports Python's syntax.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rzsb6mdGtbu66BG1EaF6Zz

* test(shared): add a re-export closure scaling guard to CI

No bench covered `buildReexportClosures` at all. Until #2864 its input was
TypeScript barrel files — a handful of shallow edges — and it admitted only
`reexport` and `wildcard` drafts. It now admits every module-level Python
`from m import x`, measured ~20x more edges on the CPython stdlib and cyclic
SCCs where there were none. The pass went from "rarely runs" to "runs over
the whole named import graph" with nothing watching it.

The regression this guards has already happened once: fc919ad6 removed
`MAX_REEXPORT_DEPTH`, which was correct for shallow barrels and stayed
invisible for as long as the input stayed shallow.

The depth arm is an EXACT structural assertion — build a chain far past the
cap, assert the longest emitted `transitiveVia` is exactly `MAX_VIA_LENGTH`.
It started as a `depth_ratio` timing arm and that was a bad gate: sampled
five times capped it scored 2.71-3.52 and three times uncapped 5.87-7.65, so
the ranges nearly touch and one uncapped run came in UNDER budget. A gate
that passes a third of the time on a broken build is worse than none, because
it gets read as evidence. The structural form fails 3/3 with 401 vs 32.

`width_ms` stays a timing arm with a deliberately loose budget, because a
structural check cannot see a constant factor: restoring a per-lookup linear
scan of `localDefs` leaves every array length untouched while making every
real analyze slower.

Both arms drive `finalize` through INDEXED hooks. Reusing the unit tests'
`defaultHooks` is the trap — its `resolveImportTarget` does `files.some(...)`
per import, which is O(imports x files) in the FIXTURE and swamps the pass so
completely that removing the cap measures as no change at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rzsb6mdGtbu66BG1EaF6Zz

* fix(cache): bump SCHEMA_BUMP 53 -> 60 for ParsedImport.reexportsName

`reexportsName` is a new field on `ParsedImport`, and `parsedfile-store.ts`
serializes the whole `ParsedFile` generically — so it is part of the cached
shape even though it is not a capture, which is the easy-to-miss variant of
the rule `parse-cache.ts` states as a MUST. (The `@import.publishes` marker
added alongside it moves the capture output too, so this qualifies twice; the
python captures golden confirms the drift.)

Without the bump, a warm `parsedfile-cache` replays pre-fix `ParsedImport`s
carrying no flag, `isNamedReexport`'s strict `=== true` takes the old path,
and the entire fix is a SILENT NO-OP on incremental analyze while every
cold-run test passes. It lands hardest on `__init__.py` — the rarest-changing,
highest-cache-hit files in a Python repo, i.e. exactly the target. A published
npm release invalidates via `GITNEXUS_PKG_VERSION`; dev trees, main-HEAD
installs and CI with a restored cache dir do not.

60, not 54, because the value has to clear every in-flight claim rather than
just origin/main: main is at 53 while open PR #2899 claims 54 and #2891 claims
59. Five exact clashes are recorded in the ledger, and the pin test cannot
detect a tie — both sides assert the same number and both pass. RE-CHECK
against origin/main immediately before merging.

Also documents the divergence between `pythonFileExportsName` and the
re-export closure. That predicate answers "does this package expose X?" from
`localDefs` alone, so with `pkg/__init__.py: from .impl import log`,
`pkg/impl.py: def log` and a same-named `pkg/log.py`, `from pkg import log`
still targets the submodule and the closure is never consulted — for exactly
the case it was built for.

Deliberately NOT fixed by reusing the flag, which is the obvious three-line
change and is WRONG: `reexportsName` is also set for `from . import log`,
where CPython binds `pkg.log` to the MODULE, not a name (verified on 3.11
against the `from .impl import log` form, which binds the function). Returning
true there would kill the correct namespace edge. Separating the two needs the
re-export's own resolved target — i.e. re-entering `resolvePythonImportTarget`
from a different `fromFile` — and that classification is the subject of open
issue #2882, so it belongs with that fix. Not a regression: both halves behave
exactly as they did before #2864.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rzsb6mdGtbu66BG1EaF6Zz

* test(python): re-baseline the scope-capture fingerprint for @import.publishes

CI's `bench/python-scope/measure.mjs --check` failed on capture fingerprint
drift. Intentional: the module-level marker added for `reexportsName` is a new
synthetic capture, and that guard hashes `tag|text|range` over every
`emitPythonScopeCaptures` output.

Attributed before re-baselining rather than after. Reverting ONLY the
`@import.publishes` emission — nothing else — restores the previous hash
a0da3e7c exactly, so the whole drift is that one marker. `capture_groups_fp`
is 3246 either way and `scaling_ratio` stays ~1.0, so no capture group
appeared or vanished and the pass is still linear.

The other nine bench guards were run rather than assumed: scope-capture,
callable-value-flow, finalize-reexport, cpp-qualified-ns,
kotlin-import-target, receiver-resolution, scope-emission, import-target and
cfg all pass. The benchmarks job runs under `-e`, so this failure masked
whatever followed it — worth checking the rest before pushing a one-line
baseline change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rzsb6mdGtbu66BG1EaF6Zz

---------

Co-authored-by: Carter LaSalle <carterlasalle@gmail.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:21:06 +01:00
DuduPhudu
fa31a7d824
fix: close the nine follow-up review findings from #2856 (routes, receiver typing, truncation honesty) (#2899)
* fix(typescript): a type parameter shadows a declared type of the same name (W2-8)

First item of wave 2, promised to the reviewer on #2856.

`export function unwrap<Result>(value: Result): Result` names the PARAMETER, not
the `interface Result` beside it — tsc resolves both annotations to the
parameter. The type-reference capture that makes a contract answerable ("what
breaks if I remove this field?") had no notion of a parameter binding, so every
annotation mentioning `Result` inside `unwrap` minted a `USES` edge into the
interface, at the same confidence as a real consumer and indistinguishable from
one. Measured on the new fixture: `unwrap` produced TWO false edges while the
genuine consumer produced one.

Blast radius is every generic whose parameter name collides with a declared
type, and the colliding names are ordinary choices for both: `Result`, `Key`,
`Value`, `Item`, `Node`, `Options`, `Config`, `Props`, `State`, `Response`.

TWO HALVES, and the first is why upstream's fix could not reach this. #2833
introduced `bindsTypeParameter` for the CALL-receiver path, where a workspace
`class T` was answering for `<T>`. Reusing it here changed nothing at first, and
the reason is its own documented contract: `@declaration.type-parameters` was
captured for class/interface declarations ONLY, so a generic FUNCTION recorded
no parameter list and the predicate correctly returned false — absence is not
evidence. The data was missing, not the logic. So:

  - TYPESCRIPT_SCOPE_QUERY now captures type parameters on `function_declaration`,
    `generator_function_declaration` and `type_alias_declaration`;
  - the graph bridge consults `bindsTypeParameter` before emitting `USES`.

Both are load-bearing — removing either one fails the fixture.

The fixture carries two controls, because the obvious wrong fix is to stop
emitting: a genuine consumer of the interface must still link, and a generic
whose parameter does NOT collide must still link its real reference. Both are
asserted, and the "genuine consumer" case is asserted FIRST so the absences
below it cannot pass vacuously.

SCHEMA_BUMP 53 -> 54: parse-time capture change. A warm cache replays defs with
no parameter list, so the guard reads nothing and the feature is inert while
looking implemented.

Capture fingerprint re-baselined with justification. NO NEW CAPTURE NAME —
diffing the capture-name sets against the wave-1 branch returns empty; the tag
existed and now fires on more declarations. capture_groups_fp 2338 -> 2371,
fixture_count 151 -> 152, scaling 1.06 < 1.5, and JavaScript's fingerprint does
not move at all, which is the check that this is the TS declaration rules rather
than something broader.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): close the four false-success paths in the graph-write-collapse guard (W2-6)

Second wave-2 item, promised on #2856. All four were reported; all four
reproduced by reading the code they name.

(a) A SAME-COMMIT RE-RUN REPORTED SUCCESS FOREVER. Every other meta-driven
    trigger — schema fingerprint, PDG mode, runner identity, CJK segmentation,
    embedding dims — has a block that forces a rebuild before the
    `alreadyUpToDate` fast path. `graphWriteCollapsed` had none; `grep -rn` found
    writes and no reads. So the one state meaning "most of your edges are gone"
    was the one state that repaired itself only if the user happened to pass
    `--force`. Now forces a full rebuild, and forcing is right rather than merely
    re-running: the persisted graph disagrees with what the pipeline produced, so
    an incremental pass over unchanged files would write nothing and re-stamp the
    same broken index as fresh.

(b) AN INCREMENTAL RE-RUN ERASED THE STAMP. `saveMeta` is a full atomic
    overwrite, and the field was spread in only when the CURRENT run had a
    verdict. `undefined` meant two different things at that site — "full run, no
    collapse" (a positive all-clear) and "incremental write, not comparable" (no
    opinion) — so the second case silently dropped `graph-write-collapsed` from
    meta.json while the edges were still missing. Now three-way: stamp on
    detection, CLEAR on a healthy full run, CARRY FORWARD when there is no
    verdict. That is the shape `branch: branchLabel ?? existingMeta?.branch` two
    lines away had all along.

(c) THE SERVER PATH NEVER CONSUMED IT. `analyze-worker-ipc.ts` projects the field
    "so a server-side caller sees the same degraded outcome the CLI does" — but
    nothing read it, so the comment described an intention and every collapsed
    run reported `complete` to the UI and to every API consumer. Now reports
    `failed` with the counts and the remedy, matching the CLI, which prints
    `Repository indexed INCOMPLETELY` and exits non-zero. A consumer that reads
    "complete" will query the index and get confident wrong answers.

(d) --pdg ROWS MASKED TOTAL STRUCTURAL LOSS. `expected` counts the in-memory
    graph plus the streamed STRUCTURAL manifest; the streamed PDG layers never
    enter `graph.relationshipCount`. But `persisted` was `stats.edges`, a count
    of EVERY `CodeRelation` row, and PDG writes into that same table. With 1,000
    structural edges expected and 4,000 PDG rows persisted, losing every
    structural edge still read `persisted = 4000`, cleared the ratio, and stayed
    silent — on exactly the large repos `--pdg` is used for.

    Worth recording that the OBVIOUS fix does not work. Padding `expected` with
    the PDG rows makes the two universes match but leaves the ratio judging a
    minority population: 4,000 of 5,000 still clears 0.5. I wrote that first, and
    the test I wrote to prove it failed. Only comparing structural against
    structural asks the question the check exists to ask, so `getLbugStats` gains
    a `structuralEdges` count excluding `PDG_EDGE_TYPES`. `TAINT_PATH` is
    deliberately NOT in that set — it is a whole-program Function→Function edge
    persisted by the normal emit, so it is structural and stays counted on both
    sides.

    `index-freshness-graph-collapse.test.ts` had pinned the masking as correct
    (`detectGraphWriteCollapse(1000, 4000)` → undefined, "PDG layers write into
    the same table, so persisted > expected is normal"). True about the table,
    and it licensed the hole. Replaced with the case that matters and a note on
    why the fix is at the caller.

The new `structuralEdges` assertion in `lbug-core-adapter` is there because the
failure mode is silent: the query sits in a try/catch that yields `undefined`,
and `undefined` makes the collapse check decline to compare — so a typo in the
Cypher would throw nothing, fail nothing, and switch the guard off. Verified
against a real LadybugDB and mutation-checked by breaking the query.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(processes): make process selection insertion-order invariant (W2-5)

Third wave-2 item. Reproduced before fixing: two equal three-step flows with
`maxProcesses: 1` select `handleAlpha`; inserting the identical nodes and CALLS
edges in reverse select `handleBeta`. Same repository, same commit, a different
persisted graph — so a filesystem that enumerates differently, or an incremental
run that reorders assembly, silently changes what the tool reports.

Four sorts ranked by score or length alone and returned 0 on a tie.
`Array.prototype.sort` is stable, so a 0 preserves INPUT order, which traces
back to `graph.iterNodes()`. Under `maxProcesses` capping that decided which
`Process` and `STEP_IN_PROCESS` nodes were persisted at all. Each now falls
through to a totally-ordered, content-derived key — node id for entry points,
the joined path for traces.

WHAT IS ACTUALLY VERIFIED, stated precisely because "four fixes" would overclaim:

  - the ENTRY-POINT sort is individually mutation-verified;
  - the two DEDUP sorts are collectively mutation-verified;
  - the TRACE-RANK tiebreak is NOT individually observable, and the source says
    so. The dedup sorts already impose a total order on the list that reaches
    it, so removing it alone fails nothing. Kept as defence in depth: it cannot
    misbehave — it only makes an already-deterministic order explicit — and it
    is what stops a change to dedup ordering from silently re-opening this.

Finding that out took two fixtures. The first (three chains, three entry points)
is separated by the entry-point sort before trace ranking is reached, so it never
exercises the trace comparator at all; the second gives ONE entry point two
equal-length branches to different terminals, which is the only shape where the
trace comparator decides. Both are kept — they gate different sites.

The invariance tests assert the INVARIANT rather than any single sort, so they
cover all four sites and any future one without needing to know where they are.
Three assertions: same selection under a cap, identical set uncapped, and
identical ORDER — the last because order is what the cap consumes, so a set-only
assertion would pass while the defect persisted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(impact): UNKNOWN dominates a mixed candidate set, instead of reporting the known floor (W2-4)

Fourth wave-2 item. The all-UNKNOWN branch here was reasoned about carefully and
is correct — its comment even names the two ways a set can be all-UNKNOWN. The
MIXED case fell straight through it.

`RISK_ORDER` is `['LOW','MEDIUM','HIGH','CRITICAL']` and has no `UNKNOWN` entry,
so `indexOf('UNKNOWN')` is -1 and an UNKNOWN candidate can never win the reduce.
An ambiguous name with one caller-less candidate (UNKNOWN, per the round-1 fix)
beside one single-caller candidate (LOW) reported `maxRisk: 'LOW'` — a confident
floor over a set containing an interpretation nobody measured. That is the same
false-safe the all-UNKNOWN branch exists to prevent, one case over, and it
surfaced in the UI as "Max blast radius N (LOW risk)".

`maxRisk` answers "how bad could this be?", and an unresolved candidate could be
CRITICAL — so any UNKNOWN in the set makes the aggregate UNKNOWN. Narrowing it
that way would normally cost information, so the measured part travels alongside
as `knownMaxRisk`, present only when the two differ: absent on a fully-resolved
set, where it would duplicate `maxRisk`, and absent on a fully-unknown one, where
there is no measured part. A reader gets "at least LOW among what resolved, and
one interpretation could not be walked at all", which is strictly more than
either value alone. The human-readable message says the same thing.

The seed gained a mixed pair because the existing one could not reach this: both
its twins are caller-less, so it only ever exercises the all-UNKNOWN branch —
which is precisely why the gap survived a round of review. Three assertions,
both halves mutation-verified.

`eval-server.ts` needs no change: it renders `result.maxRisk ?? 'UNKNOWN'`, so it
now shows UNKNOWN where it previously showed the floor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(routes): track ternary polarity in dispatch guards, so a selected verb cannot be inverted (W2-9)

`if ((req.method === 'GET' ? false : true) && pathname === '/api/i')` emitted
`GET /api/i` — the one method that branch guarantees the request does NOT have.
A ternary SELECTS between its arms, so a verb inside one is not reached merely
because the whole condition is truthy, but `findVerbInSubtree` descended into
both arms and returned the first verb it saw. Same inversion `!` produced before
d4dcba8c, one level up.

Handled by folding the ternary where an arm is a boolean literal, which is what
collapses the selection into a conjunction:

    c ? A : false  ==  c && A     both hold, so search both
    c ? false : B  ==  !c && B    c must NOT hold, so search it at flipped parity
    c ? true : B   ==  c || B     a disjunction guarantees neither operand
    c ? A : true   ==  !c || A    likewise

Two non-literal arms leave the verb chosen by an unknown condition, so the
ternary guarantees nothing. Refusing every ternary would also have fixed the
reported bug, but three of the four shapes measured were ALREADY correct and
would have silently lost their verb; they are pinned now.

A second defect in the same walk, found while reproducing: the `!` rule was
keyed on PRESENCE, returning null at the first negation it saw, while
`isNegatedContext` two functions above states the rule is PARITY and says so
outright — `!!x` is `x`. So `!!(req.method === 'GET')` dropped a verb the source
states plainly. The existing double-negation test covered the PATH position,
where the parity walk already ran, and so never saw it. The verb walk now tracks
parity too, and the two agree.

Verb-less, not route-less: the path comparison is untouched evidence that the
branch serves that path, so an inverted verb becomes a missing verb rather than
a missing route.

SCHEMA_BUMP 54 -> 55. Routes are emitted at parse time and replayed verbatim
from a warm cache, so without the bump an already-indexed repo keeps serving the
inverted verb and the fix looks inert. Free against origin/main (48).

Every rule mutation-checked: removing the ternary dispatch, either literal-arm
rule, the negated-ternary guard, or the parity walk each fails exactly the tests
that claim it. One assertion I wrote survived all five mutations and was removed
rather than kept.

Not a recall win on crypto-trading-bot, which contains neither shape — this is
precision insurance for dispatchers that do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(routes): report every method a dispatch guard serves, not just the first (R3-8 part 1)

`if ((req.method === 'GET' || req.method === 'POST') && bundlesMatch)` is two
routes. The verb walk returned the FIRST verb it found, so `route_map` presented
a two-method route as GET-only and `impact` on the POST path found nothing.
Taken verbatim from the reporting repo's researchRunRoutes.js.

`governingVerb` -> `governingVerbs`, returning a list; `findVerbInSubtree` and
`verbFromTernary` likewise. A guard with several verbs emits one route per verb
via the new `pushPerVerb` — they share a path and a handler but not a method,
and `(method, url)` is the key every downstream consumer dedups and looks up on.

A disjunction yields ALL its verbs or NONE, which also fixes an over-attribution
the first-match rule had:

    req.method === 'GET' || req.method === 'POST'   ->  GET, POST
    req.method === 'GET' || isAdmin                 ->  no verb

The second is reached for ANY method when `isAdmin` holds. Reporting `GET` — as
first-match did — describes a route open to everything as single-method, which
is the direction this module treats as more expensive than saying nothing.
Negated, `!(A || B)` is `!A && !B`, so it excludes verbs rather than offering
them and yields none.

Generic descent deliberately stays FIRST-match rather than unioning across
children: an arbitrary node says nothing about how its children combine, and two
verbs found under one are far more likely unrelated than alternatives. `||` is
the one construct that genuinely means "either of these".

Pinned against regression: the pre-existing rule that distributes ONE verb
across an OR of PATHS must not start multiplying methods, and switch arms
inherit the full method set.

SCHEMA_BUMP 55 -> 56. Routes are parse-time output replayed verbatim from a warm
cache. Free against origin/main (48).

Four mutations, each failing exactly the tests that claim it: removing the
disjunction dispatch, dropping the all-operands rule, allowing a disjunction at
odd parity, and emitting only the first verb.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(routes): read `.match()` dispatch, and the capturing wildcard it needs (R3-8 part 2)

`RE.test(pathname)` and `pathname.match(RE)` are the same test with the operands
swapped. Only `.test` was read, which is why 28 of the reporting repo's 75 routes
still named the shared route table as their handler rather than the module that
serves them: those modules dispatch with `.match`.

THE CAPTURING WILDCARD, which is the part that made the rest inert.
`regexToRoutePath` accepted `[^/]+` and refused `([^/]+)` — `(` fell through to
the metacharacter bail. So the non-capturing form translated and the capturing
form produced nothing, and every existing test passed because every existing
test used the non-capturing form. The tests were written against the
implementation rather than against the corpus, and the reporting repo contains
no non-capturing path wildcard at all: a dispatcher captures the segment because
it needs the id. This alone also repairs the already-shipped `.test` rule.
A capture around anything that is NOT one segment still bails — `(.+)` spans
slashes — and the alternation is balanced, so `([^/]+` unclosed is not a match.

`.match` differs from `.test` in one way that matters: its result is USED, so it
is almost always BOUND, and the verb then lives in a later `if`:

    const runMatch = pathname.match(/^\/api\/research-runs\/([^/]+)$/)
    if (req.method === 'GET' && runMatch) { … }

Reading the verb off the CALL would report every one of those verb-less. So a
bound match records `name -> path` and the route is emitted where the binding is
TESTED, once per test site — one binding tested for GET and for PUT is two
routes. A reference counts only in a truthiness position (`&&`/`||` operand, or
a whole `if` condition), which is what separates `if (m && …)` from `m[1]`: a
read of the captured segment says nothing about dispatch and would otherwise
mint a duplicate route per use of the id. A binding never tested still emits one
verb-less route — the code did compute an anchored match against the path.

Regexes named by a same-file const resolve too (`pathname.match(POSITION_REPLAY_RE)`),
with the same ambiguity refusal the string-constant map uses: bound twice to
different patterns means dropped, because a half-right regex is a wrong route.

SCHEMA_BUMP 56 -> 57. Free against origin/main (48).

Nine mutations, each failing exactly the tests that claim it. TWO of my own
tests initially survived their mutation and were rewritten, not kept:
- the non-path-receiver case had no path token anywhere in the fixture, so
  PATH_TOKEN_HINT skipped the file and the assertion was satisfied by a file
  that was never examined;
- the negation case used `!m`, which never reaches the negation check at all —
  a `unary_expression` parent is not a truthiness position to begin with. The
  shape that exercises it is `!(req.method === 'GET' && m)`.
A declaration-site skip written alongside them proved unreachable for the same
reason and was removed rather than left to imply a hazard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(processes): report what the detection ceilings dropped, instead of logging it at debug (W2-3)

`processProcesses` has five ceilings - the entry-point trace quota, the
per-entry trace budget, `maxTraceDepth`, `maxBranching` and `maxProcesses` -
and every one of them fired silently. The result came back looking whole and no
consumer could tell it was a sample. The code's own comment already said so:

    // A silently truncating cap reads as "this is everything", which is the
    // same class of confident-empty answer this work is about.

and then only called `logger.debug`. A log nobody has enabled is not a
disclosure.

`stats.truncation` is additive, so every existing consumer of `totalProcesses` /
`crossCommunityCount` / `avgStepCount` / `entryPointsFound` is unchanged. It
carries one boolean to branch on plus a counter per ceiling, kept SEPARATE
rather than summed because they mean different things: unexplored entry points
mean whole flows are missing, while a depth-capped trace means a flow is present
but shorter than it really is.

`processesDropped` counts against the DEDUPED population, not the raw trace
list - the gap between those two is deduplication doing its job, and counting it
as truncation would report a permanent non-zero on every healthy repo.

`truncated` is DERIVED from the counters rather than set at each site, so a
ceiling added later only has to increment its own counter to be reported.

Surfaced at `warn` and NOT gated on `isDev`: "823 flows" printed without it
reads as the complete set, which is the confident-empty failure wearing its
other face - a confident-COMPLETE one. The debug line stays for the per-entry
detail it carries.

Seven mutations, each failing exactly the tests that claim it, including BOTH
directions of the flag: hardcoding `truncated` false fails the four positive
cases, and hardcoding it true fails the nothing-was-truncated case, which is
asserted first precisely so the positives cannot pass vacuously. The
`walksCutByBudget` fixture gives every node exactly `maxBranching` callees so it
asserts its own counter and not a neighbour's.

Also fixes a defect this work exposed: 10b0c7a1 (W2-5) embedded a RAW NUL BYTE
in `trace.join(...)` instead of the backslash-u escape the rest of the repo
uses. It behaves identically at runtime, but `file` reports the source as
`data`, and grep, git diff and code search treat it as binary - several greps
against this file silently returned nothing while I was reading it. main was
clean here; two other files carry the same raw byte from before this branch and
are left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): resolve members through a MEMBER-CALL producer's return shape (W2-1)

    const svc = new SignalService()
    const r = svc.make()
    return r.secretFlag        // <- no edge

`return-shape-members` types `r` to the producer that made it, but a member call
binds the spelling `svc.make`, and slicing that to its last segment leaves
`make` — a METHOD, never a callable binding in scope. The producer lookup failed
and the pass declined.

The limit shipped documented as needing inter-procedural receiver typing. It does
not. Measured on a fixture, the pipeline had already done the hard part:

  - `readMake -> Method:...SignalService.make#0` already resolves as an ordinary
    CALLS edge, so the receiver is already typed; and
  - `Property:...SignalService.make.secretFlag@N:C` already exists, because R3-4
    anchors a returned literal's keys to the METHOD that returns them, not only
    to free functions.

Both halves were present and unjoined — the same shape as R3-5 itself.

ADDITIVE, not a reroute. The new branch sits inside `if (producerFile ===
undefined)`, so it can only fire where the callable lookup already declined;
every reference that resolved before resolves identically, by construction
rather than by test.

Nothing new is inferred. The receiver is typed by the SAME predicate that typed
`r`, and it must itself resolve to a class — a receiver that cannot be typed
still declines, so `make.<member>` is never matched by name across the graph.
That fabrication is what the existing guards exist to stop and they all carry
over unchanged: the owner must resolve, its file must match the candidate's, and
`ownFilePaths` keeps the polyglot class registry from walking a JS read into a
Java field.

The owner segment is TWO parts for a method (`SignalService.make`) and one for a
free function (`makeSignal`), which is exactly how R3-4 qualifies each. That is
what separates two methods of one class returning the same key name from each
other AND from a free function of that name — the fixture gives `secretFlag`
three owners so a wrong resolution is detectable rather than a coin flip that
happens to look right.

Four mutations, each failing exactly the two tests that claim it: removing the
fallback, using the method alone as the owner segment, taking the producer file
from the reading file instead of the owner class, and dropping the
receiver-type requirement. 3,408 resolver tests pass, including
`polyglot-property-isolation`, which is the one this could plausibly break.

No SCHEMA_BUMP: this is a resolution pass over ParsedFiles, not parse-time
output, so a warm cache replays the same input and produces the new edges.

Measured on crypto-trading-bot: ZERO new edges, byte-identical at 62,158. Its
170 `const x = new Y()` bindings are overwhelmingly built-ins (Map, Set,
Promise, S3Client) rather than workspace classes whose methods return object
literals — it is a module-style JS codebase. Correctness fix for class-shaped
code, not a recall win on this corpus, and it should not be presented as one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): type a bare parameter from what its callers pass (W2-2)

    function readSpike(spike) { return spike.wickRatio }

had nothing to type `spike` from, so the read fell through to the 0.5 name tier.
That is the standing limit of R3-5 and, measured, by far the largest: 11,012 of
13,672 property edges on the reporting repo (81%) rest on that name guess.

The two facts needed were already extracted, for a different consumer. For JS
and TS among others, `callable-flow-captures` synthesizes:

    formal    owner=readSpike  binding=spike  parameter-index=0
    argument  source=s  parameter-index=0  direct-callee-name=readSpike

Joining them on (callee, parameterIndex) says which cell reaches which
parameter, and the argument's own binding is typed by the same
`findReceiverTypeBinding` a directly-bound receiver already uses. So the
parameter inherits the producer and `spike.wickRatio` resolves as evidence
rather than inference.

No new capture, no parse-time change, NO SCHEMA_BUMP. And deliberately not a
change to the callable-value-flow solver that owns these sites: that pass is
guarded by a fingerprint CORRECTNESS gate plus a timing budget, so this reads
the same facts and computes its own map.

AMBIGUITY DECLINES. A parameter whose callers pass different producers resolves
to nothing. Picking one would fabricate at the 0.9 PRECISE tier, which no
`minConfidence` floor can filter out — the same reason `buildConstantMap` drops
an ambiguous constant instead of taking the first.

Keyed by the formal's (scope, name), not by a definition id. The first attempt
used a def and measured `paramDef=NONE`: a parameter is not reachable through
`findValueBindingInScope` (its predicate is `isOwnableValueLabel`, which lists
Const/Variable/Property/Static because it exists for OWNERSHIP registration, and
a parameter is owned by nothing) and it is not a `local` binding either. The
formal site already states the scope its parameter binds in, which is enough.

Formals carry their DECLARING FILE in the key, so two same-named functions in
different files cannot answer for each other — dropping it makes both go
ambiguous and both readers silently lose their edge.

COVERAGE, counted rather than assumed. The synthesis skips an argument that is
itself a call result (an explicit `continue` in `callable-flow-captures`), so
`f(makeSignal())` emits no argument site and only the bound spelling
`const s = makeSignal(); f(s)` is served. That looked fatal until measured: in
the reporting repo, bare-identifier arguments outnumber call-result arguments
2,563 to 50 — 51:1. Extending the shared, benched capture synthesis for the 2%
case is not worth its risk.

Four mutations, each failing exactly the tests that claim it: keeping the first
producer instead of declining on conflict, dropping the read-site lookup,
matching a formal at index 0 regardless of the argument's index, and dropping
the declaring file from the formal key. Two of those could not be caught by the
first fixture at all — it had a single parameter and a single consumer file — so
the fixture gained a two-parameter callee and a same-named twin in a second file
before they were meaningful. The test helper also had to start filtering by
source FILE, or two different `readSpike` symbols merged into one count.

Measured on crypto-trading-bot: 36 reads left the 0.5 name-guess tier. 26 became
precise 0.9 edges (return-shape reads 1,130 -> 1,156, which is the whole delta),
and 10 became honest absences — the receiver was typed, the producer's shape was
known, and the member is NOT on it, so the site is claimed as disproved rather
than left for the name fallback to invent an answer for.

That is ~0.3% of the 11,012, and it should be reported as such. The 81% figure
is the size of the PROBLEM, not of this fix: the shape requires a bound
argument, a producer that returns an object literal, and a parameter read as a
receiver, and that intersection is narrow. The remaining name-tier reads are
mostly receivers no workspace producer types at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(processes,ci): anchor trace subsumption, cover the sink wiring, stop one bench guard hiding the rest (#2894, #2896, #2895)

Three follow-ups reported against #2856 after it merged. Each was reproduced
before it was fixed.

#2894 — trace subsumption matched mid-identifier.

`deduplicateTraces` decided whether one trace is a sub-path of another with an
UNANCHORED `String.includes`, so a match could begin in the middle of a node id:

    'X->AA->B'.includes('A->B')   ->   true

and `A -> B` was discarded as redundant against a chain `A` is not a step of at
all. Reproduced directly against the function before fixing.

Padding both keys with the separator makes `includes` match whole steps only.
Reported as measured-inert and that holds — the collision needs one node id to
be a strict suffix of another at a `->` boundary, which real ids
(`Function:<path>:<name>`) do not produce. Fixed anyway because the predicate
did not mean what the surrounding code says it means, in a function whose entire
job is deciding what to delete, and nothing pinned it.

`deduplicateTraces` is exported for the test, matching how `traceFromEntryPoint`
and `buildSinkFunctionSet` are already reached. The tests use bare ids because
the shape cannot be built from realistic ones — which is exactly why nothing
caught it. Alongside the regression case, two tests pin that GENUINE subsumption
still happens, prefix and suffix, so the fix cannot degenerate into "subsume
nothing" and pass the first test trivially. Mutation-checked: reverting the
padding fails the mid-identifier test and only that one.

The encoding assumes `->` never appears IN a node id; a C++ `operator->` would
defeat the join regardless of padding. Out of scope, but the assumption is now
written down where the join happens.

#2896 — the sink wiring was only ever exercised through its fail-open catch.

`processesPhase` reads `allFetchCalls` / `allORMQueries` off the parse output
inside a try/catch that falls open to "no sinks", and every phase-level test
omitted `parse` — so all of them took the CATCH branch and the success path had
no coverage. `getPhaseOutput` is a raw `as T` cast, so a field rename would make
the phase detect zero sinks while every test still passed, because zero sinks is
what they already assert.

The new test asserts the one thing only the success path can produce: a flow
ENDING at the sink while a longer chain continues past it. Its control is the
same graph with no `parse` dep, which must NOT produce that terminal — without
the control the assertion could pass for an unrelated reason. Also asserts
`processesPhase.deps` contains `parse`, so the read and the declaration cannot
diverge, and that a parse output missing those fields still fails open rather
than losing every process.

Mutation-checked, including the exact drift scenario reported: renaming
`allFetchCalls` at the read site, dropping `parse` from `deps`, and passing no
sinks to `processProcesses` each fail exactly the test that claims them.

#2895 — a failing bench guard aborted the job and masked every later guard.

Every step in the benchmarks job was fail-fast, so the first failing `--check`
aborted it and the rest reported `skipped`, which reads identically to "nothing
to do". Audited over 13 runs on #2856: the job succeeded zero times and the last
two guards executed zero times for the life of the PR, while two reviews read
the checks summary and saw nothing wrong. Both guards did in fact pass — that
was luck, not verification.

`if: ${{ !cancelled() }}` on all ten steps after the first, so one stale
baseline reports one red step instead of hiding nine. `!cancelled()` rather than
`always()` so an explicit cancel still stops the job instead of running seven
minutes of benchmarks nobody is waiting for.

The two steps easiest to miss are covered: `Receiver-resolution drop guards`,
whose `run:` sits twenty lines below its `name:` behind a long comment, and the
final `Cross-language pipeline benchmarks` step, which is not a `--check` and so
falls outside any grep for one — and is one of the two that never ran.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(parse): capture a fetch call site even when its URL is not a literal (#2897)

The `fetch` rule required the argument to be a string or template literal:

    arguments: (arguments
      [(string (string_fragment) @route.url)
       (template_string) @route.template_url])

so `fetch(url)` with a variable matched nothing at all. Measured across this
repository's own TypeScript sources: **44 of 47 fetch calls pass a variable**, so
94% produced no site.

That is what makes R3-6 look inert. The sink set is built entirely from
`allFetchCalls` / `allORMQueries`, so a function performing an outward call
through a computed URL was never a sink, no flow could terminate there, and the
sink-first ranking rule never changed an ordering. The feature was fine; the
signal underneath it was almost always empty.

The URL alternation is now OPTIONAL, so one match covers both shapes. The R3-6
sink set needs only WHERE the program reaches outward, not where to.

Route linking is untouched, by construction rather than by hope:
`processNextjsFetchRoutes` normalizes the URL first and skips anything that
yields nothing, so a URL-less entry cannot mint a FETCHES edge. Verified on this
repo — FETCHES went 8 -> 9 across the change, i.e. the widening added sink sites
without inventing route edges, which was the one real risk here.

Tested in BOTH JavaScript and TypeScript, since the rule is duplicated in each
query block and fixing one would have left the other blind:

  - a variable argument is captured, with no URL   <- the regression case
  - a computed argument (`fetch(buildUrl(), {...})`) likewise
  - a literal URL is still captured WITH its URL   <- route linking depends on it
  - a template URL likewise
  - exactly ONE site per call — an optional alternation must not make a literal
    match twice, which would double-count the site and could mint two edges
  - `prefetch('/x')` is still not a fetch

Mutation-checked: restoring the mandatory alternation fails six of the twelve,
three in each language.

SCHEMA_BUMP 57 -> 58. Parse-time capture output is replayed verbatim from a warm
cache, so without the bump an already-indexed repo keeps its empty sink set and
the fix looks inert — which is the failure this constant exists to prevent, and
would have reproduced the very symptom being fixed.

Not addressed here, and worth stating: this widens `fetch` only. The reporter's
broader point stands — anything keyed on FETCHES / QUERIES is only as good as
the extraction underneath it, and the ORM side has not been measured. A guard
that fails when a corpus known to contain outward calls yields zero sites is the
right follow-up; this change makes such a guard meaningful rather than
tautological.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(bench): re-baseline receiver-resolution for the two fixtures this PR adds

`receiver-resolution --check` failed on:

    countArm.totalDropsAllKinds: 140 -> 148
    countArm.bySiteKind.write:    11 ->  19

Investigated before touching the baseline, because a guard that exists to catch
unexplained movement should not be silenced by an unverified story.

WHAT IT IS: the count arm runs the real pipeline over a corpus that includes
`test/fixtures/lang-resolution/`, and this PR adds two fixtures there —
`member-call-producer` (W2-1) and `parameter-producer` (W2-2). Each returns an
object literal with two keys, and a producer writing its own returned key is a
write site the receiver recorder logs. Four each, eight total.

Attributed by dumping the individual drops rather than reading the aggregate:

    member-call-producer/src/producer.js   secretFlag, wickRatio  (2 lines) = 4
    parameter-producer/src/producer.js     source, wickRatio      (2 lines) = 4

The eleven drops already in the baseline are all `javascript-object-properties`
fixtures of exactly the same shape, so the new ones are not a new KIND of drop —
they are more of one the baseline already records. This is the first case the
guard's own failure message names: "a fixture was added".

WHAT IT IS NOT: `callDrops` — THE gate number, and `call`-only by deliberate
design because reads and writes "would inflate it" — is unchanged at 102. `read`
drops unchanged at 27. The SHAPE ARM shows no drift at all: no receiver spelling
moved between RESOLVES / VISIBLE-GAP / INVISIBLE-GAP, so no resolution
regressed.

HOW IT WAS ISOLATED, since the first attempt was misleading and the record is
worth having: reverting `return-shape-members.ts` alone did NOT reproduce it and
pointed away from W2-1/W2-2. Only a commit-level bisect was trustworthy —
`origin/main` OK, W2-8 OK, W2-3 OK, then W2-1 +4 and W2-2 +4, which matches the
fixture count exactly. A file-level revert leaves the fixtures in the tree, and
the fixtures are the cause.

The update is two numbers. Nothing else in the baseline moves.

Worth noting where this failure became visible at all: under the fail-fast
benchmarks job it would have aborted the run and shown the five guards after it
as `skipped`. It is legible here because #2895 — fixed in this same PR — now
lets every later guard run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): stop `--pdg` runs reporting a healthy index as INCOMPLETE

Every `gitnexus analyze --pdg` reported a graph-write collapse and exited 1 on an
index where every row had persisted. Reported by a user hitting it on a real
repo; introduced by this PR's own W2-6(d).

    Repository indexed INCOMPLETELY
    the pipeline produced 200,501 relationships but only 64,764 are readable

The index was complete: 200,190 rows present, 109,905 PDG and the rest
structural, all queryable.

WHAT WENT WRONG. W2-6(d) made the persisted side count STRUCTURAL rows only —
correct, and the reason is in its own comment: PDG writes into the same table, so
counting everything let PDG surplus mask real structural loss. But the expected
side kept using `graphEmitManifest.totalRows`, and that is a BUFFER-POOL SIZE
HINT which counts every streamed row. PDG streams through that same sink, so the
check compared a structural-plus-PDG expectation against a structural
measurement. On any repo with a PDG layer that is a guaranteed false collapse.

It compounds rather than merely misreporting: the run stamps
`graph-write-collapsed`, and W2-6(a)'s rebuild trigger — added alongside it —
forces a full re-analyze next run, which collapses again. A permanent rebuild
loop, on an index that was never damaged, at ~100s a cycle.

MEASURED RATHER THAN ASSUMED, because the first attempt was wrong. I first
subtracted PDG edges RESIDENT in `graph.relationshipCount`, rebuilt, re-ran the
failing command and got byte-identical numbers. Instrumenting the three terms
showed why:

    relationshipCount=20,825  graphManifestTotalRows=179,676
    pdgEmitManifest=absent    residentPdgInGraph=0

PDG is not resident in the graph AND has no separate manifest — it streams
through the ordinary `GraphEmitSink`. The reverted attempt is not in this diff.

THE FIX. A pair key cannot separate them: it is `From|To` NODE LABELS, and a CFG
edge shares `Function|Function` with CALLS. Only the write path sees
`relationship.type`, so the sink now counts a `structuralRows` subtotal there and
publishes it on the manifest. `totalRows` is unchanged — it still sizes the
buffer pool, which is what it was for.

WHY THIS SHIPPED UNCAUGHT, and what changed about that. The wiring test kept a
LOCAL MIRROR of the expected-count expression "because the production expression
is inline in a 3000-line function". A mirror cannot catch a term the original got
wrong. That expression is now an exported
`computeExpectedStructuralRelationships` which production calls and the test
imports.

It also takes the MANIFEST rather than a pre-selected number, deliberately: the
defect was choosing the wrong FIELD, and a numeric parameter leaves that choice
at a call site no unit test can reach. Verified — with the helper taking a
number, reverting to `totalRows` failed nothing; taking the manifest, the same
revert fails four tests.

Verified end to end on the reported command: `analyze --force --embeddings 0
--pdg` now exits 0 with "indexed successfully", 86,963 nodes / 200,217 edges, and
the run clears the stale collapse stamp.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(routes): scope match bindings, and intersect ternary conjunctions

Two ways the dispatch-guard walk minted a route that does not exist — the one
thing this module's header says is worse than missing one.

MATCH BINDINGS WERE KEYED BY BARE NAME, FILE-WIDE. `collectFromMatchBindings`
walked from `tree.rootNode` and resolved `matchBindings.get(node.text)` at every
identifier in a truthiness position, so a same-named binding in ANOTHER function
answered for it. The poison check only fired on a second REGEX match with a
different URL; a non-match binding never entered `collectFromRegexDispatch`, so
nothing refused it. Reproduced:

    function handleReplay(req, res) {
      const m = pathname.match(/^\/api\/live\/positions\/([^/]+)\/replay$/);
      if (req.method === 'GET' && m) { … }
    }
    function handleSettings(req, res) {
      const m = req.headers['x-mode'];        // unrelated value, same name
      if (req.method === 'DELETE' && m) { … }
    }

    GET    /api/live/positions/{param1}/replay  handler=handleReplay    correct
    DELETE /api/live/positions/{param1}/replay  handler=handleSettings  FABRICATED

Wrong in method, handler and line. `m`, `match`, `result` are the ordinary names
here. Two ways the truth was then lost: the fabricated route is VERBED, so
`reconcileDispatchGuardRoutes` kept it and dropped the true verb-less one — the
#2856 `/api/report` shape, through the channel this series added — and `tested`
was name-keyed too, so the tail loop suppressed the real binding's own honest
verb-less emit before reconciliation ever ran.

`matchBindings` and `tested` are now keyed on (enclosing function, name).
`enclosingFunction` is extracted from the walk `enclosingHandlerName` already
did, so there is one function-boundary mechanism, not two. A second declarator
for a key refuses it, and an assignment refuses the name in its own scope and
every enclosing one. `buildRegexConstantMap` refuses a name rebound to anything
that is not a regex literal, closing `let RE = /…/; RE = buildDynamic(req)` and
the `new RegExp(prefix + '/x')` twin.

A use resolves only within its own function. Resolving outward would need a
complete declaration model — params, imports, catch bindings — and a miss there
fabricates exactly the route this fixes. Declining costs the verb, not the path.

THE TERNARY TOOK FIRST-MATCH WHERE THE ALGEBRA IS INTERSECTION. The docblock
proves `c ? A : false ≡ c && A` and says "both hold, so search both", but
`firstNonEmpty` returned one operand's set unintersected:

    (req.method === 'GET' || req.method === 'POST')
      ? (req.method === 'POST' || req.method === 'PUT')
      : false                                    emitted GET and POST
                                                 only POST is reachable
    req.method === 'GET' ? req.method === 'POST' : false
                                                 emitted GET, unsatisfiable

`intersectVerbs` replaces it for both conjunction shapes. An empty side still
yields to the other — "names no method" is not "admits none", which is what the
`isAdmin && POST` fallthrough is for — but two non-empty sides intersect, and an
empty intersection is an unsatisfiable guard that yields no verb.

Both changes strictly REMOVE routes, so SCHEMA_BUMP 58 -> 59: routes are
parse-time output replayed verbatim from a warm cache, and without the bump an
indexed repo keeps serving the fabricated verbed route while the fix looks
implemented.

10 tests added, 9 of which fail without the change. All 86 existing assertions
pass unchanged; none was weakened.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(scope-resolution): bind a type parameter only inside the scope it opened

W2-8 captured `@declaration.type-parameters` on EVERY `type_alias_declaration`,
but an alias becomes a SCOPE only when its value is an `object_type`
(`typescript/query.ts:149`). For a union, array, conditional, mapped, tuple or
function alias there is no scope, so the def — now carrying `typeParameters` —
attached to the innermost enclosing scope, which is the MODULE. And
`typeParameterNamesInScope` folds each scope's set from its PARENT'S, so the
name landed in every scope in the file. The `USES` guard then deleted every edge
whose target had that simple name:

    export interface Result { ok: boolean }
    export type Maybe<Result> = Result | null      // one ordinary line
    export function readResult(r: Result) { … }    // its USES edge is DELETED

Silent data loss, in the edge class whose whole purpose is answering "what
breaks if I remove this field?". Measured: adding two scope-less generic aliases
emptied the fixture of USES edges entirely.

The existing fixture could not see it — it wrote `type Box<Result> = { held: Result }`,
the ONE alias form that opens a scope.

`typeParameterNamesInScope` now reads a def's `typeParameters` only when that
declaration OPENED the scope owning it: `scope.kind !== 'Module'` and the def-id
position equals the scope range start, via the canonical `definitionIdPosition`
rather than slicing the id. That is the same alignment test `pickCallerCallableDef`
uses to tell a closure from a nested function, and it is language-neutral — it
also covers `function f() { type W<Result> = Result[] }`, which a module-scope-only
stopgap would miss.

Every language populating the capture was audited (ts, java, csharp, kotlin,
rust, cpp): all anchor it on a declaration that IS a scope node, including C++
where the capture rides `template_declaration` but the anchor is the inner
`class_specifier`. Go uses a separate sidecar. The TypeScript non-object alias
was the only mismatch in the codebase. `query.ts` is untouched.

THE GUARD ALSO SAT AT THE WRONG LAYER, which forced three defects at once. It
keyed on `edgeType === 'USES'` — and `mapReferenceKindToEdgeType` maps THREE
kinds there, `type-reference`, `value-ref` (#2437) and `macro` (#1934) — and,
because `Reference` carries no spelled name, substituted the resolved def's name
via `simpleNameOfDefId`. So `import { Result as ApiResult }` inside
`function unwrap<Result>()` deleted a REAL edge, while a namespace-qualified
target (`Host.Result`) kept a FALSE one, and a positional `@row:col` suffix broke
the last-colon parse outright.

Moved to `lookupForSite`'s `case 'type-reference'` in `resolve-references.ts`,
which has the spelled `site.name` and the reference kind in hand. One line closes
all three, deletes `simpleNameOfDefId` — a byte-identical duplicate of
`simpleNameOfGraphId` — and removes the only `graph-bridge/` -> `scope/` import
in that directory.

Honest scope: all three sub-defects are real in the code but none is observable
end-to-end today (`value-ref` never reaches this path; TypeScript emits no
cross-file USES for a type annotation at all — a separate pre-existing gap). Those
arms are labelled forward guards in the test rather than claimed as repros.

Fixture grows 1 file -> 4; 3 of 9 assertions fail without the change. The
scope-capture TypeScript fingerprint moves for FIXTURE-CORPUS GROWTH ONLY, with
per-file accounting that sums to the delta and JavaScript unchanged as the
control — see the `_rebaselined_` key.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(scope-resolution): refuse an ambiguous formal, stop the walk at the nearest binding

Two ways W2-2 typed a parameter from the wrong caller, both at the PRECISE 0.9
tier — above every `minConfidence` floor, so nothing downstream can filter them.

`formals` WAS LAST-WRITE-WINS. The key is (filePath, ownerName, parameterIndex),
`ownerName` is a bare identifier, and `emitFormalFacts` emits one site per
parameter of EVERY function collected, nested functions and class methods
included. A plain `.set` let two same-named callables in one file collide — a
free `parse` and a nested `parse`, a free `apply` and `Runner.apply` — so the
last one visited won, fabricating an edge on the loser and leaving the genuine
consumer untyped. The file's own comment covers only the cross-FILE axis.

The correct shape was thirty lines below, in the `producers` map, which does
`producers.delete(cell); conflicted.add(cell)`. `formals` now refuses the same
way: a key claimed by two DIFFERENT parameters is deleted and recorded, so a
third same-named formal cannot re-claim it. Re-stating the same cell is not a
disagreement, so a benign duplicate capture cannot poison a real key.

THE SCOPE WALK CLIMBED PAST A NEARER BINDING. The docblock claimed it stops at
the first scope carrying the name, but it consulted only `parameterProducers` —
a shadowing `const`, a catch binding or an arrow parameter is not in that map, so
the walk went straight past it to the enclosing formal:

    function readSpike(spike) { … items.map((spike) => spike.wickRatio) … }

typed the ARRAY ELEMENT from the outer parameter. `parameterProducerFor` now
stops at the first scope that binds the name AT ALL — reading the scope's own
tables, the same channels and the same reasoning as the sibling
`isNamespaceNameShadowed` — and then stops at a Function boundary. That boundary
is what covers the anonymous arrow: `collectFunctions` drops a callable it cannot
name, so an anonymous arrow emits no formal site and its scope looks empty while
in fact rebinding the name. The cost — a closure genuinely reading an enclosing
parameter now declines — is documented as the deliberate trade.

No cycle guard, deliberately and with the reason stated: both constructions of
`indexes.scopeTree` validate through `buildScopeTree`, which enforces strict
parent-contains-child ranges, so a cycle needs a scope strictly containing
itself. A per-site Set on every read/write site in the repo would defend against
a state the builder rejects.

5 fixtures, 5 assertions; 4 fail without the change and the control passes both
ways. Still uncovered and not faked: a `for (const x of …)` binder shadow — the
binder lives in the loop header, so JS emits no scope to stop at and no Function
boundary intervenes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(analyze): measure one population in every config, split the stamp on the verdict

`1b41c9df6` fixed the collapse check for the STREAMED configuration by giving
the sink a `structuralRows` subtotal. It does not cover the other one.

`resolveStreamGraphEmit` and `resolveStreamPdgEmit` both open with a
`force === true` gate, so a run without `--force` streams nothing: there is no
manifest, `structuralRows ?? 0` contributes 0, and
`scope-resolution/pipeline/run.ts:1222` (`input.pdgEmitSink ?? graph`) writes PDG
into the ordinary in-memory graph, where `relationshipCount` counts it. And
`isIncremental` requires an existing meta, so a FIRST run is a full write and the
check runs. A first-time `gitnexus analyze --pdg` on a fresh repo therefore
compared structural+PDG against structural and exited non-zero with
"Repository indexed INCOMPLETELY" on a healthy index.

MEASURED, not assumed — `runScopeResolution({ pdg: true })` with no sink:

    pdgEmitSink        = absent (non-force shape)
    relationshipCount  = 1
    residentPdgInGraph = 1
    byType             = [["CFG",1]]

The prior `residentPdgInGraph=0` was taken on a `--force` run, where
`input.graph` IS the sink; it never spoke to this case. `graph-collapse-wiring.test.ts`
had pinned the gap, asserting a PDG-inclusive in-memory count was a valid
structural expectation.

`countStructuralRelationships(graph)` filters `PDG_EDGE_TYPES` over
`forEachRelationshipFields` — the same predicate the sink uses for
`structuralRows` and the adapter for `structuralEdges` — so all three terms
measure one population in every configuration. Declining whenever
`pdg && !streaming` was rejected: that is the DEFAULT PDG shape, so the guard
would be off for every non-force run including the only full write most users
ever do. An unscannable graph (mocked pipelines) yields NaN, the same fact the
old `undefined + rows` produced and one `detectGraphWriteCollapse` already
documents as expected input.

THE THREE-WAY STAMP WAS A TWO-WAY. The comment enumerated collapse -> stamp,
healthy -> clear, no verdict -> carry forward, but the code split on the WRITE
MODE. `graphWriteCollapsed` is undefined for two different reasons, and one of
them is "the structural query threw" — so on a full run where the count could
not be READ, the code took "healthy, clear it" and erased a stamp recording real
edge loss. Run 3 then printed "Already up to date" forever: the exact failure the
comment says it fixed, reachable through the new code's own `catch {}`.

`detectGraphWriteCollapse` now returns `'collapsed' | 'healthy' | 'unmeasurable'`
with a reason, and `selectPersistedCollapseStamp` is a pure exported function
production calls. Two boundaries worth naming: `expected === 0` is unmeasurable
(its own docstring calls it "could not report a total"), but the small-repo
exemption and a cleared ratio are HEALTHY — both counts were taken. Making the
exemption a non-verdict would leave a stamp unclearable on any repo that shrank
below 100 edges, relocating the wedge rather than fixing it.

`getLbugStats` now reports `structuralEdgesError` and warns, and `run-analyze`
falls back to `stats.edges` only when the run had no PDG layer, where the two are
equal by construction. With `--pdg` on there is no substitute, so the absence
becomes an explicit unmeasurable verdict — which preserves the stamp.

13 tests added; 8 fail without the change. The integration suite now seeds a CFG
row and asserts `edges` moves while `structuralEdges` does not — the exclusion
filter was previously unexercised, its own comment conceding "structural == total
here".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(server): check the collapse before publishing the index

W2-6 marked a collapsed run's job `failed`, but the check ran INSIDE
`.then(() => backend.init())` — after the publish. `LocalBackend.init()` is the
publish step: it refreshes the registry and atomically swaps the in-memory repo
map every MCP tool and HTTP route resolves through, and its `validate` pass
prunes only entries whose metadata is provably gone, so it can publish but never
quarantine. The known-incomplete database was therefore live and queryable before
the job was ever marked failed — the job status was a label on a published index,
not a gate. The pre-existing comment two lines above says so outright: "the repo
is actually queryable when the client receives the SSE complete event."

`backend-client.ts` routes `failed` to `onError` and never calls `onComplete`, so
the UI showed an error toast while every query against that repo answered from
the incomplete graph — precisely the confident-wrong-answers failure this guard
exists to prevent.

The collapse branch now returns before publishing; the healthy path publishes via
a nested `backend.init()` so the trailing `.catch` still converts init failures
into the same message. `closeDbHandle()` runs on both paths — it is eviction, not
publication, and the worker rewrote the DB files regardless of outcome, so
skipping it would leave a stale pre-rewrite handle.

Honest limit, stated in the error string rather than overclaimed: this keeps a
FIRST-TIME analyze unpublished, which is the UI's main flow. On re-analysis of an
already-published repo the existing map entry survives and points at the same
storagePath. A real quarantine needs an un-register hook on `LocalBackend`, which
does not exist today — follow-up.

`'partial'` was considered and rejected on evidence: it is not a status. It is an
embedding-specific detail object in the `updateJob` allowlist; the status union
excludes it. Adding it would make `isTerminalJobStatus` false, so `sse-progress`
never writes a terminal frame and never calls `res.end()` — the stream hangs
open — while `backend-client` falls through to `onMessage` and `api.ts` spins the
full hold-queue timeout. `failed` at least terminates.

The failure branch also now sets `repoName`, which only the success path did.

First tests this file has ever had: 4, of which 2 fail without the change. They
assert the ORDERING, not just the final status, and build the worker message by
calling the production `projectAnalyzeResultForIpc` so a field rename breaks the
test instead of silently disabling the branch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(processes): count the entry-point cap, and make the disclosure proportionate

W2-3 added a truncation disclosure and then missed the largest ceiling it was
written to report. `findEntryPoints` ends `.slice(0, 200)` and
`entryPointsUnexplored` counted against the POST-slice list, so candidates
201..N were invisible — while the derivation docblock claimed "a new ceiling
added later cannot be forgotten here". An existing one was. On this repo's own
corpus the new counter reads 780 of 980 candidates never ranked in.

`entryPointCandidatesDropped` reports the pre-slice count, folded into
`truncated`, with `ENTRY_POINT_CANDIDATE_LIMIT` extracted and `findEntryPoints`
taking the same optional out-parameter `traceFromEntryPoint` already uses. Its
return contract is unchanged.

THE WARN FIRED ON EVERY RUN. At the shipped defaults — only `maxProcesses` is
overridden — `calleesDropped` fires for any function with 5+ callees and
`tracesDepthCapped` for any chain deeper than 10, so an ungated `logger.warn`
was constant background noise, and a warning that always fires is one nobody
reads. The split is the module's own, from the `ProcessTruncationStats` docblock:
"unexplored entry points mean whole flows are missing, while a depth-capped trace
means a flow is present but shorter than it really is."

So `warn` iff whole flows are absent — candidates dropped, entry points never
traced, or flows dropped at `maxProcesses` — and `debug` for a run truncated only
in depth or breadth. `stats.truncation` still carries all six counters; the
machine-readable channel is unchanged, only the log level moves.
`entryPointCandidatesDropped` stays in the loud set deliberately: it is the only
ceiling that GROWS with repo size, while the other two can only fire while
`maxProcesses` is small enough to bind, so gating on those alone would go silent
on exactly the large repos where 200-of-several-thousand is the thinnest sample.
The message leads with the ratio so the line carries a fact, not an alarm.

THREE COMPARATORS ALLOCATED PER COMPARISON, in the function whose own comment
explains the hoist that removed this shape (`deep_chain` 1233 -> 102 ms).
Measured here: +99 ms once per analyze at 80k functions — small, because `n` is
capped at 200 entry points x a 12-trace budget = 2,400 traces regardless of repo
size. Worth fixing anyway: 23,851 comparisons cost 70,524 joins.

One shared `sortByDepthThenPath` (Schwartzian, key built once per trace) now
serves all three sites, and `rankedByInterest` additionally hoists the `isSink`
test that ran twice per comparison. It also settles a separator inconsistency:
`deduplicateByEndpoints` joined on a SPACE while `traceOrderKey` used NUL, and
node ids embed file paths, so two different traces could produce the same key and
the tiebreak fell back to the insertion order it exists to remove — the same
hazard this series' own `->`-padding fix addresses. Order identity is pinned by a
seeded 200-trace corpus asserting the new sort equals the old one exactly.

11 tests added, 9 failing without the change, including two end-to-end
insertion-order arms. The W2-5 determinism block is unregressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

* fix(docs): restore the agent guidance, and put it in the generator that deleted it

Commit `9e602aef0` — whose message is entirely about the fetch capture — also
regenerated the machine-managed `<!-- gitnexus:start -->` block from a local
non-`--pdg` index, deleting from both AGENTS.md and CLAUDE.md:

  - the whole `MUST treat risk: UNKNOWN as unresolved, not as low` bullet
  - the `pdg_query({mode:"controls"/"flows"})` bullet
  - the `mode: "pdg"` text on the impact bullet
  - `…never read UNKNOWN as an all-clear…` from Never Do

and regressing the stats 248612/565510/918 -> 42853/135955/758. All four document
SHIPPED features: `pdg_query` at `mcp/tools.ts:675`, dispatched at
`local-backend.ts:2233`; `mode: "pdg"` at `tools.ts:448`; `riskNote` at eight
sites.

It matters more than a docs nit because the SAME series makes `UNKNOWN` dominate
a mixed candidate set (`local-backend.ts:6058`) — correct, and it makes UNKNOWN
far more common. The surviving rule only warns on HIGH/CRITICAL, so a set
measuring CRITICAL now reports UNKNOWN and that rule no longer fires, while the
rule that covered the gap was deleted in the same commit range, from all three
files agents actually read.

ROOT CAUSE, which is why restoring the files alone would not have held.
`cli/ai-context.ts` is the template. The `pdg_query` and `mode: "pdg"` text IS in
it, correctly `hasPdg`-gated — a non-PDG analyze SHOULD drop those. The
`risk: UNKNOWN` rules were never in the template at all: they had been hand-added
INSIDE the machine-managed region, so every `gitnexus analyze` on any repo
silently deleted them. This was the second occurrence; #2856's `8f8261021` was
the first. Both lines are now generated unconditionally — they describe impact's
risk semantics, which are not PDG-dependent — so regeneration restores them
instead of removing them.

AGENTS.md and CLAUDE.md are byte-identical to origin/main again, and the fixed
template reproduces that block exactly for `hasPdg: true` plus the real stats.
`.claude/skills/gitnexus-guide/SKILL.md` regains the "Inline staleness signal"
section for a live feature (`local-backend.ts:921`, `:1017-1024`, `:1995`); the
npm mirror's lack of it is pre-existing drift and is left alone, so the new sync
guard is scoped to the canonical and plugin copies.

Guards added, both demonstrated failing against the unrestored files: the managed
block must contain the UNKNOWN policy and its Always-Do/Never-Do bullet counts
must not fall below a floor, and `generateGitNexusContent` must render both lines
for `hasPdg` true AND false while keeping `pdg_query` gated. The existing
fragment lists could never have caught this — they assert presence, and this was
a deletion.

One deliberate loosening, called out rather than buried: the restored text pushes
`ai-context.test.ts`'s block-size ratio past 0.55, so it moves to 0.65. That test
argues against exactly this nudge-the-number pattern. The defence is that the
wording is origin/main's own and the 0.55 budget was calibrated against a block
already missing it; trimming shipped guidance to fit a budget would be the wrong
direction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZu3G5USo5Rs9myaaxVWjK

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-09 11:44:52 +01:00
Gergő Magyar
78ecce1b92
perf(import-target): index the workspace once per run for go/csharp/dart/ruby (#2898)
* perf(import-target): index the workspace once per run for go/csharp/dart/ruby

Four import-target resolvers answered their lookups with a full
`allFilePaths` scan per import, making resolution O(imports x files):

- go (#2877): `findRootPackageFiles` / `findAllFilesInPkgDir`, the latter
  once per path segment on the GOPATH fallback. Most Go imports are
  external, so the whole cascade ran to completion before returning null.
- csharp (#2878): the no-csproj leg took the raw Set past the memoized
  index the csproj leg was already using - up to eight passes for a
  four-segment `using`.
- dart (#2879): one scan per candidate path, and for an external package
  both candidates miss, so both always ran to completion.
- ruby (#2880): a complete `buildSuffixIndex` rebuilt and discarded per
  `require` - every require paid to index every file in the repo.

Each now reads an index memoized on the `allFilePaths` Set identity, the
shape `getPythonFileIndex` (#1918) and csharp's own `getWorkspaceFileIndex`
(#1881) already used. Two shared modules back them:

- `workspace-file-index.ts`: normalized list + `SuffixIndex` + a
  normalized->raw map, for csharp and ruby.
- `package-dir-index.ts`: "which files live directly inside a directory
  ending with <path>", for go and csharp. Candidates are bucketed by the
  directory's last segment rather than by indexing every directory suffix,
  which would cost O(files x depth) entries at kernel scale (#2649).

Behaviour is unchanged, including the tie-breaks that are expressed only
through Set-iteration order and `indexOf` positions: the go root leg stays
sorted and its package leg stays unsorted, the first-occurrence rule that
excludes a directory nested inside a same-named directory is preserved,
csharp's whole-path match still beats an earlier suffix match, and dart
still tries `lib/<rel>` fully before bare `<rel>` and matches raw paths.

Verified two ways. `import-target-index-parity.test.ts` keeps verbatim
copies of the pre-change implementations and diffs against them over a
deterministic corpus plus hand-built layouts for each tie-break; six
mutations of the new code were confirmed to fail it. Separately, the bench
corpus produces byte-identical fingerprints against the pre-change
resolvers at both 400 and 1600 files.

`bench/import-target/measure.mjs` gates both arms in CI: per-language
output fingerprints, a scaling budget (measured 0.98-1.12 here, 3.32-4.10
against the pre-change scans), and the corpus shape, so the corpus cannot
be shrunk below the size the scaling arm needs and still print PASS.

Closes #2877
Closes #2878
Closes #2879
Closes #2880

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCptYZWRgnnJ821rzebQyj

* perf(import-target): cover kotlin, and add depth + absolute-cost arms

#2872 landed the same index hoist for Kotlin while this branch was open.
Fold it into the shared measures so all five resolvers are gated on one
corpus, and adopt the two arms that PR's review proved a scaling ratio
alone cannot carry.

- `bench/import-target/measure.mjs` gains a kotlin arm: a Gradle-shaped
  corpus with per-module source roots over one package namespace, `.kt`
  and `.kts` stems, a nested same-name package directory, and a share of
  wildcard `.*` imports so the package fan-out tier — the only tier whose
  output is order-bearing — is inside the fingerprint.

- `depth_ratio`: deep arm at a FIXED file count with ~6x the path
  components. `scaling_ratio` divides the file count out, so it is
  scale-invariant and structurally cannot see a cost that grows with path
  depth instead, and `buildSuffixIndex` (C#, Ruby) and Kotlin's
  `suffixByStem` each emit one entry per component. Measured: go 0.98,
  dart 0.88 (depth-free indexes), ruby 1.48, kotlin 2.20, csharp 3.45 —
  which is why the budget is per language. One global budget would have
  to sit at 5.0 and would let Dart go 0.88 -> 4.9 unnoticed.

- `small_ms_ceiling`: an absolute bound at 4x the measured arm, because a
  constant-factor regression that grows both scale arms equally passes
  every ratio.

- The deep arm must resolve exactly what the small arm resolves. Padding
  was supposed to change depth and nothing else; a deep arm that stopped
  resolving would be timing the null path.

The five fingerprints are unchanged by this commit - verified against the
previous baseline before rewriting it, so adding the kotlin arm and the
deep scale did not perturb the four languages' output.

Kotlin joins the Set-iteration counter in
`import-target-index-parity.test.ts` too. Its own guard
(`kotlin-import-index-reuse.test.ts`) counts index BUILDS, which a scan
added beside a reused index does not move.

That counter is also the only DETERMINISTIC guard against a reintroduced
scan, and this commit documents why rather than pretending otherwise: a
full workspace scan on 1-in-32 imports was measured to pass every timing
arm here (dart, 1.458 scaling against a 1.8 budget, 1.736 ms against a
4 ms ceiling) while the counter reads 14 instead of 1. Tightening the
ceilings toward the noise floor to chase that case would only buy flaky
CI.

`bench/kotlin-import-target/` stays: it fingerprints both file-set
iteration orders and probes the four-tier cascade shape by shape, neither
of which this corpus does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCptYZWRgnnJ821rzebQyj

* perf(import-target): merge matching package dirs in one pass

`filesDirectlyInPkgDir` re-spread its accumulator once per matching
directory, costing O(files x dirs^2) copies per import. On a Go monorepo
where many services carry the same package directory (`svcN/internal/pkg`,
which Go's GOPATH cascade queries by two-segment tail) that made the index
SLOWER than the scan it replaced: 13.4x at 1600 matching directories.

Append into one array and sort once. Measured against a verbatim copy of
the pre-change scan, output byte-identical at every k:

  k=200   1400 files   old 0.126 ms   was 0.169 ms   now 0.042 ms
  k=800   5600 files   old 0.457 ms   was 3.002 ms   now 0.185 ms
  k=1600 11200 files   old 0.960 ms   was 12.890 ms  now 0.232 ms

The index now beats the scan by 2.5-4.1x on this shape instead of losing
to it by up to 13x.

Also drop the min-`ord` comparison in `firstFileDirectlyInPkgDir`: the
build loop appends a directory to its last-segment bucket the moment it
accepts that directory's first file, so bucket order already IS ascending
first-file-`ord` order and the first hit is the minimum. Differentially
verified at 0 divergences. The invariant, and the build-loop edits that
would silently break it, are now recorded at the early return.

Type the index containers as deeply readonly so Go's deliberate
`[...rootFiles].sort()` copy is compile-enforced rather than
comment-enforced, and correct the header's claim that a polyglot repo
"never pays" -- only the stored index is per-language.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013rLy781K8E3rRHFGEe1khJ

* test(import-target): guard index reuse at the adapter boundary

`workspace-file-index.ts` documented the hazard as "a defensive
`new Set(allFilePaths)` in an ADAPTER" -- the bug #1918 shipped -- and
named the unit parity test as the guard. It is not: that test imports the
resolvers directly, while production reaches them through
`<lang>ScopeResolver.resolveImportTarget`. Inserting the copy at
`go/scope-resolver.ts:31`, `csharp:35`, `dart:189` and `ruby:268` left the
parity test 28/28 green and `measure.mjs --check` PASS in all four cases.
Kotlin and Python already had adapter-level guards; go/csharp/dart/ruby
had none.

Add `test/integration/<lang>-import-index-reuse.test.ts` for the four,
mirroring the Kotlin/Python precedent: resolve through the scope resolver,
assert the file set is traversed once (twice for C#, which builds two
indexes), and pair every count with a result assertion so a count of 1
cannot be the count of an adapter that resolves nothing. Each was proven
to fail under the copy it exists to catch:

  go     expected 600 to be 1     dart   expected 600 to be 1
  ruby   expected 400 to be 1     csharp expected 600 to be 2

`CountingSet` moves to `test/helpers/counting-file-set.ts` and now counts
`forEach`, `values`, `keys` and `entries` as well as `[Symbol.iterator]`.
It missed a rescan spelled `allFilePaths.forEach(...)` entirely; with the
overrides that mutation reads 14 instead of 1.

Four fixtures that pinned the guard next door, each now shown to kill its
mutation:
- the Dart "matched RAW" case used a forward-slash target, so the basename
  bucket missed before the raw comparison was reached and it asserted
  `null === null`. A positive twin carrying the backslash in the TARGET
  catches both half-mutations.
- no C# or Ruby target addressed the corpus's `win\dir\thing` file, so
  deleting the backslash normalization in `workspace-file-index.ts` passed
  both gates. Now 4 failures.
- `normToRaw`'s first-wins rule had no normalization twin in any corpus.
- the Go nested-package fixture was decided by the `endsWith` half and
  never reached the first-occurrence branch its title names; addressing
  the directory as a single segment makes it reach it.

The parity test's own docstring no longer claims the scan count is a
complete census -- it names the three materialized arrays it cannot see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013rLy781K8E3rRHFGEe1khJ

* test(import-target): assert every scale, add collide and retained-heap arms

Three holes in the gate this PR ships as its own proof.

1. `--check` computed three fingerprints per language, stored all three,
   and compared one. `DEEP_PAD = 16 -> 0` deleted the entire depth arm and
   still printed PASS, because depth padding is count-neutral by design so
   no asserted number moved. Assert `fingerprint` per scale, and assert
   `deep.fingerprint !== small.fingerprint` so the padding's EFFECT is
   pinned, not just its output.

2. The corpus minted per-index directory names (`src/pkg${d}`,
   `src/Ns${d}`, `lib/feature${d}`), so max last-segment bucket and max
   matching dirs were both 1 -- and bucket cardinality is the only
   non-constant term the index has. The `dirCount > 1` merge branch had
   never executed in any arm. Add a `collide` arm on shared-leaf layouts
   with identical files/imports/resolved counts; it reaches 9,269
   multi-directory merges per run, up to 34 directories at once. Go and
   C#/Dart legitimately score above the linear budget there and get their
   own; Ruby and Kotlin stay at 1.8 because their keyed maps are
   collision-immune and that immunity is the assertion.

3. No arm measured memory, while the C# no-csproj leg newly retains an
   O(files x depth) suffix index. Add a retained-heap arm on the
   `bench/cfg` pattern, including its loud failure when `--expose-gc` is
   missing rather than a silent skip. Measured at 32k files:
   csharp 73.62 MiB, ruby 55.26 MiB. Ceiling is 1.5x, NOT the 4x the
   timing arms use -- the measurement is byte-stable to 0.00085% across
   processes, so 4x would be throwing away the gate. `_arms_note` records
   why, so nobody harmonises it back.

`depth_ratio`, added by this PR, flaked ~1-in-20: go peaked at 1.748 and
dart at 2.043 against a 1.6 budget, both ratios of two sub-3 ms minima.
Fixed at the estimator, not the threshold -- REPS 5 -> 15, matching
`bench/cfg`, `schema-pairs` and `callable-value-flow` (5 was the lowest in
the repo; the sibling `kotlin-import-target` uses 7, which was not enough
here). 22/22 PASS, every arm now at 70-78% of its budget with a <=1.26x
swing. No budget was widened; the distributions are recorded in
`_arms_note` so the headroom is visibly earned.

Three copies of the same overclaim corrected: the parity test NARROWS the
1-in-32 blind spot, it does not close it -- it watches the Set while the
resolvers hold materialized arrays. `_floor` no longer claims its ratios
"match" the issues' (different corpora, both quadratic).

The step moves to the END of the benchmarks job and runs with
`--expose-gc`. A failing step aborts every step after it (#2895), so the
newest, least-proven gate must not sit ahead of eight established ones.

All five output fingerprints are byte-identical to before this session --
the proof that every change here was behaviour-preserving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013rLy781K8E3rRHFGEe1khJ

* docs(import-target): point the reuse contract at the guard that guards it

`workspace-file-index.ts` told callers the unit parity test guards the
adapter-copy hazard. It does not -- it never crosses the adapter. Name
both layers and say which catches what: the per-language
`test/integration/<lang>-import-index-reuse.test.ts` files at the adapter
boundary, the parity test for a rescan reintroduced inside a resolver.

The C# namespace-dir index comment named `findDirectChild`, which this PR
deleted; it feeds `firstFileDirectlyInPkgDir` now.

Drop `GoResolveContext`, dead since the legacy call-resolution DAG was
removed in #942 -- zero importers, and `gitnexus`'s package.json declares
no `main`, `exports` or `types`, so it is not a published surface.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013rLy781K8E3rRHFGEe1khJ

* refactor(import-target): quality pass over the review-round changes

Four cleanup lanes (reuse / simplification / efficiency / altitude) over the
previous four commits. No behaviour change anywhere: all 25 bench cells
(5 languages x 5 arms) are byte-identical on files, imports, resolved,
distinct_outcomes and fingerprint, re-verified after each individual edit.

**Restores a fast path the last commit lost.** Fixing the O(k^2) accumulator
made the SINGLE-directory case — the overwhelmingly common one — copy the
bucket where the original aliased it: measured 1.11x slower at 4 files/dir
rising to 1.72x at 128. Holding the first bucket by reference and promoting to
an accumulator only when a second directory appears is 0.65-0.97x of the
previous code at dirCount=1 and parity at dirCount=64. 176-case differential,
0 divergences.

**`sortedRootFiles` accessor.** `rootFiles` was the only index container read
directly from outside the module. `readonly` is erased at runtime and
`Array.isArray` widens it back, so the copy rule now lives with the code that
owns the invariant instead of at the call site. No `Object.freeze`: V8's
PACKED_FROZEN_ELEMENTS read cost lands on the hot `matchingDirs` path.

**One shared arm for the four reuse guards.** The distinct-file-set test was
copy-pasted four ways, 33-38 identical lines each, and this repo's own helpers
(`mini-repo.ts`, `scope-model.ts`) document extracting at the SECOND verbatim
consumer. `expectDistinctFileSetsGetOwnIndex` takes what actually varies; its
`expected` type excludes `null` so the pairing rule cannot be reinstated as a
hole. The per-language first and third arms stay duplicated on purpose —
corpora and payload shapes genuinely differ. Re-proven: all four still fail
under an adapter-inserted `new Set(allFilePaths)`.

**Bench.** `dirsFor` shared by the two functions that must agree on directory
fan-out (they mint and address the same files). `SCALES` derived from the arm
table, so a future arm cannot be measured, printed and silently never asserted.
Five timing checks with one shape collapsed to a table — the trailing sentence
had already drifted into four wordings. `uniqueTarget`/`collideTarget` as flat
functions, mirroring the `uniqueDir`/`collideDir` split rather than nesting a
second axis four ternaries deep. One `identityPass` replaces two untimed full
resolution passes per cell: -371 ms median.

**CI step moved back where it belongs.** It was parked last "until #2895
lands", but that reasoning was backwards twice over: the flake that motivated
it was fixed at the estimator in the previous commit, and #2895's own audit
measured the last slot as executing zero times in 13 runs. It sits with the
other resolver-index guards; #2899 carries the `if: !cancelled()` that fixes
step masking for every step at once.

Filed rather than fixed here: #2908 (java and cobol still scan the workspace
per import, same shape as #2877-#2880, neither memoized), #2909 (make index
reuse a contract test over SCOPE_RESOLVERS on one instrument).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 09:59:37 +01:00
glier
c6b24162d9
perf(kotlin): index import resolution instead of scanning per import (#2872)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Skill copy sync / shipped skills drift guard (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* perf(kotlin): index import resolution instead of scanning per import

`resolveKotlinImportTarget` walked the entire workspace on every import.
Its four tiers — exact/suffix, directory child, package fan-out and
progressive prefix strip — each ran `for (const raw of allFilePaths)` with a
`replace(/\\/g, '/')` and several string scans per entry, and they are tried
in cascade, so one unresolved import cost two to four full passes.

Across a repository with tens of thousands of Kotlin files that is
O(imports x files): on the order of 10^10 string operations on a single
thread. It does not look like a hot loop from the outside - analyze sits at
exactly 1.00 core with a completely flat heap and emits nothing for hours,
because every allocation is a short-lived string and nothing accumulates to
hint at progress. Small repositories hide it entirely: at a few hundred files
each pass is free.

Three maps, built once per `allFilePaths` Set and memoized on its identity,
make each tier O(1): stem -> path for the exact tier, every component-suffix
of the stem for the suffix tier, and directory -> direct children for both the
fan-out and the first-child fallback. Cost becomes O(files) once plus O(1) per
import. This mirrors the existing Python index (`getPythonFileIndex`), down to
the WeakMap keying and the build counter.

Semantics are unchanged, including the parts the scans expressed only through
iteration order:

  - an exact match anywhere beats a suffix match found earlier, because the
    scan returned on the first exact hit but merely remembered the first
    suffix hit;
  - "first match" stays first in set-iteration order, so both stem maps keep
    the earliest path inserted for a key;
  - a directory-name match still honours the scan's `startsWith`-then-`indexOf`
    rule, which only ever considered the FIRST occurrence of `/dir/`. A path
    like `data/src/main/kotlin/com/example/data/Repo.kt` is therefore still
    NOT a child of `data`. That is arguably wrong, but fixing it here would
    silently move edges in every Kotlin repository; it belongs in its own
    change with its own fixtures.

That claim is gated, not asserted. `bench/kotlin-import-target` fingerprints
every `fromFile | targetRaw -> result` triple over an exhaustive branch matrix
plus a deterministic fuzz, each file set resolved in BOTH iteration orders
because that is the only place the tie-breaks above are expressed. The
committed baseline is the value the PRE-INDEX implementation produces: both
implementations print
5ad605c179081505705ff7698a09dbdbdc4831080af6d9fdec5499cc6bce28ee over the same
20074 cases, 11612 of them non-null, and anyone can re-run it by pointing the
harness's module specifier at the old file.

Its second arm is the scaling ratio, `(t_large/t_small)/(1600/400)` over a
synthetic Kotlin monorepo whose imports are ~40% unresolvable — only a miss
drives all four tiers, which is where the scan was worst. The index measures
0.99 (8.0 ms / 31.7 ms); the implementation it replaces measures 3.737
(2207.8 ms / 33003.5 ms) on that same corpus, so the budget of 1.6 separates
them by a wide margin. Take the absolute times as an order of magnitude only
(~276x, ~1041x): the floor arm was run once cold because best-of-seven against
a quadratic implementation costs minutes, while the index arm is the usual
best-of-seven. The ratios are the comparable pair. Both arms run in the
existing always-on `benchmarks (GITNEXUS_BENCH)` job, next to the C++ guard
from #2788 and the Python one from #1918.

Two unit-level guards sit alongside it: a parity test pinning the curated
cases, and an integration test asserting the index is built once across many
imports — the adapter must pass the Set through, since a defensive copy would
hand a fresh WeakMap key per call and restore the old behaviour (the same trap
Python hit in PR #1918).

Two other providers have the same defect and are left alone here, having no
repository at hand to verify a change against:

  - `go/import-target.ts`: `findRootPackageFiles` and `findAllFilesInPkgDir`
    scan unmemoized, and the GOPATH fallback calls the latter once per path
    segment but the last, so a single import can trigger several full passes;
  - `dart/import-target.ts`: the `package:` branch scans once per candidate
    path — `lib/<rel>` and bare `<rel>` — and `resolveRelative` scans again in
    its suffix fallback, also unmemoized.

`csharp/import-target.ts` is a partial case worth noting: it already builds a
memoized `getWorkspaceFileIndex`, but that is reached only when a `.csproj` is
found; the no-csproj path hands the raw Set to `resolveDirectMatch` and
`resolveByProgressiveStripping`, which scan past it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(kotlin): close the blind axes in the import-resolution gate

Review of #2872 found the weak part was the gate, not the resolver: four
plausible follow-up mutations passed `--check` with a byte-identical
fingerprint, `cases` AND `non_null`. Each is now caught, and each was
re-checked against the mutation it exists to stop.

  - The hashed record carried `order | fromFile | targetRaw | result` but not
    the FILE SET, so a corpus edit that swapped the workspace under a case
    while leaving its result string alone was invisible. Leaving the resolver
    untouched and editing only the corpus, two documented load-bearing cases
    could be gutted — the "exact beats an earlier suffix" case losing its
    competing file, the repeated-directory negative case losing its file
    entirely — with the gate green. The file set is now part of the record, and
    that same edit now moves the fingerprint.
  - The corpus capped path depth at 8 components and packages at 16 files,
    which are precisely the two axes the loops this change added run on. It now
    carries 11- and 13-component paths, queries against suffix keys deeper than
    seven segments, a 40-file package, and a fuzz that spans both. Verified:
    capping suffix-key depth at 7, skipping the `dirChildren` suffix loop above
    depth 8, and capping a bucket at 17 entries each now move the fingerprint,
    where all three previously passed.
  - `non_null` was reported but never asserted; it is asserted beside `cases`.
    That closes only the "resolves nothing at all" hole — it stayed 11612 under
    all three code mutations above and under the corpus edit — so it is a
    companion to the two fixes above, not a substitute for either.
  - A ratio cannot see a constant factor, and a file-count ratio cannot see a
    depth cost. `--check` now also asserts a DEPTH ratio (file count fixed,
    paths 24 components against 8) and an absolute ceiling on the small arm: a
    full workspace scan reintroduced on 1-in-32 imports scores 1.490, inside
    the scaling budget, while running 2.8x slower.

The baseline is re-derived, not adjusted: the pre-index implementation and the
index both print
ebf1790bf1d42dad483a51f2cbdeb2351e493b9e8236e4eedeef592dd81e2c5c over the new
20106-case corpus, 13256 of them non-null.

Both test suites were shown to be non-load-bearing and now are:

  - the parity test's repeated-directory case put `data` at the LEADING
    segment, so the `startsWith` guard fired and the `indexOf` rule its own
    comment describes was never reached — a resolver with that check relaxed to
    `>= 0` passed all 18 cases. A mid-path case now pins it, and a backslash
    fan-out case pins `norm.lastIndexOf` against `raw.lastIndexOf`, which was
    also bench-only. Both mutations now fail the unit suite.
  - the index-reuse test discarded all 200 return values, so a build count of 1
    was equally true of an adapter that had stopped resolving anything. It now
    asserts results, and its docstring premise is corrected: every one of its
    imports hit the tier-1 suffix lookup and none reached the fan-out it
    claimed to exercise. Half now genuinely do. The `undefined as never` casts
    and the `?.` are gone — both trailing parameters are optional and the
    member is required.

Resolver changes, all output-identical against the differential above:

  - `dirChildren` buckets are frozen once built. `findKotlinPackageFiles` hands
    a bucket straight out of the index, and the `readonly string[]` return type
    does not survive the caller: the finalize pass normalizes with
    `Array.isArray(t) ? t : [t]`, and `isArray`'s `arg is any[]` predicate
    widens the true branch, so `tsc --strict` accepts a `.sort()` there. A
    downstream sort would permanently reorder the cached bucket and flip the
    first-child tier for every later import in the run.
  - `stripped` is computed only after tier 1 misses, with `lastIndexOf`/`slice`
    instead of `split`/`slice`/`join`. Measured -20% small arm, -21% large arm.
  - `KOTLIN_EXTENSIONS` now comes from the existing `import-resolvers/jvm.ts`
    export instead of a fourth inlined copy.
  - A note on why the shared `buildSuffixIndex` is not reused, with the four
    probes that diverge, and the measured basename-bucket comparison — the one
    place this was less documented than the Python precedent it follows, and
    the question the Go/Dart/C# follow-ups will each face.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-08 09:31:39 +00:00
DuduPhudu
223ac7010d
feat: close reported graph blind spots in reference resolution, analyze and storage (#2856)
* fix(mcp): report UNKNOWN risk when an upstream impact walk finds no callers

`risk: LOW` asserts "safe to change" — a claim ABOUT callers. An upstream
walk that resolved none has nothing to base it on: the symbol may be
genuinely unused, or reached only through a reference class the index does
not record (a property access on a plain object, a bare-identifier read of a
module-scope const). Seeding LOW from an empty result is the false-safe
signal `anyKnownRisk` already refuses to emit on the ambiguous-candidate
path, and that #2687 removed by making an undetermined impactedCount `null`
rather than `0`.

Zero-caller upstream results now report risk UNKNOWN with a riskNote saying
absence of edges is not evidence of disuse. Downstream is untouched: an
empty downstream walk reports resolved callees, not safety.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(javascript): emit ACCESSES for bare-identifier reads of module-scope consts

A constant read only as a bare identifier — `Math.max(LIMIT, n)`, a default
parameter value, `return LIMIT` — minted no reference site at all, because
JS captured only `@reference.read.member`, which requires a receiver a bare
identifier does not have. So "who uses this constant?", the question behind
every dead-code trim and constants refactor, answered with a confident zero
in both directions.

The rest of the machinery was already in place: `FIELD_KINDS` accepts
`Const`, the scope query already declares it via `@declaration.const`, and
`read` maps to ACCESSES for any resolved target. This adds the missing
capture in VALUE POSITIONS ONLY (call arguments, default-parameter values,
return statements) — a blanket `(identifier)` rule would mint a site for
every token in the file, which is unaffordable at repo scale and would keep
alive the block-local symbols `pruneLocalSymbols` exists to drop.

Cross-file readers are NOT yet covered: the site exists and a call through
the same import statement resolves, but a value-kind def does not link
across the import edge. Recorded as a todo with the investigation.

PARSE_CACHE_VERSION bumped 44 -> 45: this is parse-time capture emission, so
a warm cache replays the pre-change capture set and the new edges never
appear — observed directly, a full `analyze --force` produced a
byte-identical graph until the cache was cleared by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(javascript): pin A1/A5 plain-object property acceptance criteria

Fixture plus todo specs for the four shapes plain-object property access has
to answer: object-literal keys indexed as Property nodes, a read through the
holding variable, a property WRITE, and a read through an untyped param.

Records the investigation so the work is resumable: the parse-query pattern
scoped to literals bound to a variable matches correctly (verified against
the raw JAVASCRIPT_QUERIES), but no Property node reaches the graph and
local-symbol-pruner is not the cause — it drops only Const/Variable/Static.
The remaining gate is in the parse worker's node-creation path.

No production code — specs only, so the suite stays green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(javascript): index object-literal keys of a named object as Property nodes

Idiomatic JS models configuration as an object literal, not a class, but
Property definition nodes existed only for DECLARED CLASS FIELDS. A config
field therefore had no symbol at all: `context({name: 'exitMinAtrMult'})`
answered "not found" for a field read and written throughout a live code
path, and ACCESSES had no target to point at.

Both halves are added for keys of a literal BOUND TO A VARIABLE — the parse
query mints the graph node, the scope query mints the def the resolver can
aim at. Unbound literals are deliberately excluded: an inline call argument
or a JSX prop bag is call-site data, not a named surface other code
references, so a node per key there would add volume without adding an
answerable question.

This lands the definition-node half only. The ACCESSES edges still require
receiver resolution — typing the const that holds the literal to the
literal's scope for the precise case, and name-based matching at reduced
confidence for the untyped-param (option bag) case. Both are recorded as
todos with the mechanism each needs.

Also records a trap that cost a wrong conclusion: under vitest the parse
worker runs the BUILT dist code (parse-impl resolves parse-worker.js, absent
under src/, and falls back to dist), so parse-query changes are invisible to
tests until `npm run build`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(cache): move the SCHEMA_BUMP pin to 45

The pin is the guard that makes two branches claiming one cache-schema
number fail loudly instead of silently serving each other's entries, so a
bump is only half-done until the pin moves with it. The bump itself landed
with the JavaScript bare-identifier captures; this is the other half.

Caught by the guard working exactly as designed — the suite failed with
"expected 45 to be 44" rather than letting a mismatched pair through.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): resolve plain-object property access by unique name

Idiomatic JS reads configuration off an object whose receiver cannot be
typed — an options bag passed as a parameter, a destructured handle, an
imported literal. No precise pass resolves those, so a field read and
written across a live code path produced no ACCESSES edge at all and "who
reads this setting?" answered a confident zero.

A last-resort pass runs after every precise pass and sees only what they
left behind. For each still-unresolved read/write site it asks whether
exactly ONE Property in the workspace carries that name. If so the read
almost certainly means it. If two or more do, nothing is emitted and the
site is COUNTED as ambiguous — a guess between them would be a coin flip,
and a wrong edge in the pre-edit safety gate is worse than a missing one.

Uniqueness is the right gate because it recovers exactly the names worth
recovering: distinctive domain fields (exitMinAtrMult, bookNotionalUsdt)
are unique in a repo and resolve, while generic keys (id, name, data) are
not and are skipped — which is where name matching would over-connect.

Bounded four ways:
- Confidence 0.5, the global tier, with the inference named in the reason,
  so a consumer can filter inferences without losing scope-resolved edges.
- Never second-guesses a precise result: sites already resolved are
  excluded, because first-write-wins stops a duplicate but NOT a second
  edge to a different target.
- Honors `fieldFallbackOnMethodLookup`. A statically-typed language opts
  out of name matching precisely because it over-connects; inferring an
  ACCESSES edge by name is the same claim and must obey the same opt-out.
- Requires an explicit receiver — a bare identifier is not a property
  access, and matching one by name would link a local to an unrelated key.

Indexes graph nodes rather than scope defs because an object-literal key
mints a Property NODE but no scope-resolution DEF: `localDefs` and
`scope.bindings` are both empty for exactly the population this serves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(analyze): record a collapsed graph write instead of reporting fresh

The dangerous half of a broken refresh: metadata IS written, so the index
reads as fresh, hooks re-arm, and every tool answers from a graph missing
most of its edges — indistinguishable from a codebase that genuinely has no
such relationships. Reported in the field as edges collapsing 23009 -> 2170
and as a CodeRelation table that never materialized.

`analyze` now compares the relationship count the pipeline PRODUCED against
what the DB hands back after the write. Both numbers are already in scope at
the same point, so the shortfall is provable rather than inferred — no
comparison against the previous index, which cannot distinguish a failed
write from a repo that legitimately shrank. A missing relation table needs
no special case: it reads back as a persisted count of zero.

On a collapse the run records `graphWriteCollapsed` in metadata, which
`getIndexIncompleteReasons` turns into `graph-write-collapsed` so status and
the MCP resources report the index INCOMPLETE rather than fresh.

A ratio, not equality: some relationship types do not round-trip one-for-one
and `--pdg` writes MORE rows into the same table, so demanding equality
would fire on healthy runs. Only a collapse is a defect. Fail-safe when the
expected count is unavailable — an implementation that offloads
relationships out of memory may not be able to report a total, and a false
"your index is broken" is worse than a missed one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ingestion): qualify object-literal Property ids by their owning object

Two config objects in one file that share a key name generated the same
`Property:<file>:<key>` id and COLLAPSED INTO ONE node, so two distinct
settings became a single symbol. Worse, the merged name then looked
workspace-unique to name inference, which happily resolved reads of it to a
node representing both — a wrong edge in the pre-edit safety gate, which is
precisely what the unique-name pass is bounded to avoid.

`objectLiteralOwnerInfo` already existed for exactly this ("so two
constructors in one file that both define `bar` stay distinct nodes") but
was gated to `Method`. `Property` now opts in.

`findObjectLiteralBindingInfo` returns `ownerName` only when asked. Its
`Method` ids must stay byte-identical — qualifying them would rewrite every
object-literal method id in every indexed repo — while object-literal KEYS,
indexed only since A1/A5, have no such history to preserve.

Found by a test written for the ambiguity path rather than by review: the
suite reported one node where two were expected, and an edge where none
should exist. Both are now pinned, along with the detection boundaries of
the B2 collapse check, which was previously an untestable inline expression
and is now a pure function.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(typescript): index type aliases and shape members as symbols

A TS frontend models its API contracts as `type X = { … }` and `interface`,
so a field on one is exactly what "who breaks if I remove this?" is asked
about. Three gaps made that unanswerable, all in the TypeScript queries:

1. No `type_alias_declaration` -> `@definition.type`, so an alias minted NO
   NODE AT ALL and a context() lookup on an exported contract type answered
   "Symbol not found". TypeScript was the ONLY language missing this — Rust
   (type_item), Kotlin (type_alias), Swift (typealias_declaration) and Dart
   all emit it. The alias was declared for scope resolution but never became
   a graph symbol.
2. No `property_signature` in the parse query, so INTERFACE members minted no
   Property nodes either — the upstream report's "class/interface index fine"
   holds only for the type, not its fields.
3. No `property_signature` in the scope query, so even with nodes present the
   resolver had no member declaration to aim at. Its sibling
   `method_signature` -> `@declaration.method` already existed; only
   properties were missing.

Interface bodies and object-type aliases both spell members as
property_signature, so one pattern per query covers both shapes.

Lands the SYMBOLS, not yet the ACCESSES edges: the shape is already a
class-like scope and now has member declarations, but no edge forms — the
remaining link is owner/type-binding, recorded as todos with the diagnosis.
Note TypeScript sets fieldFallbackOnMethodLookup:false, so unlike JavaScript
there is deliberately no name-based fallback here; the precise path is the
only route by design.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(golden): accept interface members in the mini-repo snapshot

Drift is entirely the new TypeScript shape-member indexing: the fixture's
three interfaces (ValidationResult 2, DbRecord 3, LogEntry 3) contribute
exactly 8 Property nodes, each with exactly one HAS_PROPERTY owner edge.

Verified before regenerating rather than after: every pre-existing count is
untouched (CALLS 9, IMPORTS 12, DEFINES 16, HAS_METHOD 1, MEMBER_OF 12,
STEP_IN_PROCESS 12), so nothing was rewired — the digest moved only because
8 edges were added. The fixture's inline `return { valid: false, … }`
literals correctly produced nothing, confirming the object-literal rule
stays scoped to variable-bound literals.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): never report a collapse from a non-numeric count

The B2 check reported healthy runs as total graph-write collapses. A
non-numeric `expected` (a graph implementation reporting no total, a
lightweight pipeline result) does not skip the guards — it INVERTS them:
`undefined < 100` is false, so the small-repo exemption never fires, and
`0 >= undefined * 0.5` is `0 >= NaN`, also false, so the ratio check
"passes" as well. Both bounds silently evaporate and every such run is
flagged.

That is precisely the failure this check was written to catch, reproduced
inside the check itself: an unmeasurable quantity treated as a measured
zero. Both sides are now validated as finite numbers before any comparison.

`persisted` is also passed as UNKNOWN rather than zero when the DB was not
demonstrably readable: `getLbugStats` flattens "no connection", "query
threw" and "empty table" all into `edges: 0`, so `stats.nodes > 0` is used
as independent evidence the read happened at all.

Caught by the existing run-analyze suites, not by the new unit tests — those
exercised the pure function with well-formed numbers and were blind to the
integration's actual inputs. Both cases are now pinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(typescript): make object-type aliases own their members

A TS object-type alias declares the same `property_signature` members as the
interface beside it and answers the same question, but was not a member
owner: its fields were minted with bare ids and no owner edge, so two
aliases in one file sharing a field name collapsed onto one node, while the
identical interface resolved normally.

`type_alias_declaration` joins CLASS_CONTAINER_TYPES (and
CONTAINER_TYPE_TO_LABEL, as that set's invariant requires — a container
missing there gets orphaned member edges or a wrong owner label). Aliases
with no object type (`type Id = string`) declare no members, so they own
nothing and are unaffected.

This also lands the INTERFACE field -> consumer edges, verified on the
mini-repo fixture rather than only on a purpose-built one: `saveToDb` now
links to `ValidationResult.value`, and `formatLogEntry` to `LogEntry.level`
and `LogEntry.message` — three real contract-field reads that previously had
no graph path at all. Golden updated: +3 ACCESSES, no node changes.

The ALIAS field -> consumer edge is still not linked and is recorded as a
todo with the exact blocker: resolving a receiver typed as the alias needs
the NAME to resolve to a class-like def, and `isClassLike` is
Class|Interface|Struct|Record|Enum|Trait. That predicate is read from ~12
sites including MRO and heritage, and every language mints TypeAlias, so
widening it would enrol aliases in linearizations where they do not belong.
Widening only the scope index was tried and reverted — the type-name walkers
gate on it independently, so it fixed nothing and left dead code. That needs
a deliberate "shape-like" concept, not more call-site widening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(test): record the traced diagnosis for the unlinked alias field edge

Traced to the end rather than left as "needs investigation", so the next
attempt starts from facts:

  1. Graph side is COMPLETE and symmetric with the interface —
     Property:...:LiveModeConfig.bookSlots is owner-qualified and carries
     HAS_PROPERTY.
  2. Resolution DOES reach resolveClassBindingForName('LiveModeConfig')
     (instrumented) and misses.
  3. It misses because the module scope binds LiveModeIface:Interface,
     renderAlias, renderIface — and not LiveModeConfig. The alias has no
     binding on the receiver's scope chain at all.
  4. The TS scope query tags aliases @declaration.type, but normalizeNodeLabel
     accepts only typealias / type_alias and has no "type" case, so it returns
     undefined. Kotlin and Dart use @declaration.type_alias; TypeScript is
     alone on the dead tag.
  5. Retagging is NECESSARY BUT NOT SUFFICIENT — tried, and the binding still
     does not appear, so a second gate exists in how a declaration anchored on
     a node that is ALSO a @scope.class anchor is attached: the alias appears
     to bind inside its own scope rather than hoisting to Module, where
     interface_declaration evidently does hoist.

An isShapeLike predicate (the nominal-vs-structural split: shapes declare
members, nominal types participate in MRO) plus a mirrored
findShapeBindingInScope were built and REVERTED along with the retag. With no
binding on the chain they never fire, and shipping inert widening is worse
than shipping none — the same standard applied to the earlier scope-index
attempt. The design is recorded here; it is worth doing once step 5 is fixed,
and it also unblocks Rust's parked union_item, which the MEMBER_OWNER_NODE_TYPES
comment documents as the same gap in another language.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): resolve cross-file value references, skip block-locals

Two halves of the same question, "who uses this constant?".

CROSS-FILE. `resolveReferenceSites` runs against the registries and, as its
own comment says, "imports live in finalized bindings the registries can't
see" — which is why free CALLS need `emitFreeCallFallback`. Reads had no
counterpart, so `import { LIMIT }` followed by a bare use resolved to nothing
while a CALL through the very same import statement resolved fine. This adds
the read/write counterpart, reusing `findValueBindingInScope` (which walks the
FINALIZED chain) rather than inventing a lookup. Confidence 0.9: the import
names the def, so this is precise resolution, not inference.

BLOCK-LOCALS. Bare-identifier capture also matches a read of a block-local
`const`, and an edge to one keeps alive exactly the inert locals
`pruneLocalSymbols` exists to drop — a pruned node becomes a retained node
plus an edge, in every function of every indexed repo. Emission now takes the
set of value defs bound at MODULE scope and drops ACCESSES to
Const/Variable/Static outside it. The cross-file pass carries the same
guarantee structurally: a def in another file cannot be a block-local of this
one, so it skips same-file hits entirely.

The block-local leak was already shipped in the intra-file A2 commit and was
found only because a test was written for the guard rather than the feature —
the same way the object-literal id collision surfaced.

Verified on the full resolver matrix: 3172 tests, golden unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(lbug): diagnose a vanished staging CSV instead of surfacing a Binder error

A forced rebuild could fail with "COPY failed for File: Binder exception: No
file found that matches the pattern .gitnexus/csv/file.csv" and then an ENOENT
on .gitnexus/csv/rel_Folder_File.csv — two engine-level messages that name
neither a cause nor a remedy, which is where several field reports end.

Only tables with rows > 0 enter the COPY manifest (csv-generator.ts), so an
absent file was WRITTEN during this run and removed since. Both COPY loops now
preflight and say exactly that, with the row count, both causes the reports
point at (a second `gitnexus analyze` on the same repo — they share
.gitnexus/csv — or an external cleanup of .gitnexus/), and the action to take.

Scope note, deliberately narrow: this does not attempt to fix WAL corruption
or checkpoint rotation. Those already have detection and recovery hints
(isWalCorruptionError, WAL_RECOVERY_SUGGESTION, the configurable
wal-checkpoint-threshold), and the ~6000 lines added to lbug/ + storage/ since
v1.6.9 — index-lock.ts most of all, which serializes writers and plausibly
closes the concurrent-run class outright — postdate every report in the
window. Guessing at unreproducible durability faults would be speculation;
making the one failure with NO handling legible is not.

An existing overlap test induced this exact scenario (a manifest entry
pointing at a missing csv) and asserted on the engine's wording. Its intent —
that a node-COPY failure is rethrown at the FK barrier rather than swallowed —
is unchanged and still asserted; only the message it matches moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): split shape-like from class-like, linking alias fields

Completes A4: a field on a TypeScript object-type alias now links to the code
that reads it, the last unanswerable half of "who breaks if I remove this?"
for a TS frontend that models contracts as `type X = { … }`.

`isClassLike` answered two questions that only coincide for classes:
  1. does this declare MEMBERS I can look up?   — a SHAPE (structural)
  2. does this participate in inheritance / MRO? — a NOMINAL TYPE
An object-type alias is (1) and emphatically not (2) — it has no supertypes
and no place in a linearization. Widening `isClassLike` to buy (1) would have
enrolled every language's aliases (Rust type_item, Kotlin/Swift/Dart
typealias, C typedef) into MRO and heritage, so the two questions now get two
predicates. Call sites split by which they ask, and their names already said
which: `resolveInheritanceBaseInScope` and `resolveQualifiedInheritanceBase`
keep `isClassLike`; receiver typing and member OWNERSHIP take `isShapeLike`.

Three parts, each necessary and none sufficient alone:
- `findShapeBindingInScope`, mirroring `findValueBindingInScope`'s established
  relationship to `findClassBindingInScope` (same walker, different accepted
  def-type), consulted only AFTER the class lookup misses so a class of the
  same name always wins.
- `populateClassOwnedMembers` uses it, so alias members get an `ownerId` and
  are registered under the alias. Without this the receiver resolved to the
  alias and then found no members under it.
- The TS scope query tags aliases `@declaration.type_alias`, not
  `@declaration.type`: `normalizeNodeLabel` accepts typealias / type_alias and
  has no "type" case, so the old tag mapped to NO label and TypeScript aliases
  produced no scope-resolution def at all. Kotlin and Dart already spelled it
  this way; TypeScript alone was on the dead tag.

An earlier attempt concluded a further "scope-attachment gate" existed. That
was wrong and is worth recording: scope extraction runs in the parse WORKER,
which loads built `dist`, so the retag was never executed. Rebuilt, the alias
hoists to Module scope exactly as the interface does. Same trap as the parse
query — `src` edits to anything the worker runs are invisible until
`npm run build`.

Typedef and Union stay out of `isShapeLike` deliberately: they belong
conceptually (the union_item note on MEMBER_OWNER_NODE_TYPES records the same
gap) but neither is wired as a member container, so including them would widen
a predicate nothing exercises.

Verified on the full resolver matrix: 3173 tests, golden unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(typescript): pin the type-alias capture to a tag that maps to a label

The capture test asserted `@declaration.type`, the tag that
`normalizeNodeLabel` does not recognize (it accepts typealias / type_alias and
has no "type" case). So the test passed for as long as the tag was broken: it
checked only that the capture FIRED, never that it resolved to anything, while
TypeScript aliases produced no scope-resolution def at all.

Updated to the working tag and given a second assertion that the derived kind
string is one the label mapper accepts — the property that actually matters,
and the one whose absence let a dead tag sit pinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(lbug): declare TypeAlias member pairs so analyze does not abort

Making object-type aliases member owners emits HAS_PROPERTY from a `TypeAlias`,
and the relation schema declared no such pair. The emit therefore threw
`UndeclaredRelationPairError` and the ENTIRE analyze died on any repo
containing `type X = { ... }` — a hard stop, not a dropped edge. Found by
running the analyzer over a real 16k-node TypeScript repo, not by a test.

`Method` is declared alongside `Property`: a member written
`type Handler = { onClick(): void }` is a method_signature and would fail in
exactly the same way.

Why every existing test missed it: the resolver suites build an in-memory
graph via `runPipelineFromRepo` and never write to LadybugDB, so the schema
constraint was never exercised. `structural-pair-coverage.test.ts` is the one
suite that does run the emitters against the declared pairs — and its own
docstring names the gap: coverage is bounded by NON_BRIDGE_CORPUS, "a new
structural emitter should land with an entry here". This adds that entry,
pinning TypeAlias|Property and Interface|Property as sentinels.

Verified the guard is not vacuous: removing the pair again makes the suite
fail with undeclaredPairs: ["TypeAlias|Property"].

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(processes): trace depth-first so multi-hop flows are detected

D1 ("query ranks frontend components above the backend module that owns the
concept") and D2 ("processes is dominated by trivial mechanical chains") are
the same defect, and neither is about ranking or selection.

The walk stops after a fixed NUMBER of traces, so traversal order decides which
traces those are. Breadth-first reaches every shallow terminal before any deep
one, so the quota filled with the shortest paths in the graph and the walk
stopped — `maxTraceDepth: 10` was never approached. Measured on a real repo
before the fix: of 300 processes NONE exceeded 7 steps and 90% were 3-4. A
multi-hop business flow (signal → order → exit) therefore had no process that
could represent it, and `query` could only rank the mechanical pairs that did
exist. Step 4 of the caller already sorts by length and dedupes by endpoint —
it was always asking for the deepest traces this walk could give it.

Depth-first descends to a terminal first, so the same quota is spent on paths
worth keeping. Cost is unchanged: same budget, same cycle guard, same depth
ceiling — only the order differs.

Measured on the same 16k-node repo, same build and flags, BFS vs DFS (an
earlier comparison was discarded as confounded — it crossed builds and --pdg):

  steps   6-8:  50 → 168   (3.4x)
  totals:      844 → 806

and the reported query moved from `LiveSetupView → Cn` (a React component) to
`ReconcilePositions → IsTpInProfit / WithHeld / ShouldNotify` — server-side
exit management, which is what was asked for.

`traceFromEntryPoint` is exported for the test. Traversal order is unobservable
through `processProcesses`: `findEntryPoints` supplies several starting points,
so a deep chain is traced from inside it whatever the order does. A test at
that level passes under BOTH traversals — the first version of this test did
exactly that and guarded nothing. Driving the walk directly, it fails under
breadth-first with "expected 3 to be greater than 3".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(test): correct a stale status note left behind by a later fix

The A1/A5 header still said "edge resolution REMAINING ... neither is
implemented". Both shapes resolve — the typeable receiver precisely, the
untyped one by workspace-unique name — and the tests below assert exactly that,
so the note contradicted the file it sat on.

It was accurate when written and went stale when the work continued past it.
Left as-is it would tell a reviewer that a landed feature is missing.

The TRAP note is kept: the parse worker still runs built dist under vitest, and
that is still the trap it describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): index literals behind identity-preserving wrappers

`export const INERT_EXIT_CONTRACT = Object.freeze({ ... })` minted no
`Property` node for any of its keys. The object-literal rule matches
`variable_declarator > value: (object)` as a DIRECT child, and freezing puts a
call expression in between — so the shape whose fields are most worth querying
was the one shape the rule could not see. Freezing a config object is how JS
publishes an immutable contract, which is why this reads as a confident zero
on exactly the fields a reader cares about.

The allowlist is three functions, not "any call". `Object.freeze`, `seal` and
`preventExtensions` RETURN THE ARGUMENT THEY WERE GIVEN, which is what makes
the literal's keys members of the bound name. For `const x = compute({ a: 1 })`
the literal is an argument and `x` holds compute's return value, so attributing
`a` to `x` would be a fabrication.

Two negative controls, because the obvious one is vacuous: a bare-identifier
callee is rejected structurally and would pass with no allowlist at all, so the
assertion that actually pins the predicate uses `Object.entries` — identical
shape, differing only by name. Verified load-bearing by adding `entries` to the
allowlist and watching that test alone fail.

SCHEMA_BUMP 46 -> 47: parse-time emission, so a warm cache replays the pre-fix
capture set. Observed as a false negative first — `analyze --force` returned
the old node set until the on-disk cache was removed by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): narrow multi-candidate property names by scope

Workspace uniqueness was the wrong denominator. Measured on the reporting
repo: `exitMinAtrMult` has 26 `Property` definitions — 16 in one-off
`scripts/`, 7 in the frontend, one in a test, and exactly ONE in the backend
that reads it. Every backend read was refused because of competitors the
reader cannot see. The gate was not too permissive or too strict, it was
scope-blind.

A name with several definitions is now narrowed before being abandoned:
same-file first, then files the reading file directly imports, using the
finalized import graph rather than a path-shape heuristic. Exactly one
survivor at the first non-empty tier resolves; anything else stays refused.
A tier holding several candidates stops the walk instead of falling through —
local evidence that is itself ambiguous still contradicts reaching further out.

Confidence stays 0.5 at every tier. Narrowing changes which candidate is
chosen, not the kind of claim: it is still a name match, and the round-1
contract is that filtering on confidence drops all name inference at once.
The reason string now names the tier that fired.

Ambiguity reporting goes from a count to the actual names (capped), because a
count says a gap exists while the names say which fields are unanswerable.

Measured on that repo, backend readers of `exitMinAtrMult` go 0 -> 24 and
total readers 9 -> 45, including the two call sites in
`oppositeSignalExitManager.js` the report singled out. Both narrowing tests
were mutation-checked by dropping the import evidence and confirming they, and
only they, fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): capture destructured parameter keys as property reads

`function exit({ exitMinAtrMult = 0 })` reads that property off whatever the
caller passes, exactly as `cfg.exitMinAtrMult` would. It never appears in a
member_expression, so it had no reference site at all — and this is the shape
the function that IMPLEMENTS a behaviour uses, so the most relevant reader was
the one systematically missing from "who reads this setting?".

Uses a distinct `@reference.read.destructured` anchor rather than
`@reference.read.member`. The latter is filtered emit-side to matches with a
member_expression ancestor, because calls and writes share its shape, and a
destructuring pattern has none — reusing the tag would have been silently
dropped by that filter. The `read.` head already maps to a read kind, so no
mapping change is needed.

Scoped to formal_parameters. A destructuring binding elsewhere
(`const { x } = require('m')`) is frequently an import rather than a field
read, and minting a property read there would attribute module bindings to
unrelated same-named keys.

All three cases (default value, bare shorthand, renamed key) mutation-checked
by removing the patterns and confirming those three tests, and only those,
fail. The renamed case also asserts the edge points at the KEY and that the
local alias mints nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): link type consumers to the type they name

An exported contract type owned its members after round 1 and still answered
`incoming: {}`, so "what breaks if I remove this field?" — the question a
contract type exists to answer — had no edge to walk. Measured on the
reporting repo: all 324 TypeAlias nodes AND every Interface node had DEFINES
as their only incoming edge.

Two independent causes, and the second is why the first was not enough.

TypeScript captured no type references at all — only cpp and csharp did — so
an annotation naming a declared type minted no reference site. Added for
annotations, generic arguments and `as` assertions, anchored to those contexts
rather than a bare `(type_identifier)`, which would also match the name in
`type X = …` and make every declaration a consumer of itself.

That alone fixed interfaces and left aliases still empty. `TypeAlias` was
missing from `LINKABLE_LABELS`, so alias graph nodes were never indexed in
`nodeLookup` and `resolveDefGraphId` could not bridge a def to its node — the
edge was dropped AFTER a successful lookup. `CLASS_KINDS` has always listed
TypeAlias and the ClassRegistry returned the def correctly, which is what made
this read as a resolution failure; instrumenting the lookup showed it
returning the right def all along and moved the search one table over. Exactly
the bug already documented two entries above it for Trait.

Fixes every language that spells an alias this way — TypeScript, Kotlin, Dart
and Rust all emit `@declaration.type_alias`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): capture record construction as property writes

The read side answered well after the narrowing work while "who SETS this
field?" still missed the code that stamps the value. A record built inline —
`return { exitContract: { exitMinAtrMult: settings.x } }` — is bound to no
variable, so it minted no definition and its keys referenced nothing.

Modelled as WRITE REFERENCES, deliberately not definitions. The round-1 rule
already mints Property nodes for literals bound to a variable; minting more for
anonymous records would add same-named competitors to the very name-narrowing
that makes these fields resolvable — measured at 26 competing definitions for
one field on the reporting repo, which is what made every backend read
unanswerable in the first place. A construction site is a USE of a field, not
another declaration of it.

Two positions only: nested under a key, and returned. Both are records with a
name attached (the key, or the function). An inline call argument
(`doThing({ id: 1 })`) stays excluded for the same reason round 1 excluded it
from definitions — it is call-site data, not a named surface — and is asserted
as such.

The enclosing literal is the receiver and it is anonymous, so these route
through the same narrowing and the same refusal-to-guess as every other
untyped receiver.

Verified on the reporting repo: `entryPlan.js` went from no rows to
`selectExitEnvelope` as a writer of `exitMinAtrMult`. Both captures
mutation-checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(processes): select round-robin by terminal so the list is not one flow repeated

Ranking was `sort by length` alone, so the top of the list was one behaviour
described many ways: eleven of the top fourteen processes on the reporting
repo were four entry points crossed with three terminals of the SAME
date-window utility cluster. Genuine call chains, but a reader learns one
thing from fourteen entries, and the repo's own domain flows sat below them.

Selection now round-robins across TERMINALS, deepest first. Depth still orders
within a terminal and still leads the list; what changes is that no terminal
takes a second slot until every other has had a first.

Keying on the entry point was tried first and made it worse — many files
declare a `main`, so each was a distinct entry that round-robin then awarded
its own slot, and `Main -> AlignWindowEnd` went from one row to eight. The
repetition was never in where a flow starts.

Measured on that repo: distinct terminals in the top 20 went 3 -> 20, and its
domain flows (`ReconcilePositions -> ...`) moved into the top 4%.

Two things this deliberately does not claim. The reported cause — ranking
rewarding fan-in, promoting chains ending in widely-called helpers — measured
FALSE: those terminals have one caller each (`alignWindowStart` 1,
`validateSymbol` 1). A fan-in discount was implemented against that hypothesis,
measured, and reverted for moving nothing. And a business flow still cannot be
a process in its own right: the walk only emits at a leaf, at max depth, or on
a cycle, so a flow whose meaningful endpoint calls onward survives only as
whatever leaf it bottoms out in. Both are recorded in the code so neither
reads as settled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(structural-pairs): pin the type-annotation USES pair

R2-2 emits USES INTO a `TypeAlias`, so the pair is `Function|TypeAlias` — a
different table from the `TypeAlias|Property` entry added in round 1, and one
that entry stays green without. `TypeAlias` is on the eleven-table list this
suite exists for, and an undeclared pair does not degrade: it throws
`UndeclaredRelationPairError` and kills the entire analyze on any repo
containing an annotated type. Every resolver suite still passes, because they
build an in-memory graph and never write to the DB.

That exact failure shipped once in this PR already. Two emitters into the same
label, each with its own way to reach a released build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): build the module-level set before the out-of-core seal

Review blocker. Under `GITNEXUS_DISK_SCOPE_INDEX=1` the seal replaces every
ParsedFile with a scope-STRIPPED copy, and the block-local filter's set was
built after it — so it walked `scopes: []` for every file, came out empty, and
the filter read that as "no def is module-level" and dropped EVERY
`Const`/`Variable`/`Static` ACCESSES edge in the repo. All languages, all
files, including the module-scope-const edges this PR exists to add. Nothing
threw and nothing logged, on the path the largest repos take: the exact
confident-empty answer the PR is about.

Built above the seal now, from `parsedFiles`, and passed as `undefined` rather
than an empty set when no scope was inspectable — an empty set is a legitimate
answer ("this repo has no module-level value defs") and must not be
indistinguishable from "could not look". Fails open; the block-local exclusion
is still asserted under the seal, since that is correctness rather than
optimization.

Also widens module level past `kind === 'Module'`. A `Namespace` scope (TS
`namespace`, Rust `mod`, C++/C# `namespace`) holds importable values too, and
treating its consts as function-locals dropped their reads. Included only when
the whole chain to the root is Module/Namespace, so a namespace declared inside
a function body stays local — asserted both ways.

That fixture then failed for a third reason: `@reference.read.identifier`
existed ONLY in the JavaScript query, so A2 did not work for TypeScript at all.
Added there, and both languages widened to `variable_declarator value:` and
`binary_expression` operands — the gaps review named between what A2 claimed
and what it matched.

Nothing covered `GITNEXUS_DISK_SCOPE_INDEX`. The new parity test asserts the
seal changes no edge, and was verified against an emulation of the original
bug: same-file readers vanish and only the cross-file reader survives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(typescript): anchor property_signature to declared shapes

Review blocker, and it reproduces end to end. `property_signature` occurs in
EVERY object_type in the TS grammar, not only in an interface body or an
alias's object type, so inline parameter types, inline return types and nested
object types all matched — and the enclosing-container walk hung each one off
the nearest class, interface or alias. Measured against the unanchored rule,
all four appeared as members of shapes that do not have them:

  Property:contracts.ts:Svc.inlineParamOnlyKey
  Property:contracts.ts:Repo.inlineQueryOnlyKey
  Property:contracts.ts:NestedConfig.nestedOnlyKey
  Property:contracts.ts:buildInline.inlineReturnOnlyKey@46:33

When the inline member shares a name with a real one — `run(opts: { retries:
number })` inside a class that declares `retries` — `addNode` is
first-write-wins and the two distinct symbols merge onto one node, so every
context()/impact()/rename() answer about that field describes the merge. The
sibling JS object-literal rule in this same PR is anchored for exactly this
reason; this is the TypeScript half of the same fix.

`(A (B))` matches DIRECT children, so nested object types are excluded by the
same anchor rather than by a second rule.

The first version of these tests was VACUOUS and is recorded here because the
reason generalizes: a collision and a correct exclusion both leave exactly one
node behind, so counting ids cannot distinguish them. Every inline member in
the fixture is now uniquely named, which is the only thing that discriminates —
verified by restoring the unanchored rule and watching exactly those four
assertions fail. A fifth test asserts anchoring costs no real member.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): correct the numbers feeding the graph-write-collapse guard

Review blocker. The predicate itself held under adversarial probing; every
defect was in what it was handed and what happened after it fired.

(a) `expected` was wrong twice. Under `GraphEmitSink` streaming the bulk types
leave the heap at parse time and never enter `relationshipCount`, so the count
understated the real volume by most of it and the ratio passed trivially —
on `force === true` runs, which include crash recovery AND the
`analyze --force` retry this check's own warning tells the operator to run.
Adds the manifest totals, the same correction the buffer-pool hint in this file
already makes for the same reason. Separately, an incremental run persists only
the changed subgraph while both counts are whole-scope: a 10,000-edge index
that lost 200 replacements reads 9,800 and is certified complete. The check is
skipped on that path rather than answered wrongly.

(b) A throwing edge count became a measured zero. `getLbugStats` initialised
its total to 0 and ran the query in a swallowing catch, so WAL/lock contention
during finalize — documented on this exact call — reported a healthy index as a
total collapse. It now returns `number | undefined`, and the caller requires
both a readable node count and a defined edge count.

(c) A total loss was exempted for being small. The min-edges rule tested
`expected` before looking at `persisted` at all, so `expected = 99,
persisted = 0` — every edge gone — stayed fresh and reported success. Total
loss is now decided first. The existing test asserted the defect; it now
asserts a PARTIAL shortfall, which is the case the exemption was written for.

(d) A detected collapse reported success and exited 0. It is different in kind
from the other incomplete reasons: those describe a run that did what it said
and left work for later, this one means most of your edges are gone and every
query answers a confident empty. The CLI now prints INCOMPLETE with the counts
and sets a non-zero exit code, and the flag crosses IPC so the worker cannot
send a clean `complete` either.

Nothing exercised this wiring — only the pure helper. Adds tests for all four,
each written so the pre-fix arithmetic fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): keep unique-name property inference inside one language

The pass indexed `Property` nodes from the whole shared graph. Per-language
gating decides whether it RUNS for a language; it never restricted which nodes
could be TARGETS. So the only carrier of a name could be in another language
entirely, and a read here resolved to it on name uniqueness alone — no owner,
no file, no call path.

Reproduced: a Java class declaring `private int loyaltyPointsBalance` and a JS
`cfg.loyaltyPointsBalance` on an untyped parameter produced an ACCESSES edge
from the JS function to the Java private field. Confidence does not mitigate
it, because `minConfidence` defaults to 0 — the tier is only a filter for
consumers who ask for one.

Candidates are now restricted to files in the language's own `parsedFiles`,
which is a precise restriction rather than a heuristic and needs no new node
property.

Every other fixture in the suite is single-language, so this could not be
caught anywhere by construction. The new fixture is deliberately polyglot and
asserts both halves: no cross-language edge, and a same-language unique name
still resolves.

Known and not addressed here: the index is still O(total graph nodes) and is
rebuilt once per qualifying language, the per-language whole-graph-scan pattern
`phase.ts` hoisted out for `sharedNodeLookup`. Hoisting it belongs with that
machinery rather than in this fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(processes): explore siblings in source order, log the exhausted budget

`slice(0, maxBranching)` selected the FIRST N callees while `pop()` explored
them LAST-first, so the trace budget went to the last-declared branch. For
`main() { init(); loadConfig(); run(); shutdown(); }` the walk spends itself on
`shutdown` and can drop `init` — the earliest steps of a flow, which is the
opposite of what a process describes. Selecting first-N and exploring last-first
was simply inconsistent; pushing in reverse makes the stack pop in source order.

Measured on the reporting repo, this costs depth: 6-8 step processes go 168 ->
146 of 816. Still roughly three times the pre-PR baseline of 50, and the right
trade — a deep branch is no longer reached by accident of being declared last.

The remaining limit is the BUDGET, not the traversal: with a fixed quota a deep
branch declared after enough shallow ones is not reached at all. That is now
asserted in both directions rather than left implicit, and the walk logs when it
stops with branches unexplored — a silently truncating cap reads as "this is
everything", the same confident-empty answer this work is about, and the repo
already sets that precedent for `dispatchFanoutSkipped`.

Removes the second depth test, which was vacuous: the note twelve lines above
it already said a `processProcesses`-level depth assertion passes under BOTH
traversals, and measured it does — breadth-first yields the same deepest
stepCount of 8, so it passed with the production change reverted. Traversal
order is asserted against `traceFromEntryPoint` directly; what is observable at
the pipeline level is which traces survive selection, which the diversity tests
cover.

Also renames `queue` to `stack` and corrects the BFS references in the module
docstring and the function's own JSDoc, which is what an IDE hover shows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(impact): carry riskNote onto ambiguous candidates and separate UNKNOWN's two meanings

Two problems on the ambiguous fan-out, which builds its own candidate object
rather than returning the single-symbol shape.

The narrowed type had no `riskNote` field and never read one, so a candidate
that resolved and found no callers reported `risk: UNKNOWN` with nothing
attached — losing the entire point of the change on the path where the reader
has the least context, since the name is ambiguous there by definition.

And `UNKNOWN` used to mean exactly one thing on this path: the probe threw. The
zero-caller branch gives it a second meaning, so an all-UNKNOWN fan-out could
no longer be told apart from a broken one. Candidates now carry `probeFailed`,
and the comment asserting the old reading is corrected.

Also aligns `gitnexus-web`, which review flagged as giving a different verdict
for the same symbol. That surface answers in prose rather than an enum, and its
message said the symbol "appears to be unused (not called by anything)" — the
identical false certainty in words. It now carries the same MEANING rather than
the same field. Downstream wording is unchanged: no outgoing dependencies
really is a fact about the symbol itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: replace assertions that cannot fail

Four from review, each satisfied by the defect it was meant to catch.

`new Set(props).size === 2` over two different literal strings can only ever
be 2, so it could not detect the node merge its title promises — that is a
difference in COUNT, now asserted on the raw array.

The ambiguity test asserted only an empty edge set, which is satisfied equally
by "the gate fired" and "the name was never looked up". It now also requires
the ambiguity counter to have moved.

`Interface|Property` was listed as a structural-pair sentinel beside
`TypeAlias|Property`, but both its labels are in the SCOPE_BRIDGE cross-product
so the pair is generated by construction and the sentinel cannot fail. Dropped
rather than left reading as coverage; `TypeAlias|Property` is the load-bearing
one.

`TypeAlias|Method` was declared in the schema with no fixture emitting it — a
declared pair no emitter exercises is indistinguishable from a missing one
until an analyze aborts on a real repo. Adds a method-shaped alias member, and
the suite requires sentinels to actually appear, so it is not vacuous.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document the new incomplete reason, the UNKNOWN verdict and the id churn

Review found the code changes landed without the guidance around them, and an
agent following this repo's own rules would have been told the wrong thing.

`graph-write-collapsed` joined `INDEX_INCOMPLETE_REASONS` with no Sign block
and no recovery section, while the precedent it cites
(`embedding-checkpoint-pending`) has both — so `gitnexus status` would surface a
new string naming silent wrong answers with nothing explaining trigger or
remedy. Added to RUNBOOK and GUARDRAILS, including why this reason alone also
fails the exit code.

`AGENTS.md` said MUST warn on HIGH or CRITICAL and never mentioned UNKNOWN, and
the shipped impact skill's risk table had no UNKNOWN row and still implied
few-callers ⇒ LOW. An agent obeying those rules literally sees `risk: UNKNOWN`
and proceeds, which negates the change the verdict exists to make. Both copies
of both skills updated.

`MIGRATION.md` now records that process ids do not survive this release —
positional ids plus depth-first tracing, source-order siblings and round-robin
selection mean `proc_7_handle` is a different flow afterwards. Bounded honestly:
nothing in-repo joins on a raw process id, so it is index churn, not a broken
consumer.

`ARCHITECTURE.md`'s scope-resolution stage list gains the two new stages.
The guide skill's node list gains `Property` and `TypeAlias` — the two node
types this work most prominently creates.

Also, on the pair-CSV preflight review asked to confirm: the hard abort IS
deliberate, because a fallback recovering zero rows is the confident-empty
failure this work targets. But the transient the message itself names — a second
concurrent analyze sharing `.gitnexus/csv` — is a race, so the check now
re-looks three times over ~150ms before declaring the file gone. Long enough to
ride out a rename, far too short to mask a file that is genuinely missing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: drop redundant TypeAlias pairs and keep bare identifiers off class members

Two regressions the full suite caught after the review fixes, both real.

`schema-pair-coverage` failed with eleven hand-declared pairs that a rule now
generates. Adding `TypeAlias` to `LINKABLE_LABELS` — needed so
`resolveDefGraphId` can bridge an alias def to its node — also makes it a
SCOPE_BRIDGE source and target, so the cross-product produces `File|TypeAlias`,
`TypeAlias|Property` and nine others that round 1 had declared by hand. Removed;
the invariant is that no pair is both generated and hand-declared.

This also changes what the structural-pair sentinel means, and the comment is
corrected rather than left overstating it: `TypeAlias|Property` is no longer
load-bearing because the label is off the generated grid — it is load-bearing
because it now depends on `TypeAlias` being IN `LINKABLE_LABELS`. Remove it and
the pair stops being generated while the hand declaration is gone, which is the
same state that silently breaks alias consumer edges.

`block-scope-shadowing` failed because a bare identifier resolved to a class
`Property`. `class Box { baseUrl = '...'; pick() { const baseUrl = ...; return
baseUrl; } }` linked the block-local read to `Box.baseUrl`, duplicating the
legitimate `this.baseUrl` edge. A bare identifier is not a member access: with
no receiver there is no object whose property it could be, and in JS/TS a field
read needs `this.`. Receiver-less read/write sites no longer accept `Property`
hits; callables stay reachable, so `cb = save` naming a top-level function is
unaffected.

That defect PREDATES this branch's TypeScript captures — JavaScript has emitted
bare-identifier reads since A2 and no class fixture exercised the shadow. The
TS parity added here is what surfaced it.

Golden snapshot regenerated after verifying the drift line by line: exactly
+5 USES from type annotations in the mini-repo, every pre-existing count
unchanged, so nothing was rewired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(scope-resolution): share the property-name index across language passes

Review follow-up. `indexPropertyNodesByName` scanned every node in the graph
and was rebuilt inside each qualifying language pass, reintroducing exactly the
pattern `phase.ts` hoisted out for `sharedNodeLookup` — whose comment records
why it matters: "the previous per-language rebuild burned that CPU+heap N times
and, on a huge repo, a tiny language's full-graph copy overlapped the next
language's — a real contributor to the scope-resolution memory peak."

Built once in `phase.ts` beside `sharedNodeLookup` and `sharedFnNodeIndex`, and
threaded through the same `prebuilt*` seam, so tests and isolated calls still
build their own.

Sharing is only safe because the per-language restriction MOVED rather than
disappeared: the shared index is whole-graph, and candidates are filtered to
the language's own files at lookup time. That also fixes a subtlety the
per-language build had backwards — the cap now applies to the FILTERED set, so
a name carried by forty properties across a polyglot monorepo but only two in
the language being resolved is still answerable, where a global cap would have
refused it.

The tri-state at the lookup boundary is deliberate and the three outcomes are
not interchangeable: no property of this name in this language (nothing to say,
and NOT an ambiguity), too many to choose between (reportable), or a list to
narrow.

Caught mid-change by the polyglot fixture: an intermediate state shared the
index without moving the filter, and the cross-language edge came straight
back. That test earning its keep twice is the reason it exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): report when a field's only anchor is another language

Round 3, found OUT-OF-SAMPLE — six field names appearing in no prior report, so
nothing here was tuned against them. All six answered 0 backend ACCESSES while
their definitions sat in `apps/research-dashboard/**`: TypeScript only. The
in-sample set scored 5/5 and the out-of-sample set 0/6, and the gap is entirely
this.

Per-language inference (`3c5eadc7`) is right and stays. What was wrong is that
declining is INVISIBLE: an empty result for a field anchored only in TypeScript
is byte-identical to an empty result for a field nobody reads. One says "look
in the other language or grep"; the other says "delete it". That is the same
confident-empty failure this series exists to remove, one surface over — and
this time the missing fact is about the ANALYZER's reach rather than the code.

Declines are now counted and named, with the languages the anchors actually
live in, kept SEPARATE from ambiguity because the remedies differ: ambiguity
wants better receiver typing, this wants an anchor in the reading language.
Collapsing them would tell a reader the wrong thing to do. A non-zero count
warns at analyze time regardless of dev mode.

The facts are published as `PipelineResult.propertyInference`, which they had
to be for any of this to be testable — and that exposed a second defect. The
round-2 ambiguity assertion, which I told the reviewer of #2856 I had
strengthened, read its stat off a `scopeResolution` field that does not exist
on PipelineResult: the `if (undefined) return` guard swallowed it and the test
passed with the production code deleted. Both that assertion and the new ones
now read the published field, and the guard is an assertion rather than an
escape. Verified by deleting the counter and watching them fail.

Reported by the same round-3 method note that caught it: verifying a fix
against the cases it was written for only proves those cases pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(context): explain an empty property result caused by a cross-language anchor

The other half of R3-1. The analyze pass now knows which fields it declined to
link because every definition of the name lives in another language; this puts
that fact where it is actually read.

`context()` on such a field previously returned an incoming list byte-identical
to a genuinely unread field. The two demand opposite actions — "look in the
other language, or grep" versus "delete it" — so the difference has to travel
with the answer:

  unresolved: property reads of this name were NOT linked: every definition of
              it is typescript, and name inference does not cross languages.
              An empty or short incoming list here is not evidence the field is
              unused — confirm with a text search, or give it an anchor in the
              reading language.
  anchorLanguages: ['typescript']

Carried through repo meta because the graph cannot answer it: the unlinked
reads mint no edge and no node, so the only record is the pass that declined
them.

Keyed on the NAME, not on the resolved label. Gating on `=== 'Property'` was
tried first and is wrong — the label reads `''` on this path for a plain
Property node, so the gate silently suppressed the entire feature while every
test still passed. Caught by asserting the field is DEFINED rather than
guarding on it, which is the same anti-pattern that made two earlier
assertions vacuous. The meta list only ever contains property names, so
matching the name is itself the type check.

Cached per (index, indexedAt): `ensureInitialized` deliberately avoids a
per-call `loadMeta` because every tool routes through it, so this re-reads
exactly when a re-analyze could have changed the answer and never otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): report declined property reads for opt-out languages too

Generalizing R3-1 rather than waiting for it to be re-reported in the other
direction. The reported case was a JavaScript read whose only anchor was
TypeScript; the mirror — a TypeScript read anchored only in JavaScript — was
still silent, because a language that sets `fieldFallbackOnMethodLookup: false`
had the whole pass skipped, and skipping emission also skipped REPORTING.

Detection is not inference. Counting what could not be linked asserts nothing
about what it means, so `reportOnly` runs the pass for its facts while emitting
no edge, and the opt-out keeps protecting exactly what it protected before.

Two things this turned up that a single-instance fix would have missed:

The cross-language fixture could NOT prove `reportOnly` is load-bearing — the
per-language candidate filter already blocks those edges, so the assertion
passed with the flag forced off. The case that discriminates is a SAME-language
TypeScript read that name inference could legitimately link and the opt-out
forbids; forcing the flag off there emits `readsTsOnly -> tsOnlyBudget`, which
is the violation.

Getting to that case surfaced a sibling gap, recorded but NOT fixed here: the
object-literal `Property` rule is JavaScript-only, so `const CONFIG = { ... }`
in a `.ts` file mints no node and its keys are invisible. The first draft of
this fixture used exactly that shape and could not discriminate for that reason.
It is the TypeScript half of R2-1a and wants its own change, not a rider on
this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(typescript): index object-literal keys, as JavaScript already did

The sibling recorded in `0c5a4f64` and deliberately left out of it. Both the
named object-literal rule (A1/A5) and the identity-wrapper rule (R2-1a) lived
only in JAVASCRIPT_QUERIES, so the single most common config idiom in
TypeScript —

    export const tsRuntimeConfig = { tsConfigRetries: 3 };

— minted no node for any key. `context()` answered "Symbol not found" and a
precise read through the holding variable had nothing to resolve to.

TypeScript sets `fieldFallbackOnMethodLookup: false`, so these gain no
name-based inference. What they gain is the PRECISE path, which is the route
TypeScript is meant to use: `tsRuntimeConfig.tsConfigRetries` has a typeable
receiver and now resolves. A read through an untyped receiver stays unresolved
and, since `0c5a4f64`, is reported as such rather than answering an empty set.

Scoped exactly as the JavaScript rules are — bound to a variable, and for the
wrapper only the three functions that return the argument they were given —
with the same `Object.entries` negative control pinning the allowlist.

Found by fixture, not by report: the first draft of the `reportOnly` test used
a TS `const CONFIG = { ... }` as its discriminator and could not discriminate,
because the shape mints nothing. That is the whole argument for sweeping a
class instead of waiting for each instance to be filed.

SCHEMA_BUMP 47 -> 48: parse-time, so a warm cache replays ParsedFiles carrying
none of these matches and the keys stay invisible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): anchor anonymous returned object literals to their function

The last gap round 3 named, and the dominant shape in idiomatic JS: 437
`return {` sites in a single backend directory of the reporting repo, including
the ~25-field payload of its entire signal pipeline. The literal binds to
nothing, so its keys could not even be named — "who reads wickRatio?" had no
symbol to ask about.

The enclosing FUNCTION is the owner: the literal is that function's return
shape, a contract its callers consume. Keys qualify as `<function>.<key>`, so
two functions returning the same name stay two shapes rather than one merged
symbol, and multiple returns in one function stay distinct by position.

RECONCILING THIS WITH R2-1b, which deliberately modelled returned keys as WRITES
to avoid adding same-named competitors to narrowing. These are definitions, but
narrowing now ranks DECLARED anchors — named literals, class fields, interface
and alias members — strictly above return shapes. A name that already resolved
keeps resolving to what it resolved to before, so the competitor problem R2-1b
was avoiding cannot come back. Mutation-checked: dropping that ranking breaks
five pre-existing R2 resolutions.

That also required an R2-1b assertion to change, and the change is a
strengthening rather than a concession. It asserted `toHaveLength(1)` — no new
definition — as a proxy for "adding definitions must not move an existing
answer". The proxy is now false while the property still holds, so the property
itself is asserted directly.

No `HAS_PROPERTY` edge from the function: that would be a `Function|Property`
relation pair the schema does not declare, and an undeclared pair does not
degrade — it throws and kills the whole analyze. That already shipped once in
this PR.

Two things found by dumping rather than assuming, both fixed here:

SHORTHAND keys were not matched at all. `return { symbol, interval, score }` is
the commonest spelling and the reporting repo's own payload is mostly this form,
but tree-sitter models it as `shorthand_property_identifier`, which `(pair)`
does not match. Caught by dumping the golden fixture and seeing a literal
returning `{ level, message, timestamp: Date.now() }` had indexed only
`timestamp`. Now covered in return position AND in the variable-bound rule,
which had the same gap.

Provenance was flagged by owner-presence, which mislabelled the anonymous case:
a callback's return shape yields no name to qualify by, so it looked like a
DECLARED anchor and would have outranked real declarations. Flagged by position
now — a different question from whether a name could be derived.

SCHEMA_BUMP 48 -> 49. Within one PR the version only has to differ from main's,
but a build stamped 48 was installed and used to analyze before these captures
existed, so caches stamped 48 carry none of them — the intermediate-build hazard
this ledger already records for 33/34.

Golden regenerated after verifying the drift: exactly +10 Property and +10
DEFINES, every pre-existing count unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): rank production anchors above test fixtures

Found by testing R3-4 on the reporting repo instead of on its fixtures. Anchoring
returned literals took `wickRatio` from 6 definitions to 13 — and backend reads
still resolved to nothing, because SEVEN of the new JavaScript anchors compete
and four of them are in `tests/`. A test constructs throwaway shapes carrying
production field names; a read in shipped code cannot mean one of them.

Applied before the declared/return-shape split, because "is this the shipped
program" is the stronger signal — a declaration inside a test fixture is still a
test fixture. Skipped when the READER is itself a test, since a read there
legitimately means the test's own shape.

The first version of this test was vacuous and the mutation check caught it: the
reader sat in the same file as the production anchor, so the same-file tier
resolved it whether or not this tier existed. The reader now lives in a file
that imports neither anchor, which leaves production-vs-test as the only thing
that can decide.

Honest about what this does NOT do: it narrows `wickRatio` from seven candidates
to three, and three functions in different files each returning that field is
GENUINELY ambiguous — refusing is correct, and the ambiguity is now counted and
named rather than silent. The reported question ("who reads wickRatio?") is
answerable only where one producer exists; where several do, the honest answer
is the list of producers, which R3-4 made nameable for the first time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(scope-resolution): resolve members through a call result's return shape

The question three rounds of reports could not answer, and the one narrowing
must refuse by design: a field produced by SEVERAL functions. A read of
`spike.wickRatio` could mean any producer, so name inference correctly declines
and no amount of tier-tuning changes that. It needs evidence, not inference.

The evidence existed in two halves that had never been joined. The call-result
type binding (`const alert = formatSpikeAlert(row)` binds `alert` to a TypeRef
whose rawName is the callee) predates all of this work; it simply had nothing to
resolve to when the callee returned an anonymous literal, because an anonymous
literal named nothing. R3-4 gave it a name. Joining them:

    const alert = formatSpikeAlert(row);
    alert.wickRatio   ->   Property:...:formatSpikeAlert.wickRatio

Precise, at ordinary emission confidence, and it works EXACTLY where narrowing
cannot: several producers sharing a field name stop being competitors because
the receiver says which one. Runs before the name fallback and claims its sites,
so a precise answer is never second-guessed by a name match.

Measured on the reporting repo: 1,410 precise edges, and all six fields round 3
verified OUT-OF-SAMPLE go from 0 backend readers to 7, 11, 10, 7, 6 and 14.
Round 3 scored 0/6 on that set; this is 6/6.

The bound is asserted, not just documented: a read off a BARE PARAMETER has no
binding here, because typing it needs the caller's type to flow in — that is
inter-procedural and genuinely larger. Those reads still fall through to name
inference and are still reported when it declines. The fixture has two producers
sharing a field name precisely so the test cannot pass by name matching, and
mutation-checking the owner lookup fails it.

No SCHEMA_BUMP: this is scope resolution, not parse-time capture, so a warm
cache already carries everything it reads. Noted in the ledger because the
reflex on this branch has been to bump, and an unnecessary bump costs every user
a full re-parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "return-shape anchoring" (R3-4/R3-5): it degrades query

Reverts af5eec5c, c764847a and 4f93f32e. The capability was real and measured —
all six fields round 3 verified OUT-OF-SAMPLE went from 0 backend readers to
7/11/10/7/6/14, 0/6 to 6/6, via 1,410 precise return-shape edges. It is reverted
anyway, because it costs more than it buys in its current form.

`cli-limit-e2e` caught it. Bisected to af5eec5c: on the mini-repo fixture,
`query('message')` returned two processes before and NONE after. The mechanism
is not window displacement — that hypothesis was tested with a partition that
kept function-local property keys from taking window slots, and it changed
nothing. Indexing the keys of every returned literal adds many nodes whose names
are ordinary words, which moves the BM25 CORPUS statistics: "message" gets less
discriminating, and `createLogEntry` — the callable that actually carries the
processes — stops ranking at all. A corpus-level effect is not repairable by a
tie-break.

Trading a regression in `query`, one of the core tools, for coverage in
`context` is the wrong trade, and shipping it because the number was good would
be the same mistake this PR spent three rounds removing: a confident answer that
is worse than the honest one.

What the work established, and what re-landing needs:

  - The mechanism is right. Joining the existing call-result type binding to a
    named return shape resolves `alert.wickRatio` by EVIDENCE, which is why it
    succeeds exactly where name inference must refuse.
  - The cost is search dilution, and it needs to be measured on BM25 ranking
    BEFORE the capture lands — not discovered by a downstream e2e test.
  - The likely shape of the fix is keeping return-shape keys out of the text
    search corpus while keeping them in the graph, which needs persisted
    provenance rather than the in-memory flag used here.

Kept: everything through 8972d223, which is verified green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(search): give the index a notion of DETAIL symbols, and re-land R3-4/R3-5

Reverts the revert. The return-shape work was correct and measured — 1,410
precise edges, and all six fields round 3 verified out-of-sample going 0/6 to
6/6 — and it was dropped for a regression that was really a MISSING LAYER: the
search index had no way to say "this symbol is queryable but is not a concept a
text search should surface on its own".

Indexing the keys of anonymous returned literals adds many nodes whose names are
ordinary words (`message`, `value`, `timestamp`). Without that notion they
compete on equal terms in FTS, push the CALLABLES named after the same concept
past the search's row cap, and `query('message')` returned two processes before
and none after.

The layer, rather than a workaround:

  - `Property.isDetail`, persisted. A Property-only column, which that table
    already precedents with `declaredType`, set where the key is minted.
  - `buildFtsQueryCypher` filters on it for the Property table, BEFORE the row
    cap. That placement is the whole point: rows crowded out never reach the
    caller, so no downstream re-ranking can recover them. Two downstream fixes
    were tried first — a tie-break and a partition of the merge window — and
    recovered nothing, which is what located the real seam.
  - `IS NULL`-tolerant, so an index written before the column existed still
    answers instead of returning nothing.

Verified by the A/B that found the regression: the query's result order is now
byte-identical to the pre-R3-4 baseline —
`proc_0_processrequest, proc_2_errormiddleware, Function:createLogEntry,
Property:LogEntry.message` — with the return-shape coverage retained.

The determinism guard then caught prose in the new DDL comment containing the
token this repo scans for, which would have read as an unordered query. Reworded;
that suite is doing exactly its job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(processes): let a flow end where the program reaches outward

The item three rounds kept circling. A trace was only emitted at a node with NO
outgoing calls, so a real flow — scan, score, arm, PLACE THE ORDER — is always a
PREFIX of some longer chain that runs on into date helpers, and could never be a
process in its own right. Ranking could not fix that; the flow was never a
candidate to rank.

What blocked it was signal granularity, and the fix is the layer that was
missing rather than a heuristic. GitNexus already knew where the program reaches
outward: the parse phase collects fetch calls and ORM queries carrying
`filePath` + `lineNumber`. Those facts only ever produced FILE-level edges
(`File -[FETCHES]-> Route`), which cannot end a trace — every function in a file
containing one would qualify. Attributing each site to the function whose range
CONTAINS it turns the same facts into the function-level signal the walk needs:
no new extraction, no new relation pair, no schema change. Innermost wins, so a
closure that performs the call is the sink rather than the function spanning it.

Three touch points, and the second is the one that makes or breaks it:

  - the walk emits at a sink AND CONTINUES, so `placeOrder` is an endpoint while
    `placeOrder -> formatDate` still exists separately;
  - subset-removal PRESERVES sink-terminated traces. A sink flow is by
    definition a prefix of the chain that runs past it, so emitting one at the
    walk and deleting it one step later would have been a no-op. Mutation-
    checked: removing this preservation fails all three sink tests, including
    the one asserting the sink is reached at all;
  - selection ranks sink-terminated above leaf-terminated, then by depth.

`processes` now declares `parse` as a dependency. It historically avoided that
on the grounds the dependency was spurious for a progress counter — it is no
longer spurious, so it is declared rather than reached for implicitly, and the
read fails open so a pipeline without that output detects no sinks instead of
losing every process.

Bounded honestly: this fires where fetch/ORM extraction fires. On the reporting
repo it will do nothing until route detection handles hand-rolled dispatchers,
since that codebase routes with `pathname === '/api/...'` on raw node:http and
produces zero Route nodes — a separate gap, and the next one worth closing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(processes): the comment above the sink ranking still described it as unreachable

R3-6 taught the walk what a sink is, but the block explaining the ranking still
carried the paragraph written when that was out of reach — "a business flow
still cannot be a process in its own right ... fixing that means teaching the
walk what a sink is" — sitting directly above the code that does exactly that.
A reader arriving at `rankedByInterest` would take the limitation as current.

The measured-false fan-in finding stays; it is still true and still worth not
re-deriving. What replaces the stale half is the bound that IS current: sinks
fire where fetch/ORM extraction fires, so a codebase whose outward calls are not
detected as such still sees leaf-terminated traces only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(routes): read a route that is declared by a comparison, not by a framework

`route_map` on the reporting repo returned

    {"routes": [], "total": 0, "message": "No routes found in this project."}

for a codebase with SEVENTEEN route modules, an `apiRouteTable.js`, and 113
path comparisons. Not a partial answer — a statement about the code, and a
false one. Same confident-empty class as the rest of this branch, except here
it takes out a whole tool.

Four route-discovery paths existed — filesystem convention, single-file
framework route, cross-file framework route, decorator — and every one of them
needs a FRAMEWORK to declare the route. A raw `node:http` server declares it
the only way the language offers:

    if (req.method === 'GET' && pathname === '/api/live/portfolio') { … }

A path, a verb, and a handler. Nothing in the pipeline could read it.

The failure modes are not symmetric, so the rules are weighted accordingly: a
route this misses is a coverage limit, a route it invents is `route_map`
asserting something false. A comparison therefore qualifies only against a
demonstrable request path (`pathname`, `*.pathname`, `req.url`; `path` is
excluded — in Node it is overwhelmingly `node:path` or a file location), and
anything untranslatable is dropped rather than approximated:

  - `pathname.startsWith('/api/')` is a namespace test; minting `/api` would
    claim a route nobody serves;
  - a bare `pathname === '/'` with no verb is more often the static-file
    normalisation branch (`pathname === '/' ? '/index.html' : pathname`) than a
    route — WITH a verb the intent is unambiguous, so that form IS taken;
  - an anchored regex converts only when its body is a literal path plus
    single-segment wildcards, so `/^\/api\/research-runs\/[^/]+$/` becomes
    `/api/research-runs/{param1}` while an optional group or an alternation
    bails.

Three things went in that nobody reported, each found by measuring rather than
by a second report.

`switch (pathname) { case '/api/x': }` is the same dispatch in different
syntax, and waiting for a bug report per shape is how a graph stays permanently
one idiom behind the code it indexes.

The reconciliation had to move up a level. The reporting repo keeps its path
table (`isKnownApiPath`) in one module and its handlers in sixteen others, so a
per-file rule sees each half separately and lists every route twice — once
verb-less with the table as its "handler", once properly. Measured: 22 of the
first 94 routes were that shadow. Only the whole registry can tell them apart,
so the rule lives in the routes phase and touches dispatch-guard routes only —
a framework route without a verb is method-agnostic BY DECLARATION (a Django
function view, a Laravel resource), a fact rather than a weaker observation.

And a path composed from a constant needed folding. One of those seventeen
modules writes every one of its routes as `` `${autoTradeBasePath}/rules` ``,
where the base is an alias of a module-level literal. Refusing that lost the
whole file — and lost it INVISIBLY, since a module with unfoldable paths and a
module with no routes are the same empty answer. Same-file only, literals only,
one alias hop, and it refuses on ambiguity: a name declared twice with
different values is dropped rather than guessed, because a partially-folded
path is a wrong route and a wrong route is the failure this module exists to
avoid.

Wiring is a LanguageProvider hook, not a language check in shared code.
`extractDecoratorRoutes` was already the general "route from this file's own
AST" channel rather than a decorator-only one — express routes have flowed
through it as `decorator-express.get` for a while — so the transport, the
`(method, url)` dedup and the handler-symbol resolution all apply unchanged.
`ExtractedDecoratorRoute.source` carries the one thing that genuinely differs:
a decorator route is DECLARED, a dispatch-guard route is INFERRED. The walk is
gated behind a substring pre-filter so it costs nothing on files that cannot
produce a route, and the gate is sound by construction — every rule reaches a
route only through `isPathExpression`, which needs one of exactly those tokens.

SCHEMA_BUMP 49 -> 51, two entries. Decorator routes are worker output carried
in the parse cache, so a warm cache replays results predating the extractor and
`route_map` stays empty — the symptom this fixes, wearing the mask of "the
extractor does not work". The second bump is the v34 hazard tripping again: a
build stamped 50 had already been used to analyze before folding existed, so
caches stamped 50 carry the unfolded route set. Caught by measuring — the
post-folding run came back suspiciously fast and would have reported the
pre-folding number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): ask whether a value def is FUNCTION-LOCAL, not whether it is module-level

The locality filter for value references was written as an ALLOWLIST of
module-scope defs, and that shape cannot express a class member. A value def has
three homes, not two: module level, a function body, and a CLASS body. Java and
C# fields and Python class attributes live in the third, so an allowlist keyed on
"module level" excludes every one of them by construction.

The guard written to make that safe could not fire either. The set arms whenever
a Module scope is FOUND, and Java has module scopes while having no module-level
values at all — so for Java it armed permanently empty, which is exactly the
state the guard exists to distinguish from "there genuinely are none".

Inverting it removes the class. A blocklist of defs positively identified as
function-local fails safe: a Java field, a Python class attribute, or a language
whose scopes could not be inspected is emitted rather than dropped. That also
retires the arming flag — an empty blocklist and an uninspected one mean the same
thing, and both mean "emit". The failure mode moves from "silently deletes an
edge class" to "retains an inert local", which is the right direction for a tool
whose stated principle is that a confident empty answer is the worst outcome.

MEASURED, because the review that prompted this reported it as a P0 deleting
every Java/C#/Python field ACCESSES edge, and that half does not reproduce.
Instrumenting the bridge over `java-write-access` shows ZERO value-ACCESSES
candidates reaching the filter: Java field references resolve to a `Property`
target and `isValueDefinitionLabel` covers only Const/Static/Variable, so the
filter is never consulted there. Pipeline-level edge sets are byte-identical with
the filter forced on and forced off, across four shapes — Java cross-file field
writes, Java cross-file constant reads, Java bare same-class constant reads, and
a Python module-constant/class-attribute mix. The defect is real and latent; the
blast radius is not. Fixed anyway, because the predicate asks the wrong question
and the next change that makes the bridge the sole emitter would ship the
deletion for real.

New `value-ref-locality.test.ts` pins the invariant triple — local dropped,
module-scope kept, class member kept — by TARGET rather than by `reason`. The
per-language suites filter on `rel.reason === 'read'|'write'` while the bridge
stamps `scope-resolution: read|write`, so they are blind to bridge-side change in
both directions. The file states plainly which half gates the mechanism (JS,
mutation-verified) and which gates only the outcome (Java, because the mechanism
is unreachable there), so it cannot be mistaken for a stronger gate than it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(docs): restore the agent guidance a generated-block refresh deleted

Commit 8f8261021's message is entirely about cross-language anchor reporting; it
also regenerated the `gitnexus:start` block in AGENTS.md and CLAUDE.md against a
LOCAL, non-PDG index and swept six documentation/config files along with it. The
review caught this and it is correct. Restored:

  - the index stats, which regressed 248612 symbols / 565510 relationships /
    918 flows -> 29969 / 118986 / 762 — my machine's index described as the
    project's;
  - the whole `pdg_query` bullet and the PDG half of the impact bullet, while
    both capabilities remain live in `mcp/tools.ts` and `local-backend.ts`;
  - the "Inline staleness signal" section in the guide skill, content that never
    left `origin/main` and that this branch had no reason to touch;
  - `.mcp.json`, which had moved from `npx -y gitnexus@latest mcp` to a bare
    `gitnexus` — a fresh clone with no global install gets a dead MCP server.

The worst of it is self-inflicted in a specific way worth naming: commit
411cac9b9, four hours earlier on this same branch, ADDED the instruction telling
agents not to read `risk: UNKNOWN` as an all-clear. The refresh deleted it. So
the branch shipped a new UNKNOWN verdict and simultaneously removed the guidance
for reading it — the exact false-safe this PR exists to remove, reintroduced one
layer up in the docs.

Re-applied that guidance, and found the drift is wider than reported. The review
noted the `.claude/` copy contradicting the plugin mirror; in fact the UNKNOWN
block was present in ONE of five shipped distributions. `gitnexus/skills/` (the
npm package), `gitnexus-cursor-integration/`, and `.agents/` were missing it too,
so every non-Claude consumer of this skill had the old table.

`shipped-skills-sync.test.ts` passed 54/54 through all of that. Its byte-identical
check covers only the plan/work/review/lfg family, and the standard skills are
guarded solely by per-skill fragment lists — so a fragment nobody listed is a
fragment nothing protects. Added the UNKNOWN fragments to that list, plus a
`copies.length > 1` assertion so an empty copy list cannot make the loop vacuous.
Verified it fails against the pre-fix tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): require the return-shape producer to RESOLVE, not merely to name-match

Review finding 2, reached independently by three Claude lanes and two Codex
legs, and reproduced here. `emitReturnShapeMemberAccesses` took the receiver's
type binding, then filtered a WHOLE-GRAPH property index with `idNamesMember` —
a textual match on the node id. Any node whose id happened to read
`<producer>.<member>` qualified, in any file and any language, and it emitted at
the 0.9 PRECISE tier where a `minConfidence` floor cannot filter it out. The
sibling unique-name pass was given a per-language restriction for exactly this
hazard; this pass consumed the same shared index with none.

Three guards, catching different shapes:

  - the producer must RESOLVE to a definition (`findCallableBindingInScope` — a
    CALLABLE lookup: the producer is the function whose return shape owns the
    member, and it resolves through finalized import bindings so a producer in
    another file still yields its own file);
  - the member must live in that definition's file;
  - that file must belong to the language being resolved.

The third is not redundant with the second, which is the part worth recording.
A receiver typed by CONSTRUCTION (`const bound = new Loyalty()`) resolves through
the shared class registry, which is polyglot — so the producer resolves into
`Loyalty.java`, its members legitimately live in that same file, and file
equality waves the cross-language edge straight through.

Also fixes the sibling P2: a site where the receiver IS typed to a producer that
owns no such member now claims the site. That branch is the strongest negative
evidence the pipeline can produce, and letting it fall through meant the 0.5 name
fallback answered a question the precise pass had just DISPROVED — measured,
linking a read to an unrelated same-named key in another file.

`polyglot-property-isolation` gains the bound-receiver arm the review asked for,
and it is the right arm: the pre-existing case has an untyped receiver and so
only ever exercised the unique-name pass, while one extra token routes an
identical read through this one. Mutation-verified — restoring the pre-fix
matching makes exactly the new leak assertion fail. The first version of that arm
was silently vacuous (it introduced a JS key of the same name, which destroyed
the fixture's Java-only premise), which is why it now asserts on the TARGET FILE
rather than on the absence of a name.

KNOWN LIMIT, stated rather than papered over: a member-call producer
(`const r = svc.make()`) binds `svc.make`, which resolves to no callable, so this
pass now declines it. Codex B3 raised that converse case and it is real. Fixing
it means typing `svc` and then finding `make` on that type — a larger piece of
work, queued for the follow-up PR. Declining is the correct interim behaviour:
the alternative is matching `make.<member>` by name across the graph, which is
the fabrication this commit removes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): resolve the import map by point lookup so the seal cannot empty it

Review finding 4, reproduced end-to-end by two lanes: the same commit and the
same repo produced a DIFFERENT graph depending on `GITNEXUS_DISK_SCOPE_INDEX`.

`buildDirectImportMap` built `scopeToFile` by walking `parsed.scopes`. The
out-of-core seal replaces `emitParsedFiles` with a scope-STRIPPED copy — that is
its documented contract, scopes are reachable only via `scopeTree.getScope`
afterwards — so under the seal the map came out empty, every `directImports`
lookup returned undefined, and tier-2 narrowing died repo-wide.

The reporting is the worse half. The loss surfaced as `ambiguous`, which means
"several candidates and the pass refused to choose". The truth was "the evidence
was discarded one function earlier". A reader acting on that would go looking for
better receiver typing to fix a problem that was not there.

This is the SECOND consumer of `parsed.scopes` on this branch to hit the seal.
The first was hoisted above it. This one is converted to the point lookup
instead, which is the stronger fix: a point lookup survives the seal by contract,
so there is no ordering left for a future edit to get wrong.

The parity assertion that would have caught it now exists. The sealed harness in
`javascript-const-references` already ran the fixture both ways, but every
assertion in it pinned ONE field's readers — which is exactly how a second
instance slipped in, since no assertion happened to cover a narrowed name. It now
also compares the WHOLE ACCESSES edge set between the two runs, as a sorted diff
so a failure names the edges that moved, with a non-empty guard so two empty sets
cannot compare equal and assert nothing. Mutation-verified: forcing the map empty
fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(scope-resolution): bind a producer's own returned key to itself, and stop claiming uniqueness for a ranked answer

Review finding 3, accepting the two defects it demonstrates and declining the
remedy it proposes. Both halves are mutation-verified.

1. A SITE INSIDE ITS OWN RETURN SHAPE NOW BINDS TO ITS OWN KEY.

   `export function buildB(row) { return { tickIntervalMs: row.b } }` writes the
   key that IS `buildB.tickIntervalMs`. Ranking declared anchors above return
   shapes is correct for a READ through a receiver, but applied to this site it
   handed the write to a same-named module const that `buildB` never touches —
   a wrong edge — while the node the key actually defines was left with no
   writer at all. Both halves wrong from one rule applied to the wrong shape.

   Checked before every other rule, because it is evidence rather than ranking:
   the owner qualifier on the candidate id and the enclosing callable are the
   same symbol. Nothing outranks that.

2. THE TIER NO LONGER LIES.

   `workspace-unique` is a claim that exactly one node in the workspace carries
   the name — a fact about the graph, and the label a reader trusts most. An
   answer reached by FILTERING (tests down-ranked, return shapes down-ranked)
   is a weaker claim, and it was reported under the same label. The edge is
   unchanged; what it is allowed to say about itself is not. `narrowed` now
   counts these correctly too, since it keys off the tier.

WHAT I AM NOT DOING, and why. The review proposes dropping the same-file and
imported-file tiers "and keeping only genuine workspace-uniqueness". That would
revert the measured R2 result taking backend readers of `exitMinAtrMult` from
0 to 24. Workspace uniqueness was already measured too strict on that repo: the
field carries 26 Property definitions — 16 in one-off scripts, 7 in the
frontend, one in a test, and exactly one in the backend that reads it. Strict
uniqueness declines all 24.

The alternative suggestion — require the receiver to bind to the owning object —
has the same effect by another route: the population this pass exists for is the
untyped option bag, whose receiver binds to nothing. Requiring a binding turns
the pass off for its own use case. So the two demonstrated defects are fixed and
the capability around them is kept, at half confidence, naming its inference in
the reason string, and honoured only where `fieldFallbackOnMethodLookup` allows.

The R3-5 precision test needed rescoping rather than relaxing: it asserted that
EVERY edge to the contested field is a precise return-shape edge, which the
producer's own (correct, name-tier) write now violates. It asserts the reader
edges are precise and the producer's write binds to its own key — two different
claims reached two different ways, which is what the code now models.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(bench): re-baseline the JS/TS scope-capture fingerprints for this branch's capture additions

The `Cross-language scope-capture fingerprint + scaling guards` CI step was
failing on TypeScript and JavaScript, and it had been failing for the whole PR —
the branch changed both SCOPE queries without ever updating the guard's
baseline. It only surfaced now because a merge conflict had prevented CI from
running at all, so nothing reported it.

Re-baselined per the file's own instruction ("re-baseline intentionally on a
legitimate capture change"), and verified first rather than rubber-stamped. The
capture-name sets in both scope queries, diffed against `origin/main`:

  TypeScript  + @reference.read.identifier      (A2, bare-identifier reads)
              + @reference.type                 (R2-2, type references)
  JavaScript  + @reference.read.identifier      (A2)
              + @reference.read.destructured    (R2-1c)
              + @reference.write.property-key   (R2-1b)

Nothing removed on either side. A pure superset is the check that no EXISTING
capture moved — which is the failure mode a fingerprint guard exists to catch,
and the reason to look before regenerating.

Consistent everywhere else too: `capture_groups_small`/`_large` are unchanged
(4503/14403) because those measure the SYNTHETIC scaling source this branch does
not touch, so only the fixture-corpus number moves — 2097 -> 2338 across 21 new
lang-resolution fixtures, 146 -> 151 files. Scaling stayed linear and inside
budget (typescript 1.116, javascript 1.010, both < 1.5), so the added rules cost
no super-linear time. Prior and new hashes are recorded in the baseline note, as
every previous entry in that file does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(bench): re-baseline the receiver-resolution drop guard for the new WRITE site kind

Second of the two bench guards that had been failing for the whole PR without
anyone seeing it — CI could not run while the branch was conflicted, so both
went unreported until the merge cleared.

The drift is a new site KIND, not a movement in an existing one:

    totalDropsAllKinds  129 -> 140
    bySiteKind          {call: 102, read: 27}
                     -> {call: 102, read: 27, write: 11}

`call` and `read` are byte-identical, which is the check that matters. This
branch added write-site captures the corpus never had — `@reference.write.
property-key` (R2-1b record construction) and the destructured-read rules — so
write sites reach receiver resolution for the first time, and 11 of them have a
receiver that does not resolve. A drop is the honest outcome for those; the
alternative is the name-inferred guess this series spent three rounds bounding.

Verified it is NOT caused by this session's review fixes before re-baselining:
removing the `memberNotOnShape` site-claim added in 69047086 and re-running gives
the identical 129 -> 140 / write: 11 drift, so the movement predates today and
belongs to the capture work, exactly as the arithmetic above says.

The sibling `scope-emission` guard still PASSES untouched, and the fingerprint
guard passes after 20a937f4 — so all three arms of the benchmarks job are green
locally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(routes): track boolean polarity in dispatch guards, so a negated condition cannot invent a route

Reproduced exactly as reported. `dispatch-guard.ts` refuses to inherit a verb
from an `if` whose `else` branch holds the comparison — the module's own doc
comment explains why: that branch runs precisely when the condition did NOT
hold, so attributing it is backwards. `!` is the same fact written as an
operator, and it was not handled. A stated invariant with half an
implementation, which is worse than an absent one, because the comment reads as
though it were covered.

Measured against the real extractor before fixing:

    if (!(pathname === '/api/admin'))                  ->  '' /api/admin   INVENTED
    if (!(req.method === 'GET') && pathname === '/x')  ->  GET /x          INVERTED
    if (!(req.method === 'POST' && pathname === '/w')) ->  POST /w         BOTH

And the review is right that this is not additive-only. Driven through the real
pipeline with a policy module that serves nothing plus a one-line route table,
the invented `GET /api/report` collected into `verbedUrls` and
`reconcileDispatchGuardRoutes` then EVICTED the true verb-less route for that
path. A false route deleted a real one. After the fix that repo yields exactly
one route, verb-less, path intact.

Parity, not presence: `!!x` is `x`, so counting negations and testing the parity
is the only rule that keeps a doubly-negated guard working. A negated VERB drops
to verb-less rather than dropping the route — `!(method === 'GET')` means every
method except GET, which no single value expresses, while the path evidence is
untouched. Applies to the regex arm too; `!/^\/api\/x$/.test(pathname)` had the
identical hole.

Deliberately NOT keeping the `statement_block` break from the suggested patch.
It is unreachable — the `!` in `if (!cond) { … }` lives in the condition, a
SIBLING of the block, never an ancestor of anything inside it, and the only
shape that puts a `!` above a block is an IIFE, which the function-boundary stop
catches first. Unreachable in the UNSAFE direction, too: breaking early
under-counts negations, and an under-count reads a negated guard as positive and
invents the route. Verified by mutation — with the break present, deleting it
fails nothing; the other three guards each fail a test when removed.

Six new cases, all previously absent (`grep -c '!(' ` over both test files was 0,
and the only negation covered was `!==`, the form that already worked).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(bench): re-baseline the emit-persistence byte-identity fingerprint for the isDetail column

The third bench guard this branch left red, and the one the earlier
rebaseline pass missed: the `benchmarks (GITNEXUS_BENCH)` job has never
succeeded once in eleven attempts, and since step 11 aborts the job, the
two steps after it — the streaming PDG-emit guard and the cross-language
pipeline benchmarks — have never executed at all.

    [emit-persistence --check] FAIL: byte-identity fingerprint drift
      (got 4ee15e74…, expected 69e9182a…)

Cause is this branch's own `isDetail` BOOLEAN on the Property table
(PROPERTY_SCHEMA), which `streamAllCSVsToDisk` writes as one more header
field and one more cell per Property row.

Verified header-only rather than regenerated on faith. Dumping every CSV
the bench emits on both `origin/main` and this branch and diffing them
per file (name, byte length, sha256): the file set is identical at 35
CSVs, 34 of the 35 are byte-identical, and the sole difference is
property.csv growing 68 -> 77 bytes as the header gains `,isDetail`. The
synthetic graph mints no Property nodes, so not one data row moved —
which is the thing this fingerprint exists to catch. Both timing gates
were green throughout (scaling_ratio 0.783 against a 1.8 budget,
elapsed_ms_large 229ms against the 1000ms backstop), so no throughput
claim is being rebaselined away.

Justification recorded in a `_rebaselined_<reason>` key, the convention
bench/scope-capture/baselines.json already sets, and the note now says so
explicitly so the next regeneration records its reasoning too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCptYZWRgnnJ821rzebQyj

* perf(processes): build each trace key once, not once per comparison

`deduplicateTraces` held its `join('->')` inside the `some()` callback, so
every already-kept trace had its key rebuilt from scratch against every
candidate: O(T*U) joins of O(depth * id-length) characters. The
allocation, not the substring scan, is what the pass spends its time on.

Nothing about breadth-first search made that safe. It only hid the cost by
keeping traces short — measured on this repo the walk averaged 4.3 steps
before D1 and 9.4 after, which roughly doubles both the number of
surviving traces and the length of every key, so the same quadratic that
was affordable under BFS is about six times the work under DFS. That is
the whole of the slowdown D1 was carrying; the depth-first walk itself is
cheaper than the queue it replaced (`pop()` against an O(frontier)
`shift()`), and its frontier is bounded by depth rather than by breadth.

Hoisting the join into a `uniqueKeys` array removes the multiplication.
Measured back to back on one host, 5 reps, 25k callables, production sink
path (main -> this branch before -> this branch after):

    deep_chain      876.8ms -> 1233.1ms -> 101.9ms
    mixed_cycles    731.4ms -> 1130.6ms -> 132.8ms
    shallow_wide    572.5ms ->  531.8ms ->  49.6ms

and on the real gitnexus/src corpus (11,490 symbols) process detection
goes 204ms -> 89ms against main, having been slower than main before.

Output is unchanged, which is the property that matters here: swapping the
file back and forth and diffing every non-timing field across all sixteen
shape x scale x sink-variant configurations gives no difference, and the
real corpus returns the same 936 processes / 4,648 steps either way. Sink
keys are pushed alongside the traces they belong to, so the comparison set
is the same set it always was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCptYZWRgnnJ821rzebQyj

* fix(processes): type the parse-output read as ParseOutput

The R3-6 sink read declared its own structural shape for the parse output
instead of naming `ParseOutput`, which made it the only one of the five
parse consumers in the repo not bound to the real type — cross-file.ts,
orm.ts, routes.ts and tools.ts all pass the type argument.

`getPhaseOutput` is a raw `as T` cast, so a local shape checks nothing at
runtime and only severs the compile-time link: renaming `allFetchCalls` on
`ParseOutput` would still compile here and silently detect zero sinks
forever. Verified with a real `tsc --noEmit --strict` run over exactly that
rename — the typed consumers error, this one did not. The runtime `.filter`
stays, since it is the only thing actually guarding the cast.

Also brings the phase docblock back in line with the deps array, which was
missing `structure` (pre-existing) and `parse` (added by this branch), and
records the two parse fields the phase now reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCptYZWRgnnJ821rzebQyj

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-08-08 09:58:14 +01:00
azizur100389
7ac0c86165
fix(scope-resolution): link Record graph nodes (#2871)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(scope-resolution): link Record graph nodes

Register Record definitions and caller anchors so Java and C# record targets and initializer sources resolve to canonical nodes. Keep generated LadybugDB relation pairs and the production benchmark baseline synchronized.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bench): harden schema-pair production gate

Derive the production count from executable DDL, fail closed when its budget is missing, and independently pin Record-to-Property coverage. Keep benchmark evidence machine-scoped and correct stale schema counts.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-07 18:33:27 +01:00
Gergő Magyar
997fc05b83
fix(resolution): resolve calls through a generic-typed field receiver in every language (#2833) (#2855)
* test(resolution): pin generic-typed field receivers across languages (#2833)

A field whose declared type carries a type argument (`repo: Repo<User>`)
emits zero CALLS edges — not a truncated chain, not an edge to the
interface declaration, nothing. This adds the cross-language matrix that
measures it, modelled on the #2807 inferred-field matrix: every language
runs the same two calls, one through a generic-typed field and one
through a non-generic control field, and each language is compared
against its OWN control row rather than an absolute edge count.

Measured state, pinned here as `known-gap` so the file is green on main
and flipping a row is a visible edit:

  affected    TypeScript, C#, C++, Python
  unaffected  Java, Kotlin, Go, Rust, Swift, Dart

The unaffected six erase type arguments at interpret time (Java's
`stripGeneric`, F41 #1928; Swift likewise). TypeScript, C# and Python
instead run a container ALLOW-LIST that returns the type ARGUMENT, so a
user-defined `Repo<User>` survives verbatim into a lookup that binds
nothing.

The `ts-local-vs-field` case is the bug in one file: `viaLocal` and
`viaParam` both resolve for the identical type, and only `viaField`
loses every edge — a bare name reaches Case 4 and its generic-aware
lookup, a dotted field receiver does not.

Negative controls pin what erasure must NOT do: an unbounded type
parameter denotes no declaration, and a C++ explicit specialization is a
different class from its primary template. The `Box2<T>` row pins a
PRE-EXISTING false edge (a workspace class named `T`) so it cannot later
be mistaken for fallout from this work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* refactor(resolution): move resolveClassBindingForName to the shared walkers (#2833)

Pure relocation, no behaviour change: the generic-aware class lookup
moves from `passes/receiver-bound-calls.ts` to `scope/walkers.ts`, beside
the bare `findClassBindingInScope` it wraps. Its two existing callers —
`classifyReceiverOrigin` and Case 4 — import it from the new home and are
otherwise untouched.

The move is required rather than cosmetic: `receiver-bound-calls.ts`
already imports from `compound-receiver.ts`, so having the compound
receiver call into the pass would close an import cycle. `walkers.ts` is
the shared floor both already depend on.

Verified behaviour-neutral: the #2833 matrix is 44/44 identical before
and after, across all fifteen fixtures.

detect_changes attributes `resolveInheritanceBaseInScope`,
`resolveQualifiedInheritanceBase` and `EMPTY_BINDINGS` to this commit;
those are line-shift artifacts of inserting a function above them, and
their bodies are byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(resolution): type generic field receivers through the generic-aware lookup (#2833)

A field receiver is spelled `this.repo` — dotted — so it types through
the receiver-chain fold and the text cascade, both of which reach
`findClassBindingInScope`. That function has no notion of type arguments,
so a field declared `Repo<User>` resolved to nothing and the call site
emitted NO edge at all: not the interface declaration, not the
implementation fan-out, nothing. A local or parameter of the identical
type is a bare name, reaches Case 4 and its generic-aware
`resolveClassBindingForName`, and resolved fine. The bug was the
asymmetry, not the generics.

Three receiver-typing lookups now call the generic-aware helper instead:
`typeOfMemberOnClass`'s primary and module-hoist branches, and the
cascade's bare-identifier type-binding read. Every other one of the 38
`findClassBindingInScope` call sites is untouched — its own docstring
records that widening it globally suppresses the `?? otherResolver(...)`
fallbacks two dozen callers rely on, which would retarget inheritance
edges, and impact rates it CRITICAL with 12 direct dependents.

Order matters and is preserved: the helper tries the exact name, then an
arity- and token-exact match against `def.templateArguments`, and only
then falls back to the base name. Erasing first would collapse a C++
explicit specialization onto its primary template — `Vec<bool>` really is
a different class. A bare type parameter carries no type arguments, so it
never enters the generic branch and cannot be erased into a class that
happens to share its name.

Measured: TypeScript and C# generic-typed fields now emit exactly what
their non-generic control rows emit, primary plus interface-dispatch
fan-out. Java, Kotlin, Go, Rust, Swift and Dart are byte-identical. Both
type-parameter negative controls are unchanged. C++ and Python are still
open and stay pinned as known-gaps — they fail for different reasons and
get their own commits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(cpp,python): bind generic-typed member fields so their calls resolve (#2833)

Completes #2833 for the two languages the shared resolution change could
not reach. Each failed for its own reason, and both were found by
measurement rather than assumed.

C++ — a CAPTURE gap, not a resolution one. All three `field_declaration`
type-binding rules required `type: (type_identifier)`, so a member
declared `Repo<User> repo;` is a `template_type` and matched none of
them: the field got no type binding at all, and every call through it
lost its edge in both the bare and `this->` spellings. A LOCAL of the
identical type resolved the whole time, because the local declaration
rules gained their `template_type` variant long ago. Three mirrored
rules close it, one per declarator shape (plain, pointer, reference).
Written as separate patterns rather than one alternation: a node-type
alternation in a field position is a tree-sitter 0.21 hazard this repo
has been bitten by before.

Python — the bracket spelling never entered the generic branch. Its
`stripGeneric` is a container allow-list over `[...]` that returns the
type ARGUMENT (`list[User]` to `User`), so a user-defined `Repo[User]`
matched nothing and survived verbatim, and the shared lookup's generic
branch is gated on `<`. It now reduces a subscripted type neither
allow-list claims to its base name — the same rule Java and Swift
already apply to `<...>`. Deliberately the LAST resort: a container must
reach its own rule first, or `list[User]` would type the receiver as the
container and retarget every call in a for-loop chain. The as-written
spelling survives on `TypeRef.declaredSpelling`, which is what the fold's
index step reads.

Both are parse-time and land in the cached ParsedFile, so SCHEMA_BUMP
goes 45 -> 46 with its pin test. Verified free against origin/main; the
ledger in that file records three prior EXACT clashes, so re-check again
immediately before merge.

The matrix now covers the spellings real code writes, all measured: a
nullable generic, a bounded wildcard, a raw type, a nested generic and a
multi-argument one. None needed work beyond the shared lookup, which is
the evidence that base-name erasure is the right primitive. The C++
specialization control now asserts what it was written for:
`Vec<bool>.save` and `Vec.save` are DIFFERENT target ids, so the
arity/token match still wins over erasure.

scope-capture is byte-identical for cpp and c, so no rebaseline — the
bench corpus contains no generic-typed member field, which is worth its
own coverage issue.

Two pre-existing gaps were measured and are deliberately NOT fixed here,
because in both cases the language's own non-generic CONTROL row fails
identically: C++ `this->field.m()` emits nothing, and JavaScript/PHP
docblock-declared field types bind nothing at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(python): do not reduce containers or typing special forms to a base name (#2833)

Review finding on this branch's own Python change, caught by probing the
interpreter directly rather than by reading it.

The base-name reduction was reached by FALLTHROUGH: "neither container rule
matched" was treated as "not a container". It is not, and two measured
shapes proved it:

  dict[str, list[User]]   ->  dict          (was: the annotation, intact)
  Dict[str, Repo[User]]   ->  Dict
  Callable[[int], User]   ->  Callable
  Literal["a"]            ->  Literal
  Union[A, B]             ->  Union
  tuple[int, ...]         ->  tuple

The dict rule's value group cannot span a nested `]`, so a nested value
declines and falls through — and the dict rule's own comment says that
shape is deliberately "left for a downstream strip pass". Collapsing it to
`dict` destroyed the value type instead. The typing SPECIAL FORMS are worse:
`Callable`, `Literal`, `Annotated` and `Union` are not classes, and reducing
them to a bare name binds any workspace class that happens to share it —
a fabricated edge, which is strictly worse than the missing edge #2833 set
out to fix, and those names are ordinary enough for a real codebase to
declare.

Reduction is now guarded by an explicit deny set covering the containers the
two allow-lists already own and the typing special forms. Everything named
there keeps its as-written text and resolves exactly as it did before #2833.

`arr[0]` also reduces to `arr` in isolation, but that is unreachable and is
now documented as such: every Python `@type-binding.type` capture is a
`(type)`, `(identifier)`, `(attribute)` or `(dotted_name)` node, so a
subscripted VALUE expression never reaches the interpreter.

Pinned by a new unit test that asserts all four groups — user generic
reduces, container reduces to its ELEMENT, declined container shape stays
intact, special form untouched. Reverting the deny set fails three of its
five cases.

Also corrects `resolveClassBindingForName`'s docstring, which this branch
had made false: it claimed only `classifyReceiverOrigin` passes the
decoration stripper, while the three receiver-typing lookups in
compound-receiver.ts now pass it too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(resolution): rank base-name candidates lexically and refuse arg-pinned defs (#2833)

Review of #2855 found that this PR turned a MISSING C++ edge into a
CONFIDENTLY WRONG one — the direction this subsystem calls unrecoverable.

`resolveClassBindingForName` ended with an unguarded base-name fallback
that returned the first same-named class the scope chain reached. A C++
primary template carries `templateArguments === undefined`, so it can
never satisfy the exact-args branch, and every non-specialized
instantiation fell through to that fallback. Measured through the real
pipeline: with the primary forward-declared and the specialization
defined first, `Vec<int> vi; vi.save()` emitted `Vec<bool>::save`.
Declaring the primary first gave the correct target — selection was
SOURCE-ORDER DEPENDENT. Two more triggers behaved the same way: a
partial specialization (`Vec<int*>` against `Vec<T*>`), and lexical
shadowing between a global `Box<bool>` and a namespaced `N::Box<bool>`.

Two changes, neither of which is any of the three remediations the
review proposed — each was rejected on measured evidence:

- Exact-argument matching is now LEXICAL-FIRST. Candidates come from the
  scope chain, and the workspace-wide qualified-name bucket is consulted
  only when the chain produced no exact match, so cross-file
  specializations still bind.
- The base-name route refuses a definition that pinned its own template
  arguments: if the fallback's answer carries `templateArguments`, the
  visible candidates are re-decided with those removed — exactly one, or
  decline.

Why not the filed options. "If specializations exist and none matches
exactly, return undefined" deletes a green committed row
(`neg-cpp-specialization/runInt` legitimately resolves to the primary).
"Resolve all defs for the base name, return only on exactly one" deletes
a working edge for C# `partial class Repo<T>` split across files — two
unspecialized defs under one name is legitimate, and
`QualifiedNameIndex`'s own docstring names that case. Preferring the
primary alone fixes nothing about shadowing, which is a ranking bug.

The guard is expressed as `carriesOwnTemplateArguments`, not as
"specialization", so shared pipeline code still names no language
(AGENTS.md R6). It can only fire where a declared name carries concrete
arguments — measured `undefined` for `class Repo<T>` in TypeScript and
C# and for a C++ primary template — so the blast radius is bounded to
C++-style specializations.

Partial-specialization SELECTION is deliberately not implemented:
choosing `Vec<T*>` for `Vec<int*>` needs template-argument deduction,
which is a semantics expansion and cannot live in language-neutral
shared code. The source-order dependence is what is fixed; the answer is
now deterministically the primary.

Also in this commit: dropped an unreachable `?? []` (QualifiedNameIndex
returns a frozen empty array on miss by contract) whose comment was
wrong on both clauses; made the docstring true about argument ERASURE
being what widens what binds, rather than only the decoration stripper;
and corrected a stale pointer that still placed
`resolveClassBindingForName` in `receiver-bound-calls`.

`findClassBindingInScope` itself is untouched — 38 call sites, CRITICAL.

Verified: matrix 56/56, cpp.test.ts 334, unit scope-resolution 1505.
Mutation proof: reverting this file fails the three trigger cases and
passes the non-regression cases; restoring it passes all five.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(python): close the deny-set drift axis by case-folding, not by vigilance (#2833)

Review of #2855 found `NOT_A_USER_GENERIC` was a closed list over an
open universe: four review lanes each escaped it with a DIFFERENT set of
names. `Deque` was the sharpest — its lowercase twin `deque` was already
listed, so the omission was an internal inconsistency rather than a
judgement call, and with a workspace `class Deque` present
`self.dq: Deque[User]` fabricated a `Deque.appendleft` edge.

The structural cause is PEP 585: nearly every container has two
spellings differing only in case (`deque`/`typing.Deque`,
`frozenset`/`FrozenSet`). Exact matching forced every pair to be listed
twice, so any half-pair was a silent escape. The deny lookup is now
CASE-FOLDED, which closes that axis by construction — `Deque` becomes
impossible rather than remembered.

`SINGLE_ARG_CONTAINERS` and `MAPPING_CONTAINERS` are now the single
source of truth: they build the two container regexes (verified
byte-identical `.source` and `.flags`, so zero behaviour change) and
feed the property test. The deny set is re-scoped to a closed, auditable
universe — the documented Python stdlib type-system surface — and grew
39 -> 65 concepts: the `collections.abc` views, `contextlib` managers,
`re.Pattern`/`Match`, the `IO` family, ordinary-named stdlib generics
(`Queue`, `Task`, `Future`, `PathLike`), the remaining typing special
forms, and the generic machinery (`Generic`, `Protocol`, `TypeVar`...).

Third-party generics (`Mapped`, `QuerySet`, `Model`) are deliberately
NOT added and are pinned as a decision: that universe is open,
enumerating it only chases the last escape, and declining `Model` would
cost real edges in the many projects that declare one.

The review's suggested property test — derive the names from the
`single`/`dict` regex sources — would NOT have caught `Deque`: `deque`
appears in neither regex, only in the deny set. Both properties are
implemented, since they catch different drift.

The unit test was also TAUTOLOGICAL: it asserted members OF the deny
set, so it structurally could not detect an omission. It now asserts
case-fold closure and PEP 585 alias coverage, and the capture fixture
drops its `as unknown as` cast for the fully-typed helper pattern the
sibling `java-interpret.test.ts` already uses.

Still at interpret time, so no further SCHEMA_BUMP (already 45 -> 46).
Proving the base is a class the FILE can see — the real fix for the
remaining exposure, since `findClassBindingInScope` binds any name with
exactly one workspace def regardless of scope or imports — is a
follow-up, not reachable from this file.

Mutation proof: restoring HEAD's deny-set contents and exact-match
lookup fails four assertions including the `Deque` pair, with the
pre-existing guard rows still passing; restoring gives 125/125.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(cpp): capture qualified generic member fields, and make the bench gate see them (#2833)

Review of #2855 found that the three `field_declaration` rules this PR
added only matched a DIRECT `template_type`, so the common real-world
spelling still bound nothing: `std::vector<Item> items;`,
`ns::Repo<User> r;` and `std::unique_ptr<Repo> p;` parse as a
`qualified_identifier` WRAPPING a `template_type`. "C++ fixed" was
overstated.

Six new patterns — three declarator shapes (plain, pointer, reference)
by two qualifier depths — written as separate patterns rather than one
alternation, keeping the tree-sitter 0.21 field-position discipline the
existing rules follow.

The design choice was measured, not assumed. Codex suggested preserving
the full qualified spelling and normalizing `::`; preserving resolves
NOTHING, because `findClassBindingInScope`'s dotted-tail fallback splits
on `.` while C++ writes `::`, and `ns::Repo` is not an index key either
(C++ emits no `@declaration.qualified_name`). Measured: `ns::Repo<User>`
resolves to nothing, `ns.Repo<User>` resolves to `Repo`. Since a
tree-sitter capture is a NODE and not synthesized text, the only lever
is which node to capture — so `@type-binding.type` goes on the INNER
`template_type`, dropping the qualifier and landing on the same
single-match-or-decline path the bare spelling already takes.

Qualifier depth 3+ (`a:🅱️:c::Repo<User>`) remains uncaptured. Stated as
a limit and pinned by a test row, not claimed as fixed.

The bench blindness the review identified is also closed. The
`scope-capture` C++ corpus contained ZERO template-typed member fields —
confirmed a fourth way by applying six demonstrably behaviour-changing
patterns and getting a byte-identical fingerprint. The corpus now
carries generic and qualified-generic members, and the gate is load
bearing for the first time: three states that all hashed to 856d02f3
before now differ (pre-#2833 0e7cbda7, +this PR's 3 rules de07d8b5,
+these 6 rules bd47c82d). Rebaselined for cpp only; c is unchanged.
Histogram diff: only 5 tags move with the fields, each by exactly +40
(20 entities x 2), and every `@reference.*` count is unchanged.

Over-match is preserved: 20 shapes still produce no field capture,
including the 8 original method/pointer/reference/function-pointer/
using/typedef/friend/operator forms plus their `std::`- and
`a:🅱️:`-qualified variants.

Not fixed here, deliberately: NON-generic qualified fields
(`ns::Address addr;`, `std::string name;`) still capture nothing.
Closing that needs six more patterns and would newly bind every
`std::string`/`std::mutex` member repo-wide, changing edges far outside
#2833. Separate issue.

The template-template-parameter hazard the review filed against these
rules is NOT capture-side: a tree-sitter query has no scope knowledge,
so it cannot know `Map` is bound by the enclosing `template <...>`
header, and the PRE-EXISTING `type: (type_identifier)` rule already
captures a bare `T item;` and erases it the same way. It is handled by
the lexical ranking in `walkers.ts` in this series.

Mutation proof: reverting this file fails 9 of 32 assertions (all eight
qualified spellings return no capture) while every over-match negative
still passes; restoring gives ALL PASS. Bench `--check` passes for all
15 languages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* test(resolution): pin specialization order, shadowing and the untested spellings (#2833)

Grows the generic-field matrix 56 -> 114 tests, closing every coverage
gap the #2855 review named and turning the fix-agents' scratch evidence
into permanent rows.

The rows that discriminate against the resolver fix (they fail if
`walkers.ts` is reverted):

- C++ specialization must not depend on DECLARATION ORDER: the
  forward-declared-primary/specialization-first arrangement must land on
  the primary, same as the mirror arrangement. Plus a cross-case
  property asserting the two independently built fixtures agree.
- Partial specialization is deterministic in both orders. The note says
  explicitly that selecting `Vec<T*>` would need argument deduction and
  that flipping this row later is a deliberate expansion, not a
  regression fix.
- Lexical shadowing: the namespace-local `N::Box<bool>` wins for a field
  inside `N`, and the global specialization wins at global scope.

The NON-REGRESSION rows are load-bearing — they are why two of the three
proposed remediations were rejected: cross-file C++ specialization
binding, and C# `partial class Repo<T>` split across two files with the
field in a third (two legitimate unspecialized defs under one name).

Coverage the review found missing: C++ pointer and reference generic
fields (two of this PR's three original rules had ZERO coverage); all
six qualified patterns plus the depth-3 boundary pinned as empty;
TS/C# multi-arg container collision; an anti-vacuity sibling for
`neg-bounded-type-parameter`; Swift/Dart rows restructured so the
ANNOTATION is the only possible source (the old rows gave the field an
initializer of the same generic type and could not tell which resolved);
and cross-file, inheritance/MRO, import-alias, static-member and the
TypeScript module-hoist branch.

Six things were measured and pinned AS MEASURED rather than asserted as
wishes, each flagged in its row note: a static/class-level member emits
nothing for generic AND non-generic alike (a static gap, not a generics
one); a cross-file C++ primary template does not bind while the
cross-file specialization does; `std::unique_ptr<Payload>` types to
`unique_ptr` rather than `Payload` (smart-pointer transparency is not
applied on the qualified path); two same-named C++ specializations in
one file collapse to one node id; and the container-name collision
(`Map<string, User>` binding a workspace `class Map`) is recorded as
INTENDED, since the annotation does name that class.

The `new Set(...)` dedup was kept rather than narrowed: a per-case
surplus-edge sweep measured ZERO duplicate edges anywhere in this file,
Swift included, so the quirk that justified a blanket dedup does not
reproduce. The sweep now pins zero surplus per case, so a real
double-emit fails instead of being absorbed.

The file is deliberately NOT split: four assertions compare cases
against each other, cost is linear in cases, and the 1,800,000 ms
`beforeAll` is kept because the same run measured 271-428 s depending on
host load — a tighter bound converts contention into a red suite. The
reasoning is recorded in the file header.

Also corrects the SCHEMA_BUMP pin-test title, which still said (#2766).

Mutation proof: reverting `walkers.ts` fails exactly the five order and
shadowing assertions and passes the other 109; restoring gives 114/114.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* feat(resolution): capture declared type parameters so a type variable is not a class (#2833)

Three review findings were blocked on one missing fact. `templateArguments`
records the arguments a declaration was written AGAINST (`struct Vec<bool>`);
nothing recorded the parameter list a declaration DECLARES (`template <class T>`,
`class Box<T extends Repo>`). So the resolver could not tell a type variable
from a class, and:

- `class Box2<T> { t: T }` beside a workspace `class T` emitted a FALSE edge
  `run2 -> T.foo`. `T` carries no type arguments, so it never entered the
  generic branch — the plain lookup simply bound a same-named class. The
  lexical grounding added elsewhere in this series cannot help, because
  `export class T` IS lexically bound.
- `class Box<T extends Repo> { t: T }` resolved to nothing: no recorded bound
  to resolve through.
- A full specialization `template<> struct Vec<T*>` and a partial
  `template<class T> struct Vec<T*>` were byte-identical (`['T*']`).

`SymbolDefinition.typeParameters` now records `{ name, bound? }` in declaration
order (substitution is positional). `bound` is kept verbatim and un-split, so
`Repo & Closeable` stays whole; ABSENT means UNKNOWN, never "unbounded", which
is what keeps unconverted languages behaving exactly as before.

Transport is the raw parameter-list node via `@declaration.type-parameters`,
read by a language-neutral parser that recognizes TOKENS, not languages:
`extends`/`:` introduce a bound, the name is the trailing identifier, so
`class T`, `typename T`, `in T`, `out T`, `reified T` and `class... Ts` are one
rule. Populated for TypeScript, C++, Java, Kotlin, C# and Rust. JavaScript, C,
COBOL, PHP and Ruby have no declared type parameters to capture; Go and Python
spell them with SQUARE brackets, which this parser deliberately rejects as
ambiguous against subscript and array spellings (Go already has a working
main-thread sidecar in this series); Dart and Swift are straightforward
follow-ups.

Two latent hazards found and closed on the way:

- The new capture was not in `KNOWN_SUB_TAGS`, so it could out-span its own
  declaration and become the anchor — silently DROPPING the whole class def.
- A templated C++ struct matches both the standalone and `template_declaration`
  patterns, minting two defs under one id, and only one twin could see the
  parameter list. `buildDefIndex` is first-write-wins, so MATCH ORDER decided
  whether `Vec` remembered `T`. A narrow duplicate-declaration backfill gives
  both twins the list.

Also fixed by its own test: a Rust lifetime `'a` parsed as a parameter named
`a`, which would have shadowed a real class.

Parse-time output lands in the cached ParsedFile, so SCHEMA_BUMP goes 46 -> 47.
Re-checked against origin/main at write time: main is on 45; 46 was taken by
this same branch, and a warm cache stamped 46 carries ParsedFiles with no
`typeParameters` at all.

The csharp and rust capture goldens were regenerated with the tests' own
documented `UPDATE_GOLDEN=1`; only digests moved, no captureGroups.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(resolution): ground erased base names, and stop a class name from being enough (#2833)

The review's central risk was that this PR converts MISSING edges into
CONFIDENTLY WRONG ones. Base-name erasure (`Repo<User>` -> `Repo`,
`Repo[User]` -> `Repo`) bound through a workspace-wide qualified-name
fallback that consults NO scope, NO import and NO module — it bound any name
with exactly one workspace def. That is why a Python `Mapped[User]` could bind
an unrelated `class Mapped`, and why the language deny lists were papering
over an open universe.

`resolveErasedBaseName` now admits an erased base on one of four grounds,
strongest first: the scope chain binds it; the declaration is in the SAME
FILE; the index proves the name is a template family; or the file binds no
cross-file class at all, so its silence is no evidence. The last ground fails
toward permissive on purpose — every way it can be wrong costs a wrong edge
that already existed, never a working one.

Two measurements drove that design and refuted the simpler rule. A C++
`#include` materializes NO binding whatever, and C# resolves cross-namespace
without `using` through the index — so a pure "require lexical grounding" rule
would have deleted every cross-file C++ generic member. Both are now pinned.

Python erases at CAPTURE time, so by resolution there is no `<` and the
grounded route was never entered. `erasedTypeApplication` rebuilds the
application from `TypeRef.declaredSpelling` — strictly: the raw name must be
the base and the argument list the whole balanced remainder, so `User[]`,
`vector<Item>` and `Repo<User>?` decline and behave exactly as before.

Closing it took finding FOUR emitters, not one. Three were in Case 4; the
fourth was `emitReferencesViaLookup` re-emitting the refused edge from the
pre-resolved reference index, which needed the site marked handled with a
recorded `receiver-unresolved`. A fifth lived in the text cascade: a declined
fold falls THROUGH by design, and the cascade held its own ungrounded copy of
the member-typing lookup. This file typed a receiver from a `TypeRef` in five
places and the PR had wired three; all five now go through one
`classOfDeclaredType`.

Also here, from the same review:

- Type parameters no longer bind a same-named class (uses the new
  `typeParameters`), and a BOUNDED parameter resolves through its bound.
- A cross-file C++ PRIMARY template now binds: a ranking bug, not a capture
  one — the index fallback needs exactly one candidate and `Vec` held two, so
  removing the argument-pinned declaration leaves one.
- `this->field.m()` resolved to nothing for generic AND non-generic alike. A
  language that declares `this` IS the enclosing class
  (`resolveThisViaEnclosingClass`) synthesizes no `this` typeBinding, so a
  chain whose BASE is `this` could never seed its head. Reading the provider
  flag keeps the rule language-free.
- Class-level (static) member receivers emit nothing in TypeScript and Kotlin
  — for the non-generic control too. Case 6 types them from the DEF side
  (`isStatic` + `declaredType` on the field node), which needs no capture
  change; the target lookup stays the ordinary instance walk, so a static
  field HOLDING an instance still binds an instance method and a genuine
  static call is untouched.

Partial-specialization SELECTION is deliberately not implemented: it needs
argument deduction against a parameter list, and full C++ partial ordering is
a real algorithm with no measured driving case. The discriminator now exists
if someone wants it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* fix(cpp,js,php,go): close the remaining per-language generic-field gaps (#2833)

Four language gaps the review measured, each with a different cause.

**C++ qualified member fields.** `std::vector<Item> items;`, `ns::Repo<User> r;`
and `ns::Address addr;` captured NOTHING: every field rule required the type
node to BE a `type_identifier` or `template_type`, and a qualified member type
is neither — tree-sitter wraps both in a `qualified_identifier`. Three
depth-agnostic rules (one per declarator shape) now match the outer node, which
also REMOVES the depth boundary rather than raising it: depths 1-4 capture,
generic and non-generic alike.

Preserving the qualifier resolves nothing — measured: `ns::Repo<User>` binds
neither way, because the dotted-tail fallback splits on `.` while C++ writes
`::`, and `ns::Repo` is not an index key. Since a capture is a NODE and not
synthesized text, the qualifier is dropped in `interpret.ts` by a top-level-only
`::` split, so `std::vector<std::string>` reduces to `vector<std::string>`, not
`string`.

Measured cost of the non-generic half, which was the reason to hesitate: field
captures go 8 -> 32 across the C++ bench corpus, but the resolution-level census
over those 13 repos is 32 CALLS edges before and 32 after, BYTE-IDENTICAL. It
fabricates only where a workspace class shares a std name (`class string` beside
`std::string name;`), which is the same accepted policy the already-landed
qualified-generic rules carry, pinned in the matrix as intended.

**JavaScript `@type {Repo<User>}` and PHP `@var Repo<User>`.** Neither bound a
field type — and neither did the NON-generic control, so this was a docblock gap
rather than a generics one. PHP needed TWO captures, not one: with only the type
binding, `$this->repo->save()` resolved until a second class declared `save` and
then went unresolved, because narrowing a same-named method needs the receiver's
member owned. Generics do NOT come free in PHP — `normalizePhpType('Repo<User>')`
returns `'User'` by the container-element convention, so passing the raw spelling
through would have emitted `User::save`; type arguments are erased at capture
instead. In JavaScript they DO come free, verified byte-identical to the
TypeScript control. Both decline what they cannot prove: arrays, `list<User>`,
unions, `Promise`/`Array` wrappers (via an exported predicate rather than a
copied name list), statics, and any property that already has a native type.

**Go generic interfaces.** `UserRepo` genuinely DOES implement `Repo[User]` —
the spec says a generic type must be instantiated, that instantiation
substitutes type arguments and yields a new non-generic type, and that a type
implements an interface when it is in its type set. So the old behaviour was a
FALSE NEGATIVE and the matrix note calling it "already correct" was wrong.
Satisfaction is now checked against POSITIONALLY SUBSTITUTED method sets, so
`Repo[Order]` does not match a `Save(x User)` implementor — substitution, not
erasure. #2829's exact method-set model is untouched: pointer receivers still
follow MS(*T), unexported names stay package-scoped, the declaration's own
method set is still checked first, and the harvest is gated so a repo with no
generic interface never runs it. `go.test.ts` is unchanged at 296 passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* test(resolution): pin every fix from the review, 114 -> 155 rows (#2833)

Eight rows in this matrix pinned gaps that the fixes in this series close, so
each asserted the opposite of the new truth. All eight are flipped, and the
prose describing them as open gaps is corrected. Nine new cases cover the fixes
that would otherwise have shipped unpinned.

Flipped, each measured: the type-parameter FALSE edge (`run2`) is gone; a
bounded parameter now resolves through its bound with fan-out; the cross-file
C++ primary binds; the C++ qualifier depth boundary is removed rather than
raised; Go gains its two structural implementors and JOINS the paired sweep,
which had quietly excluded it — that exclusion was the taxonomy admitting a bug;
and both static-member rows resolve.

Added: JS `@type` and PHP `@var` docblock fields with three PHP declines; a
Kotlin `companion object` receiver (given an INTERFACE control so the paired
sweep can check it, which `ts-reach-shapes` cannot — its two sides are not
count-comparable); the Python third-party grounding refusal plus the ground that
still ADMITS, so an empty row can never be read as "erased names never resolve";
the four mirrors that would break if grounding were tightened (same-file and
imported Python, a C++ `#include`, C# cross-namespace without `using`); C++
qualified non-generic fields including the fabrication policy and its absence
case; `this->field.m()` for generic and non-generic with bare controls; and a Go
negative proving substitution is positional, not erasure.

Three shapes are pinned AS MEASURED with notes saying they are deliberate limits
so nobody "fixes" them by accident: C++ partial-specialization selection is
deterministically the primary (real selection needs argument deduction);
`std::unique_ptr<T>` types to the pointer, not the pointee (`.` and `->` are
indistinguishable to the resolver, so transparency would trade a recoverable
miss for a confident wrong edge); and two same-named C++ specializations in one
file collapse to one node id, which is why the shadowing fixture uses two files.

One row pins a REMAINING wrong edge rather than hiding it: `m.inner.ping()` on
a `Mapped[User]` head still binds the unrelated workspace class, while the
one-segment-shallower `m.save(u)` correctly declines. The obvious one-line guard
was written and MEASURED not to close it, so the surviving route is elsewhere
and wants its own diagnosis — a broader refusal would change chain-head
resolution for every language without pinning the shape it is meant to fix.

`bench/scope-capture` is rebaselined for the six languages whose captures moved,
regenerated from a fresh measurement rather than pasted; `--check` passes for all
15.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* perf(resolution): remove three measured hot-path regressions this series added (#2833)

A quality pass over the #2833 series found three performance defects it had
introduced, all measured, plus dead code and stale docs from six agents having
appended to the same files across four rounds. No behaviour change: the
resolver suite is identical before and after, and every scope-capture
fingerprint is byte-identical.

**An accidental quadratic in Go instantiation harvesting.** `collectGoInstantiations`
calls `record()` for every type binding and every declared, return and parameter
type in every Go file, and the `includes('[')` gate does not filter Go's most
common types — `map[string]string`, `[]map[string]*v1.Pod` and
`map[string]map[string]int` all produce a `map` candidate. Each false base then
failed a full scope-chain walk and fell through to a LINEAR SCAN OF EVERY
INTERFACE IN THE PROGRAM, with no dedupe on the spelling, so the same
`map[string]string` written 10,000 times paid 10,000 scans. Now a
qualified-name index built in `buildDetectionIndexes` (one probe, ambiguity
semantics preserved exactly) plus a per-scope base memo:

    8,000 interfaces / 80,000 spellings:  6,662 ms -> 104 ms   (64x)

`resolveEmbeddedInterface` held a byte-identical copy of that scan and now
shares the helper. `GoInstantiation` was a single-field wrapper and collapses
to the array it wrapped; its two parallel maps fold into one whose inner key IS
the dedupe. `candidateStructIdsFor` was rebuilt per instantiation although
every substituted method set has the same key set — hoisted, and materialized,
because one branch returned a live iterator that would have yielded nothing on
a second pass.

**`scanForCrossFileClass` asked a name-keyed question that needs no name key.**
It answered "does this file bind any cross-file class" by probing every
accessible namespace once PER NAME. It now iterates the channels directly,
taking whichever side is smaller so a large namespace table cannot reintroduce
the product. Predicate and early exit preserved:

    5,000 module names x 1,000 namespaces:  159.0 ms -> 1.2 ms   (132x)

**A duplicated scope walk on every generic receiver.** `resolveClassBindingForName`
computed the lexical candidate list, then `resolveErasedBaseName` recomputed
the identical `findAllBindingsInScope`. Computed once and passed:

    receiver at depth 8:  5,617 ns -> 3,091 ns   (-45%)

**A whole extra AST traversal per JavaScript and PHP file.** The docblock
synthesis passes each added a full tree walk to find one node kind — the ninth
in the JS emitter, the third in PHP. `node.namedChildren` materializes a
wrapper array across the N-API boundary for every node, so one added pass cost
1.9x what parsing the entire file costs. Folded into the existing walks as one
more node kind; capture output is byte-identical and every fingerprint is
unchanged. Total emit time per file drops 4-7%.

Hygiene, all verified stale rather than assumed:

- `receiverOriginOpts` passed `resolveThisViaEnclosingClass`, which
  `classifyReceiverOrigin` never reads — the "both hooks" comment above it is
  true again.
- The `stripDecoration` docstring's caller roll-call claimed the only
  edge-emitting caller "emits no edge and can only change a diagnostic label".
  Case 6 passes it and does emit edges. Replaced the roll-call with the rule;
  six rounds each appending a name to a list is how it went wrong.
- A Python comment described the resolution-time grounding as a follow-up that
  "this parse-time pass cannot do" — it landed in this same branch and is
  pinned by `py-erased-grounding`.
- `classOfDeclaredType` took a `scopeId` all five callers derived from the
  `TypeRef` they also passed. Dropped, so "these five are the same call" is
  enforced rather than asserted.
- Three exports with no consumer outside their own file.
- PHP had three copies of one preceding-comment sibling walk and two regexes
  for one tag, so a fix to either reader of `@var` would land on one and not
  the other — the symptom being a field typed differently from its own foreach
  element type. One walk, one regex.

Tests: the new matrix leaked a fixture repo per case; it now carries the
sibling suite's `cleanupTempDirSync` and the Windows EBUSY reasoning that goes
with it. `PAIRED` was a second hand-maintained list and 19 of 41 cases had
silently fallen out of it — it is derived from the cases now, with a new
assertion that each case is either swept as a pair or carries a written reason
it is not. That recovered one genuine omission (`php-typed-property`).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KtNfG6EPn738Y51AYs7wDp

* test(bench): rebaseline receiver-resolution for the #2833 this-> fix

The `Receiver-resolution drop guards` CI step failed on this branch:

  shapeArm.cpp.fieldReceiverCall:  "INVISIBLE-GAP" -> "RESOLVES"
  shapeArm.cpp.decoratedFieldType: "INVISIBLE-GAP" -> "RESOLVES"

Both are the intended improvement. The guard is exact-match by design —
the drop count cannot move without a deliberate rebaseline, and the
rebaseline path demands the movement be explained — so this records the
two shape flips and leaves the call-drop count arm untouched.

BASELINE.md still claimed `this->repo.save()` and `this->repo->save()`
were INVISIBLE-GAP. That is now false: the `resolveThisViaEnclosingClass`
head seed added in this PR resolves both. Also notes what the control
established — this was never a generics gap, since the non-generic
control failed identically before the fix.

* docs(parse-cache): narrow the SCHEMA_BUMP ledger to what the bump delivers

The ledger claimed a warm cache would make "the whole fix ... a silent
no-op on every incremental analyze". That overstates the constant. The
bump invalidates the PARSE half; whether the re-parsed captures reach the
graph is gated separately and does not move:

  - `isIncremental` (core/run-analyze.ts) tests `!options.force`, an
    existing meta, `!schemaFingerprintMismatch(...)`, feature parity,
    non-empty `fileHashes` and a git repo. SCHEMA_BUMP is in none of them.
  - the incremental branch writes back only `hashDiff.toWrite` and logs
    the rest as "unchanged file rows preserved".
  - SCHEMA_FINGERPRINT hashes node/relation DDL, untouched here, so it is
    byte-identical and moves nothing either.

So an incremental analyze re-parses an unchanged file correctly but keeps
its existing rows; the new edges land on the next full rebuild. That is
the pre-existing contract for every capture change, not a regression in
this PR — but the comment should not promise more than it delivers.

Comment only; no behavior change. SCHEMA_BUMP stays 48.

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:14:13 +01:00
Gergő Magyar
a033b04c46
fix(go): scope and define each type_spec, not the type_declaration (#2837) (#2843) 2026-08-06 00:55:48 +01:00
Gergő Magyar
aaa78f9590
fix(scope-resolution): fan out interface dispatch from Case 3b receivers (#2832) (#2842)
* fix(scope-resolution): fan out interface dispatch from Case 3b receivers (#2832)

Case 3b (chain-typebinding) folds a receiver through the same
`resolveCompoundReceiverClass` call and the same `[owner, ...mroFor(owner)]`
walk Case 0 uses, but emitted its edge without calling
`emitInterfaceDispatchFor`. When that fold landed on an Interface the site got
one edge to the interface's bodiless declaration and none to any
implementation — the defect #2813 reported for field receivers, in the half
#2829 did not cover.

The gap was a property of how a receiver was SPELLED rather than of what it
resolved to. `d.repo.save()` contains a dot, so it took Case 0 and fanned out;
binding the identical field to a local first — `const r = d.repo; r.save()` —
made the receiver a bare name with a dotted typeBinding, which is Case 3b, and
lost every implementation edge.

`ownerDef` is the receiver's own folded type, matching Case 0's `currentClass`
and Case 4's `ownerDef`, not the owner of the member the MRO walk settled on:
a receiver folding to a concrete class that merely inherits an interface method
must not fan out, because its runtime type is that class. The closure
self-gates on `ownerDef.type !== 'Interface'`, so the call is inert for every
concrete receiver and needs no language check. Confidence is the 0.85 literal
this case's own primary emits, so dispatch edges never outrank the edge they
hang off; Case 4's site.kind-dependent value has no counterpart here because
Case 3b's primary does not vary that way.

The new fixture pins the route as well as the fix. `const r = d.repo` reaches
Case 3b and nothing else can take the site: Case 0 needs a `.`/`(` in the
receiver name or a minted receiver chain, and `encodeReceiverChain` returns
undefined for the empty step list a bare identifier produces; Case 4 excludes
itself on the dot. Before the fix the primary assertion passed while the
fan-out came back empty — the exact "reached Case 3b and stopped at the
declaration" signature.

Resolution-side only: this changes what the resolver produces, not how it is
stored, so no SCHEMA_BUMP applies. An existing index must be re-analyzed to
show the new edges.

Follow-up from #2829.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(scope-resolution): record Case 3b's interface-dispatch fan-out in I4 (#2832)

Invariant I4 documented the fan-out as something "Cases 0 and 4 both perform"
and spelled out Case 0.5's exclusion, while saying nothing about Case 3b —
which is what made 3b's missing fan-out an undocumented asymmetry rather than
a deliberate exclusion someone could defend or point at.

With the fan-out added, Case 0.5 is the only case that folds or walks to a
receiver type without dispatching to implementations, and its exclusion is
gated behind `resolveThisViaEnclosingClass`. Saying so explicitly keeps the
next reader from having to re-derive which cases fan out by reading the pass.

Comment-only; `detect-changes --scope staged` reports no graph change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(scope-resolution): add the concrete-implementor control for Case 3b (#2832)

The Case 3b fan-out shipped with one negative control — a chain folding to
PlainCache, a class that implements nothing. That proves only the weak claim:
no interface anywhere near the site, no fan-out.

Add the stronger negative. SqlRepo implements Repo, so an interface IS in
scope and `save` is a name Repo declares, yet the receiver's folded type is
the concrete class and nothing may fan out. This is the control that fails if
a later change fans out from the interface a member is DECLARED in rather than
from the receiver's own folded type.

The comment says what the control cannot do, too: it cannot catch "member
owner passed instead of folded type" in TypeScript, because an implementing
class always declares the member itself, so the MRO walk never settles on the
interface's bodiless declaration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NdZtWXQJUGB1ZGLH2YNw3o

* docs(scope-resolution): correct four overclaims the review found (#2832)

A multi-lane review of this PR reproduced, against the real pipeline, that
several claims in the comments and one test name assert more than the code
delivers. No behavior changes here — only the text, and one test rename.

1. The test comment gave the WRONG REASON why the concrete-implementor control
   cannot catch "member owner passed instead of folded type". It said an
   implementing class always declares the member itself; `class C extends Base
   implements I {}` is valid TypeScript and inherits it. The real reason is that
   TypeScript's MRO chain never contains an implemented interface, so the walk
   cannot settle on an interface declaration for a concrete receiver. The
   mutation IS expressible where a concrete class inherits a `default` interface
   method (Java, Kotlin) — reproduced during review — so this is language-scoped,
   not inherent, and a follow-up fixture is tracked.

2. "fans out to every implementation of the folded interface" certified a
   completeness that does not exist. TypeScript emits heritage edges for
   `class_declaration` only (languages/typescript/captures.ts:749, stated in its
   own docstring at :732-733), so `abstract class X implements I` and `interface
   B extends A` produce no heritage edge and still dead-end on the bodiless
   declaration. Renamed to name the shape actually covered, with a KNOWN GAP
   note. The gap is in the capture layer and predates this fan-out.

3. Invariant I4 said Case 0.5 is the ONLY case that resolves a receiver type
   without fanning out. Cases 3 and 5 do too, by direct lookup rather than a fold
   or MRO walk. The sentence now says which distinction it means and states the
   reachability argument (no known language reaches Case 3 with an Interface —
   every one that could strips the namespace qualifier first, sending it to
   Case 4) instead of implying a completed audit.

4. The gate's rationale claimed the `ownerDef.type !== 'Interface'` test is right
   for every non-Interface receiver. An abstract-class receiver also dead-ends on
   a declaration-only member and does not fan out. Noted, with why widening the
   gate belongs to Cases 0 and 4 across all languages rather than to #2832.

Also completes the module-level case ladder, which still credited the fan-out to
Case 0 alone and omitted it from the Case 3b entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NdZtWXQJUGB1ZGLH2YNw3o

* docs(scope-resolution): name Case 2 in the I4 exclusion list too (#2832)

The first pass at this correction listed Cases 3 and 5 as the other cases that
resolve a receiver type without fanning out, and was itself incomplete: Case 2
also walks an MRO and its binding admits `Interface`. It is excluded for a
different reason than 3 and 5 — its receiver IS the type name, so the site is
static dispatch and a fan-out would be wrong, whereas 3 and 5 resolve by direct
lookup rather than a fold or MRO walk. Say both rather than enumerate one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NdZtWXQJUGB1ZGLH2YNw3o

* fix(scope-resolution): close the two gaps the #2842 review left open

Both were pre-existing and reached by Cases 0 and 4 as well; the Case 3b
fan-out only widened the population of sites that hit them. Researched against
the real TypeScript compiler and the language service before choosing
semantics, plus how comparable tools draw the same lines.

1. THE FAN-OUT COULD TARGET A STATIC MEMBER

`class C implements I { static save() {} }` does not satisfy `I` — TypeScript
rejects it as TS2420, "Property 'save' is missing" — so an edge from an
`I`-typed receiver to a static member names a target dispatch can never
produce. The closure picked targets with `pickOverload`, which applies no
static filter, while the surrounding cases pick their own primary with
`pickFirstNonStaticOnly`: the speculative edges were picked with weaker rules
than the certain edge they hang off. A same-name static+instance pair also made
`pickOverload` return OVERLOAD_AMBIGUOUS, suppressing the CORRECT edge too, so
this was a false negative as well as a false positive.

Every comparable tool draws this line: tsserver partitions static from instance
results, clangd gates on `isVirtual()` (C++ forbids virtual statics), jdtls
filters abstract-or-static, and class-hierarchy analysis expands only VIRTUAL
call sites.

The guard prefers `provider.isStaticOnly` where a language declares it and
falls back to the graph node's `isStatic`. That order is load-bearing, not
stylistic: the method extractor derives `isStatic` from the OWNER type as well
as the member (`staticOwnerTypes`), and the JVM config lists
`object_declaration` — so reading the flag first would delete Kotlin `object`
implementations, which are singleton INSTANCES and genuinely reachable. Kotlin
is the only hook implementor and marks exactly the companion-promoted set;
Ruby's `singleton_class` (`def self.foo`) is correctly filtered by the
fallback.

2. TYPESCRIPT HERITAGE WAS CLASS-ONLY

`interface B extends A` and `abstract class X implements I` emitted no heritage
edge at all, so the subtype closure had nothing to descend and both shapes
dead-ended on a bodiless declaration — including the very example the closure's
own docstring cites as the reason it exists. Since Case 3b's dotted-alias
binding survives qualifier-stripping only in TS/JS, this was the language that
actually reaches the new path.

The two shapes reach their bases differently: an abstract class carries the
same `class_heritage` child a concrete one does, while an interface's bases
hang off `extends_type_clause` directly. That clause's `type` field is
`multiple: true`, so `childForFieldName('type')` would silently drop `C` from
`interface B extends A, C` — hence iterating named children.

Deliberately NOT structural matching. TypeScript is structurally typed, so a
class satisfies an interface without `implements`, but tsc's own navigation is
declaration-only and says why: "users are typically only interested in explicit
implementations... The type checker doesn't let us make the distinction between
structurally compatible implementations and explicit implementations, so we
must use the AST." scip-typescript reached the same design independently. gopls
does match structurally, but only because Go has no `implements` keyword to
prefer.

Abstract declarations are still walked THROUGH rather than targeted — the rule
everywhere is "does it have a body?", which is what `isDeclarationOnly`
already tests.

VERSIONING. The capture change is parse-time, so a v43 warm cache would serve
entries missing the new matches: SCHEMA_BUMP 43 -> 44 with its pin test moved
in the same commit, verified against origin/main at a857f4c5a (still 43).
Rebaselined only the `typescript` scope-capture fingerprint, justified by a
capture-name histogram diff over the same 145-file corpus: the only deltas are
@reference.inherits 17 -> 20 and its paired @reference.name 245 -> 248, emitted
together by `emitTsInheritanceBase`. Every other capture count is byte-identical
and javascript is unchanged, the language having no interfaces.

Tests: the fan-out now covers a static-shadowing subclass, a concrete class
below an abstract intermediate, and an implementor of an extending interface.
Resolvers 3176 passed; all five bench gates pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NdZtWXQJUGB1ZGLH2YNw3o

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:50:08 +01:00