Commit graph

1743 commits

Author SHA1 Message Date
gitnexus-release-bot[bot]
3dc2d883c4 release: v1.6.10-rc.150 2026-08-02 20:18:10 +00:00
azizur100389
797e4ef8f6
fix(ai-context): document CLI graph fallbacks (#2803)
* fix(ai-context): document CLI graph fallbacks

Teach generated GitNexus guidance to pair mandatory MCP graph checks with repo-scoped CLI fallbacks so agents can keep working when MCP is unavailable.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style(ai-context): satisfy Prettier

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-02 20:52:44 +01:00
Gergő Magyar
990d79ba8c
fix(mcp): make impact/context reproducible — deterministic ordering on every capped query (#2787) (#2796) 2026-08-02 17:03:15 +00:00
Gergő Magyar
010a7d806a
fix(schema): declare the full scope-resolution relation cross product (#2792) (#2793)
* fix(schema): declare the full scope-resolution relation cross product (#2792)

`RELATION_SCHEMA` was hand-listed, and every prior fix added only the
FROM/TO pair named in a crash report — `Const→Method` in #2769, the
Swift/Rust member pairs before it. So `analyze` kept aborting at
`assertDeclaredPair` on the next codebase whose edges happened to land on
a different pair; #2792 reports `Class→Variable` on Java.

Audit the surface instead of the symptom. `buildGraphNodeLookup` skips
any node whose label is not in `isLinkableLabel`, so the lookup holds
only linkable-labelled nodes — and both endpoints of every graph-bridge
edge resolve through that lookup. The emittable surface is therefore
exactly:

  FROM  LINKABLE_LABELS + File   (the module-level caller fallback)
  TO    LINKABLE_LABELS + CALL_TARGET_TYPES

`isCallerAnchorLabel` is a strict subset of linkable and contributes
nothing on top. `CALL_TARGET_TYPES` contributes `Delegate`, which
`tryEmitEdgeWithExplicitTargetId` can emit without going through the
lookup at all.

Generate that 14x14 block into the DDL rather than listing it: 223 -> 322
declared pairs, and no future pair from these sets can be missing by
construction. The containment/inheritance/DI/route/cluster/PDG pairs stay
hand-declared — no single predicate describes them.

Both label sets live in the ingestion layer, which `core/lbug` must not
import, so schema.ts carries twin lists. test/unit/schema-pair-coverage.ts
derives the requirement from the originals and fails CI when either set
grows without the pairs landing here — the piecemeal loop this fix ends.

Measured before widening: at 322 pairs the cost is inside noise
(1.09s vs 1.12s per 300 anchored queries on a 32-table DB), but the full
32x32 cross product is ~1.8x on untyped-endpoint anchored queries. The
audited subset is the right scope, not "declare everything".

INCREMENTAL_SCHEMA_VERSION 34 -> 35: LadybugDB fixes endpoint pairs when
the rel table is created, so a pre-v35 database physically cannot store
these edges.

Closes #2792

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(schema): declare the non-bridge structural pairs COBOL and Vue emit

The generated scope-resolution block closed the half of RELATION_SCHEMA a
label predicate can describe. The hand-declared half was still stale: with
#2791's Function->Variable fix applied, `analyze` continued to abort on this
repo's own test/fixtures/lang-resolution with

  Relationship label pair Module→Property is not declared

A full sweep (assertDeclaredPair patched to log-and-skip, run over the whole
fixture corpus) found 13 undeclared pairs over 106 edges. This branch already
covered 3 of them via the cross product; the remaining 10 come from emitters
outside the graph bridge:

  - cobol-processor.ts mints Module / Namespace / Record / Property /
    CodeElement and wires them with CONTAINS, CALLS and ACCESSES (9 pairs)
  - vue-sfc-extractor.ts emits BINDS_EVENT_HANDLER from a handler Function to
    the child component's File, the only edge whose target is a File (1 pair)

CodeElement, Namespace, Record and File are in neither scope-bridge label set,
so neither the generated block nor schema-pair-coverage.test.ts can reach them.

Adds test/integration/structural-pair-coverage.test.ts, which derives the
requirement from a corpus instead of a predicate: it runs the real pipeline
over the non-bridge fixtures and requires every FROM/TO pair they produce to
be declared. Mutation-checked — dropping `FROM Function TO File` fails it with
exactly Function|File.

Verified: cobol-app, vue-basic and php-transitive-traits now index instead of
aborting; the full lang-resolution corpus completes at 10,876 nodes / 18,517
edges; scrypster/muninndb at 0b7a4272 (the #2789 repro) completes at 20,069
nodes / 71,580 edges, matching #2791 exactly, so this supersedes that PR.

* refactor(test): simplify the structural pair coverage guard

Cleanup pass over the previous commit. No behaviour change to the schema.

- reuse `FIXTURES` and `runPipelineFromRepo` from resolvers/helpers.ts instead
  of re-deriving the fixture root and importing pipeline.js directly
- gate on `distWorkerExists()` like every other integration test that passes
  `workerUrlForTest`, so a missing dist skips rather than fails
- run the three fixtures with `it.concurrent.each`; they share nothing and the
  cost is almost all worker spawn plus grammar load, which overlaps well
  (tests phase 21-24s -> 5.6s measured)
- replace the sentinel-in-a-Set filter with a plain `.filter()` chain, matching
  the sibling unit test, and move the declared/table lookups off the per-edge
  path onto the deduped set
- move the pure string pin out of the integration tier into
  schema-pair-coverage.test.ts, where the identical construct already lives, so
  it needs no build and survives fixture deletion
- trim the schema and test prose that restated the code, and correct the
  BINDS_EVENT_HANDLER attribution: it is emitted by
  languages/vue/scope-resolver.ts, not vue-sfc-extractor.ts
- amend the v35 comment to mention the 10 structural pairs it now also stamps

Still mutation-checked: dropping `FROM Function TO File` now fails both the
integration sweep and the unit pin with exactly Function|File. 89 tests green.

* fix(schema): generate the attachment pair surface and close four analyze aborts

Review of the generated scope-bridge cross product found four `analyze`
hard-aborts still live at head, each reproduced end-to-end on the default
user path (`analyze --index-only --skip-git`):

  Method→Annotation   Spring `@Bean` + `@ConditionalOnMissingBean` (Java + Kotlin)
  Method→File         Vue Options-API `methods:` handler bound to a child event
  Namespace→Record    COBOL `DECLARATIVES` / `USE AFTER STANDARD ERROR ON <file>`
  Class→Tool          `@mcp.tool()` applied to a class

All four are pre-existing on main, and both existing guards were structurally
blind to them: the unit guard derives from LINKABLE_LABELS ∪ CALL_TARGET_TYPES
(none of Annotation/Tool/Record/File-as-target is a member) and the corpus
guard ran three fixtures that exercise none of these emitters. All 16 tests
passed while all four crashes were live.

The PR's model — "bridge endpoint × structural endpoint" — does not fit:
Namespace→Record is structural on both sides. The property that does hold is
that the ANCHOR is a lookup result, not a literal at the emit site, so the
emitter cannot constrain its label. That gives a second closed-form rule:

  DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS

DEFINITION_ANCHOR_LABELS is derived from NODE_TABLES by subtraction, so a new
node table joins automatically. 332 → 450 declared pairs.

Sized against a committed harness (gitnexus/bench/schema-pairs), real
@ladybugdb/core, identical data: 450 costs 0.93–1.05× of 332 on untyped-endpoint
anchored queries — inside noise — versus 1.22–1.43× at 641 and 2.03–2.34× at
1024. The harness reproduces the known #2792 cliff, which is what makes the 450
figure trustworthy.

Also in this change:

- Delete the 161 hand-declared pairs the rules already generate (233 → 72).
  The declared set is byte-identical at 450; those lines were load-bearing
  shadow, because the generator suppresses anything already declared
  structurally, so narrowing a rule later would silently keep pairs alive.
  A new guard fails CI if a hand-declared pair is ever re-added inside a rule.
- Import LINKABLE_LABELS / CALL_TARGET_TYPES instead of hand-copying them.
  The twins' stated justification ("the ingestion layer must not be imported
  here") is false: csv-generator.ts and lbug-adapter.ts, siblings in the same
  directory, already do, and no rule in AGENTS.md / ARCHITECTURE.md /
  CONTRIBUTING.md / GUARDRAILS.md states otherwise.
- Resolve `resolveStreamGraphEmit` after the guards that rebind `options.force`,
  not at function entry. It gates on `force`, and every freshness guard runs
  ~360 lines later, so the v34→v35 bump would have pushed every existing index
  down the non-streamed emit path — losing the #2680 memory streaming added for
  the #2649 kernel-scale OOM, for exactly the population most likely to be
  memory-constrained.
- `UndeclaredRelationPairError` now carries the relationship type, both node ids
  and the source file, with a matching CLI branch. The old message named only
  the abstract label pair, which a user could not act on. Found through the
  cause chain, since pipeline-phases/runner.ts rewraps every phase failure.
- Share one classifier (`relPairKeyFor`) across the router, both emit sinks and
  the corpus guard, which previously hand-mirrored the router's skip rule; one
  cause-chain walker in lib/utils.ts; one exported pair-matching regex.
- Corpus guard: four new fixtures reproducing the aborts, per-fixture sentinel
  pairs so a fixture that stops emitting fails loudly instead of passing
  vacuously on an empty graph.

The per-edge path stays allocation-free: the failure context is passed
positionally and the message is built only inside the throw.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

* test(bench): re-baseline the COBOL capture fingerprint for the new fixture

`bench/scope-capture` globs `lang-resolution/cobol-*`, so the
`cobol-declaratives` fixture added in 81daf370e (to reproduce the
`Namespace→Record` analyze abort) joined that corpus and shifted the
fingerprint — 14 → 15 files.

Verified corpus-only, not a capture change: with that one fixture moved
aside the fingerprint is byte-identical to the prior baseline
(d45bb091…), and 81daf370e touches no COBOL capture code. The new value
reproduces CI's reported hash exactly. Scaling 0.677 < 1.5 budget.

`bench/scope-capture/measure.mjs --check` → PASS (15 languages).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 16:26:25 +01:00
Gergő Magyar
74409a37f6
perf(cpp): index qualified namespace members once per pipeline run (#2788) (#2794)
* perf(cpp): index qualified namespace members once per pipeline run (#2788)

`resolveCppQualifiedNamespaceMember` walked every parsed file — rebuilding a
per-file `scopesById` map each time — once per qualified `ns::member()` call
site, so the scope-resolution emit phase cost O(callsites x scopes). On a
1,473-file C++ repo that was 25.3 min of a 33-min analyze, with 75% of total
self-time in this one function. Its inner `findMemberInNamespaceTransitive`
compounded it: each recursion step filtered `scopesById.values()` by parent,
O(scopes^2) per file on its own.

This is the same bug #1990 fixed in the sibling ADL path (`pickCppAdlCandidates`
-> `AdlCandidateIndex`), so it gets the same fix: a `QualifiedNsMemberIndex`
(receiver simple name -> member simple name -> callable defs) built lazily once
per `parsedFiles` identity and reset by `clearCppInlineNamespaces`, which runs
from `cppScopeResolver.loadResolutionConfig` at the start of every pass. Per
call site the work drops to two Map lookups.

Ordering is preserved exactly — file-major, `parsed.scopes` declaration order,
a namespace's own `ownedDefs` before its inline-namespace children, depth-first
— because the caller takes `allHits[0]` for the single-hit case and
`narrowOverloadCandidates` is first-wins. Non-inline nested namespaces are
still not descended into, and same-name hits across inline children still
report `'ambiguous'` (#1564).

Measured with `PROF_SCOPE_RESOLUTION=1 analyze --force --index-only` on a
synthetic corpus (`namespace ns_i { inline namespace v1 { ... } }` plus 20
`ns_j::fn()` call sites per file):

| files | emit before | emit after |
|-------|-------------|------------|
| 100   | 153ms       | 16ms       |
| 200   | 704ms       | 24ms       |
| 400   | 3,293ms     | 42ms       |
| 800   | 16,898ms    | 78ms       |

Before, doubling the file count quadrupled emit; now it doubles. At 800 files
total scope resolution goes 17.2s -> 394ms.

Output is unchanged, verified rather than assumed: a full graph dump (sorted
nodes + relationships) from a baseline build at the parent commit and from this
one are byte-identical on all 134 `cpp-*` fixtures merged into a single repo
(1573 nodes / 1997 relationships) and on the 400-file synthetic corpus.
`test/integration/resolvers/cpp.test.ts` passes 334/334.

#1990 shipped its ADL fix without a scaling gate, which is how the bug class
came straight back here, so this adds one: `bench/cpp-qualified-ns` measures
`(t_large/t_small)/(1600/400)` — 0.93-1.21 indexed versus 3.45 for the old
per-call-site scan — alongside a fingerprint over every
`receiver::member -> outcome` the corpus resolves, and CI runs it with
`--check`. `test/unit/cpp-qualified-ns-index.test.ts` covers the cache
invalidation the index introduces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(cpp): address tri-review findings on the qualified-namespace index (#2788)

Multi-engine review of this PR (Claude swarm + ce-code-review, Codex
gpt-5.6-sol swarm + ce + adversarial) returned two P1s and five smaller
findings. All are fixed here.

P1 — the index defeated the pipeline's post-language memory release.
`scope-resolution/pipeline/phase.ts` evicts each language's files and then
calls `forceGc()`, on a stated premise that "This language's ParsedFiles are
now unreachable", sizing C/C++ at ~17-20GB on the Linux kernel. The
module-level `let qualifiedNsIndexSource` falsified that: it pinned the whole
`parsedFiles` array, and the index held defs reaching into those files'
scopes, until the *next* C++ pass cleared it — which in a single analyze never
comes. C++ is 7th of 16 in SCOPE_RESOLVERS, so the set survived nine later
language passes plus emit. Replaced with a
`WeakMap<readonly ParsedFile[], QualifiedNsMemberIndex>`, the pattern already
used by `moduleScopeIndexByPass` in `cpp/file-local-linkage.ts`.
`clearCppInlineNamespaces` still swaps in a fresh WeakMap, because the index
has a second input (`inlineNamespaceScopeIds`) the key cannot observe.
Measured with `--expose-gc`: 61.2MB retained after the caller drops the array
before, 0.1MB after.

The ADL twin (`adl.ts`) has the same pattern, so the hazard predates this PR —
but `pickCppAdlCandidates` returns early before `ensureAdlIndex` on
`noAdlSites`/empty `argInfoBySite`, so it rarely arms, whereas a qualified
`ns::member()` index arms on almost every C++ workspace. Moving the ADL twin
to a WeakMap is left as a follow-up.

P1 — the new bench could not see the regression class it exists to gate.
`callSites()` drew every receiver from `ns_${...}`, so the receiver lookup
never missed; production is the opposite, since Case 1.5 in
`receiver-bound-calls.ts` is reached by every plain-identifier receiver call
and misses on most. A rescan reintroduced only on the receiver-bucket-absent
path scored 1.279 and PASSED the old bench. The corpus now mirrors production
(~1 in 5 receivers name a declared namespace) and adds a namespace reopened
across files, a same-name inline nest, a member declared at both namespace and
inline-child level, and call sites carrying a real `Callsite` so
`narrowOverloadCandidates`/`cppConversionRank`/
`isOverloadAmbiguousAfterNormalization` are inside the fingerprinted surface at
all. That same rescan now measures 4.538 and FAILS; defeating the dedup now
fails the fingerprint arm where it previously passed byte-identical. The
fingerprint moved once, deliberately, for the corpus expansion — recorded in
`_rebaseline_2788_review`, explicitly not precedent.

Also fixed:

- Unbounded recursion aborted analyze. `collectNamespaceMembers` recursed per
  inline child with no bound and threw an uncontained `RangeError` at inline
  depth 8000 (`phase.ts`'s try has a `finally`, no `catch`), and a receiver
  *miss* paid full recursion where the deleted walker skipped on a name
  mismatch. An explicit work-stack alone would only have converted that into
  an OOM at depth 6000, because the eager table was quadratic in memory too:
  for a depth-D chain it legitimately holds D(D+1)/2 entries, since `v2::foo()`
  is a valid receiver at every level. Replaced with a lazily-queried node graph
  (per-scope own-member buckets plus direct child links, resolved on demand and
  memoized per receiver+member). Build is now linear; depth 100000 costs 133ms
  where 8000 previously threw.
- "#1990 shipped without a scaling gate" was false. #1990 did ship
  `test/integration/cpp-adl-benchmark.test.ts` (f1b843838). Corrected in the
  bench header and the CI step comment. The accurate point is narrower and
  stronger: that bench asserts `callsResolved === 0`, so it never drives the
  qualified-receiver path, and it is `skipIf(!GITNEXUS_BENCH)` while the only
  step setting that variable lists neither C++ bench — so it has never run in
  CI. Wiring it in is a follow-up.
- "Ordering is load-bearing" was not a live property: `allHits[0]` is only
  reached at length 1, and both tail branches return `'ambiguous'`. Order is
  still preserved for byte-identity with the pre-#2788 walker; the comment now
  says that instead, and the test named for ordering is renamed to the
  parent/inline-child visibility it actually asserts.
- "Same contract as `ensureAdlIndex`" overstated parity — the sibling ships a
  `validateAdlSeqCoverage` guard because it reads
  `seqByNodeId.get(...) ?? 0`, which can silently collapse candidates. This
  index has no analogous defaulting read, so no guard is added; the comment now
  says why.
- Two coverage gaps closed, both mutation-verified: cross-file merge of one
  namespace reopened in two files (no existing test covered it — confirmed by
  making each file clobber the previous and watching only the new test fail),
  and the same-name inline nest whose dedup, when defeated, flips a resolved
  def to `'ambiguous'`.
- `resolveCppQualifiedNamespaceMember`'s JSDoc now names both production call
  sites, including the callsite-less `resolveAdlCandidates` path.
- Memoized candidate buckets are frozen, so a future in-place sort in
  `overload-narrowing.ts` throws instead of silently corrupting later
  resolutions now that the array is shared across call sites.
- `_scaling_note`'s "measured 0.93-1.21" band did not reproduce; it is now the
  honestly observed 1.28-1.45, with the small arm widened to ~14ms (halves the
  spread) and a triage line saying a scaling failure is a timing signal to
  re-run, unlike the deterministic fingerprint arm.

Verification: 840,000 differential probes against the pre-#2788 walker
extracted from base, 0 mismatches, plus 48,000 candidate-order comparisons,
0 mismatches. cpp resolver integration suite 334/334. Unit suite 10/10.
tsc, eslint, prettier clean. Bench --check PASS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bench): escape the NUL separator so measure.mjs stays a text file

The fingerprint key separator was written as a literal NUL byte instead of the
`\u0000` escape. Git classifies any file containing a NUL as binary, so the
whole bench showed as `Bin` with no diff on GitHub and could not be reviewed —
the same defect this branch already fixed once before the tri-review.

Escaping it is byte-for-byte equivalent at runtime (both produce U+0000), so
the committed fingerprint is unchanged and `--check` still passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(cpp): quality cleanups from a four-angle review of the #2788 series

Reuse, simplification, efficiency and altitude passes over the diff. No
resolution behaviour changes: the bench fingerprint is unchanged and the C++
resolver integration suite still passes in full.

Efficiency

- The `Object.freeze` added in the review-fix commit to close an aliasing
  residual costs 4.6x on the narrowing path — V8 moves frozen arrays to
  PACKED_FROZEN_ELEMENTS, off the fast path for the `.filter`/`.map`/`.some`
  runs `narrowOverloadCandidates` does at every multi-candidate call site.
  The hazard it guards is already a compile error (both the memo and the
  parameter are `readonly SymbolDefinition[]`), so it is now a dev-only
  tripwire. Gated on `isSemanticModelValidatorEnabled()` — the repo's opt-IN
  form used by `phase.ts` and `validate-bindings-immutability.ts` — not
  `adl.ts`'s opt-out `NODE_ENV !== 'production'`, which would keep paying the
  cost in CLI runs where `NODE_ENV` is unset. Large bench arm 100.3 -> 70.6ms.
- The index build scanned every scope in every file twice; pass 1 now collects
  `[scope, node]` pairs for pass 2 to iterate. 800k scopes 14.40 -> 8.21ms.
- `simpleNameOf` uses `lastIndexOf('.')` + `slice` instead of
  `split('.').pop()` (83 -> 24 ns/call), semantics verified byte-equivalent
  over 12 edge cases including `undefined`, `''`, `'a.'`, `'.b'` and `'a..b'`.
- The per-call `hookCtx` literal is hoisted to a module const.
- The bench's fingerprint pass resolved 960k call sites to produce 5,800
  distinct outcomes; it now dedups on the key it already builds. Bench wall
  time 2.4 -> 1.9s, fingerprint byte-identical.

Reuse

- `bucketOwnMembers` used an inlined `Function | Method | Constructor` compare;
  it now calls the canonical `isOverloadableCallable`. `graph-bridge/ids.ts`
  already carries a note that inlining this list recreated twin-list drift once.

Altitude

- `adl.ts` had the same retention defect this series just fixed next door: a
  module-level `let adlIndex` + `let adlIndexSource` strongly pinning the whole
  `parsedFiles` array until the next C++ pass, which in a single analyze never
  comes. Converted to the same `WeakMap` shape. Measured with `--expose-gc`:
  89.11MB retained after the caller drops the array before, 0.17MB after. Six
  file-local helpers now take the index as a parameter; no exported signature
  changed.
- The index's freshness depended on `clearCppInlineNamespaces()` being called
  from another file, guarded only by a warning paragraph. An epoch bumped in
  both `populateCppInlineNamespaceScopes` and the clear is now stored with the
  memo, so a missed clear degrades to a rebuild instead of a stale answer —
  confirmed by driving inline state mid-pass without the clear. Roughly line
  neutral, since it replaces most of the paragraph.
- `test/integration/cpp-adl-benchmark.test.ts` is wired into the
  `GITNEXUS_BENCH` step. It is `skipIf`-gated and was absent from that step's
  explicit file list, so #1990's ADL emit-scaling guard had never executed in
  CI. It passes; ~50s added to a 25-minute job. `cpp-pipeline-benchmark.test.ts`
  is deliberately NOT wired: it costs 115s for guards covering generic
  per-language pipeline scaling that nothing here touches.

Simplification

- Deleted a comment referencing a `inlineChildrenByParent` map that only ever
  existed inside this branch's own first commit, so "the legacy map" pointed a
  reader at code that never shipped.
- Dropped three unreachable `undefined` guards (`strict: false`, no
  `noUncheckedIndexedAccess`), keeping the load-bearing `visited` check.
- Compressed the `rootsByReceiver` doc from 17 lines to 7 — it was the longest
  comment in the file and guarded the least consequential property — the
  `validateAdlSeqCoverage` paragraph from 7 lines to 3, and turned three
  restatements of the uncaught-throw and dedup arguments into pointers.
- Test fixtures: dropped the dead `'Module'` union arm, added a one-line `ns()`
  builder for the nine hand-written scope literals, and moved the file to
  `test/unit/scope-resolution/cpp/` where every other C++ scope-resolution unit
  test lives. 403 -> 337 lines, same 10 tests, and the cross-file mutation check
  still fails exactly one test.

Not done, and why: merging this index into `AdlCandidateIndex` (they key on
different names with different inline-transparency depth — a refactor with
correctness risk, not a cleanup); `ScopeTree.getChildren` (trades in-memory
`parsed.scopes` for store hits on a hot path); a shared `cpp/` util for the five
pre-existing `simpleName` copies; sharing fixtures across the bench/test
boundary (no precedent in this repo); and an exact-count arm for the bench,
which is a gate redesign worth its own change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 15:52:33 +01:00
Gergő Magyar
911151e230
fix(resolution): resolve Go pointer-receiver calls, and report the program boundary instead of hedging (#2766) (#2782)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
2026-08-01 22:42:18 +01:00
Karl Lehenbauer
639eb04b31
fix(swift): preprocess indented conditional directives so class bodies survive parsing (#2771)
* fix(swift): preprocess indented conditional directives so class bodies survive parsing

* fix(swift): make conditional-directive blanking comment-, string- and brace-aware (#2771)

Addresses the review findings on PR #2771. The transform fired
unconditionally, which turned valid Swift into parse errors while missing the
most common shape it was written for.

- The blank/keep decision now consults `blockCommentDepth`, so `  #endif */` —
  the result of commenting out a conditional block — keeps its comment
  terminator. Previously `hasError` went raw=false -> preprocessed=true and the
  rest of the file was swallowed.
- The decision keys on the scanner's brace depth instead of indentation. A
  column-0 `#if` inside a class body is blanked (6 of 7 body shapes previously
  still lost the enclosing declaration) and an indented file-scope directive is
  not — matching what the doc comment already claimed. Bare-CR line endings,
  NBSP/ideographic indentation and a leading BOM are recognized too.
- A group is blanked only when every branch is brace-balanced. An `#if`/`#else`
  that splits a declaration header leaves one unmatched `{` once both branches
  survive, which collapsed five top-level nodes into one and gave unrelated
  types fabricated `NetworkClient.` qualified names. Such a group now degrades
  to the pre-fix behavior.
- Multiline strings honour `\"""` escapes, and a plain `"""` closes even when a
  `#` follows it, so the scanner no longer wedges in string state and silently
  stops blanking for the rest of the file.
- The pound run is counted once per position and skipped. It was quadratic:
  10.6s for one 64k-`#` line, well inside the 512 KB walker limit.
- Extended regex literals (`#/.../#`) no longer open a phantom block comment.
- Directive-free files return early, matching `stripUeMacros`.

Worker parity: `emitSwiftScopeCaptures` and `emitCppScopeCaptures` re-apply
their provider's `preprocessSource` on the parse-cache-miss path — Dart already
did this — and the embedding parse in `ensureAndParse` applies the hook as
well. Before this the worker and the scope-capture/embedding halves analyzed
different programs, turning a consistent degradation into cold-run/warm-run
non-determinism. A new parity test pins the equivalence for every provider that
defines the hook.

SCHEMA_BUMP 37 -> 38: this changes parse semantics, the chunk key hashes raw
on-disk bytes, and `preprocessSource` runs after the key is computed — so a
same-package-version warm cache would replay pre-fix Swift results verbatim,
including across `--force`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(ingestion): apply preprocessSource once in the scope bridge (#2771)

Follow-up cleanup on the review fixes. The previous commit re-applied each
provider's `preprocessSource` inside `emitSwiftScopeCaptures` and
`emitCppScopeCaptures`, mirroring what Dart already did — three copies of the
same rule, and a contract that asked every future emitter to remember it.

`extractParsedFile` is the single funnel every `emitScopeCaptures` caller
passes through (parse worker, scope-resolution run, Vue script extraction), and
it already receives the provider. Applying the hook there on the cache-miss
path covers all three languages and every future one, names no language in
shared code, and drops Dart's unconditional transform on the cache-hit path.
Verified the three emitters use `sourceText` for nothing but the parse, so the
substitution is output-identical — which the parity test asserts directly.

Also from the cleanup pass:

- the parity test derives its language list from the provider registry, so a
  new provider adopting the hook fails until it adds a fixture
- `ensureAndParse` resolves the provider from the language it already computed,
  instead of a second extension table (`getProviderForFile`)
- the preprocessor returns `sourceText` unchanged when no group was blanked,
  which is the common case for files whose only directives are top-level
- `split(/(\r\n|\n|\r)/)` replaces the hand-rolled line splitter, and the
  per-group brace bookkeeping is two scalars instead of an array
- the hint regex is derived from the line regex so the two cannot drift
- unit assertions compare the WHOLE preprocessed file against the expected
  blanking, replacing per-line spot checks; the pipeline tests share one
  `runFixture` helper and `getNodesForFile` in the resolver test helpers
- `LanguageProvider.preprocessSource` documents the real call sites and says
  plainly that the set is not closed — `populateRangeBindings` still hands
  language helpers raw text

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-01 20:10:58 +00:00
azizur100389
d268f351d3
fix(group): preserve manifest-only impact crossings (#2784)
* fix(group): preserve manifest-only impact crossings

Keep proven manifest cross-repo hits when the far endpoint has no concrete graph symbol, avoiding a guaranteed failed UID fan-out.

* fix(group): verify manifest-only neighbor repos

Keep manifest-only crossings from bypassing neighbor repository resolution so unavailable repos still surface as truncated fan-out.

* fix(group): distinguish boundary-only impact crossings

Keep manifest-only boundaries visible without treating unattempted fan-out as completed impact or escalating risk, and cover service scope, deduplication, and real bridge persistence.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-01 20:24:06 +01:00
dependabot[bot]
a6aae8142d
chore(deps)(deps): bump brace-expansion from 5.0.7 to 5.0.9 in /gitnexus (#2786)
Bumps [brace-expansion](https://github.com/juliangruber/brace-expansion) from 5.0.7 to 5.0.9.
- [Release notes](https://github.com/juliangruber/brace-expansion/releases)
- [Commits](https://github.com/juliangruber/brace-expansion/compare/v5.0.7...v5.0.9)

---
updated-dependencies:
- dependency-name: brace-expansion
  dependency-version: 5.0.9
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-01 18:00:21 +00:00
ivangegovdve-sudo
565287528d
fix(ci): stop CI Report dying silently when the tests job fails (#2728)
* fix(ci): stop CI Report dying silently when the tests job fails

The "Build report" step in ci-report.yml runs under
`bash --noprofile --norc -e -o pipefail`. It located its inputs with

    UNIT_SUMMARY=$(find "$DIR/test-reports" -name ... 2>/dev/null | head -1)

`coverage-merge` in ci-tests.yml is `needs: tests` with no `if: always()`,
so any failing shard skips it and the `test-reports` artifact is never
uploaded. `find` then runs against a directory that does not exist and
exits 1; `-o pipefail` carries that status through `| head -1`, the
command substitution hands it to the assignment, and `-e` kills the step.

The death is invisible: `2>/dev/null` discards find's error and the whole
report is built into `$GITHUB_OUTPUT`, so the step logs nothing and just
reports "Process completed with exit code 1". "Comment on PR" is then
skipped, so the CI Report workflow fails and posts nothing on exactly the
PRs whose tests failed — when the report is most useful. The
"Coverage data unavailable" fallback already existed for this case but
was unreachable, because the script died ~160 lines before it.

Route the four lookups through a `find_first` helper that returns empty
when the root is absent. Verified by extracting the step body and running
it against both artifact layouts: with `test-reports` present the output
is byte-identical to the previous script (1335 bytes), and with it absent
the step now exits 0 and emits the coverage-unavailable report instead of
exiting 1 with an empty $GITHUB_OUTPUT.

Observed on 32 of the last 100 failed runs; correlation with the tests
job's conclusion was 6/6 failure and 4/4 success in the sampled runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): let the prebuild assertion report a missing .node

`Build prebuild` deletes `$pkgdir/prebuilds` before running prebuildify,
so a run that emits nothing without failing leaves `find` searching a path
that no longer exists.  Under the step's `shell: bash` (`-e -o pipefail`)
that `find` exits 1 and kills the step before the `test -n "$out"` guard
below it — the guard written to explain exactly this case never runs, and
the job dies with a bare "Process completed with exit code 1".

Same shape as the `ci-report.yml` fix in this PR: a lookup that exits
non-zero on an absent root pre-empts the fallback beneath it.  `|| true`
hands the empty result to the guard, which still fails the build, now with
`::error::prebuildify produced no .node`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-08-01 18:31:49 +01:00
MyShining
1147646518
feat(spring): model AOP transactions, caching, and security (#2783)
* feat(spring): model AOP advice and proxy behavior

* fix(spring): address AOP review findings

---------

Co-authored-by: Shining <xuenning@qiyi.com>
2026-08-01 17:22:12 +01:00
ChunxueLi
99291891b7
feat: make MAX_CALLABLE_VALUE_TARGETS configurable via env (#2725)
* feat(scope-resolution): make MAX_CALLABLE_VALUE_TARGETS configurable via env

The branch's original commit was a whole-file snapshot taken at a stale base
and never touched the constant, so the env read was missing and the branch's
own test failed. Implemented here, matching the sibling
GITNEXUS_MAX_PROPERTY_DISPATCH_FANOUT knob.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(callable-value-flow): add env override tests

* docs(callable-value-flow): document GITNEXUS_MAX_CALLABLE_VALUE_TARGETS env

Adds a Troubleshooting subsection to README.md and a commented entry to
gitnexus/.env.example for the new per-callable-site dispatch-target cap
(default 32), following the maintainer's review request to document the
knob alongside its implementation.

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ubuntu <ubuntu@localhost.localdomain>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-08-01 12:42:36 +01:00
Karl Lehenbauer
bebb1d2367
fix(schema): declare Swift member-containment pairs in CONTAINS DDL (#2769)
* fix(schema): declare Swift member-containment pairs in CONTAINS DDL

* fix(schema): declare remaining Rust impl/trait and JS/TS object-literal HAS_METHOD pairs; guard streamed emit sinks against undeclared pairs (PR #2769 review)

* refactor(schema): share one declared-pairs constant across router and sinks

DECLARED_REL_PAIRS was being computed independently in three places
(csv-generator.ts, graph-emit-sink.ts, pdg-emit-sink.ts) from the same
static RELATION_SCHEMA parse. Export the existing constant from
csv-generator.ts (already imported by both sinks) instead.

assertDeclaredPair now takes the pre-built pairKey rather than the two
labels, since every caller (RelPairRouter.route, both sinks' addRelationship)
needs that same key immediately after for its own Map/stream lookup on the
per-streamed-edge hot path — avoids rebuilding the template string twice
per edge.

Also drops two schema.test.ts assertions that duplicated coverage already
in the more narrowly-named regression tests below them, and trims the v32
ladder comment to point at assertDeclaredPair's docstring instead of
re-explaining the same failure mechanism.

* fix(schema): use replaceAll for the pair-arrow error message (CodeQL)

.replace(str, ...) only touches the first match; CodeQL flags that as
incomplete string escaping regardless of the caller's invariant that
pairKey contains exactly one '|'. replaceAll is equivalent here and
silences the alert.

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-08-01 12:35:37 +01:00
ChunxueLi
c238085676
feat: make MAX_PROPERTY_DISPATCH_FANOUT configurable via env (#2726)
* feat(scope-resolution): make MAX_PROPERTY_DISPATCH_FANOUT configurable via env

* test(property-dispatch): add env override tests

* docs(scope-resolution): document GITNEXUS_MAX_PROPERTY_DISPATCH_FANOUT in README and .env.example

Add a dedicated troubleshooting subsection and .env.example entry for the
GITNEXUS_MAX_PROPERTY_DISPATCH_FANOUT environment variable, matching the
format used by the sibling GITNEXUS_MAX_CALLABLE_VALUE_TARGETS knob.

Closes maintainer request: "Could you please document this in the readme
plus the .env.example?"

---------

Co-authored-by: Ubuntu <ubuntu@localhost.localdomain>
2026-08-01 12:18:07 +01:00
Void Freud
0f78179e7f
fix(jvm): enforce proximity-bounded sibling injection (#2732)
* fix(jvm): bound same-package sibling injection

* docs(jvm): document sibling injection cap

* fix(jvm): mark truncated sibling sets incomplete and bound the merge

Review follow-ups on the sibling injection cap (#2732):

- The cap silently produced a third visibility state. Before it, a file was
  either fully visible (package under 500 files) or fully incomplete; once
  `injectedIds.size` hit the cap, real siblings were dropped while
  `isVisibilityIncomplete` still returned `false`. That flag gates wildcard
  attribution in seven Spring passes (`bean-candidates.ts:199` and the java/
  kotlin bean-metadata, conditionals, config-bindings and DI resolvers), so
  201-500-file packages — exactly this cap's population — resolved wildcard
  annotations against a truncated sibling set with no log signal. Truncation
  now marks the file incomplete and analyze warns once with the affected file
  count.

- The cap only bounded `bindingAugmentations`; the two `typeBindings` merges
  below it still absorbed every sibling, so a class excluded from the binding
  set could still steer receiver/variable type inference through
  `scope.typeBindings`. Both halves now use the same bounded sibling set, and
  the merge iterates that set directly rather than filtering a full rescan, so
  the cap bounds the work as well as the result.

- Path segments are split once per bucket instead of on every pairwise
  proximity comparison — that comparison runs O(files²) per package.

- `JvmPackageFact` was re-declared locally instead of imported from
  `package-facts.js`, where the canonical declaration still serves both
  languages' facades and capture side-channels. Nothing kept the copies in
  sync. Restored the import.

- README/.env.example: `GITNEXUS_MAX_INJECTED_SIBLINGS` does not lift the
  fixed 500-file package skip (including at `0`), and truncation disables
  wildcard attribution for the affected files. Both are now stated.

* test(jvm): restore the language-facade coverage and pin the cap's behaviour

The cap rewrite replaced the per-language harness with generic fixtures,
dropping the Java/Kotlin capture-side-channel and facade coverage (package fact
extraction, the 500-file skip, fail-closed on a file that produced no
ParsedFile) and leaving a proximity fixture whose candidates were already in
order — so it could not tell a working sort from plain truncation of the input.

Restores that harness and adds cap-specific cases on top, driven through the
shared JVM factory. The fixture interleaves near and distant siblings, so the
retained set is only reachable by a working proximity sort. Covers: the exact
capped set, truncation marking the file visibility-incomplete, type bindings
bounded by the same sibling set, the unbounded `0` override staying complete,
and the documented default of 200 applying when the variable is unset.

Each new case fails against the pre-fix implementation.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-08-01 09:38:10 +01:00
Void Freud
7ee0df9e55
fix: serialize global registry transactions across processes (#2716)
* fix: serialize global registry mutations

* fix: serialize global registry mutations

* fix: keep registry reads lock-free

* test(registry): document cross-platform lock coverage

* fix(registry): isolate the registry lock's namespace, timeout and diagnostics

Review follow-ups on the global registry lock (#2716):

- The lock took `getGlobalDir()` itself, which is byte-identical to a repo's
  index slot when that repo is rooted at the user's home directory (a real
  dotfiles layout). `runFullAnalysis` holds the per-repo lock across its whole
  pipeline and `acquireIndexLock` is not reentrant, so `registerRepo` /
  `adoptFlatBranchLabel` self-deadlocked until the wait ceiling and then failed
  the analyze. The registry now locks a private `<globalDir>/registry-lock`
  namespace no index slot can ever resolve to.

- The lock inherited the index lock's 10-minute default timeout, sized for
  multi-minute analyze runs. `gitnexus augment` — documented to cold-start in
  under 500ms and shelled out from editor tool-use hooks — reaches it through
  `listRegisteredRepos({ validate: true })`. Registry transactions are
  sub-second, so they now get their own 5s ceiling.

- Contention was silent: no `log`/`onWaitStart` was wired, and the primitive's
  own texts attribute a wait to "another gitnexus analyze", which misnames a
  registry holder. A registry-specific line is emitted on wait start instead.

- On timeout the transaction now proceeds unlocked with a warning rather than
  throwing. The lost-update race it guards was unguarded before this branch, so
  degrading to the old best-effort behaviour beats failing `analyze`/`list`/
  `index` outright — none of which wrap these calls in a handler — on a wedged
  lock.

- `adoptFlatBranchLabel`'s recursive `fs.rm` no longer runs inside the lock;
  only the closing re-read/mutate/write does, mirroring `clean.ts`, which
  deletes the branch directory before calling the locked `removeBranchIndex`.
  A slow delete no longer blocks every registry operation on the machine.

* test(registry): cover the remaining locked mutators and the colliding layout

Three of the five functions the registry lock wraps had no overlap coverage, so
a future narrowing of the lock would go unnoticed. Adds:

- overlapping `removeBranchIndex` calls on two branches of one entry,
- an overlapping `unregisterRepo` / `registerRepo` pair on distinct repos,
- a registration issued while an index lock is held on the global directory,
  which reproduces the home-rooted self-contention the lock namespace fix
  addresses.

Each fails without the corresponding fix: the two overlap tests lose an update
when `withRegistryLock` is bypassed, and the collision test sees the wait
announcement and the degraded-write warning once the lock namespace is reverted
to `getGlobalDir()`. The collision test asserts on those log records rather
than on elapsed time, so it stays deterministic on a slow runner.

* perf(registry): keep the validation walk out of the registry lock

`listRegisteredRepos({ validate: true })` held the global lock across its
read-only validation walk — an `fs.access` per entry, slow on a network mount
or a large registry — even though the common case prunes nothing and writes
nothing. That is the same lock `gitnexus augment` takes on every editor tool
call, so unrelated registry work serialized behind a walk that never touched
the file.

The walk now runs unlocked; the lock is taken only when an entry is provably
gone, and the prune is applied to a snapshot re-read inside it, so a
registration that lands during the walk is no longer clobbered by a stale
write. Same shape as the `adoptFlatBranchLabel` split.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-08-01 08:47:14 +01:00
Gergő Magyar
064832f50c
fix(mcp): resolve false FTS-missing warnings in the query tool (#2773)
* fix(mcp): surface resolved repo/branch/indexed-at in the FTS-degraded warning

Turns the generic "FTS indexes missing" message into a diagnostic
that reveals what this MCP session actually resolved, so a CLI/MCP
mismatch or stale-connection theory is visible in the warning text
itself instead of requiring a separate debugging round-trip (#2767).

* fix(mcp): stop swallowing real FTS query errors behind the missing-index message

queryFTSViaExecutor previously collapsed 'index genuinely missing' and
'a real query/connection error occurred' into the same silent null,
so a real failure could masquerade as the generic FTS-degraded
message with no diagnostic trail — even when it happened on only
some of the per-table queries while others succeeded. Classifies the
failure (mirroring queryFTS's own check for this exact cypher call),
always logs a non-benign error server-side regardless of overall
outcome, and surfaces it (redacted) in the client warning only when
every table failed (#2767).

* fix(mcp): give --repair-fts a dedicated freshness signal for warm readers

--repair-fts intentionally never restamps indexedAt (it doesn't
regenerate the graph), so a long-lived MCP session's pool staleness
check had no explicit signal that a repair happened, only the
incidental file-identity delta. Reuses the existing (forensic-only)
capabilities.fts.status field: repair-fts now stamps just that
sub-field (everything else byte-identical), and ensureInitialized
compares it as a third, independent reinit trigger alongside the
existing stamp/identity checks, seeded at cold init too so a fresh
process's first warm check doesn't false-trigger (#2767).

* test(mcp): warm session picks up an out-of-band --repair-fts rebuild (#2767)

New end-to-end integration test: a real writable LadybugDB session
builds an index WITHOUT FTS, a real LocalBackend observes 'FTS
indexes missing' through the real pool, a separate writable session
performs the exact repair-fts writes (real createSearchFTSIndexes +
the #2767 capability-only meta stamp), and the SAME still-warm
backend re-queries successfully without a restart — closing the one
end-to-end gap no existing test covered.

Running this against the real engine surfaced a second real message
shape for a missing FTS index ("doesn't have an index with name X",
not just "does not exist") that the U2 classifier didn't recognize —
fixed classifyFtsQueryError to match both, with a regression test
pinning the exact observed string.

* fix(review): address code-review findings on the #2767 FTS fix

- Anchor classifyFtsQueryError to the exception class (mirroring
  isBenignDropFtsIndexError) instead of a bare substring search, so a
  real, differently-classed error that happens to echo the benign
  phrase in its body (e.g. an echoed user query) can't be
  misclassified as a benign missing-index (adversarial review).
- Re-read the on-disk meta immediately before the --repair-fts
  capability stamp write instead of reusing the pre-rebuild snapshot,
  so a concurrent writer (e.g. the HTTP server's background embedding
  checkpoint job) landing mid-repair isn't silently reverted.
- Surface a client-facing partial-result warning (mirroring the
  existing enrichmentDegraded convention) when some FTS tables
  succeed but at least one hits a real error, instead of only logging
  it server-side.
- Update RepoMeta.capabilities' stale 'no programmatic readers'
  docstring now that ensureInitialized reads capabilities.fts.status.
- Widen the warm-session integration test's polling deadline for more
  margin over the production 5s staleness-check throttle.

* fix(ci): drop the cold-init loadMeta call ensureInitialized never needed

It stole the mocked loadMeta call an unrelated upstream PDG test
depends on (test/integration/impact-pdg-statement-precise.test.ts
queues a single mockResolvedValueOnce for its own PDG-config read;
the extra call consumed that slot before the PDG code ran, so it
fell through to the mock's null default and epistemic came back
undefined instead of 'pdg-intra-procedural'). Cold init now leaves
lastObservedFtsStatus unseeded — the cost is at most one redundant
initLbug call on the first warm check, which no-ops via a single
fs.stat when nothing actually changed, not a real reopen.

* fix(review): address tri-review findings on the #2767 FTS fix

Fixes two P1s (misleading repair-fts advice on real query errors;
embedding-checkpoint job silently reverting the capabilities.fts stamp
for up to its 30-minute lifetime), five P2/P3s (stale indexedAt in
warnings, extension-unavailable noise, mismatched log severity, a
table-missing vs index-missing conflation confirmed against a live
LadybugDB, and a reinit-watermark latching bug), and the four residual
items already self-disclosed in this PR's description (shared FTS
error classifier, consolidated per-pool observed-state map, a
redactPaths whitespace gap, and an isolated ftsCapsChanged test).

A /simplify pass afterward caught one more real bug: the
extension-unavailable short-circuit only guarded the MCP pool path,
so the CLI-path fix above it started surfacing spurious non-benign
errors for the same expected degraded state the pool path stays
silent on — now both paths agree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-01 08:42:51 +01:00
azizur100389
84f584449d
fix(python): resolve classes through module imports (#2770) 2026-08-01 06:02:47 +01:00
dependabot[bot]
51095c19f8
chore(deps): bump actions/setup-python from 6.3.0 to 7.0.0 (#2757)
Some checks are pending
Gitleaks / gitleaks (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 6.3.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](ece7cb06ca...5fda3b95a4)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-31 13:05:18 +01:00
azizur100389
454d383416
fix(ingestion): join multi-line closure bindings on initializer startLine (#2735) (#2762)
* fix(ingestion): join multi-line closure bindings on initializer startLine

Graph-node captures sit on the outer binding wrapper while scope-resolution
anchors on the inner callable; the line-only position join missed when those
split across lines and fail-closed dropped the real CALLS edge (#2735).

* fix(ingestion): unwrap Ruby call+block for multi-line lambda joins

Cover Kotlin/Ruby/Dart multi-line closure CALLS in integration tests, and
dig Ruby's call/block field so do-end bindings join on the block start line.

* style(ingestion): format closure join changes

* fix(ingestion): make closure position join language agnostic
2026-07-31 11:58:47 +01:00
dependabot[bot]
d99d828a52
chore(deps)(deps): bump express-rate-limit in /gitnexus (#2764)
Bumps [express-rate-limit](https://github.com/express-rate-limit/express-rate-limit) from 8.6.0 to 8.6.1.
- [Release notes](https://github.com/express-rate-limit/express-rate-limit/releases)
- [Commits](https://github.com/express-rate-limit/express-rate-limit/compare/v8.6.0...v8.6.1)

---
updated-dependencies:
- dependency-name: express-rate-limit
  dependency-version: 8.6.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-31 10:42:37 +01:00
dependabot[bot]
e0dc0c2d5e
chore(deps): bump release-drafter/release-drafter from 7.5.1 to 7.6.0 (#2756)
Bumps [release-drafter/release-drafter](https://github.com/release-drafter/release-drafter) from 7.5.1 to 7.6.0.
- [Release notes](https://github.com/release-drafter/release-drafter/releases)
- [Commits](4d75298e00...eada3c96a6)

---
updated-dependencies:
- dependency-name: release-drafter/release-drafter
  dependency-version: 7.6.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-31 10:34:51 +01:00
dependabot[bot]
909f2f85b6
chore(deps): bump the codeql-action group across 1 directory with 3 updates (#2755)
Bumps the codeql-action group with 3 updates in the / directory: [github/codeql-action/init](https://github.com/github/codeql-action), [github/codeql-action/analyze](https://github.com/github/codeql-action) and [github/codeql-action/upload-sarif](https://github.com/github/codeql-action).


Updates `github/codeql-action/init` from 4.37.0 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](99df26d4f1...e4fba868fa)

Updates `github/codeql-action/analyze` from 4.37.0 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](99df26d4f1...e4fba868fa)

Updates `github/codeql-action/upload-sarif` from 4.37.0 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](99df26d4f1...e4fba868fa)

---
updated-dependencies:
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/init
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-31 10:34:35 +01:00
dependabot[bot]
e8e572fbff
chore(deps)(deps): bump i18next from 26.3.0 to 26.3.6 in /gitnexus-web (#2751)
Bumps [i18next](https://github.com/i18next/i18next) from 26.3.0 to 26.3.6.
- [Release notes](https://github.com/i18next/i18next/releases)
- [Changelog](https://github.com/i18next/i18next/blob/master/CHANGELOG.md)
- [Commits](https://github.com/i18next/i18next/compare/v26.3.0...v26.3.6)

---
updated-dependencies:
- dependency-name: i18next
  dependency-version: 26.3.6
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-31 10:34:16 +01:00
MyShining
de84ad6297
feat(spring): index @Bean factories and @Resource injection (#2740)
* feat(spring): index Bean factories and Resource injection

* fix(spring): address Bean and Resource review findings

* refactor(lbug): keep relation pair parsing in router

* test(lbug): preserve schema exports in WAL mocks

* test(cache): align schema bump pin

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-07-31 10:33:42 +01:00
dependabot[bot]
5ed9617ff4
chore(deps)(deps): bump @langchain/core in /gitnexus-web (#2749)
Bumps [@langchain/core](https://github.com/langchain-ai/langchainjs) from 1.2.2 to 1.2.3.
- [Release notes](https://github.com/langchain-ai/langchainjs/releases)
- [Commits](https://github.com/langchain-ai/langchainjs/compare/@langchain/core@1.2.2...@langchain/core@1.2.3)

---
updated-dependencies:
- dependency-name: "@langchain/core"
  dependency-version: 1.2.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-07-31 08:59:03 +00:00
dependabot[bot]
c1ee62854a
chore(deps)(deps-dev): bump @babel/types in /gitnexus-web (#2750)
Bumps [@babel/types](https://github.com/babel/babel/tree/HEAD/packages/babel-types) from 8.0.0 to 8.0.4.
- [Release notes](https://github.com/babel/babel/releases)
- [Changelog](https://github.com/babel/babel/blob/main/CHANGELOG.md)
- [Commits](https://github.com/babel/babel/commits/v8.0.4/packages/babel-types)

---
updated-dependencies:
- dependency-name: "@babel/types"
  dependency-version: 8.0.4
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-31 08:32:45 +00:00
dependabot[bot]
59ea1ce2c8
chore(deps)(deps-dev): bump @types/node in /gitnexus (#2763)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 26.1.1 to 26.1.2.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 26.1.2
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-31 07:58:11 +01:00
dependabot[bot]
c0f2eb594e
chore(deps)(deps-dev): bump @vitejs/plugin-react in /gitnexus-web (#2753)
Bumps [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) from 6.0.2 to 6.0.4.
- [Release notes](https://github.com/vitejs/vite-plugin-react/releases)
- [Changelog](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md)
- [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.0.4/packages/plugin-react)

---
updated-dependencies:
- dependency-name: "@vitejs/plugin-react"
  dependency-version: 6.0.4
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-31 07:57:50 +01:00
dependabot[bot]
7890798192
chore(deps)(deps-dev): bump @types/node in /gitnexus-web (#2752)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.9.5 to 26.0.1.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 26.0.1
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-31 07:57:31 +01:00
Gergő Magyar
27ab37c432
feat(resolution): type receiver chains from AST structure across all 14 languages (#2708) + epistemic lower-bound (#2744) (#2747) 2026-07-31 07:12:57 +01:00
Gergő Magyar
9c24e3459e
fix(rust): qualify items by their enclosing mod chain so same-named ones stay distinct (#2742) (#2745)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(rust): let the qualified-call filter see inline modules

The negative filter added in #2741 builds its set of known module names from
FILE PATHS, so an inline `mod x { … }` — which appears in no path — was absent
from it. Every module-qualified call into an inline module was therefore
rejected before any candidate channel ran, which is a hole in that optimisation
rather than in the resolution logic it guards.

The per-pass index now unions the file-derived names with inline module names
taken from the scope model: a `mod` declaration binds a `Namespace` def locally
in the declaring scope, and that binding is the only place an inline module's
name exists. Collected in the same walk that already builds the module → scope
map, so it costs no extra pass.

Found while fixing #2742, where a correctly resolved call into `mod inner { … }`
still could not reach its target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(scope-resolution): try the namespace-prefixed node key before the bare one

`resolveDefGraphId` looked up the plain `qualifiedName` key first and only
retried with the `namespacePrefix`-qualified key afterwards. For the defs that
carry a prefix the qualified name is a bare TAIL, so the plain key happily
matched a same-named item at a different namespace depth in the same file and
returned it before the more specific retry was ever reached.

The namespace-prefixed key is strictly the more specific of the two, so it is
now tried first. Where no such node exists the lookup falls through to exactly
the previous order, which keeps the #1982 behaviour this retry was added for.

Without this, a call into `mod inner { fn dispatch }` resolved to the correct
definition and then mapped it onto the crate-root `fn dispatch` node — the
self-loop #2742 describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): qualify items by their enclosing mod chain so same-named ones stay distinct (#2742)

Node identity is `<label>:<file>:<qualifiedName>` and carried no module path, so
an inline `mod inner { fn dispatch }` and a crate-root `fn dispatch` in the same
file collapsed onto `Function:<file>:dispatch`, first-wins. Resolution already
picked the right definition — the target simply was not representable, so a
correct resolution still rendered as a self-loop and `impact` reported the real
callee as unreached.

The mechanism already existed: `qualifyRustImplTargetByModScope` has walked
`mod_item` ancestors for impl targets since #1982. Generalised to
`qualifyByEnclosingModScope` and applied to free items, so
`mod inner { fn dispatch }` becomes `Function:<file>:inner.dispatch`. Keyed
purely on the `mod_item` node type, exactly as the impl qualifier already was,
so it is a no-op for every language whose grammar has no such node.

Two constraints found by tests rather than by reading, both now encoded:

  - The helper normalised `::` to `.` unconditionally. With no enclosing `mod`
    that rewrote a top-level `impl a::Inner` from `a::Inner` to `a.Inner` and
    moved its node id away from the one the HAS_METHOD owner edge emits,
    breaking the #1975 scoped-impl ownership. It now returns raw text untouched
    when there are no mod segments, which also makes the change strictly
    additive for every id that has no enclosing module.

  - Qualification is scoped to items with no enclosing class/impl. A method
    already carries its owner's name, and that owner's id is mod-scoped by the
    impl qualifier, so qualifying the method again breaks the same byte-for-byte
    agreement. Same-named methods on same-named types in sibling modules
    therefore still collapse — a narrower residual than the free-item case fixed
    here, and one belonging to the owner edge rather than to this path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(storage): bump schema versions for the mod-qualified Rust node ids (#2742)

`INCREMENTAL_SCHEMA_VERSION` 24 -> 25 and `SCHEMA_BUMP` 31 -> 32.

Node ids change for every Rust item inside any `mod` block, and
`#[cfg(test)] mod tests` makes that close to every Rust repository. A pre-v25
index therefore holds ids an incremental top-up cannot reconcile — the old nodes
would simply be stranded — so the reuse gate has to force a full re-analyze. The
qualified name is computed in the parse worker, so a warm parse cache would
likewise replay the old unqualified ids and keep the collapse.

This branch originally claimed v24; #2708 took that number and merged first, so
it is renumbered to v25 here. That is exactly the collision the v29 note in
parse-cache.ts warns about, and re-checking against origin/main at rebase time
rather than at branch time is what caught it. #2708 did not touch `SCHEMA_BUMP`,
so 32 is free.

The version-pin test moves with the bump by design, including the new pre-v25
row in the reuse-gate table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): stop mod-qualifying container ids while their owner edges stay bare (#2745 review)

#2742 re-keyed Rust node ids by the enclosing `mod` chain. The mint moved; the
owner-edge anchor did not. `findEnclosingClassInfo` mints a member's owner id
from the container's BARE `nameNode.text` and only follows a qualified shape when
the provider sets `classExtractor.qualifiedNodeId`, which Rust does not.

So every `struct` / `trait` / `enum` / `impl` declared directly inside a `mod`
got a node id that none of its member edges pointed at. Five lines of idiomatic
Rust were enough:

    pub mod engine { pub struct Config { pub retries: usize } }

    NODE     Struct:src/lib.rs:engine.Config
    DANGLING HAS_PROPERTY Struct:src/lib.rs:Config -> Property:src/lib.rs:Config.retries

The rows are discarded by the IGNORE_ERRORS COPY retry, so the struct silently
lost every field. A trait impl inside a `mod` additionally dropped its
METHOD_IMPLEMENTS edge outright.

The same gap put `impl a::Inner` inside a `mod` back on the #1975 rake that
`qualifyByEnclosingModScope`'s own docblock warns about. The impl-target branch
deliberately fires only for an UNSCOPED `type_identifier`; the new gate had no
such restriction and picked up the scoped targets that branch had just excluded,
minting `Impl:<file>:outer.a.Inner` against an anchor still reading
`Impl:<file>🅰️:Inner`.

The member side was already excluded via `!enclosingClassInfo`. This adds the
owner side, gated on `MEMBER_OWNER_NODE_TYPES` — derived from
`CLASS_CONTAINER_TYPES`, which is already the single source of "this node type
owns member edges" and already carries an INVARIANT note binding it to
`CONTAINER_TYPE_TO_LABEL`. A language adding a container therefore cannot gain a
mismatched id shape here without also failing that invariant. Keyed purely on
tree-sitter node types, so no language name enters shared ingestion.

`union_item` is listed too: its fields are captured as Property but it is not a
recognized owner, so they carry no HAS_PROPERTY edge and cannot dangle — it is
here so a union's id keeps the same shape as the struct beside it.

Containers still collapse across sibling modules, exactly as before this fix.
That residual belongs to the owner edge, and is not worked around here.

Regression tests use the UNFILTERED `findDanglingEdges(result)`. Every other
dangling assertion in `rust.test.ts` passes `['HAS_METHOD']`, which is precisely
why the HAS_PROPERTY breakage shipped with a green suite. They assert the NODE
id rather than only the edge's anchor, because the anchor was already bare while
the bug was live — an edge-only assertion passes in both builds. All four fail
when the new gate clause alone is reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* fix(rust): resolve modules nested inside an inline mod (#2745 review)

The #2730 self-loop survived one `mod` deeper:

    pub mod outer {
        pub mod tools { pub fn dispatch() {} }
        pub fn dispatch() { tools::dispatch(); }
    }

    CALLS outer.dispatch -> outer.dispatch          <- the #2730 symptom
    NODE  outer.tools.dispatch                      <- correct target, unlinked

Two gates were blind to a nested inline module, so the hook refused and the shared
lexical tier bound the call to the enclosing same-name `dispatch`:

`knownModuleNames` was collected by walking `moduleScopeByFile`, which maps a file
to its ROOT `Module` scope only. A `mod` nested inside an inline `mod` binds in the
parent module's scope, so the walk saw depth-1 inline modules and missed every
nested one — `tools` never entered the set and the negative filter rejected the
qualifier before any candidate ran.

`declaresSubmodule` had the same root-only assumption, so even with the name known
the candidate `outer::tools` was never yielded.

Both now read the def index. Names come from every `Namespace` def; inline module
PATHS are derived from the members' `namespacePrefix` rather than from the `mod`
defs, because a `mod` def carries no nesting information of its own — inside
`mod outer { mod tools { … } }` the inner def is `qualifiedName: 'tools'` with NO
`namespacePrefix`, while every def within it is stamped `outer.tools`. A
`Namespace` scope also owns its OWN def rather than its children's, so the scope
tree cannot answer this either: the `mod outer` scope lists `outer`, never `tools`.

Restricted to non-empty prefixes, so this stays a DECLARATION check. Including
file-derived modules would let an undeclared or `cfg`-gated file on disk outrank a
real `use` binding — the regression #2741's review already fixed once. File-backed
submodules therefore keep going through the binding check.

A module with no defs at all is absent from the set, which is harmless: it has no
member for a qualified call to resolve to.

Cost is one pass over an already-resident def index, memoized per resolution pass
on the existing WeakMap — the same order of work as the binding walk it replaces,
and it subsumes it. `isLocalNamespaceBinding` was going to single-source the
duplicated "locally declared submodule" predicate the review flagged; deriving
paths from members removed the second copy outright instead.

Regression fixture covers depth 2 and depth 3, so the fix is depth-agnostic rather
than depth-2 special-cased.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* fix(rust): let an imported type outrank a same-named module (#2745 review)

Widening the negative filter to inline `mod` names let a type-qualified call
through whenever a module happened to share the type's name. The
crate-root-relative candidate then captured it:

    // src/lib.rs
    pub mod Buffer { pub fn with_capacity() -> usize { 111 } }
    // src/b.rs
    use crate::c::Buffer;                        // the real target lives in c.rs
    pub fn call() -> usize { Buffer::with_capacity() }

    base: (no CALLS edge — unresolved)
    PR:   CALLS b::call -> Function:src/lib.rs:Buffer.with_capacity   <- fabricated

`ids.ts` states the doctrine this broke: a missing edge is the correct failure
direction for a graph whose consumers include `impact`; a fabricated caller is not.
The base produced the missing edge and the PR produced the fabricated one.

That third candidate is the loosest of the three — a guess at a crate-root-relative
path the caller never wrote, kept for 2015-edition style. In Rust 2018 a bare first
segment resolves in the CALLER's module, so a local binding for that segment
settles the question: it is now skipped when the head names anything non-module in
the caller's own module. Candidates 1 and 2 are untouched, and they run first, so
the legitimate `use crate::tools;` path is unaffected.

The binding lookup goes through `lookupBindingsAt`. A first attempt read
`Scope.bindings` directly and the guard never fired: a `use` binding is finalize
OUTPUT and absent from the scope's own local table, which is exactly the
imported-type case being guarded. Contract I8 in `contract/scope-resolver.ts`
requires that channel anyway.

The regression test asserts the forbidden TARGET rather than an empty edge set, and
separately asserts the module member still exists as a node — otherwise the test
would pass just as well if the call went unresolved for some unrelated reason, or
if the module node disappeared entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* fix(scope-resolution): try the namespace-prefixed name on the TAGGED keys too (#2745 review)

`resolveDefGraphId` gained a namespace-prefixed retry for the plain qualified key
in this PR, but the five tagged keys above it — template constraints, parameter
types, parameter shape, arity, template arguments — kept composing from the bare
`qualifiedName`.

For a namespace- or `mod`-qualified def those keys are simply dead:
`node-lookup.ts` registers them under the QUALIFIED name (`inner.dispatch#0`)
while this side built `dispatch#0`. The keys exist to separate overloads, so a
mod-scoped overload set was relying on whichever later key happened to catch it.

Verified as a miss rather than a mis-hit before changing anything — an end-to-end
run with a crate-root decoy of the same name and arity binds correctly — so this
is hygiene, not a live bug. Worth doing while the code is open rather than leaving
five keys dead and the behaviour dependent on fallback order.

Both name forms now go through one `lookupTagged` helper, most specific first, so
a sixth tagged key cannot be added with the bare form only. That also removes the
five hand-repeated `qualifiedKey(...)` / `nodeLookup.get(...)` pairs.

Also pins the C++ `EXTENDS` retarget this PR's reorder produces.
`cpp-two-phase-dependent-base-cross-ns-deep` declares a global `Inner` decoy
alongside `ns:🅰️🅱️:Inner`; the base's `qualifiedName` is a bare `Inner` with the
path on `namespacePrefix`, so only the prefixed key separates them, and only if it
runs first. The improvement was riding unasserted in a Rust-scoped PR.

The captures golden covers every `rust-*` fixture, so the three fixtures added by
this review series drift it; regenerated with UPDATE_GOLDEN=1.

Verified: 785 tests across cpp / csharp / rust resolvers and the
callable-id-lockstep unit test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* fix(rust): keep a mod declared inside a fn from hoisting above the callable (#2745 review)

`fn wrapper() { mod helper { fn dispatch } }` minted
`Function:<file>:helper.wrapper.dispatch@2:8` — the mod segment composed OUTSIDE
the enclosing-callable prefix, inverting the real nesting.

Nothing dangled: the `@line:col` suffix already makes a function-local callable's
id unique, which is also why the mod segment adds no identity in this position. The
path simply read as a lie about the source. It is now skipped rather than
reordered — interleaving two qualifier passes to fix the order would be real
machinery for a shape whose ids are already unique.

Also folds in the three documentation and structure findings from the same review:

- The 4-clause gate is extracted to a named `qualifiesByEnclosingModScope`, matching
  the two conditions directly above it in the same function, which were already
  named consts.
- `qualifyByEnclosingModScope`'s docblock documented only the impl-target contract
  even though the generalized name has had a second, looser caller since #2742. It
  now states both, and says which gate belongs to which — that gap is what let the
  #1975 scoped-impl regression through in the first place.
- The "cheap rejection BEFORE any index work" comment was no longer true:
  `passIndexFor` walks the def index on its first call in a pass. Corrected rather
  than left to mislead the next reader into thinking the filter is free. What it
  still buys — skipping the per-site candidate search, the part that scales with
  the workspace — is stated instead.

Verified: 279 tests across the Rust resolver suite and the Rust scope-resolution
unit tests. Captures golden regenerated for the extended fixture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* test(storage): move main's SCHEMA_BUMP pin to 32 for the mod-qualified ids (#2745 review)

`#2736` added a pin asserting `SCHEMA_BUMP === 31` on main, which arrived on this
branch through the merge of main while `0062a5c2` had already bumped the constant
to 32. Neither side conflicted textually — the pin and the constant live in
different files — so the merge was clean and the test failed instead.

That is the pin working as designed: it exists so a bump cannot ride along
unnoticed, and this is the fifth time a SCHEMA_BUMP collision has been caught by a
guard rather than by review. Updated to 32 with the reason recorded inline.

`INCREMENTAL_SCHEMA_VERSION` needs no second bump: 25 was introduced by this
unmerged branch, so no released index carries it, and its own pin in
`call-summary-schema-version.test.ts` is already consistent.

Verified: 119 tests across the parse-cache, schema-version, incremental-orchestration
and the two identity suites that arrived with the merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

* test(bench): rebaseline the Rust capture fingerprint for the three new fixtures (#2745 review)

CI caught what I missed:

    [scope-capture --check] FAIL: rust: capture fingerprint drift
      (got 05acbaca..., expected 90fda086...)   fixture_count 202

`bench/scope-capture` fingerprints the whole `rust-*` fixture corpus, so the three
fixtures added by this review series drift it. I rebaselined the
`rust-captures-golden` snapshot and stopped there — a new fixture is a call site of
BOTH, and updating only one is how this reached CI red.

This is the same class as PR #2743's headline finding, from the other direction: an
id-shape change makes every synthetic corpus a call site, and the author fixed the
unit-test fixture and missed the bench. Here it is a fixture-count change rather
than an id-shape change, and the review that flagged the #2743 lead as "REFUTED,
bench/ has no Rust node-id corpus" was right about node ids and wrong about the
corpus fingerprint. Noted for the next author in the baseline entry itself.

Verified as pure corpus growth rather than a capture-logic shift: removing ONLY the
three new fixture directories and re-running reproduces the prior fingerprint
exactly (196 fixtures, capture_groups_fp 3432), and restoring them gives the new
one (202, 3556). `emitRustScopeCaptures` is untouched by this series. Scaling 1.022
local / 1.057 CI, well inside the 1.5 budget.

`bench/python-scope` globs `python-*` only and is unaffected; no other bench walks
the Rust corpus. `--check` now PASSes for all 15 languages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GvVxWt1ShhEP8CiMj3b6Z

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 17:08:00 +01:00
azizur100389
e723f3c2ee
fix(scope-resolution): parse def coordinates after file paths (#2743)
* fix(scope-resolution): parse def coordinates after file paths

Anchor coordinate parsing to the known file path so coordinate-like path fragments and private symbol names cannot corrupt closure attribution.

* fix(bench): use production definition ids
2026-07-30 07:34:32 +01:00
azizur100389
bee3e82ab2
test(scope-resolution): guard closure identity invariants (#2748) 2026-07-30 05:34:21 +01:00
Gergő Magyar
bc76ba2f25
fix(resolution): type inline constructor receivers in every spelling (#2708) (#2737)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(resolution): resolve constructor-expression receivers (#2708)

`Service(db).do_work()` emitted no CALLS edge, so the caller was missing
from `impact(direction: "upstream")` and `context()` while the two-step
spelling of the same call (`s = Service(db)` then `s.do_work()`) resolved.

The receiver reaches `resolveCompoundReceiverClass` intact — Case 0 in
`receiver-bound-calls` routes it there because the text contains `(`. The
free-call branch then only knew one shape: a function whose return-type
binding names a class. A class has no return-type binding, so `Service`
resolved to nothing and the member call was dropped.

Handle the constructor shape: in languages that construct without a `new`
keyword (Python, Kotlin, Swift, Scala) a free call naming a class IS a
constructor call, so the expression's type is that class. The existing
return-type path still runs first and wins, keeping this strictly
additive — `new`-keyword languages never reach the new line because their
receiver text keeps the keyword (`new Service(db)`), which matches no
class binding.

Verified on the issue's 4-file repro: `route_inline` now emits
`CALLS → Service.do_work` and `impactedCount` goes 1 → 2.

Note the issue's second ask — degrading `epistemic` to `lower-bound` when
a receiver goes unresolved — is NOT addressed here.
`computeEpistemicBoundary` keys only on the target's own heritage edges
and runs at query time against the index, while unresolved references
live in an in-memory `resolutionOutcomes[]` that is never persisted. That
needs unresolved-receiver counts in the index first, so it is left for a
follow-up.

Tests: new `python-inline-constructor-receiver` fixture plus three
integration cases (inline resolves, two-step still resolves, no
cross-class fan-out). Two of the three fail without the source change.
Full `test/integration/resolvers` suite passes (2928 tests) — the fix is
shared across every language, so no-regression coverage matters more than
the new cases. Python captures golden regenerated: additions only, no
existing digest changed, confirming capture output is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* refactor(resolution): state the construction rule once, cover every spelling (#2708)

The first commit fixed `Service(db).do_work()` by special-casing a bare
class-name callee inside the free-call branch of the compound receiver
resolver. That was the right rule in the wrong place: it covered one
surface syntax out of three, and asserted rather than declared which
languages it applied to.

Probing the same shape across languages showed the bug is wider:

  | spelling               | languages          | dropped before? |
  |------------------------|--------------------|-----------------|
  | `Service(db).m()`      | Python             | yes             |
  | `new Service(db).m()`  | JS/TS, Java, C#    | yes             |
  | `Service.new.m()`      | Ruby               | yes             |
  | both forms             | PHP, Swift, Dart,  | no — already    |
  |                        | Kotlin             | resolved        |

So the rule is stated once — "constructing a class yields an instance of
that class" — and the per-language surface syntax is declared through a
new `ScopeResolver.constructionSyntax` hook, matching how this file
already gates language-varying behaviour (`stripReceiverCastExpressions`,
`hoistTypeBindingsToModule`). Shared pipeline code names no language.

  - `bare: true`      — Python
  - `keyword: 'new'`  — JS/TS, Java, C#
  - `selector: 'new'` — Ruby, including the parenthesis-less `Service.new`
    spelling that reaches the chain walker rather than the call branch

Opt-in is per-language for two reasons. Correctness: `bare` would mistype
`stat(&st).field` in C, where a struct and a function may share a name.
Evidence: PHP, Swift, Dart and Kotlin resolve this shape already, so they
stay unwired instead of carrying a declaration that changes nothing —
each verified by diffing analyzer output between builds with and without
the change, not assumed.

The keyword gate also keeps a bare factory call honest: in a `new`
language, `makeOther(db).doWork()` still resolves through the factory's
return type and is never read as constructing a same-named class.

Tests: TypeScript fixture (inline `new`, a plain `.js` file for the
javascript provider, two-step, and the factory guard) and a Ruby fixture
(`Service.new` with and without an argument list, plus two-step). With
the source change stashed, the inline cases fail and the factory/two-step
cases still pass. The Python cases from the first commit are unchanged.

No Kotlin fixture: its cases passed without the change, so they would
document coverage this commit does not provide.

Full `test/integration/resolvers` + `test/unit/scope-resolution`: 4234
passed, 1 skipped. Ruby captures golden regenerated — additions only, no
existing digest changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(resolution): only treat a construction selector as construction on the class itself (#2708)

The `selector: 'new'` rule fired on any receiver whose type was class-like,
which is true both when the receiver IS the class constant (`Factory.new`) and
when it is a value of that class (`factory.new`). `isClassLike(...)` cannot
tell those apart, so an instance receiver took the construction path too and
skipped the member lookup that should have run.

That replaced a CORRECT edge with a wrong one. Measured against the base build
on a class defining an instance method `new` returning a `Product`:

  factory = Factory.new; factory.new.run
    before this PR:  Product#run   (correct)
    after  this PR:  Factory#run   (wrong)

Track whether resolution currently sits on the class constant or on a value of
that class, and apply the selector rule only to the former. The head of a chain
is a class constant only when it resolved straight to a class binding rather
than through a typeBinding; every hop past it yields a value, so the flag
clears. The `obj.method()` branch derives the same fact from whether `objExpr`
is a bare name resolving to that class.

`Factory.new.run` keeps the behaviour this PR introduced (Factory#run), which
is itself a fix over the base build's Product#run.

KNOWN LIMITATION, now documented on the contract field and asserted by a test
so a future change to it is deliberate: a class-level override
(`def self.new` returning another type) is still read as construction. The
scope model records no staticness per member, so `def new` and `def self.new`
are indistinguishable at this layer; separating them needs the language
provider to record staticness first. An earlier attempt to use
`TypeRef.source` as a proxy was abandoned after tracing showed Ruby records
body-inferred return types as `return-annotation` too, so it does not
discriminate.

Tests: `ruby-construction-selector` fixture pins all three shapes — class
constant, instance receiver, and the documented class-level-override
limitation. Ruby resolver suites: 185 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(resolution): resolve generic construction receivers (#2708)

`new Box<string>().unwrap()` reached the class lookup as `Box<string>`, which
names no class binding, so the member edge was still dropped while the
non-generic spelling resolved. `new Foo<T>()` is ordinary in all three
keyword-wired languages, so the fix covered a materially narrower slice of
real code than intended.

Retry the lookup on the base name via `stripTemplateArguments` — the same
normalization `resolveClassBindingForName` already applies to typed receivers
in the sibling `receiver-bound-calls` pass. The exact-name lookup still runs
first, so a class whose name legitimately contains `<` is unaffected.

Measured on the probe that first showed the gap:

  before: | viaGeneric | Class:src/box.ts:Box |            (construction edge only)
  after:  | viaGeneric | Method:src/box.ts:Box.get#0 |     (member edge resolved)

Tests: `viaGenericCtor` added to the typescript-inline-constructor-receiver
fixture, asserting both the target file and that the resolved id is `Box`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(resolution): resolve construction in the chain-head position (#2708)

`new Service(db).inner.deep()` emitted only the construction edge. The chain
walker seeds its starting class from the head segment, which arrives as
`new Service(db)` and reduces via `stripCallParens` to `new Service` — no
binding and no class of that name, so the walk was never seeded and every
segment after it resolved to nothing.

Seed the head through the same construction rule the call branch already uses.
A constructed value is an instance, so the class-constant flag from the
previous commit correctly stays false — `new Factory().new` does not get the
selector treatment.

The gap was asymmetric across the languages this PR wires: Python's bare form
strips to a plain `Service` and was already seeded, so only the keyword
languages were affected.

Tests: `viaChainHead` added to the typescript-inline-constructor-receiver
fixture. Note the fixture annotates `readonly inner: Inner` explicitly —
with an unannotated initializer the walk stops at the field, which is
field-type inference and a separate concern from head seeding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(resolution): match the construction keyword by token, not by one space (#2708)

The keyword form was matched with `startsWith(`${keyword} `)`, so only a
single space separated `new` from the type. Any other trivia the source used
— a tab, a line break — failed the match and the member-call edge was lost.

Match the keyword as a whole token followed by one or more whitespace
characters instead. `newService()` still fails the match, which is the point:
it is an ordinary call, not a construction, and must keep resolving through
its own return type.

The keyword is escaped before it enters the pattern. It comes from a language
provider rather than from user input, but a keyword containing a regex
metacharacter would otherwise build a silently wrong pattern.

Tests: tab-separated and newline-separated `new` added to the
typescript-inline-constructor-receiver fixture. Note these cases only survive
because `gitnexus/test/fixtures/` is listed in the repo-root `.prettierignore`
— running prettier from inside `gitnexus/` does not pick that file up and
normalizes the tab away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(resolution): resolve qualified construction callees (#2708)

`new ns.Service().doWork()` emitted only the construction edge. The call
branch splits the callee at its last `.` before construction is considered,
so a qualified type name was routed into `obj.method()` resolution as if
`ns` were a receiver and `Service` a member.

A keyword-marked expression is never a member call, so resolve it as
construction before the split. The callee lookup now also handles a dotted
name: an unambiguous `qualifiedNames` match first, then the trailing simple
name, mirroring how receiver resolution elsewhere in this pass degrades.

Measured:

  before: | viaQualified | Class:src/svc.ts:Service |            (construction only)
  after:  | viaQualified | Method:src/svc.ts:Service.doWork#0 |

Bare-form qualified construction (Python `models.User(db).save()`) is NOT
addressed here: that shape currently emits no edges at all, including no
construction edge, so it is a namespace-import resolution gap upstream of
this pass rather than a construction-typing one.

Tests: `viaQualifiedCtor` added to the typescript-inline-constructor-receiver
fixture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* fix(java): drop the unreachable constructionSyntax declaration (#2708)

Java was wired `{ keyword: 'new' }`, and the PR described it as one of the
languages that needed the fix. Measuring both ways shows it never did: Java
resolves `new Svc().doWork()` identically with and without the change,
because `java/captures.ts` (#2564) already rewrites an
`object_creation_expression` receiver to the constructed type's simple name,
so the raw `new Svc()` text never reaches this resolver.

The decisive evidence is generics: Java resolves `new Box<User>().doWork()`,
which the keyword path could not do before the template-argument fix earlier
in this series — the resolution demonstrably comes from the capture rewrite,
not from here.

Removing the declaration rather than leaving it as defensive configuration:
an unreachable per-language opt-in reads as coverage that does not exist, and
the contract now records why Java is excluded so the omission is not mistaken
for an oversight.

Verified after removal: the Java probe still resolves both the inline and
two-step spellings, and the Java resolver suites pass (252 passed, 1 skipped).

An earlier coordinator measurement in this review claimed Java WAS broken on
base; that comparison was invalid (the "without fix" build had not been
rebuilt). Corrected here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* refactor(resolution): state the selector rule once and derive its option type (#2708)

Two follow-ups from review, no behaviour change (643 resolver tests pass
unchanged before and after):

The `Class.new` selector rule was written out twice — in the `obj.method()`
branch and again in the chain walker — against differently named locals,
while the construction helper's own doc comment claimed the rule was stated
in exactly one place. Both sites ask the identical question, so they now call
one `isConstructionSelectorHop` predicate, and the doc comment says what is
actually true.

`ResolveCompoundReceiverOptions.constructionSyntax` re-declared the contract's
object shape by hand. It was the file's first object-shaped duplicate, and
because the value arrives as a non-literal variable, TypeScript's excess
property check would not fire: a sub-field added to the contract later would
type-check and then be silently ignored here. It is now derived with
`ScopeResolver['constructionSyntax']`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* test(resolution): cover the C# construction path and pin the wiring inventory (#2708)

Three coverage gaps from review, no behaviour change.

C# had no fixture despite being the only keyword-wired language whose
behaviour genuinely depends on the construction rule — measured absent on base
and present on head. `csharp-inline-constructor-receiver` covers the inline
spelling, the two-step spelling, and a static factory that must keep resolving
through its return type rather than being read as construction.

The TypeScript two-step assertion checked only `toContain('Service')`, and the
same fixture defines `LegacyService` — `'LegacyService'.includes('Service')` is
true, so the assertion could not distinguish the two targets. It now pins
`targetFilePath` the way its sibling assertions already do.

Nothing guarded the deliberate opt-in set, so an accidental wiring of a
language that already resolves the shape, or a silent loss of one that needs
it, would pass the whole suite. `construction-syntax-wiring.test.ts` pins the
inventory in both directions: exactly which languages declare
`constructionSyntax` and with which spelling, and that java/php/swift/dart/
kotlin stay unwired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* chore(storage): bump INCREMENTAL_SCHEMA_VERSION to 23 for the #2708 edge changes

This series changes which CALLS edges are emitted for source whose CONTENT has
not changed — inline constructor receivers that previously emitted nothing now
resolve, and the Ruby selector fix moves one edge back to the member it always
belonged to. That is precisely the class of change the version-history block in
this file requires a bump for, and the reuse gate is a strict equality on the
persisted stamp.

Without it, every existing v22 index passes the gate on the next `analyze` —
or is served by the same-commit "already up to date" fast path — and keeps
returning the pre-fix graph for unchanged files. `impact(direction: "upstream")`
and `context()` would go on omitting the very callers #2708 is about, with no
warning, until something unrelated forced a full re-analyze. The fix would
have shipped without reaching anyone who already had an index.

Precedent is unbroken across the recent resolution PRs: #2723 → v22,
#2699 → v21, #2695 → v20, #2563 → v14, each with its own rationale paragraph.
This adds v23 in the same form.

The pinned assertion in call-summary-schema-version.test.ts moves with it, as
that test documents it is designed to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* chore(bench): re-baseline the fixture-corpus fingerprints for #2708

Both bench harnesses fingerprint an entire fixture corpus by directory prefix
(`bench/python-scope/measure.mjs:38`, `bench/scope-capture/measure.mjs:76`), so
every fixture directory this series adds moves a committed baseline. Neither
script writes the baseline itself — running without `--check` only prints, and
the file is edited deliberately, which is what its own comment asks for.

Regenerated, last in the series so the fixture set was final:

  bench/python-scope/baseline-fingerprint.txt   36e29abc… -> f120df92…
  bench/scope-capture/baselines.json  ruby       070e4e11… -> fea3edf8…
                                      typescript 281e9548… -> cad25be9…
                                      csharp     e05dc274… -> 05a85bae…

CI only ever reported the python drift, because the benchmarks job runs the
python step first and aborts there; the cross-language step never ran. Both
were verified locally after the update:

  [measure --check] PASS (capture fingerprint + scaling)
  [import-target-fingerprint --check] PASS (resolver fingerprint)
  [scope-capture --check] PASS (15 languages)

The `csharp` and `ruby` entries moved because of the fixtures added earlier in
this series, not the original ones — a reminder that this baseline moves with
any fixture addition, not just the one that first triggered it.

Captures goldens regenerated alongside (csharp, ruby); both additive only, no
existing digest changed. The python golden did not move: no `python-*` fixture
was added after its last regeneration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqnsJK3Cnbu3bjdZbzgMMP

* Update tests for passesReuseGate function

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 16:19:53 +01:00
Gergő Magyar
df06529950
fix(rust): resolve module-qualified calls against the module tree (#2730) (#2741)
* fix(rust): resolve module-qualified calls against the module tree (#2730)

A Rust call written with a path (`tools::dispatch(..)`) was captured with only
its tail identifier, making it indistinguishable from a bare `dispatch(..)`.
The scope-chain walk then resolved the bare name lexically and bound it to
whatever `dispatch` was nearest — which, for the common wrapper idiom

    fn dispatch(..) -> ToolOutcome { tools::dispatch(..) }

is the wrapper itself. The graph gained a self-loop, the real cross-module edge
never existed, and `impact` reported the callee as unreached: the issue's
repository showed its central tool dispatcher as `risk: LOW` with 0 affected
processes and both "callers" being `#[cfg(test)]` functions, while still
labelling the result `epistemic: "exact"`.

Resolve paths the way rustc does, over the module tree rather than the
filesystem:

  - `mod_item` now emits `@declaration.namespace`, so a Rust module is a named
    definition rather than an anonymous scope region. This mirrors the existing
    C++ `namespace_definition` capture and lets the shared `tagNamespacePrefixes`
    pass stamp members with their enclosing module path — that pass needed no
    changes to start working for Rust.
  - `module-path.ts` reconstructs the other half of the tree: crate roots are
    directories holding `main.rs`/`lib.rs`, and a file's module path is its
    location below that root. A definition's module is its file's module plus
    any enclosing `mod` blocks.
  - `crate::`, `self::` and `super::` are prefix transforms on the calling
    module, not reasons to stop resolving.
  - The final path segment is looked up as a member of the resolved module,
    including members it only re-exports. A `pub use` creates no binding on the
    re-exporting module's own scope, so re-exports are followed through that
    module's import edges.

Resolution runs ahead of the implicit-`this` and scope-chain tiers, so an
explicit path outranks a lexical shadow, and returns undefined on an unknown
module, a missing member or a tie — leaving the existing chain untouched. The
new `ScopeResolver.resolveQualifiedFreeCall` hook is optional and unset for
every other language, so this is additive.

Fixes the reported case (direct callers 2 -> 3, impacted 2 -> 6, the Agent
module now visible) plus multi-segment paths, `super::` paths and `pub use`
facades, each of which previously produced a wrong edge.

Known limitation, pre-existing and unchanged by this commit: an inline
`mod inner { fn dispatch }` and a crate-root `fn dispatch` in the same file
collapse to one graph node, because node identity is `<file>:<qualifiedName>`
and does not carry the module path. That is a separate defect requiring
module-path-qualified node ids and an incremental-schema migration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* test(rust): rebaseline the scope-capture fingerprint for the module-tree captures

`mod_item` now emits `@declaration.namespace` and scoped call sites carry
`@reference.qualified-name`. Both are additive, so every bench fixture holding a
`mod` block or a `Foo::bar()` call gains capture groups, and the corpus grew by
the three `rust-2730-*` fixtures.

Only the Rust fingerprint moves. The other 14 languages are byte-identical,
which is the intended blast radius for a language-local capture change.
Scaling stays linear at 1.043, well inside the 1.5 budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): carry crate identity in qualified module paths (#2741 review H1)

A module was identified by its path segments below a crate root, so
`crates/alpha/src/tools.rs` and `crates/beta/src/tools.rs` were the same module.
A cargo workspace routinely gives several members the same internal module name
— `util`, `error`, `config`, `types` are near-universal — and that made
qualified resolution do one of two wrong things:

  - where only one member defined the called name, the call bound ACROSS crates;
  - where both defined it, the lookup saw two candidates, refused, and handed the
    site back to the lexical walk that emits the same-name self-loop. The fix for
    #2730 therefore switched itself off in exactly the workspace layouts it was
    written for, and #2730's own reported reproduction repository is multi-crate.

A module is now `{ crateRoot, segments }` and `sameModule` compares both. Rust
has no implicit cross-crate paths — reaching another crate requires naming it —
so two modules in different crates are never the same module. Anchored paths
(`crate::`, `self::`, `super::`) resolve inside the caller's own crate and
inherit its root.

Covered by a two-member workspace fixture where both crates define
`tools::dispatch` behind a same-name wrapper, plus unit tests for the path
arithmetic itself, including the branches no fixture reaches (a file under no
crate root, a `super::` chain walking above the crate root).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): count only module members when resolving a qualified call (#2741 review H3)

Module membership was inferred from the file path alone, so any callable in the
right file counted as a member of the module. A `fn` nested inside another `fn`
has the same `filePath`, the same bare `qualifiedName` and no owner, making it
indistinguishable from a module-level item:

    pub fn dispatch() -> usize { 3 }              // the real member
    pub fn wrapper() -> usize {
        fn dispatch() -> usize { 99 }             // counted as a second member
        dispatch()
    }

Two candidates tie, the lookup refuses, and the call falls back to the lexical
walk that emits the same-name self-loop — so an unrelated local helper anywhere
in a module silently reinstated #2730 for every qualified call into it.

The scope model already draws the line exactly: a module-level item is bound
with `origin: 'local'` in its module's own scope, a function-local item binds in
the enclosing Block, and an `impl`/trait method binds in the Class scope.
Membership is now that binding lookup rather than a path comparison.

Inline-`mod` members bind in their Namespace scope rather than the file's Module
scope, and reaching it would mean walking every child scope — faulting them back
in from disk on the out-of-core path. They keep being identified by the
`namespacePrefix` the shared tagging pass stamps on them, which a file-module
member never carries. The documented residual is a `fn` nested inside a `fn`
inside an inline `mod`, which inherits that prefix; that is strictly smaller than
before and costs a refusal, never a wrong edge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): require a use-binding to name a module, not a type (#2741 review H2)

Import resolution deliberately strips a trailing symbol segment when probing for
a file — "the last segment might be a symbol (function, struct, etc.), not a
module. Strip it and try again" (import-resolvers/rust.ts). So
`use crate::client::ClientBuilder;` also resolves to `client/mod.rs`.

The qualified-call resolver took that at face value and treated the imported
TYPE as the module `client`. Rust impl methods carry a bare `qualifiedName`, so
`ClientBuilder::new()` was then looked up among `client`'s module members and
bound to an unrelated module-level `new` — turning an unresolved site into a
false edge, which the module's own contract calls the worse outcome.

A binding now has to name the module it resolved to. The edge's
`targetExportedName` is the tail of the written path, so comparing it against the
resolved module's own tail separates the cases exactly:

    use crate::tools;                 tail `tools`         module ['tools']    accept
    use crate:🅰️:b as tools;         tail `b`             module ['a','b']    accept
    use crate::tools::{self, Ctx};    tail `tools`         module ['tools']    accept
    use crate::client::ClientBuilder; tail `ClientBuilder` module ['client']   reject

Covered by a fixture where `client/mod.rs` deliberately holds both
`impl ClientBuilder { fn new }` and a module-level `fn new`, so a regression
re-binds to the wrong one, plus a control asserting a genuine `client::new()`
module qualifier still resolves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): give src/bin targets their own crate root (#2741 review)

Cargo auto-discovers a binary target for every `src/bin/<name>.rs`. Each is a
separate crate with its own `crate::` root, and its submodules live under
`src/bin/<name>/`.

Only `main.rs` and `lib.rs` established a crate root, so those entry files were
folded into the surrounding library and given the invented module path
`bin::<name>`. That made `crate::helper()` inside a binary resolve into the
LIBRARY's `helper` — and unlike the other findings in this review, this one
downgraded an edge the lexical walk had previously resolved correctly, so it
made existing output worse rather than merely failing to improve it.

`src/bin/<name>.rs` is now its own crate root (as is the `src/bin/<name>/main.rs`
directory form), so a binary's modules and the library's modules of the same name
are no longer the same module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): only try a submodule candidate the caller actually declares (#2741 review)

The first candidate module was `callerModule ++ qualifier`, yielded before the
`use` channel and never checked against anything. That let file layout outrank a
real import: with `use crate::b;` in `src/a/mod.rs` and an undeclared — or
`cfg`-gated — `src/a/b.rs` present on disk, `b::f()` bound to the sibling file,
where rustc resolves it to `crate::b`.

A `mod` declaration, inline or file-backed, emits a `Namespace` def bound locally
in the declaring scope, so the candidate is now gated on that binding rather than
assumed. When the caller does not declare the submodule the candidate is skipped
and the `use` and crate-root channels still run, so this only removes guesses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): follow only real re-exports, and refuse on an ambiguous one (#2741 review)

Two problems in the re-export channel.

A private `use` was followed as though it re-exported. `use crate::tools::helper;`
makes `helper` visible INSIDE the module; it does not put it on the module's
public surface, so `facade::helper()` does not compile. Only `pub use` does, and
finalize already distinguishes them — `reexport` for `pub use`, `named` for a
private one. The `alias` kind is now accepted alongside `reexport`, because
`pub use x::y as name` is a re-export that was previously ignored entirely.

The lookup also took the first matching edge in file-iteration order, which is
parse-pool order. Two `cfg`-exclusive facades re-exporting the same name are
indistinguishable at this layer, so picking one baked a coin flip into the graph.
It now refuses on a genuine tie, consistent with how member lookup already
behaves.

The pre-existing limitation that only FILE modules are reachable — a `pub use`
inside an inline `mod facade { … }` has no `moduleScopeByFile` entry — is now
stated in the code. Reaching those would mean walking every child scope and
faulting the scope tree back in from disk, which is the cost that index exists to
avoid; a miss falls through to the unchanged chain rather than guessing.

The regression test deliberately makes the re-exported name globally ambiguous.
Without that, the pre-existing unique-global free-call fallback resolves the call
on its own and the assertion passes whatever this channel does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* perf(rust): stop type-qualified calls paying for module resolution (#2741 review)

The capture carrying `rawQualifiedName` matches every `scoped_identifier`
callee, so this hook was reached by `Vec::new()`, `String::from()`,
`Self::method()` and every other type-qualified call — the overwhelming majority
of `::` calls in real Rust, none of which name a module. Each one ran the full
candidate search before returning undefined, and every candidate that missed then
walked all of `workspaceIndex.moduleScopeByFile`. Total cost grew as
`qualified-call-sites x files`; two independent measurements put per-site cost at
0.117 -> 0.428 ms across 301 -> 1201 files, i.e. linear in workspace size.

Two changes:

  - The module index now carries a flat set of every module segment name in the
    workspace, and a qualifier whose head matches none of them is rejected before
    any candidate work. Measured at 0.02 us per rejected call and flat in file
    count (500 -> 8000 files), against a previously linear per-site cost.

  - Module scopes are indexed by module identity once per pass rather than
    rediscovered by scanning every file per candidate. On the out-of-core scope
    index that scan was worse than CPU: `moduleScopeByFile` fetches through
    `scopeTree.getScope`, so a full sweep could fault every module scope back in
    from disk — the pattern `workspace-index.ts` added `exportedCallableByName`
    to avoid. Given the #2649 and #1871 history this mattered before merge.

The captures golden is regenerated for the fixture files added earlier in this
series; `emitRustScopeCaptures` itself is unchanged by this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(storage): bump schema versions so the #2730 fix reaches existing indexes

Neither invalidation constant was bumped, so the fix did not reach the users who
reported the bug.

`INCREMENTAL_SCHEMA_VERSION` 22 -> 23. The incremental write set only covers
CHANGED files, so a top-up against a pre-v23 index keeps the wrong self-loop —
and keeps reporting the callee as unreached — for every unchanged Rust file. The
constant's own doc block states this rule, and the precedent is exact: v11 is the
same file (`rust/query.ts`) gaining a capture that changes CALLS edges, with the
same "force a full re-analyze" contract, and v12 is a second Rust instance.

`SCHEMA_BUMP` 30 -> 31. `@declaration.namespace` and `@reference.qualified-name`
are parse-time captures, so a warm parse cache replays the old capture set
verbatim: `rawQualifiedName` comes back undefined and no Namespace def exists to
hang a module prefix on, turning the entire resolution tier into a no-op on
unchanged files. `PARSE_CACHE_VERSION` folds in the package version, so a tagged
release would have invalidated eventually — but source, dev and CI builds at the
same version would not, and the v29 note already warns that relying on someone
else's bump is how a change ships with no invalidation at all. Re-checked against
origin/main at commit time, as that note instructs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(scope-resolution): let a language opt out of the already-namespaced guard (#2741 review)

`tagNamespacePrefixes` skips a def whose `qualifiedName` already equals, or is
prefixed by, its enclosing namespace path. That is right for C++ and C#, where
the qualified name genuinely carries the namespace.

Rust qualified names never do, so the guard fired on a coincidence: in
`mod a { pub fn a() }` the member's name equals its module's name, the prefix was
skipped, and `moduleOfDef` then reported the member as belonging to the PARENT
module. `crate:🅰️:a()` refused, and the def became indistinguishable from a
crate-root `fn a` for the module matcher.

The guard is now conditional on a `qualifiedNamesCarryNamespace` option that
defaults to the existing behaviour, and Rust opts out. The shared pass stays
language-neutral — the decision lives with the provider that knows what its own
qualified names contain.

C++ and C# resolver suites pass unchanged alongside the Rust ones (600 tests).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* fix(rust): refuse a leading :: path instead of reading it as relative (#2741 review)

A leading `::` anchors at the extern prelude: `::tools::dispatch()` names the
CRATE `tools`, not a module of the current one. The path split filtered the empty
leading segment away, which silently reinterpreted the path as relative and let
it resolve against a local module that happens to share the name.

Extern crates are outside the workspace module tree, so the qualified tier now
refuses and leaves the site to the unchanged chain.

The regression test asserts the tier does not bind into the local `tools` module,
rather than asserting no edge at all: the lexical tier still resolves the bare
tail on its own, and that behaviour is not what this change governs. Asserting an
empty edge list would have been testing a different tier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* refactor(rust): reuse the canonical callable predicate and drop dead re-exports (#2741 review)

`CALLABLE_TYPES` was a local copy of the set behind `isOverloadableCallable` in
`utils/callable-labels.ts`. Two copies of the same set drift: extending the
canonical one with a new callable kind would silently leave qualified calls of
that kind unresolved here, with nothing to catch it. Use the shared predicate.

The trailing `export { moduleOfFile, moduleOfDef }` and
`export type { ScopeResolutionIndexes }` were commented as being "for the
resolver's unit tests". No test imports them: the only importer of this module
anywhere in src or test is `rust/scope-resolver.ts`, which takes just
`resolveRustQualifiedFreeCall`. Both functions are already exported from
`module-path.ts` (where the new unit tests take them from), and
`ScopeResolutionIndexes` is canonically exported from
`model/scope-resolution-indexes.ts`. Removed rather than left as surface that
implies a contract it does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* test(rust): rebaseline the scope-capture fingerprint with the correct prior hash

The rebaseline note added with the original fix cited
`Prior 655aed01…`, which was two rebaselines stale — it predates both #2604 and
#2714. The true pre-PR value on the base commit is `7f1240b3…`. CI could not
catch it: the gate compares the live fingerprint against the stored one and never
reads the prose, so the audit chain these notes exist to provide was broken with
nothing to flag it.

The note now carries the correct prior value, and the fingerprint is regenerated
for the fixtures this review series added. Scaling 1.061, well inside the 1.5
budget; fixture_count 196; the other 14 languages remain byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

* test: move the schema-version pin to 23

`call-summary-schema-version.test.ts` asserts the exact value of
`INCREMENTAL_SCHEMA_VERSION` and enumerates which stamped versions the
incremental reuse gate accepts. It moves with every bump by design — that pin is
what stops an id- or edge-changing commit shipping without invalidation.

Updated for the bump to 23, with the pre-v23 case added to the reuse-gate table:
a v22 index predates Rust module-qualified call resolution, so every unchanged
Rust file would keep the same-name self-loop and keep reporting the real callee
as unreached.

Caught by CI rather than locally, because the earlier sweeps in this series
covered `test/integration/resolvers/` and `test/unit/scope-resolution/` only —
the pin lives outside both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J5nWP4BFeBSjbQ3uqAjtxV

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 13:41:47 +01:00
azizur100389
79ff44dfa9
fix(config): honor parts negation on Windows (#2720)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Skill copy sync / shipped skills drift guard (push) Has been cancelled
Normalize repository-relative paths before applying ignore-package rules so `.gitnexusignore` negation can override hardcoded `parts` exclusions during Windows traversal.

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-28 20:04:34 +01:00
Void Freud
7be6d29ca0
fix: require repo in multi-repo MCP tool schemas (#2717)
* fix: require repo in multi-repo MCP schemas

* style(mcp): fix server test formatting

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(mcp): cover repository schema policy

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-28 20:04:04 +01:00
Abhigyan Patwari
ee9987fdc5
Merge pull request #2719 from voidfreud/fix/configurable-embedding-timeout
fix: allow slower remote embedding responses
2026-07-29 03:46:37 +09:00
Gergő Magyar
89ea233e10
fix(js): index CommonJS exports.foo = function () {} exports (#2723) (#2729)
* fix(js): index CommonJS `exports.foo = function () {}` exports (#2723)

Functions assigned to an `exports` / `module.exports` property were not
indexed at all. On a CommonJS codebase — the dominant pre-ESM Node style
(Express, Firebase Functions) — the graph held every internal helper and
missed the entire public API: `impact({target: 'areVariablesValid'})`
answered `Target not found` for the one symbol whose blast radius mattered.

The gap had two halves, and fixing either alone leaves the feature broken:

1. `tree-sitter-queries.ts` carried `@definition.function` rules for every
   declaration form and every variable-binding closure form, but none for
   `assignment_expression` — so no `Function` node was created.

2. The scope-resolution queries (`languages/{javascript,typescript}/query.ts`)
   likewise had no `@declaration.function` for the shape. Adding only (1)
   moves `impact` from "not found" to "found, zero callers", because call
   resolution reaches a definition through the scope declaration, not
   through the graph node.

Both layers now carry the rule, for `function` / `async function` / arrow /
async arrow / generator right-hand sides, in JavaScript and TypeScript. The
receiver is pinned to `exports` / `module.exports` with `#eq?` predicates:
the general `X.foo = function () {}` shape also covers `Foo.prototype.bar`
and `this.handler`, which are member constructs with their own ownership
questions, and a broader rule would emit ownerless top-level Functions for
them. The declaration binds the bare property name into the module scope,
which is what importers see, so `const { foo } = require('./m')` matches by
name and a namespace `m.foo()` walks the module's defs.

Verified end to end: node emission for every listed form plus TS parity, and
CALLS edges for same-file `exports.foo()`, cross-file namespace `m.foo()`,
and cross-file destructured `require()`. The generator call-resolution case
was confirmed to fail against the pre-fix build before the rule landed.

Fixes #2723

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(js): stop the CJS export rule from shadowing a declared function (#2723)

Review of the previous commit caught a regression it introduced. The CJS
`@declaration.function` rules bind the exported property name into the module
scope — which is the point, since that is what importers resolve against. But
when the file ALSO declares that name lexically:

    function dup(v) { return v; }
    exports.dup = function (v) { return !v; };
    function callIt(v) { return dup(v); }

the module scope ends up holding two declarations named `dup`, the name is
ambiguous, and the resolver drops `callIt -> dup` entirely — an edge that
resolved fine before #2723. Confirmed by rebuilding both states: present at
ff86ccf1e, missing at f302916c. A silently missing caller is worse than the
gap #2723 set out to close; it is the impact-under-reporting class this repo
has been bitten by before.

The emitter now drops the CJS `@declaration.function` in exactly that case.
The lexical declaration already supplies the module-scope name, so importers
still resolve through it and intra-module resolution returns to its pre-#2723
behavior — verified by re-running the probe that found the regression.

Implemented at the established seam: a shared pure helper both capture
emitters import and apply at the existing `@declaration.function` filter,
mirroring `array-callback.ts` (#1876), which solves the same
"drop a spurious declaration emit-side" problem.

The module-scope name set is computed once per file and memoized per program
root in a WeakMap, rather than walked per export — a 1000-export CommonJS
module is precisely the shape #2723 was reported against, and the per-export
walk would be quadratic there. `tree.rootNode` was probed to confirm it
returns a stable object identity, so the memo actually hits; measured scaling
across 250/500/1000/2000 exports is linear.

Only the scope declaration is suppressed. The graph node comes from a
separate query and collapses onto the lexical declaration's node by name, so
no node is lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(js): fold the CJS export RHS forms into one pattern per receiver

Benchmarking the #2723 rules found the only reproducible cost is tree-sitter
query COMPILE, paid once per worker process when the lazy Query singleton is
built. Steady-state per-file emit cost and memory proved to sit below the
measurement noise floor, so there is nothing to win there.

The six CJS scope-query patterns per language (3 right-hand-side forms x 2
receiver forms) collapse to two, folding the RHS forms into an inner leaf
alternation. Measured over 5 runs per build, variance under 1ms:

    query compile   base      before     after
    JavaScript      42.2ms    50.6ms     45.1ms
    TypeScript     116.5ms   136.4ms    123.3ms

That recovers ~65% of the added compile cost in both grammars — about 19ms
per worker process, so ~75ms on a 4-worker analyze — and removes 45 lines of
duplicated query text.

The alternation is deliberately the INNER LEAF form. tree-sitter 0.21.1 has a
known hazard where a top-level `[...]` alternation makes sibling branches
share a single predicate bucket, silently dropping matches with no compile
error (it has bitten this repo twice: #1904, #1912). Here every predicate
sits on a capture OUTSIDE the alternation — `@_cjs.exports` / `@_cjs.module`
are on the left-hand side and bound in every branch — which is the documented
safe shape. Verified rather than assumed: a probe asserts all six receiver x
RHS combinations still bind both `@declaration.function` and
`@declaration.name` in both grammars, and that `exportz.x` / `module.other` /
`Foo.prototype.bar` / `this.handler` / aliased `exports` are still rejected —
26/26 checks, so the predicate bucket is intact.

No behavior change: the graph output on a 600-file corpus is identical
node-for-node and edge-for-edge, and the 142 JS/TS integration tests plus
1299 scope-resolution unit tests are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(js): index prototype and `this` member assignments as Methods (#2723)

Follow-up on the known limitations listed with the CJS export fix. Three of
them close here; the rest are recorded below with what they actually cost.

`Foo.prototype.bar = function () {}` is the dominant pre-ES6 method form — the
same population as the CJS exports this PR started with — and it was equally
invisible: no node at all, so `impact` could not reach a single prototype
method. `this.handler = function () {}` inside a constructor is its sibling,
and the pre-ES6 form of the closure-valued class field #2693 already models as
a Method.

Both now emit a `Method` with an owner edge:

    function Foo() {}
    Foo.prototype.bar = function (v) { return v; };
    // Method:f.js:Foo.bar,  HAS_METHOD Function:f.js:Foo -> Method:f.js:Foo.bar

The label comes from `provider.labelOverride` (Function -> Method) and the
owner from a sibling of `findObjectLiteralBindingInfo` — the helper that
already answers "this Method's owner is named by syntax, not by an enclosing
container" for object-literal methods. No new shared-code seam was invented.

Ownership resolves to what the file actually declares, so the edge points at a
node that exists: `function Foo` gives a `Function` owner, `class Foo` a
`Class` owner, and an owner the file does not declare
(`External.prototype.x = …`) claims NO owner edge rather than one pointing at
a fabricated node. A `this.x = fn` inside a class constructor needs none of
this — parse-worker resolves its owner from the enclosing class first.

Member ids qualify by owner (`Method:f.js:Foo.bar`). Without that, two
constructors in one file that each define `bar` collapse onto a single
`Method:f.js:bar` — the same identity collapse #2699 fixed for function-local
callables. Only the new prototype/`this` path qualifies, so object-literal
method ids are byte-identical to before.

Third fix, the orphan twin: `class Dup {}` plus `exports.Dup = function () {}`
emitted `Class:f:Dup` AND an unreachable `Function:f:Dup`. The scope
declaration for a shadowed CJS export is suppressed (previous commit), so the
node had nothing that could resolve to it; with a `function` of that name the
node collapsed by id anyway, but with a `class` the labels differ so it
lingered. `labelOverride` now returns null for that case and no node is
emitted.

## Still open, with measured cost

- Receiver-typed CALLS to a prototype method (`f.bar()`) do not resolve yet. A
  class method resolves because the class owns a scope the resolver attaches
  members to; a prototype assignment has no such scope, so this needs the
  scope layer to associate members with the constructor's type. Verified as a
  control that `new KlassC().meth()` does resolve, so this is specifically the
  missing half, not a general gap.
- `exports.fwd = lib.imported` still does not forward to the original
  definition — resolution/finalize-layer aliasing, reachable by no query rule.
  Note `exports.localFn = localFn` (declare-then-export, by far the more
  common idiom) ALREADY resolves and needed no work.
- Aliased `const e = exports; e.foo = fn` and module-top-level `this.x = fn`
  remain unindexed. The latter is only an export under CommonJS semantics; in
  ESM top-level `this` is undefined, so it needs a CJS gate rather than being
  applied to every `.js` file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(js): index CJS exports assigned through an alias (#2723)

`const e = exports; e.foo = function () {}` exports `foo` exactly as
`exports.foo = fn` does, but no query can express "an identifier that happens
to alias the exports object" — the receiver is only knowable per file.

So the member-assignment rules now match ANY identifier receiver and the
emitters classify. A file's module-scope aliases (`const e = exports`,
`const m = module.exports`) are collected once per program root and memoized
beside the declared-name set, so the answer costs one top-level pass no matter
how many assignments ask.

The widening is only safe because the pruning is exact, and that is the risk
worth stating plainly: without it every `obj.handler = function () {}` in every
JS/TS file would emit a spurious top-level `Function` named `handler`. Both
layers prune:

  - graph nodes, in `labelOverride`: an assignment-anchored capture that is not
    a recognised shape returns null, so no node is emitted at all;
  - scope declarations, in both capture emitters: a receiver that is not the
    exports object declares nothing at module scope.

Verified on both sides. `obj.notAnExport`, `self.alsoNot` and
`localThing.nope` produce no node and no declaration, while an aliased export
resolves cross-file through both the namespace and destructured `require()`
forms.

## Cost

Re-benchmarked, because this widens a query the previous commit had just
optimized. Query compile over 3 runs: JavaScript 46.5ms, TypeScript 123.8ms —
+1.4ms and +0.5ms against the optimized state, since dropping the `#eq?`
predicates offsets the added patterns. Steady state on the assignment-heavy
corpus (200 files x 30 member assignments, the worst case for a widened
receiver) stays inside the +-4% noise band established earlier. Heap unchanged.

One measurement artifact worth recording so it is not mistaken for a
regression later: the real-repo TS corpus went 93,586 -> 93,748 captures across
these commits. That is corpus drift, not over-matching — the benchmark walks a
sorted file list and takes the first 400, and this work added a new source file
to that tree. The repo's own TypeScript contains zero occurrences of the
`identifier.property = function` shape, so the widened rule contributes nothing
there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(js): treat module-level `this.X = fn` as a CommonJS export (#2723)

In CommonJS, module-level `this` IS `module.exports`, so

    this.handler = function (data) { … };

at the top of a `.js` file exports `handler` exactly as `exports.handler`
does. It previously produced an ownerless `Method` that no importer could
reach.

The CommonJS gate is the whole point of the change, not a detail. Top-level
`this` is `undefined` in ESM, so the same line exports nothing there — treating
it as an export would mis-index every `.mjs`, every `"type": "module"` package,
and every `.ts` that compiles to ESM. Detection is deliberately asymmetric: an
`import`/`export` statement settles the file as ESM immediately, a `require()`
call or an `exports`/`module` reference marks it CommonJS, and a file carrying
NEITHER signal is left alone — silence is not evidence of CommonJS.

`this` nesting follows the receiver rule the scope queries already encode
(#2701): an arrow does not bind `this`, so a top-level arrow's `this` is still
the module's and passes through the walk, while every other function form binds
its own receiver and stops it — that is an instance member, which keeps the
Method-plus-owner treatment from the previous commit.

Verified across all three cases rather than just the happy path: a CJS file
exports both the `function` and arrow forms and they resolve through a
cross-file destructured `require()`; an ESM file's identical line produces no
export; and a file with no module-system signal produces none either.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(js): forward CJS re-exports to the original definition (#2723)

Last of the known limitations. `exports.fwd = lib.imported` assigns an
EXISTING symbol rather than a function literal, so no definition rule reaches
it and importers of `fwd` resolved to nothing:

    const lib = require('./lib');
    exports.forwarded = lib.imported;      // via a namespace binding
    const { second } = require('./lib');
    exports.alsoForwarded = second;        // via a named binding

Both forms are now synthesized as re-export markers in the same post-query
pass that already decomposes `require()`, reusing the decomposer's existing
vocabulary rather than adding a case to it.

The kind is the whole fix, and it was established by measurement, not by
reading. Emitted first as `named-alias` — the shape the destructured
`require()` form uses — the forwarding still did not resolve: an import
binding is PRIVATE to its module, exactly as in ESM, where `import { X }`
does not re-export X. `reexport-alias` (`export { X as Y } from './m'`) is
what a CJS forwarding assignment actually is, and with it the call resolves
through the forwarding module to the original definition.

`exports.foo = localFn`, where the right-hand side is a locally DECLARED
function, is deliberately not handled here: the module scope already binds
`localFn`, importers already resolve through it (verified before writing any
code), and synthesizing a second binding would re-create the ambiguity the
shadow guard exists to prevent.

JavaScript only, matching where CJS `require()` decomposition already lives —
`typescript/captures.ts` has no require pass at all, since a `.ts` file using
CJS forwarding is vanishingly rare next to the cost of a second
implementation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(js): document the CommonJS export surface now that #2723 closed it

The known-limitation note still described `exports.X` as unmodeled and listed
the re-export edge as missing. Both are stale: the note now states which forms
declare a module-scope name, which two cases are deliberate non-cases (a
locally declared value needs no second binding; a name the module also declares
lexically is suppressed rather than made ambiguous), and that member
assignments through a receiver are Methods with an owner edge.

`module.exports = fn` — an anonymous default with no name to bind — remains
the one genuine limitation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(js): index the CommonJS default export `module.exports = fn` (#2723)

The last documented limitation. `module.exports = function () {}` exports the
whole module as a callable, so there is no property to take a name from and
nothing declared it.

Named after the file by `deriveDefaultExportHocName` — the convention this
repo already applies to anonymous default exports — so `index.js` takes its
parent directory. A NAMED function expression keeps its own name instead,
which is more informative than the file.

Two traps, both found by probing rather than reading:

  - The widened member-assignment rule already matches this shape, capturing
    the LEFT property as the name — the literal `exports`. Left alone it
    produced `Function:<file>:exports`, and for a named function expression a
    SECOND node beside the real one. The worker now overrides the captured
    name for this shape, which is why it takes precedence over `nameNode`.
  - `labelOverride` was suppressing the node entirely. That is the widening's
    safety net working as designed — an assignment-anchored capture that is
    not a recognised shape emits nothing — and this was simply a shape it had
    not been taught.

The scope declaration is synthesized in the capture emitter rather than the
query, because a tree-sitter pattern has no access to the file path the
anonymous name derives from. Without it the node would exist with nothing
resolving to it, the half-fixed state this issue already had to correct once.

`exports = fn` is deliberately NOT indexed, and there is a test pinning that:
reassigning the `exports` binding does not export anything in CommonJS, it
only breaks the alias to `module.exports`, so indexing it would invent an
export that does not exist.

## Limit worth knowing

`const m = require('./mod'); m()` resolves only when the local binding name
matches the derived name — a naming coincidence, not a mechanism. Resolving a
renamed binding (`const renamed = require('./mod'); renamed()`) needs the
finalize layer to treat a called namespace binding as the target module's
default export, which is separate work. The node itself is always emitted, so
`impact` / `context` / `rename` reach it either way — which is what #2723
asked for.

Adjacent gap found while measuring, NOT addressed here: ESM
`export default function () {}` (anonymous) is equally unindexed. Same class,
different construct, and widening to it would change behaviour for files this
issue never touched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(js): record module.exports = fn and the two remaining default-export gaps

The note still listed `module.exports = fn` as unmodeled. It is indexed now;
what remains is narrower and worth stating precisely: resolving a CALL through
a RENAMED default-export binding needs finalize-layer work, and anonymous ESM
`export default function () {}` is unindexed for the same underlying reason
but is a different construct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(js): close every finding from the #2729 tri-review

A 15-lane review (Claude swarm + ce personas, Codex gpt-5.6-sol swarm + ce +
adversarial) ran the real pipeline against both this branch and its base and
diffed the graphs. On canonical CommonJS shapes the branch was DELETING call
edges that existed at base and FABRICATING edges present in no source. A
fabricated edge is worse than the gap #2723 set out to close: it hands
`impact` a caller that does not exist.

Almost all of it reduced to two root causes.

**1. The exports receiver was identified by TEXT, with no scope lookup.**
The canonical UMD wrapper takes the exports object as a PARAMETER:

    (function (exports) { exports.publicApi = function () {}; })(this);

A text match called that a module export, invented a symbol, and — because the
invented name then collided with module scope — deleted the factory's real call
edges. The same blindness made `const helper = require('./helper');
exports.helper = fn` resolve an importer into a DIFFERENT module's function.
Receivers (and aliases) are now rejected where a parameter or enclosing local
shadows them.

**2. The shadow guard reached one of four export forms.**
It was wrong in four distinct ways: it never fired for an aliased receiver
(`root` was not forwarded), for module-level `this`, or for the default export
— each dropping a real edge, and the default-export case merging two functions
onto one node so the inner call resolved to itself. And it fired when it should
NOT have, deleting a genuine export whose name merely collided with a
non-callable variable:

    let cache = null;
    exports.cache = function (v) { cache = v; return cache; };

There is now one entry point (`cjsExportedName`) covering direct, alias, `this`
and default forms, comparing against CALLABLE declarations only.

Also fixed:

- Prototype owners bound to variables. `var Foo = function () {}` is the
  dominant pre-ES6 constructor — the population this work targets — and owner
  lookup handled only declarations, so two same-named members collapsed onto
  one unqualified node with no owner edges at all.
- TypeScript parity: the default/re-export declaration synthesis lived only in
  the JavaScript emitter, so a `.ts` file emitted the node with nothing
  declaring it. Extracted to a shared module used by both.
- Module-level `this.X = fn` in ESM or a no-signal file no longer mints an
  ownerless `Method`; `.cjs`/`.cts` and `.mjs`/`.mts` are now positive
  module-system signals where the file path is available.
- The MCP graph-schema resource documented HAS_METHOD as Class-owned only,
  while this work adds Function (constructor) owners.
- Two dead exports removed; an orphaned JSDoc reattached to the function it
  describes.
- Tests: the `exports = fn` negative test passed trivially (no JS/TS query
  matches a bare-identifier LHS at all, so it would pass with every guard
  deleted) — it now carries a positive control in the same fixture. A
  bounds-y `.some(...)` assertion was replaced per DoD.md:82. Six regressions
  added, each confirmed failing against the pre-fix build.

**Schema constants bumped LAST, deliberately.** `INCREMENTAL_SCHEMA_VERSION`
20->21 and parse-cache `SCHEMA_BUMP` 28->29, with the pin and reuse-gate tests
updated. This change alters what is emitted for source whose content has not
changed, so without the bump an existing index keeps serving the pre-fix graph
for every unchanged CommonJS file — breaching DoD.md:61. Bumping it BEFORE the
correctness fixes would have been worse: it would have propagated the fabricated
and deleted edges to every index on upgrade.

One review finding was withdrawn rather than fixed: a claimed O(n^2) memo
failure did not survive verification. Clean production-shaped measurement
(fresh parse per file, no instrumentation) shows linear scaling — 0.355, 0.221,
0.216, 0.212 ms/declaration at N=500/1000/2000/4000. The earlier
"reproduction" was an artifact of replacing `globalThis.WeakMap` to count
misses, which perturbs the identity semantics under test.

Verified: 1469 tests across 89 files, including the full scope-resolution unit
suite, the JS/TS resolver suites, closure-binding labels, const-function-twin
and the pipeline golden.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 19:45:01 +01:00
Gergő Magyar
e172aca4ee
Merge branch 'main' into fix/configurable-embedding-timeout 2026-07-28 18:56:06 +01:00
Gergő Magyar
13c77db4d9
fix(ci): stop the placeholder review, verify citations, repair once (#2733) 2026-07-28 18:51:01 +01:00
Gergő Magyar
0ce7880290
fix(scope-resolution): a closure binding is a call SOURCE in every language, and function-local values carry their own identity (closes #2699) (#2718)
* test(scope-resolution): audit the consumers of file-scoped node ids (#2699 part A)

#2699 item 4 — "audit consumers that assume file-scoped ids" — after #2695/#2714
gave function-local CALLABLES position-bearing ids. Tests and findings only; no
production change. That split is deliberate: `impact` reports
`resolveDefGraphId` at CRITICAL with 23 DIRECT dependents across 7 modules
(every language MRO builder, both Spring attachers, C++ member lookup,
tryEmitEdge, emitReferencesViaLookup, buildGraphTargetIndex, emitFreeCallFallback,
emitReceiverBoundCalls, preEmitInheritanceEdges, emitDetectedInterfaceImplementations,
phpEmitUnresolvedReceiverEdges, emitRubyMixinEdges, emitRustTraitImplEdges,
emitDartHeritageEdges), so changing that key chain is its own change, not a
rider on an audit.

A2 — detect_changes: CONCERN RESOLVED, now pinned. The worry was that an id
containing `@row:col` re-keys whenever a declaration MOVES, making every edit
look like symbol churn. It cannot: `local-backend.ts` maps diff hunks to
symbols by LINE-RANGE OVERLAP (`n.startLine`/`n.endLine`) and merely REPORTS
`n.id`. Node identity never participates in the match. New structural test
asserts the WHERE clause never gains `n.id =` or `n.id IN`, keeps the one
legitimate id-shaped predicate (the `BasicBlock:` prefix exclusion, #2082 U7),
and confirms the id is returned rather than matched. Structural in the same
idiom as `detect-changes-worktree.test.ts`, and labelled as not proving runtime
behaviour.

A1 — ANSWERED, and the answer is that #2699 is NOT fully closed by items 1-3.
The fail-closed guard is gated on `isOverloadableCallable`
(Function | Method | Constructor), so a function-local VALUE never reaches it.
Measured on a fixture: a top-level `const handler` and a function-local
`const handler` still produce ONE node, `Const:v.ts:handler`. That is the
residual half of the issue's original complaint. Pinned as a KNOWN LIMIT with
its reason (widening identity to values re-keys ~14,700 build-time nodes to
change ~800 persisted ones — the decision recorded in `parse-worker.ts`), and
deliberately NOT fixed here.

A3 — id-persisting consumers, classified:
  - detect_changes ................ SAFE (position-keyed; pinned by A2)
  - MCP impact/context/trace ...... SAFE (resolve by name/uid at query time)
  - bench fingerprints ............ SAFE (digest capture shape, not node ids)
  - rust-captures golden .......... SAFE (digests captures, not ids)
  - cfg pipeline-pdg snapshot ..... AT RISK by design — pins exact edge ids, so
    it trips whenever attribution changes. That is the gate working; #2714
    already exercised it.
  - wiki / group-contract links ... NOT id-keyed on locals (locals are never
    cross-file addressable, per the document-scoped contract of item 2).

Verified: tsc clean; 14/14 across the two touched files; `detect_changes`
reports 0 changed symbols (tests only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(php): a closure binding is a call SOURCE, not only a TARGET (#2699 part B, S1)

A call made inside a closure binding was attributed to the ENCLOSING scope, so
the closure was a call TARGET but never a call SOURCE: impact(handler,
direction:"downstream") reported nothing even though the closure calls out.

Root cause, probe-measured rather than inferred. Instrumenting
pickCallerCallableDef (graph-bridge/ids.ts) to log every rejection reason shows
the closure's own scope EXISTS and its range DOES contain the call site, but its
ownedDefs is EMPTY, so the ":94" owned-callable filter drops it and attribution
falls through to the ":97" enclosing-scope fallback.

The reason is one missing query rule. javascript/query.ts pairs the binding name
with the closure via @declaration.function anchored on the INNER arrow node, so
anchor.range equals the @scope.function range and pass2AttachDeclarations
attaches the declaration to the CLOSURE's scope. No other language had that
rule — PHP, Rust, Kotlin, Ruby and Dart all captured named function
declarations only. That single omission is the entire empty-ownedDefs cause.

This ports the rule to PHP with the same anchor discipline (@declaration.function
on the inner anonymous_function / arrow_function, NOT on the
assignment_expression wrapper). PHP needs nothing else: it already declares
(anonymous_function) and (arrow_function) as @scope.function, so the rule alone
completes it.

Measured on a fixture: `$handler = function ($x) { return target($x); }` inside
outer() now emits

  Function:src/a.php:outer.$handler@3:2 -> Function:src/a.php:target

where it previously emitted `outer -> target`.

The pinned test in closure-binding-labels.test.ts asserted the OLD, wrong
behaviour by design ("to catch that asymmetry changing in EITHER direction"), so
it is INVERTED here rather than deleted, per its own instruction. Its block
comment is corrected to record the measured root cause, including that Kotlin
and Ruby will need BOTH this rule AND a relaxed kind gate (their lambda_literal
/ do_block is @scope.block deliberately, #1757), and that Dart has no closure
scope at all.

Verification: closure-binding-labels 50/50; PHP resolver suites 221/221
(php, php-coverage, php-response-shapes). detect_changes {staged}: 1 changed
symbol (PHP_SCOPE_QUERY), 0 affected processes, risk LOW.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(rust): emit a node for a closure binding and make it a call SOURCE (#2699 part B, S3)

Rust was the one exception to #2687's "a closure bound to a name is a Function
node in every language": `let handler = || target(1);` produced NO graph node at
all, so the closure could be neither a call target nor a call source.

Needed BOTH query channels, which is the finding worth recording. Porting only
the scope-resolution rule (as S1 did for PHP) changed nothing measurable here,
because there was no node to attribute anything to:

  - languages/rust/query.ts — closure-binding declaration, @declaration.function
    on the INNER closure_expression so anchor.range aligns with the existing
    (closure_expression) @scope.function. This is what gives the closure's own
    scope a callable in ownedDefs, which is what stops pickCallerCallableDef
    falling through to the enclosing fn.
  - tree-sitter-queries.ts — @definition.function on the OUTER let_declaration.
    This emits the Function NODE that Rust never had.

Note the deliberate anchor asymmetry between the two channels: the graph-node
channel anchors the WRAPPER (matching the existing
(lexical_declaration (variable_declarator ... (arrow_function))) rule), while
the scope-resolution channel anchors the INNER closure (to align with
@scope.function). Getting these backwards silently produces either no node or
an unattributable one, so both sites carry a comment saying so.

Measured on a fixture — `let handler = || target(1);` inside outer():

  Function:src/a.rs:outer                CALLS  Function:src/a.rs:outer.handler@2:4
  Function:src/a.rs:outer.handler@2:4    CALLS  Function:src/a.rs:target

Previously the whole binding was absent and the call read as `outer -> target`.
The rule also covers `move` closures: the closure_expression node spans the
`move` keyword.

Verification: closure-binding-labels 50/50; rust.test.ts 192/192;
rust-coverage, rust-f70, rust-scope all pass; rust-captures-golden passes
UNCHANGED, so no golden regeneration was required. detect_changes {staged}:
2 changed symbols (RUST_SCOPE_QUERY, RUST_QUERIES), 0 affected processes,
risk LOW.

One caveat on the suite runs: this host times out `beforeAll` hooks at the
default 60s under load — rust.test.ts needed --hookTimeout=600000 to complete,
and a concurrent second vitest run starves worker startup entirely (every test
fails at ~5001ms). Both are host artifacts, not signal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(kotlin,ruby): a closure binding is a call SOURCE, via a Block-scope callable boundary (#2699 part B, S2)

Kotlin and Ruby anchor a closure on a Block-kind scope — Kotlin lambda_literal
and Ruby do_block/block are @scope.block DELIBERATELY (#1757 smart casts), so
they must not be re-kinded. pickCallerCallableDef gated its child-scope walk on
kind === 'Function', so a closure there could never become a call SOURCE.

Both halves are required; neither alone changes anything:

1. kotlin/query.ts and ruby/query.ts gain the closure-binding declaration rule,
   with @declaration.function on the INNER lambda_literal / block so its range
   aligns with the @scope.block range (the anchor discipline documented in
   javascript/query.ts). Without this the closure scope owns no callable def.

2. pickCallerCallableDef accepts a Block-kind child as a callable boundary when
   the scope IS that callable's body. Without this the kind gate still rejects.

The alignment test in (2) is the part worth scrutiny. Relaxing the kind gate to
accept ANY Block owning a callable would be a real regression: a nested
`fun foo()` declared inside a block is owned by that block, so a call made at
BLOCK level — outside foo — would be misattributed to foo. Comparing the def's
declaration position against the scope's start position discriminates them: for
a closure the declaration and the scope sit on the SAME node, so the positions
match; for a nested function the block starts at `{` while the def starts at the
declaration, so they do not. Existing Function-kind behaviour is untouched, so
every already-working language is unaffected by construction.

The comparison is base-safe: scope-extractor.ts builds a def id as
`def:<filePath>#<startLine>:<startCol>:<type>:<name>` from the same Range a
scope carries, so both sides share one coordinate base. This is called out in
the helper's docblock because `defStartLine` nearby documents its own output as
1-based, which invites a wrong "fix" (#2377 is exactly this class of hazard).

Ruby's call forms are restricted to lambda/proc by name: an unrestricted
(call block: (block)) would match ANY method call taking a block, so
`mapped = items.map { |i| ... }` would wrongly declare `mapped` a callable.
Verified against the parser: 3 matches (->, lambda, proc), map excluded.
Separate #eq? patterns rather than one #match? alternation, which is a known
hazard on this tree-sitter line.

Measured on fixtures:

  Kotlin  Function:src/A.kt:outer.handler@2:4  CALLS  Function:src/A.kt:target
  Ruby    Function:src/a.rb:outer.handler@4:2  CALLS  Method:src/a.rb:target#1

previously `outer -> target` and `outer#0 -> target#1`.

The pinned Kotlin test asserted the old behaviour by design and is INVERTED, not
deleted. Ruby had NO pinned case, so a new one is added rather than inverted.
The describe title no longer claimed something false ("not yet a call SOURCE"
now holds only for Dart) and was retitled.

Verification: closure-binding-labels 51/51; kotlin.test.ts, kotlin-coverage,
ruby.test.ts, ruby-scope, ruby-namespaced all pass (478 passed / 1 expected
inversion before the test was flipped). impact on pickCallerCallableDef:
CRITICAL, 191 impacted, ONE d=1 (resolveCallerGraphId) — the return contract is
unchanged, so that dependent is unaffected. detect_changes {staged}: 5 changed
symbols, 2 affected processes (both EmitReferencesViaLookup, one of them the new
ScopeIsCallableBody step), risk medium.

Dart remains the last failing language: dart/query.ts declares no
@scope.function at all, and dart/captures.ts synthesizes one only from a
declaration WITH a body node, which an expression-bodied closure lacks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(dart): give a closure binding a scope and a distinct identity (#2699 part B, S4)

Dart was the last language where a closure binding could not be a call SOURCE,
and fixing only that would have made the graph WORSE, not better. This lands
both halves together for that reason.

## The attribution half

A Dart closure had no scope at all. `dart/query.ts` declares no
@scope.function anywhere — Dart's function scopes are SYNTHESIZED in
`dart/captures.ts` from `declNode` + `findFunctionBody(declNode)`, and
`findFunctionBody` looked only at the next named SIBLING for a `function_body`.
A closure literal carries its body as a CHILD (`function_expression_body`), so
it matched nothing and no scope was produced.

`query.ts` gains the closure-binding declaration rule and `findFunctionBody`
understands the child form. Deliberately NO @scope.function is added to the
query: it would collide at identical range with the synthesized one, and
duplicate scope ids make `buildScopeTree` throw, which DROPS THE WHOLE FILE.

## The identity half, and why it is not optional

With attribution alone, two same-named closures in one file both keyed to the
bare `Function:a.dart:handler`. One node then appeared to call BOTH targets —
a CALLS edge present nowhere in the source. That is worse than the missing edge
it replaced, so S4 could not ship without this.

Root cause is not Dart-specific. `enclosingCallablePrefix` derives a SEMANTIC
relation — what encloses this callable — by SYNTACTIC ancestor walk. Dart parses
`int outer() { … }` as `function_signature` followed by `function_body` as
SIBLINGS, so the enclosing callable is never an ancestor of code inside it and
no membership set can fix that; the walk looks in the wrong direction.

This is what SCIP and real compilers avoid by construction. SCIP keeps a local
symbol opaque (`local <id>` — no name, no position, no chain) and models
containment as a SEPARATE `enclosing_symbol` field; its spec says the local/global
choice should follow ACCESSIBILITY, not the ability to name an enclosure. Dart's
own analyzer answers this from `Element.enclosingElement` in the element model,
never from AST ancestry. clang uses `name@offset` for a function-local; Kythe
uses a document-scoped VName plus a `childof` edge. Identity is positional and
opaque; enclosure is a relation.

`findSplitBodyCallableAncestor` is the narrow fix at that seam: a fallback used
ONLY when the ancestor walk finds nothing, recovering the callable from the
body's preceding sibling.

The sibling must be a BARE SIGNATURE, and that restriction is load-bearing —
"any preceding callable sibling" is WRONG and was caught regressing PHP during
this work. In `<?php function target($x) {…} $handler = function ($x) {…};` the
closure is at FILE level, so the ancestor walk correctly finds nothing, the
fallback runs, and an unrestricted version mis-qualified the file-level
`$handler` as `target.$handler`. A preceding sibling is only an ENCLOSING
callable when it cannot hold its own body.

`SPLIT_SIGNATURE_NODE_TYPES` is exactly that set and is DERIVED, not listed:
`LOCAL_SCOPE_BODY_NODE_TYPES` is already `FUNCTION_NODE_TYPES` minus the bare
signature types, so the difference between them IS the split-signature set
(`function_signature`, `method_signature` — verified at runtime). PHP's
`function_definition` carries a body and is in both, so it is excluded. No
language is named in shared code, and any future split-grammar language is
covered for free.

## Verification

Full resolver sweep — the gate that caught #2714's Rust regression — 2926
passed / 1 skipped / 0 failed across 51 files. closure-binding-labels 52/52;
dart.test.ts, dart-coverage, callable-id-lockstep, function-local-identity,
caller-identity-regression all pass (156/156 across 6 files).
impact on `enclosingCallablePrefix`: LOW, 5 impacted, 3 d=1 all inside
parse-worker. detect_changes {staged}: 5 changed symbols, 0 affected processes,
risk LOW.

Three existing Dart expectations FLIPPED rather than being deleted: Dart locals
now carry the same enclosing-callable + position identity every other language
got in #2695, so `local.dart:handler` became `local.dart:caller.handler@1:2`.
A new test pins the actual defect — two same-named closures staying DISTINCT
nodes — because the qualification assertions alone would not fail if the
fabricated edge returned.

One note for future work: an id-shape assertion here carries a call-site suffix
on indirect invocations (`…handler@3:2:5:9`) but not on direct calls. That is
the callable-value-flow pass keying its edge by invocation position, not part of
the node id.

Part B is now complete: PHP (S1), Rust (S3), Kotlin + Ruby (S2), Dart (S4).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(scope-resolution): close every deferred item on #2699 (A1 values, twin-list guard, schema bumps)

Clears the limitations this PR had been carrying rather than leaving them as
follow-ups.

## A1 — function-local VALUES now carry their own identity

This was #2699's ORIGINAL complaint and the one a callable-only gate could never
reach: a top-level `const handler` and a function-local `const handler`
collapsed onto ONE `Const:v.ts:handler`. #2695 restricted position-qualified
identity to Function|Method|Constructor because the collision that produced
wrong CALLS edges was between callables, and widening churned ids for symbols
the pruner mostly deletes. The churn is real and is accepted here deliberately.

Widening needed THREE gates aligned, not one:
  - id-building     — `parse-worker.ts` nestedCallablePrefix
  - resolution      — `ids.ts` position key
  - registration    — `node-lookup.ts` position-key registration

Missing the third would register no position key for values, so every lookup
misses and falls through silently. That is the #2714 failure mode: the caller
attaches to a node that does not exist and the edge is DROPPED, which looks like
"zero dangling edges" from outside. All three now route through ONE predicate,
`isPositionQualifiedLocalLabel`, rather than repeating the label set a third
time.

Only LOCALS move. The prefix comes from `enclosingCallablePrefix`, which returns
undefined when nothing encloses the declaration, so top-level and class-member
ids are untouched — verified by the full resolver sweep, where a leak onto class
members would have broken assertions in every language. `Property` is included
on purpose: a class field stays unqualified because the prefix walk boundaries
on class-likes, while an object-literal property inside a function is genuinely
local and would otherwise keep the old collision.

Measured: `Const:v.ts:handler` + `Const:v.ts:run.handler@3:2`, two distinct
nodes. The KNOWN LIMIT test is FLIPPED per its own former instruction ("this
test should be updated as part of it rather than deleted").

## Schema bumps — required by Part B, not just by A1

INCREMENTAL_SCHEMA_VERSION 20 -> 21, parse-cache SCHEMA_BUMP 27 -> 29.

SCHEMA_BUMP is 29, not 28, and that is the point of re-checking it against
origin/main at MERGE time rather than branch time. This branch cut at 27 and
bumped to 28; #2415 also bumped 27 -> 28 and merged first. The automated
main-merge onto this branch surfaced the collision — leaving it at 28 would have
shipped this whole change with NO parse-cache invalidation, so every warm cache
keeps replaying the pre-fix captures and ids. This is the third instance of that
collision recorded in parse-cache.ts (#2632/#2653 hit it at v21, and
#2653/#2654 hit INCREMENTAL_SCHEMA_VERSION the same way).

Part B already changed emitted node ids AND edges on files that did not
themselves change (Dart locals re-keyed, Rust gained a node it never emitted,
five languages gained closure-source attribution). A v20 index topped up
incrementally keeps serving the old attribution, and a warm parse cache replays
the old captures and ids verbatim. Shipping S1-S4 without these would have let
every existing index silently keep the pre-fix graph.

## Twin-list drift guard — the sixth instance in this family

`IMPLICIT_RECEIVERS` (gitnexus-shared lookup-core.ts) and `THIS_RECEIVERS`
(type-env.ts) spell the same concept in two packages, and nothing enforced
agreement — `$this` was added to the shared list in #2714 only because it was
already in the other. New structural test asserts set equality plus the ONE
deliberate asymmetry (`Me`, Visual Basic spelling, absent from the shared list
because no SupportedLanguages entry uses it) in BOTH directions, so re-adding it
there or dropping it here each fail loudly.

Structural rather than value-imported: both constants are module-private, and
exporting them purely to be testable would widen two public surfaces to satisfy
a test.

## Two false comments corrected

  - `lookup-core.ts` said "see the drift guard noted in #2714", implying a guard
    existed when it was only a deferred follow-up. It exists now, and the
    comment points at it.
  - `callable-id-lockstep.test.ts` claimed its regex "fails if any site
    reconstructs the id". It matches ONE template spelling; a hand-rolled
    concatenation still slips past. Now stated as a tripwire for the known
    shape, not a proof.

## Skill learnings

Four entries appended to eval/workflow_bench/learnings.jsonl from this run: the
v9fs safe-writer failure, backticks silently terminating a query template
literal (hit three times), a module-level TDZ const that passes tsc and then
presents as N file failures with ZERO failing assertions, and concurrent vitest
runs starving worker startup so a whole suite fails at ~5001ms.

## Verification

Full resolver sweep 2926 passed / 1 skipped / 0 failed (51 files) — identical to
pre-A1, which is the evidence that only locals moved. All EIGHT bench gates PASS
with fingerprints UNCHANGED, so no regeneration was needed. function-local-identity,
callable-id-lockstep, receiver-twin-list-drift and closure-binding-labels 71/71.
tsc --noEmit clean.

detect_changes {staged}: 9 changed symbols, 14 affected processes, risk HIGH —
expected, and the reason the sweep above is the gate rather than a targeted list.
Every affected process routes through `resolveDefGraphId`, the key chain Part A
measured at CRITICAL with 23 direct dependents.

Deliberately NOT done: the SCIP end state (opaque `local <id>` plus an explicit
enclosure EDGE instead of containment encoded in the id string). It is a design
direction, not a limitation of this work, and it is INCOMPATIBLE with A1 — A1
widens chain-encoded identity, that removes chain encoding entirely. Bundling
both would re-key every local twice. Written up in the research notes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* test: update two assertions the #2699 changes correctly invalidated

Both failed on CI at da3d8397 and are fixed here. Neither is a behaviour
regression; both pinned values that this PR deliberately changed.

1. this-boundary.test.ts — "a Kotlin lambda still sees the receiver"

The `this.m()` edge still exists and `this` still resolves to the enclosing
receiver, which is the ONLY property this test exists to guard (its own comment
said so: "what matters here is only that the `this.m()` edge still exists at
all"). Only the SOURCE moved, from `run` to the lambda:

  Method:K.kt:K.run#0       -> Method:K.kt:K.run.f@2:16
  Method:K.kt:K.run.f@2:16  -> Method:K.kt:K.m#0

The comment justifying the old expectation is now false and is corrected rather
than left: it said the lambda "is not its own caller anchor" because Kotlin
scopes `lambda_literal` as a BLOCK. Kotlin still scopes it as a block (#1757 is
unchanged) — what changed in S2 is that a Block-kind scope is accepted as a
caller anchor when the scope IS the callable's body.

2. call-summary-schema-version.test.ts — INCREMENTAL_SCHEMA_VERSION pin

Moves 20 -> 21 with the bump, which is the point of pinning it: a change that
alters emitted ids or edges without bumping would otherwise ship silently.

Also adds the missing reuse-gate case. `passesReuseGate(20)` now asserts FALSE —
a v20 index predates closure bindings becoming call SOURCES, the Rust node for
`let f = || …`, the Dart closure scope + enclosing-callable identity, and
position-qualified function-local values. Topping such an index up incrementally
keeps serving the old attribution, including the Dart case where two same-named
closures collapsed onto one node and asserted a CALLS edge present nowhere in
the source.

Why CI found these and local verification did not: the verification set was
`test/integration/resolvers/` plus a hand-picked list, and both failures sat
outside it — one integration test about `this` (which a caller-attribution
change obviously touches) and one unit test pinning the exact constant that was
bumped. Grepping for the changed constant, and for tests asserting closure
attribution, would have found both. All 8 suites that reference the schema
constants were then run: 177/177, no third pin.

Verification: this-boundary + call-summary-schema-version 17/17;
the 8 schema-referencing suites 177/177. detect_changes {staged}: 0 changed
symbols, 0 affected processes, risk LOW (assertion-only edits).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(scope-resolution): resolve every finding from the multi-engine review of #2699 part B

The first cut of Part B shipped four P1 defects. A two-engine review (Claude
swarm + ce personas; Codex gpt-5.6-sol swarm + ce + adversarial) found all four,
three of them because an independent engine disagreed with the authoring one.
Each is fixed here and pinned in test/integration/closure-review-findings.test.ts.

## P1-1 — a multi-line closure binding fabricated a CALLS edge

The worst of the four, because it reintroduced the exact defect class #2699
exists to remove. The two query channels anchor on DIFFERENT nodes by design
(graph-node on the outer wrapper, scope-resolution on the inner closure) and the
bridge joins them on line only. Same line, the join matches. Split across lines:

    $multi =
        function ($x) { return target($x); };

the join missed, `resolveDefGraphId` failed closed, and `resolveCallerGraphId`
then CLIMBED to the parent scope — emitting `outer -> target` although `outer`
calls nothing, while the real `outer.$multi` node sat with zero outgoing edges.

`resolveCallerGraphId` now fails closed at the owning callable instead of
climbing. If we have identified the callable that owns a call site and cannot
name its graph node, crediting an ancestor is not graceful degradation — it
invents a relationship. A missing edge is the correct failure direction for a
graph whose consumers include `impact`.

Getting there took two attempts, worth recording: the first guard keyed on the
def's qualifiedName carrying the `@line:col` local marker, but that suffix is
added by parse-worker for GRAPH NODE ids and scope-resolution defs do not have
it, so the guard never fired. `pickCallerCallableDef` now reports whether the
callable came from a child scope, and the fail-closed applies to the owning
callable either way.

## P1-2 — TS constructor parameter properties were re-keyed as locals

A REGRESSION against the base, not merely an incomplete fix. Admitting
`Property` to the position-qualified set made the enclosing-callable walk reach
the constructor's `method_definition` THROUGH the parameter list — a
LOCAL_SCOPE_BODY hit that lands before any class boundary — so
`constructor(private readonly port: Port)` produced
`Property:svc.ts:Service.constructor.port@2:14` instead of `Service.port`. That
silently empties the slot `impact`, `rename` and FTS address while the class
still asserts HAS_PROPERTY against it, and it is the Angular/NestJS DI idiom.
A real instance exists in this repo at src/core/group/service.ts:304.

parse-worker.ts already computed the correct exemption (`isFunctionLocalProperty`,
lines 2245-2257) two lines above; the new ternary discarded it. Now reused, so
the owner-edge decision and the id decision cannot disagree.

## P1-3 — Dart top-level and `final` closures were never call sources

The rule matched only `initialized_variable_definition`, Dart's FUNCTION-LOCAL
shape. A top-level `var` is `initialized_identifier` and a top-level
`final`/`const` is `static_final_declaration`; the second declarator of
`var f = ..., g = ...` is also `initialized_identifier`. None got a declaration
capture, so `findFunctionBody` never synthesized their scope.
dart/captures.ts ALREADY listed all three in bindingNodeTypes for callable-flow
— the declaration rule simply did not mirror it. It does now.

## P1-4 — Ruby `do ... end` and `Proc.new` closures were uncovered

`do ... end` is the dominant MULTI-LINE Ruby style and produces `(do_block)`;
all three patterns matched `(block)` only. The scope channel already covered
both, so these closures got a Block scope owning nothing and their calls fell
through to the enclosing method. The PR's own Ruby test used the brace form, so
it passed.

Fixing it needed BOTH channels — tree-sitter-queries.ts had no graph-node rule
for the `(call)` forms either, exactly as Rust did. Verified: brace, do/end and
Proc.new are now all sources.

## Also from the review

- The split-signature fallback could fire on VALID TypeScript: a
  `declare namespace` containing a bodyless overload made the next declaration's
  `export_statement` a sibling of a `function_signature`, so `send` became
  `internalHelper.send@2:9`. The fallback now requires the matched node to be
  the signature's BODY (a body holds statements; a declaration wrapper holds
  another signature), which separates the two without naming a grammar.
- `isCallableDef` re-spelled `Function | Method | Constructor` in the same file
  that imports `isOverloadableCallable` and calls it three times — a NEW twin
  list, in the PR whose headline is a twin-list drift guard. It now delegates.
- A partial edit had left a self-contradictory comment in parse-worker.ts
  ("Restricted to CALLABLE labels: the / Applies to VALUES as well as callables").
- eval/workflow_bench/learnings.jsonl carried a "skill": "gitnexus-plan" entry,
  but that skill's SKILL.md:347 states feedback is chat-only and forbids
  appending learnings during a planning task. Dropped; the three gitnexus-work
  entries are sanctioned and stay.
- Ruby's lambda/proc patterns tested the method NAME only, so `MyMod.lambda { }`
  was captured as a closure binding. `!receiver` now constrains them.
- Rust's closure work (S3) had ZERO test coverage anywhere — verified once by a
  throwaway fixture and never pinned. Now covered.

## Verification

Full resolver sweep plus the identity/closure suites: 3017 passed / 1 skipped /
0 failed across 58 files (up from 2997 — the new tests). This is the gate that
mattered for P1-1: failing closed instead of climbing could have silently
deleted real edges in any language, and ~2900 resolver assertions say it did
not. All EIGHT bench fingerprint gates PASS with fingerprints UNCHANGED.
tsc --noEmit clean.

detect_changes {staged}: 14 changed symbols, 18 affected processes, risk
CRITICAL — expected, since P1-1 changes the fallthrough of `resolveCallerGraphId`,
the key chain Part A measured at CRITICAL with 23 direct dependents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:25:19 +01:00
Gergő Magyar
706aa1326b
Merge branch 'main' into fix/configurable-embedding-timeout 2026-07-28 17:18:30 +01:00
Gergő Magyar
b0cacd05ee
fix(ci): stop the review agent rejecting its own graph-backed reviews (#2731)
* fix(ci): stop the review agent rejecting its own graph-backed reviews

The context-evidence gate only counted a `context` call when the call
itself passed `file_path` equal to a changed path. The review skill
teaches plain `context({name})`, so 17 of the 26 review-agent run
failures were complete, graph-backed reviews thrown away after full
model spend, with no log line saying which invariant failed.

Prove the evidence from the result instead: `status=found` plus a
`symbol.filePath` inside the repo-scoped changed-path set. Every other
check stays exactly as it was - strict JSON, orchestrator-only turns,
result ordering, duplicate tool-id rejection - and the `repo` argument
still selects the head or the merge-base path set.

Same failure inventory, smaller classes:

- rejection now logs why (in-scope, out-of-scope, sidechain, unresolved
  and off-path counts plus up to three sanitized paths), and the
  envelope error names the message count and first-message shape
- Glob/Grep leave the tool set: they were enabled through `--tools` but
  never allow-listed, so every lane call was denied and burned turns
- both pinned `npm ci` installs retry three times; one registry
  ECONNRESET killed a whole run
- the prompt matches the new contract and asks for the structured body
  even when the analysis is incomplete

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(skills): mirror the review-skill tool-set change into the shipped copies

The npm package, Claude plugin, and Cursor integration ship byte-identical
copies of .claude/skills/gitnexus-review, and the drift guard compares them.
Dropping Glob/Grep from the lane frontmatter and the SKILL.md sentence only
landed in the canonical tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): stop one junk context result discarding a proven review

Tri-review of this PR found that the previous commit fixed one spurious
rejection and created another. Widening evidence candidacy from "the call
that named a changed path" to "every orchestrator context call" also
widened the *strict-parse* surface: `contextResultProvesChangedPath`
throws rather than returning false, so a single malformed payload
anywhere in the transcript now discarded a review that an earlier call
had already proven. The MCP makes that reachable without any misbehaving
model - `GITNEXUS_MCP_DEFAULT_MAX_TOKENS=12000` truncates any context
payload over ~48 KB mid-JSON and appends a marker - and it also destroyed
docs-only runs that the `no_indexable_changed_symbols` mode exempts.

Reproduced by running the workflow's own embedded script on both trees:
a proving evidence call followed by one truncated exploratory call gave
`failure_code: null` on the base and `invalid_execution_transcript` on
the head; it is `null` again here.

- payload-shape failures are caught and counted (`malformedResults`)
  instead of thrown; transcript-structural invariants (envelope, tool
  shapes, duplicate ids, empty tool_result) still fail closed
- diagnostics gained the reasons they were blind to: errored results,
  results that arrived out of order or via a sidechain, unanswered
  in-scope calls, and malformed payloads. A rejection can no longer
  print an in-scope call with every reason at zero
- a deletion-only PR no longer registers head-scoped candidates that can
  never be satisfied: an empty eligible set is out of scope, not a result
  "outside the changed paths"
- the mandatory-body prompt clause now pairs with a required `complete`
  boolean. An incomplete analysis publishes its partial body labelled
  `incomplete_analysis` instead of passing as an accepted review
- `Agent(a,b,c)` is split into six separate `Agent(x)` rules: the pinned
  base action parses allowedTools with `.flatMap((v) => v.split(","))`
  (parse-sdk-options.ts at 3553f843), which shattered the grouped rule
  into `Agent(ci-correctness-lens`, four bare names, and
  `ci-critic-lens)` before the SDK saw it. Pre-existing and unproven at
  runtime, but the split form is correct under either reading and lets
  the header's dispatch canary actually prove something

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): require a line range for context evidence

The tri-review's adversarial lane executed `context({name: 'AGENTS.md'})`
and had the result accepted: the gate checked only that the resolved
filePath was in the changed set, so a bare File node passed for a review
of that file's contents. The trusted prescan already defines an indexable
symbol as one with startLine and endLine, so require the same here.

Pre-existing rather than introduced by this branch, but it is the same
"what counts as proof" surface the rest of this PR tightens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): close the remaining tri-review findings

Addresses every finding the tri-review left open after 0432214d and
1d9f2d75, across both engines.

Reliability and maintainability (Codex ce, ce-reliability, ce-maintainability):
- both pinned `npm ci` installs now call one shared
  `.github/scripts/npm-ci-retry.sh` instead of two near-identical 12-line
  blocks that differed only in a label
- each attempt runs under `timeout` (default 600s, overridable), so a slow
  registry can no longer crowd the model review out of the job's budget
- the helper distinguishes a timeout kill (124) from an npm rejection in
  its log

Test coverage (ce-testing, Codex swarm P3, ce-security, swarm test-ci):
- the retry helper is now exercised behaviourally with a stub npm: first-try
  success runs once, two failures recover on the third, three failures exit 1
- a non-string `symbol.filePath` is a clean reject, not a type error
- an adversarial resolved path (ESC, newline, `::set-output`, RTL override)
  is proven sanitized before it reaches the job log
- the envelope error's shape string is asserted
- an in-scope call whose result never arrives is counted, not silent
- install flags that keep the runtime inert (`--ignore-scripts`, `--prefix`,
  the lock-bound registry) are asserted against the helper they moved into

Correctness and clarity (risk-architect, ce-standards):
- the prompt now tells the model to prefer the uid form or pass file_path
  when a bare name could resolve into an unchanged file, which was the
  narrower off-path failure mode the gate rewrite left behind
- `contextResultProvesChangedPath` -> `contextResultProvesEligiblePath`,
  matching the set-membership contract its sibling was renamed for
- the transcript fixture's default no longer carries a `file_path` the gate
  ignores, which implied the opposite of the contract
- SKILL.md says "file reads" rather than naming a CLI-specific tool, per
  the CLI-neutrality rule in AGENTS.md; mirrored to all three shipped copies
- the interactive-swarm README notes the CI lanes are narrower

Publisher (ce-reliability residual, pre-existing):
- the publish job no longer gates the whole job on authorization, so a
  request rejected at normalization no longer strands the "review in
  progress" marker on the PR forever. Publication stays authorization-gated
  at the step; only the marker cleanup is unconditional.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): ship the install helper executable

The extracted helper was committed 100644, so the workflow's direct
invocation would have failed on the runner with permission denied - a
break introduced by the extraction itself, invisible to every existing
assertion. Set the mode and pin it with a test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:13:29 +01:00
Void Freud
e20b41fffc fix: allow slower remote embedding responses 2026-07-28 17:03:49 +03:00
MyShining
ff86ccf1e7
feat(spring): model profiles, conditions, and auto-configuration (#2678)
* feat(spring): model conditions and auto-configuration

* fix(spring): align auto-configuration declarations

* perf(spring): streamline auto-configuration indexing

* test(spring): move timing benchmark out of vitest

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-28 07:05:41 +01:00
Gergő Magyar
e307286d52
fix(scope-resolution): a named receiver's member never resolves lexically, + two #2695 follow-ups (#2714)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(scope-resolution): a named receiver's member never resolves lexically (#2699)

`lookupCore` Step 1 walked the lexical scope chain for every lookup, including
explicit-receiver property reads. So `options.baseUrl` could bind to an
unrelated function-local `const baseUrl` in the same file, and
`config.extractVisibility(node)` to the enclosing class's own method.

This is the residual half of the defect JS/TS block scopes narrowed in #2695.
Blocks moved nested-block locals off the chain of a reference outside the
block, which removed 114 false edges; a local declared directly in the function
body stayed on it, and no amount of extra scopes reaches that case. Fixed at
the cause instead: `recv.name` names a member of whatever `recv` denotes, so a
binding of the bare tail name in an enclosing scope is never the right answer.
Steps 2 and 3 (receiver type / owner members) are the legitimate routes.

`this` and `self` are EXEMPT, and that exemption was measured, not assumed.
Skipping Step 1 for every explicit receiver removed 711 edges on a 762-file
corpus — but 2 of those were genuine: `self.srcIx` and `self.streamedAt(...)`
after `const self = this`, reaching their own class's members through the
class-body scope. For a self-receiver the members and the lexical chain
legitimately overlap; for a named receiver they never do. Exempting the self
names keeps both true edges and still removes 709 false ones, adding none.

The removals were classified by reading source at the site, not by pattern-
matching ids — an "is the target a member of the source's owner?" heuristic
labelled 43 of them plausible and every one I then read was false:

    language = config.language;          -> the class's own `language`
    dirMap.get(...) / exactMap.get(...)  -> a sibling object-literal `get`
    return config.extractVisibility(n);  -> the class's own method (self-edge)
    writer.close();                      -> GraphEmitSink.close

Residual, deliberately kept: a `this.x` read can still bind lexically to a
same-named local. That is the price of the two true self-alias edges above.

`INCREMENTAL_SCHEMA_VERSION` 19 -> 20: a v19 index holds these false
CALLS/ACCESSES on every unchanged file and would keep serving them through the
reuse gate.

Test confirmed discriminating: it fails with the guard reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(typescript,javascript): a generator expression binding is a Function node (#2693)

`const g = function* () {}` matched none of the closure-binding definition
rules — they covered `arrow_function` and `function_expression` only — so the
binding emitted a `Const` node. `buildGraphTargetIndex` admits callable nodes
only, so `g()` resolved to nothing.

Same defect shape as the `var` case #2693 already fixed: a different grammar
node for the same construct, and the resulting graph node was not callable.

Adds the four variable-binding shapes in both languages: `const`/`let` and
`var`, each plain and exported. Purely additive — no existing pattern is
reordered or rewritten, because the #2687 pre-scan dedup is order-dependent
and collapsing the value/callable pair depends on which match wins.

Deliberately NOT covered, and the query comment says so: a generator in an
object-literal pair or a HOC wrapper still falls through anonymous. Those are
rarer, and each additional pattern is another chance to disturb the dedup.

`SCHEMA_BUMP` 26 -> 27: definition captures are parse-time, so a warm parse
cache would replay the old ones verbatim — `--force` does not clear it.

Two tests confirmed discriminating (they fail with the patterns reverted), plus
a guard that the already-working generator DECLARATION form is unaffected,
since it shares the emit path these were inserted beside.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(ingestion): keep caller attribution in lockstep with definition ids (#2699)

The definition phase appends `localIdentity` to a nested callable's own name
segment (`run.save@3:2`); `findEnclosingFunctionId` did not, so the two phases
derived different ids for the same callable. The failure mode is silent — the
caller id names a node that does not exist, so the edge is dropped rather than
reported — which is why the parse-worker docblock calls this pair a lockstep
guarantee and asks that both phases derive the prefix from one place.

The condition is now byte-identical to the definition phase's
(`nestedPrefix !== undefined`), so the two cannot diverge again.

Scope of the claim, stated plainly: no reproducing case was found, and this
changes nothing measurable on a 762-file TypeScript corpus. TS/JS resolve
callers through `resolveCallerGraphId` in the graph bridge, not this path;
`findEnclosingFunctionId` serves the `callExtractor` languages, and the
corpus does not exercise a nested callable there. The review that raised it
(P3) observed zero dangling edges, and "zero dangling" is also what silently
dropped edges look like — so this closes a documented contract rather than a
demonstrated bug, and carries no test of its own.

Rides the `SCHEMA_BUMP` 26 -> 27 in the preceding commit: caller attribution
runs in the worker, so a warm parse cache would replay the old ids.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* docs(test): correct the block-scope header that this PR made false (#2699)

Review finding (MEDIUM). The file header still described `lookupCore` Step 1 as
walking the lexical chain for EVERY lookup, and called the function-body-local
case "unchanged and still mis-resolves ... pre-existing and tracked
separately". Commit 59b892ca in this same PR falsified both, and the describe
block added ~80 lines lower in this same file asserts the opposite — a reader
scoping future work from the header would have concluded the case was still
open.

Rewritten to state what the code does: Step 1 is skipped for a NAMED explicit
receiver, the function-body case is fixed here, and the surviving residual is
that a `this`/`self` read can still bind lexically to a same-named local —
with the reason those two names are exempt (they keep the genuine
`const self = this; self.member` reads that Step 1 resolves correctly).

Also corrects a PRE-EXISTING staleness inherited from #2695 in the same
paragraph block: "the genuine bare read of that same local must still emit its
edge" describes a test that no longer exists, because TypeScript emits no
`@reference.read` for bare identifiers at all. Fixed here rather than left
adjacent to a freshly corrected sentence.

Comments only — `detect_changes` reports 0 changed symbols across 1 file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* refactor(ingestion): give the nested-callable id rule one definition (#2699)

Review finding (LOW): the lockstep change in this PR shipped without a test.
The plan called for a unit test asserting the two id-derivation phases agree.
Two things changed that plan during execution, both recorded here.

FIRST — there are THREE phases, not two. Re-verifying the plan's assumption
(`grep -n localIdentity`) found a third call site: the worker-path node-id
derivation in `processFileGroup` (parse-worker.ts:2316), whose own comment
already acknowledged the coupling. `impact` on `localIdentity` corroborates:
three direct dependents, all in the Workers module. So the invariant three
phases must agree on is now ONE function, `nestedCallableQualifiedName`, and
divergence requires deleting a call rather than editing a duplicated
expression.

SECOND — the planned `_forTest` alias seam does not work for this module.
`parse-worker.ts` posts a `ready` message to `parentPort` at module scope, so
value-importing it from a unit test throws before any test runs; the existing
unit tests that reference it use `import type` only, which erases. The rules
therefore move to a new pure module, `workers/callable-id.ts`. That is what
makes them testable at all, rather than merely commented.

Pure refactor — no id changes. Verified by the suites that assert exact node
ids (`Function:svc.ts:run.save@7:2`, `Function:c.php:run.$save@3:2`): 74/74
green, and `detect_changes` reports only the three expected symbols and the
two `processFileGroup` flows `impact` predicted.

The test pins both halves: the rule's contract, and a structural assertion
that no site has re-inlined `${prefix}.${localIdentity(...)}` — the unit
assertions alone would still pass if a fourth phase spelled the rule out by
hand, which is exactly how the divergence arose.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(scope-resolution): give PHP's `$this` the same self-receiver exemption (#2699)

Review finding (LOW). The Step-1 skip added in this PR exempts `this`/`self`,
but the receiver name arrives as the reference node's RAW SOURCE TEXT —
`extractExplicitReceiver` returns `cap.text` verbatim — so PHP's `$this->x`
presents as the string "$this" and matched neither entry. PHP was the one
supported language whose self-receiver got no exemption at all.

Measured, and the measurement is why this is framed as consistency rather
than a bug fix:

  - Corpus delta ZERO. 762-file TypeScript corpus, CALLS+ACCESSES set diff:
    13179 -> 13179, added 0, removed 0. So no INCREMENTAL_SCHEMA_VERSION bump
    (stays 20), per the plan's decision rule.
  - No PHP shape found that DISCRIMINATES. Both the simple `$this->prop` /
    `$this->helper()` shapes and a closure reading `$this->…` inside a method
    that also declares a same-named local produce byte-identical edge sets
    with `$this` present and absent — Step 2 resolves the receiver's type
    first. The added test is therefore labelled a COMPANION INVARIANT, exactly
    as the `this.baseUrl` case beside it is, and does not claim to prove the
    fix.

It is still worth making: the exemption is protective, and the 709-removed /
0-true-lost measurement that justified the narrow guard was TypeScript-only,
so PHP's safety was never established by evidence. This closes that by
construction.

Two corrections to what the plan assumed, both found by checking:

  - The plan (and my first draft of this comment) claimed the codebase had no
    precedent for handling a sigil'd receiver name. FALSE: `THIS_RECEIVERS` in
    `core/ingestion/type-env.ts:244` has always listed `$this`, and it is the
    ingestion-side twin of this very list. The precedent does not merely
    exist, it validates the approach chosen here — list the spelling as data,
    do not strip sigils.
  - That twin also lists `Me`. Deliberately NOT mirrored: no entry in
    `SupportedLanguages` is Visual Basic, so it could only ever exempt a
    variable that happens to be called `Me`.

The two lists are otherwise the same set with nothing enforcing it — a fifth
instance of the twin-list drift class this PR keeps meeting. A drift guard is
the right fix and is out of scope here; noted for follow-up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* fix(rust): resolve `Self` in scope-resolution type bindings (#2699)

CI regression, caught by `tests / ubuntu / coverage` on 13d5e738 and traced to
the named-receiver Step-1 skip earlier in this PR (59b892ca), not to the three
commits above it — verified by reverting those three and reproducing the
failure unchanged.

`test/integration/resolvers/rust.test.ts > resolves fresh.validate() inside
impl User via Self {} inference` failed: 192/192 on main, 191/192 on this
branch. The fixture calls `fresh.validate()` where `let fresh = Self { .. }`
inside `impl User` — a genuine call to `User::validate`, and a TRUE edge that
the skip deleted.

Root cause is a twin-channel disagreement, not the skip:

  - `type-extractors/rust.ts:142` substitutes `Self` -> the enclosing impl
    type into the TYPE-ENV channel via `findEnclosingImplType`.
  - `languages/rust/interpret.ts` recorded `@type-binding.type` verbatim, so
    the SCOPE-RESOLUTION channel bound `fresh: Self` — a type that does not
    exist, leaving the receiver's type unknown and Step 2 unable to resolve.

`main` passed only because Step 1 still walked the lexical chain for named
receivers: the impl scope binds `validate` by name, so the call resolved BY
ACCIDENT. Stopping that walk turned a latent gap into a lost edge. The fix
closes the gap rather than restoring the accident — `Self` is now substituted
at capture-emit time in `languages/rust/captures.ts`, where the impl node is
reachable, reusing the `findEnclosingImpl` + `syntheticCapture` idiom already
in that file.

CORRECTION to this PR's central claim. "709 removed / 0 added / 0 true edges
lost" was measured on a 762-file TYPESCRIPT corpus and stated without that
qualifier. Rust lost one true edge. The measurement stands for TypeScript; it
did not generalise, and the PR body is being updated to say so.

Scope of the breakage, measured rather than assumed: 1 failure in 2927 tests
across all 51 resolver files. Every other language — Go, Java, C#, Kotlin,
Swift, Python, PHP, Ruby, Dart, C++ — passes, which is why this is a targeted
fix and not a revert of the skip.

Re-baselined `bench/scope-capture` for RUST ONLY (655aed01 -> 7f1240b3); the
other 14 language fingerprints are byte-identical. The drift is the intended
output change and the reason is recorded in the baseline entry, per that
file's own "explain, never re-baseline to make CI green" rule.

Verified: rust resolvers 192/192; all 51 resolver files 2926 passed / 1
skipped / 0 failed; the 8 targeted suites 96/96; all 8 CI bench gates PASS;
`tsc --noEmit` clean; `detect_changes` reports one touched symbol
(`emitRustScopeCaptures`) and no affected flows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

* test(golden): refresh the Rust capture golden and the C# PDG snapshot (#2699)

The two committed artifacts CI flagged after 5f55fe46. They drifted for
OPPOSITE reasons, so each was inspected before regenerating rather than
refreshed on sight.

RUST GOLDEN — drifted because 5f55fe46 CORRECTS the output. A `Self` type
binding now records the enclosing impl's type instead of the literal `Self`,
in both the `let x = Self { .. }` and `fn new() -> Self` forms. Blast radius
verified exact: 5 fixtures drifted, all 5 contain `Self`, and every
`Self`-bearing rust fixture is among them (rust-self-struct-literal,
rust-constructor-type-inference, rust-default-constructor,
rust-method-enrichment, rust-scoped-multi-file).

C# PDG SNAPSHOT — drifted because the named-receiver Step-1 skip (59b892ca)
REMOVED A FALSE EDGE. CALLS 7 -> 6, and the edge that went is:

    Demo.Resolve.Parse@142:12#1 -> Demo.Resolve.Parse@142:12#1

a self-call, from `int Parse(string v) => int.Parse(v);`. `int.Parse(v)` is
System.Int32.Parse; the lexical chain was binding it to the enclosing local
function that happens to also be called `Parse`. Same defect class as
`writer.close()` -> GraphEmitSink.close. The snapshot's own comment says it
exists so "a future refactor that silently rewires the C-family graph trips
this gate" — it tripped correctly, and the rewiring is an improvement.

Both failures were PRE-EXISTING on this PR from 59b892ca, not from the three
commits above it — verified by reverting those and reproducing unchanged. They
went unseen because this PR's CI was never watched after its first push.

Verified after regeneration, WITHOUT update flags so they must genuinely pass:
rust-captures-golden 9/9; pipeline-pdg 31/31. The snapshot diff is 3 lines,
all inside the C# entry — no other language's snapshot moved. `detect_changes`
reports 0 changed symbols (test artifacts only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184RmD24KFJidYqpM7v3XjR

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 17:56:38 +01:00
Abhigyan Patwari
93c964609a
Merge pull request #2715 from azizur100389/azizur/md060-markdown-tables-2709
fix(ai-context): emit compact markdown tables
2026-07-27 15:45:27 +05:30
Abhigyan Patwari
652ef6842e
Merge pull request #2712 from abhigyanpatwari/dependabot/npm_and_yarn/gitnexus/tar-7.5.22
chore(deps)(deps): bump tar from 7.5.20 to 7.5.22 in /gitnexus
2026-07-27 14:09:07 +05:30