Commit graph

972 commits

Author SHA1 Message Date
Gergő Magyar
713aa92bec
Merge branch 'main' into lang/php-migration-v2 2026-05-12 16:26:10 +01:00
Harlan Zhou
a2f1b07700
fix: resolve TypeScript ESM .js extension imports to .ts source files (#1525)
* fix: resolve TypeScript ESM .js extension imports to .ts source files

TypeScript ESM requires imports to use .js extensions even when source
files are .ts (moduleResolution: node16/bundler). The import resolver
now strips JS-family extensions (.js/.jsx/.mjs/.cjs) and retries with
TS equivalents (.ts/.tsx/.mts/.cts) when the literal .js file does not
exist. This fallback only applies to TypeScript/JavaScript languages.

Also adds .mts/.cts to the EXTENSIONS list for completeness.

Fixes #1503

* fix: address review findings — normalization, edge-case tests, integration test

- Fix makeCtx to use production normalization (.replace backslash)
  instead of .toLowerCase() (Finding 3)
- Add tests for .mjs/.cjs with competing .ts/.mts siblings (Finding 1)
- Add tests for ./dir.js → dir/index.ts boundary (Finding 2)
- Add integration test verifying full pipeline CALLS edges for ESM
  .js imports (Finding 4)
- Document path alias .js limitation as known follow-up (Finding 5)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* chore: retrigger CI after bot-only tip commit

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-12 16:05:58 +01:00
Gergo Magyar
fcb0e30dd1 fix(php): PSR-4-compliant fixture + pre-existing test debt cleanup
Four tests were failing on lang/php-migration-v2 in the tests/ubuntu/coverage
job before this commit. Root causes and fixes:

1. cross-file-binding.test.ts: Consumer-Before-Provider PHP
   The fixture had BProvider.php at app/BProvider.php while declaring
   namespace App\Models — a PSR-4 violation that caused PHP's import-
   target resolver to return null for 'use function App\Models\getUser'.
   Without a resolved import target, no IMPORTS edge was emitted, the
   SCC graph had no edge between consumer and provider, and the shared
   propagateImportedReturnTypes pass never mirrored getUser's return
   type into AConsumer's scope. Result: $u (bound to 'getUser' by
   query.ts:167-169 type-binding.alias) never collapsed to User, so
   $u->save() went unresolved.

   Move BProvider.php to its PSR-4-compliant location at
   app/Models/User.php. The class+function were already in namespace
   App\Models; the rename just aligns the file path with the
   namespace prefix mapped by composer.json. After the move:
     - PSR-4 directory scan finds the file for 'use function
       App\Models\getUser' (lines 65-86 of php import-resolver).
     - IMPORTS edge AConsumer.php → app/Models/User.php emits.
     - SCC reverse-topological walk in propagateImportedReturnTypes
       mirrors getUser→User from User.php's module-scope typeBindings
       into AConsumer's, then chain-follows $u → getUser → User.
     - $u->save() resolves to User#save.

2. registry-primary-flag.test.ts: 'flipped languages' size === 1
   When PHP was added to MIGRATED_LANGUAGES at commit 69786b16 (the
   PHP scope-based resolution migration), the hard-coded opt-out list
   in this test did not include REGISTRY_PRIMARY_PHP. With PHP
   default-on and not opted out, enabled.size returned 2 (Java + PHP)
   instead of 1. Rewrite the opt-out loop to iterate
   MIGRATED_LANGUAGES dynamically so future Ring 3 additions land
   without test churn.

3. overload-narrowing.test.ts: 'falls back to the full overload list'
   Commit af9af4a9 (PR #1497 production review fixes) deliberately
   tightened narrowOverloadCandidates to stop silently rescuing
   empty filter sets when every candidate had definite bounds — the
   call is genuinely arity-incompatible (e.g. PHP variadic with
   required-prefix called with too few args). The unit test still
   asserted the old soft-rescue behavior. Rename and update the
   assertion to match the intentional new semantics; document that
   the 'anyUnknownBounds' branch in the source is structurally
   unreachable from this caller shape (an unknown-bounds candidate
   always passes the arity filter, so arityMatches is always
   non-empty when anyUnknownBounds is true).

4. registries.test.ts: 'keeps incompatible candidates (soft penalty)'
   Same root cause as #3 at the registries layer: lookup-core.ts now
   drops every candidate when all are 'incompatible' and none are
   'unknown'. The 'soft penalty' kept-with-evidence behavior survives
   only when at least one candidate's arity verdict is 'unknown' —
   that signal differentiates definite mismatch from missing metadata.
   Update the existing test to assert the new hard-rejection behavior
   and add a new test exercising the surviving soft-penalty path with
   one 'unknown' + one 'incompatible' candidate.

Verification:
- Originally-failing tests: 4 of 4 now pass (no skips).
- PHP integration suite: 205/205 in primary mode, 197+8 skipped in
  legacy mode (unchanged).
- Cross-language unit + integration: no regressions.
2026-05-12 16:03:53 +01:00
Gergő Magyar
fb093f7f07
Merge branch 'main' into lang/php-migration-v2 2026-05-12 14:04:10 +01:00
evolution
0daae93701
fix(lbug): drain checkpoint result before close (#1506)
* fix(lbug): drain checkpoint result before close

* test(lbug): cover checkpoint drain lifecycle

* fix(lbug): close query results after reads

* fix(lbug): close all stream query results

* fix(lbug): harden query result cleanup

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-12 14:03:45 +01:00
Gergő Magyar
82b51508a5
Merge branch 'main' into lang/php-migration-v2 2026-05-12 13:15:49 +01:00
Abhigyan Patwari
4fa40e9881
feat(analyze): incremental indexing (parse cache + DB writeback + scope-res short-circuit) (#1479)
* docs: incremental indexing design spec

Captures the design agreed in brainstorming on 2026-05-10:
- Transitive importer closure with public-surface-change optimization
- Git-only change detection (non-git repos: full rebuild as today)
- New default behavior; --force opts out
- New hydratePhase + loadGraphFromLbug primitive
- Iterative closure expansion with parseCache reuse
- incrementalInProgress dirty flag for crash recovery

Prior art: PR #592 (zenprocess), PR #533 (davidbeesley),
PR #1146 (azeemshaik025) — referenced and credited.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(communities): seed Leiden RNG for deterministic community detection

The vendored Leiden algorithm defaults to Math.random for tie-breaking
and randomized walks, which produces non-deterministic community
assignments and modularity values across runs on the same graph.

Pass a seeded mulberry32 RNG (LEIDEN_SEED=0xC0DE) so:
- The same graph always produces the same partition
- Modularity values are reproducible
- Equivalence tests for incremental indexing can compare community
  assignments byte-for-byte

This is foundational for the upcoming incremental-indexing feature
(see docs/superpowers/specs/2026-05-10-incremental-indexing-design.md)
where the correctness contract is incremental output ≡ full rebuild
output.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(incremental): change-detection, surface signatures, closure expansion

Three new modules supporting the incremental-indexing pipeline:

* core/incremental/git-diff.ts — getChangedFilesSinceCommit() unions
  'git diff lastCommit HEAD' (committed) with 'git status --porcelain'
  (dirty tree). Renames flattened to delete(orig) + add(new). Throws
  LastCommitMissingError when lastCommit is gone (caller falls back to
  full rebuild).

* core/incremental/surface.ts — extractSurfaceSignature() produces a
  stable hash of a file's publicly-visible symbols (functions, classes,
  methods, interfaces, types, heritage). Body-only edits → same hash.
  Signature/heritage changes → different hash. Drives the closure
  scoping optimization.

* core/incremental/closure.ts — computeImporterClosure() iterative
  fixpoint: parse each closure file, extract surface, query DB
  importers, expand. Uses a parseCache so each file is parsed once.
  Generic over TParseResult so closure logic is decoupled from the
  pipeline's parse representation.

32 unit tests across the three modules. Tests cover edge cases:
clean tree, dirty-only, mixed, renames, deletes, multi-hop cascade,
cycle termination, surface invariance, etc.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(lbug): loadGraphFromLbug, queryImporters, deleteAllCommunitiesAndProcesses

Three new primitives in lbug-adapter.ts to support incremental indexing:

* loadGraphFromLbug(graph, unchangedFilePaths) — streams all nodes for
  files in the set across every hydratable node table (excludes
  Community/Process — graph-wide, regenerated downstream). Then loads
  edges where both endpoints belong to loaded nodes, excluding
  MEMBER_OF / STEP_IN_PROCESS edges (also graph-wide).
  FilePaths chunked at 200 per query to keep statement size bounded
  on huge repos. Endpoint-level join filters by source-side filePath
  in the query, target-side checked JS-side via the loadedNodeIds set.

* queryImporters(targetFilePath) — returns DISTINCT a.filePath where
  a -[IMPORTS]-> b and b.filePath = target. Powers closure expansion:
  when a changed file's surface signature changes, all its importers
  must be re-parsed.

* deleteAllCommunitiesAndProcesses() — drops Community/Process nodes
  (and their edges via DETACH DELETE) at the start of each incremental
  run so the communities/processes phases regenerate them from the
  fully-merged graph. Required for the 'Leiden runs on full graph'
  correctness invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(pipeline): hydrate phase + parse-filter for incremental indexing

Wires the incremental-indexing infrastructure into the phase-based
pipeline. Three coordinated changes:

* New hydratePhase (deps: structure) — loads node/edge state for files
  OUTSIDE ctx.options.filesToParse from the existing LadybugDB index.
  Runs before parse so the parse phase can produce a partial graph
  while downstream phases (mro, communities, processes) still see the
  full graph. No-op in full-rebuild mode (filesToParse unset).

* PipelineOptions.filesToParse: optional ReadonlySet<string>. When
  set, parse phase filters scanned files to this set; hydrate fills
  the complement. Set by runFullAnalysis when it detects an eligible
  incremental run; never set by callers directly.

* gitnexus-shared PipelinePhase enum: 'hydrate' added so progress
  callbacks can report the new phase distinctly from 'structure'.

Phase order: scan → structure → hydrate → markdown,cobol → parse
→ routes,tools,orm → crossFile → scopeResolution → mro → communities
→ processes. Communities (Leiden) still runs on the full graph,
satisfying the correctness invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): incremental orchestrator branch + meta schema

Wires incremental indexing into runFullAnalysis. Highlights:

* RepoMeta schema extended: schemaVersion, surfaceSignatures, and
  incrementalInProgress fields. INCREMENTAL_SCHEMA_VERSION = 1.

* core/incremental/file-hash.ts — v1 surface signature: SHA-256 of file
  content. v2 will switch to a true surface-only signature (defined in
  surface.ts) so body-only edits don't expand the closure. The plumbing
  is signature-agnostic so the swap is local.

* core/incremental/orchestrator.ts — eligibility check, closure
  computation (uses file-hash as the surface signal), dirty-flag
  management, subgraph extraction, signature merge.

* run-analyze.ts adds:
  - hasDirtyTree() check on the existing 'lastCommit==HEAD' early-exit
    so an uncommitted edit triggers re-index (was a coarse equality
    check before).
  - incremental branch: try incremental first; fall through to full
    rebuild on any setup failure or eligibility miss.
  - runIncrementalBranch() — opens existing DB, deletes closure-file
    rows + Community/Process, runs pipeline with filesToParse, writes
    only the changed-subgraph back, refreshes FTS, updates meta with
    new surfaceSignatures and clears the dirty flag.
  - Full-rebuild path now populates surfaceSignatures + schemaVersion
    in meta.json so the next run is eligible for incremental.

Crash recovery: incrementalInProgress is set BEFORE any DB mutation
and cleared on success by overwriting meta.json. A crash anywhere in
between leaves the flag set, and the next analyze run forces a full
rebuild (cheapest path back to a known-good index).

v1 limitation documented: body-only edits trigger 1-hop closure
expansion (content-hash signal). True surface-only optimization is
deferred to v2 — see design doc for the integration path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): drop invalid --no-renames=false from git diff

The flag --no-renames=false isn't valid git syntax (it's parsed as a
file path). Git's default rename detection is on; removing the flag
keeps that behavior.

Caught while running an end-to-end smoke test against a small fixture
repo: incremental setup failed with 'Command failed: git diff
--name-status -z --no-renames=false ...'. After the fix, the
incremental path runs cleanly: closure is computed, hydrate phase
loads unchanged-file state from DB, parse phase only re-parses files
in closure, and the writeback updates only changed nodes/edges.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Revert v1 incremental indexing (5 commits)

Reverts the v1 design that parsed only closure files into a fresh
graph and tried to hydrate the rest from DB. Real-repo equivalence
test failed: cross-file resolution operates on partial parse data
(closure files only), so CALLS edges that resolve through unchanged
files silently fall off. Diff against full rebuild on the same
edited state: -50 nodes, -425 edges, -5 communities, -48 processes.

Architecture pivot: switch to PR #533-style content-addressed parse
cache. Pipeline parses every file (cache-served when possible),
giving cross-file resolution full data, with DB writeback then
restricted to changed-file rows.

Reverts:
  d4b9de47 fix(incremental): drop invalid --no-renames=false
  f35f7634 feat(analyze): incremental orchestrator branch + meta schema
  bc039686 feat(pipeline): hydrate phase + parse-filter
  98bb893d feat(lbug): loadGraphFromLbug, queryImporters, ...
  aa8d7ae3 feat(incremental): change-detection, surface signatures, closure

Kept:
  d9e340b0 feat(communities): seed Leiden RNG (foundational)
  8235ca36 docs: incremental indexing design spec (will be revised)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): incremental DB writeback (Option B)

Equivalence-preserving incremental analyze. The pipeline still parses
every file (correctness invariant: cross-file resolution / scope
resolution / MRO / community detection all need full graph data); the
saving comes from selectively replacing only changed-file rows in
LadybugDB instead of wiping and reloading the whole graph.

How it works:

* On every analyze, we hash all source files (SHA-256 of content) and
  store the map in meta.json.fileHashes alongside schemaVersion.
* The next run loads the prior map and diffs:
  - changed: content hash differs → file's DB rows replaced.
  - added: not in prior map → file's DB rows inserted.
  - deleted: in prior map but not on disk → file's DB rows dropped.
* If the diff is non-empty AND no --force / no schema mismatch / no
  dirty flag, take the incremental path:
  - Set incrementalInProgress dirty flag (BEFORE any DB mutation).
  - Open existing DB (no wipe).
  - deleteNodesForFile() for each changed/added/deleted file.
  - deleteAllCommunitiesAndProcesses() — Leiden regenerates these.
  - extractChangedSubgraph() from the in-memory ctx.graph: nodes whose
    filePath is in the writable set + Community + Process + edges with
    at least one endpoint in the writable set (edges entirely between
    hydrated unchanged nodes are skipped — already in DB).
  - loadGraphToLbug() on the subgraph. Unchanged-file rows in DB
    untouched.
  - Recreate FTS indexes.
  - Update meta with new fileHashes; clear dirty flag.
* Otherwise full-rebuild path runs as before.

Crash recovery: incrementalInProgress is the dirty flag. Set before
destructive ops; cleared on success. Set on next-run startup → forces
full rebuild (cheapest path back to known-good).

Other changes:
* Dirty-tree gate on the existing 'lastCommit==HEAD' early-return:
  uncommitted edits no longer slip through as 'already up to date'.
* deleteAllCommunitiesAndProcesses helper in lbug-adapter.
* Skip the embedding cache+restore cycle when willTryIncremental is
  true — embeddings stay in DB; re-inserting them would PK-conflict.

End-to-end equivalence verified on this repo (993 files, 24K nodes):
incremental run produces byte-identical {nodes, edges, clusters,
flows} to a full rebuild from the same edited state.

Speedup is currently modest (~5% on this repo) because the parse
phase still runs in full. Parse-cache integration is a separate
follow-up that composes cleanly on top of this work.

See docs/superpowers/specs/2026-05-10-incremental-indexing-design.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): chunk-level parse cache for full incremental speedup

Composes with the incremental DB writeback (commit 27f3b49d) to deliver
the major-speedup half of incremental indexing. Previously, the parse
phase ran in full on every analyze; the speedup came purely from
selective DB rewriting. With this commit the parse phase also reuses
prior tree-sitter output for chunks whose contents haven't changed.

How it works:

* Cache layer (gitnexus/src/storage/parse-cache.ts):
  - File: <repo>/.gitnexus/parse-cache.json. Versioned, atomic write.
  - Key: chunk content hash = sha256(sorted(filePath:fileContentHash
    for each file in chunk)).
  - Value: ParseWorkerResult[] (raw worker output for the chunk,
    pre-merge).
  - Granularity: per chunk (~20MB byte-budget). A change to one file
    invalidates only its chunk — typically 1 of ~50 on a 1000-file
    repo (~98% cache hit ratio on a small edit).

* Worker contract (gitnexus/src/core/ingestion/parsing-processor.ts):
  - Extracted the chunk-result merge loop into a public
    mergeChunkResults() so the same logic applies to live worker
    output AND replayed cache entries.
  - processParsingWithWorkers / processParsing accept an optional
    outRawResults out-parameter that captures worker output before
    merging — used by parse-impl to populate the cache after a miss.

* Parse phase wiring (parse-impl.ts):
  - For each chunk, compute its content hash (after reading file
    contents). Cache hit → mergeChunkResults() on cached results,
    skip the worker dispatch entirely. Cache miss → run workers
    normally, capture raw results, store under the chunk hash.
  - Cache mutations happen in-place on the ParseCache passed via
    PipelineOptions.parseCache.

* Lifecycle (run-analyze.ts):
  - loadParseCache() before pipeline runs.
  - Cache passed via runPipelineFromRepo's PipelineOptions.
  - saveParseCache() after the pipeline + DB writeback succeed.

Equivalence verified on this repo (993 files, 24K nodes):

  Cold (no cache, full work):           141.1s
  Warm cache + 1-file edit, incremental: 63.6s  ← 55% speedup
  Warm cache + 1-file edit, --force:     71.6s  ← 49% speedup

All three runs produce byte-identical {nodes, edges, clusters,
flows}. The cache survives --force (content-addressed = always
correct), so even forced rebuilds get the parse-skip benefit.

Why chunk-level rather than per-file: workers process sub-batches and
emit aggregated ParseWorkerResults. Per-file granularity would require
restructuring the worker contract; chunk-level captures most of the
practical speedup with no worker-side changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(parse-impl): smaller default chunk budget (20MB→2MB) for cache granularity

The parse cache is keyed at chunk granularity. With the previous 20MB
budget, a typical mid-size repo (e.g. this worktree at 9MB total
parseable source) fits in a single chunk — meaning ANY file change
invalidates the whole chunk and re-parses every file.

2MB default produces ~5x more chunks on the same input, so a one-file
edit invalidates ~1/N of cached chunks instead of the whole thing.
Cold-run overhead from more chunks is <5% (one extra serialization
pass per chunk).

Override via GITNEXUS_CHUNK_BYTE_BUDGET env var for benchmarking.

Measured on this repo (~9MB / 887 parseable files):
  Cold (no cache):                    143s
  Warm cache, no source changes:        2s  (early-return)
  Warm cache + 1-file edit:            81s  (~43% off cold)

Speedup is bounded by the scopeResolution phase (~58s flat regardless
of parse cache) and by GitNexus's own auto-writes during analyze
(AGENTS.md / .claude/skills/ etc. mutate between runs and invalidate
chunks containing them). Both are addressable in follow-ups.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(scope-resolution): reuse worker-produced ParsedFile + stabilize chunk order

Two compounding optimizations that drop warm-cache analyze from
~134s to ~38s on a 1000-file repo (72% faster), and cold rebuild
from ~143s to ~86s (40% faster) by short-circuiting work that was
previously re-done.

1. SCOPE-RESOLUTION: REUSE WORKER PARSEDFILE

Previously, the scope-resolution phase re-parsed every file with
tree-sitter on the main thread (~58s on a 1000-file repo) because
worker-produced tree-sitter Trees can't cross the worker MessageChannel.

But the worker ALSO produces a  artifact via
, which structured-clones fine — and it's exactly
what scope-resolution would re-derive. Threading those ParsedFiles
through the parse phase () into
 ( map) lets scope-
resolution skip its extract loop on a per-file basis.

The fast path is bounded only by  per file (cheap
graph mutation). On this repo: scopeResolution went from 58s → 5s.

2. MAP-PRESERVING PARSE-CACHE SERIALIZATION

 is a
which JSON.stringify collapses to . The first attempt at threading
parsedFiles through the parse cache crashed at runtime with
"importerModule.typeBindings is not iterable" because cached entries
came back as plain objects.

Added a JSON replacer/reviver pair in parse-cache.ts that round-trips
Map and Set instances through tagged plain objects (). Symmetric: save uses replacer, load uses reviver.

3. STABLE CHUNK ORDERING

The byte-budget chunker walked files in filesystem-scan order, which
on Windows isn't guaranteed to be stable across runs. Even with
identical source content, two scans could place files in different
chunks, shifting chunk hashes and causing 100% parse-cache misses.

Added a deterministic alphabetical sort on  before
chunking. Chunk membership is now stable across runs, so a single-file
edit invalidates exactly one chunk, not all of them.

Measured on this repo (993 files, 24K nodes):
  Cold rebuild:                        86s  (was 143s)
  Warm cache, no source changes:        3s  (early-return)
  Warm cache + 1-file edit:            38s  (was 134s)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(incremental): update spec + AGENTS.md + GUARDRAILS.md for shipped design

- Rewrite docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
  to describe the architecture that actually shipped (parse cache +
  incremental DB writeback + scope-resolution short-circuit), with the
  v1 hydrate-phase post-mortem preserved as historical context.
- AGENTS.md "Keeping the Index Fresh" section: note that incremental
  is the new default and --force is the explicit opt-out; mention
  the parse-cache file location and that it's safe to delete.
- GUARDRAILS.md Signs: add an "Index seems corrupt or incremental is
  misbehaving" entry pointing users to --force as the manual escape
  hatch (the dirty flag handles automatic recovery).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(incremental): bugbot review + CI test failures

Bugbot (PR #1479):
- Medium: pruneCache was exported but never called -> cache grew
  unbounded. Wire pruneCache into run-analyze before saveParseCache,
  using a transient usedKeys Set on ParseCache that the parse phase
  populates as it processes chunks.
- Low: willTryIncremental (pre-pipeline) and isIncremental
  (post-pipeline) could desync, silently dropping embeddings on
  mispredicted runs. Removed the prediction; the embedding cache
  now loads unconditionally when shouldLoadCache is true. The
  re-insert step gates on the actual isIncremental value to avoid
  PK-conflicts when the incremental-writeback path keeps DB rows.

CI test failures:
- cli-e2e #1169 + run-analyze.test.ts #1233: my dirty-tree gate on
  the lastCommit==HEAD early-return saw GitNexus's own auto-generated
  outputs (.claude/, .cursor/, AGENTS.md, CLAUDE.md) as dirty,
  perpetually defeating the up-to-date fast path. Extended the
  pathspec exclusion to cover all auto-gen outputs, not just
  .gitnexus/.
- ruby field-type disambig: my chunk-stability sort exposed a
  pre-existing order-dependency in Ruby cross-file resolution
  (`user.address.save -> Address#save` only resolves correctly when
  user.rb parses before address.rb in some configurations). Removed
  the sort. Filesystem ordering is stable enough in practice that
  the parse cache still hits the common case; the pre-existing
  fragility is left for a separate fix.
- pipeline-graph-golden: regenerated. Seeded Leiden RNG produces a
  partition different from the previous Math.random snapshot.
- staleness `parallel calls` was a CI timing flake; passes locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): re-insert cached embeddings on incremental path

Bugbot re-review caught: deleteNodesForFile cascades to the
CodeEmbedding table (DELETE WHERE e.nodeId STARTS WITH ...), so
changed-file embedding rows are wiped along with their nodes. The
previous fix gated re-insert on `!isIncremental`, which silently
dropped those embeddings — a regression versus the full-rebuild path's
"preserve embeddings by default" guarantee.

Remove the `!isIncremental` gate. The per-batch try/catch already
handles the unchanged-file PK-conflict case ("some may fail if node
was removed, that's fine") with the same semantics, so re-inserting
the full cached set on incremental works:

  - changed-file rows: deleted, then re-inserted from cache (preserved)
  - unchanged-file rows: still in DB, re-insert PK-conflicts and is
    silently ignored (existing rows are correct)

Cost: re-inserting ~24K embeddings on incremental when only a few
files changed — most are no-op conflicts. Bounded by batch size of
200; ~3-5s overhead. Worth it for correctness.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): address Claude+Bugbot review findings + remove design doc

Addresses CHANGES_REQUESTED review on PR #1479:

1. Remove docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
   per maintainer request.

2. BLOCKER (Claude Finding 1, Bugbot Round 3): Stale cross-file edges
   between unchanged files. extractChangedSubgraph excluded edges where
   both endpoints were unchanged-file nodes — when a barrel/re-export
   file changes, cross-file resolution may update CALLS edges between
   two unchanged files that would then be silently lost.

   Fix: 1-hop importer-closure expansion of the writable set in
   run-analyze.ts. Before deleting/rewriting rows, query DB for
   importers of every changed/deleted file and add them to the writable
   set. Their nodes get deleted+rewritten too, so cross-file's refined
   edges land in the DB. Re-added queryImporters to lbug-adapter.ts.

3. BLOCKER (Claude Finding 3): Parse cache key omitted parser version.
   After a GitNexus upgrade, the cache silently replays pre-upgrade
   ParseWorkerResults against the new schema → wrong CALLS/IMPORTS/
   scope edges with no visible signal.

   Fix: PARSE_CACHE_VERSION now embeds the gitnexus npm package
   version (read at module load via createRequire on package.json).
   Format: `${SCHEMA_BUMP}+${PKG_VERSION}` e.g. "1+1.6.4". Any release
   that bumps package.json automatically invalidates the on-disk cache.
   Mismatched versions fall through to an empty cache (next save
   overwrites with the new version baked in).

4. BLOCKER (Claude Finding 2): No automated tests for incremental
   behavior. Added 28 unit tests across 3 files:

     - incremental-file-hash.test.ts (10 tests)
       diffFileHashes classification, computeFileHash determinism,
       computeFileHashes batch / missing-file tolerance, sorted output.

     - incremental-parse-cache.test.ts (12 tests)
       computeChunkHash stability and order-independence, version
       prefix format, pruneCache, load/save round-trip on empty /
       missing / corrupt / version-mismatched files, AND a Map/Set
       round-trip test that pins the JSON replacer/reviver behaviour
       (without it, ParsedFile.scopes[*].typeBindings collapses to
       {} and downstream `.get()` / iteration throws).

     - incremental-subgraph-extract.test.ts (6 tests)
       writable-set node inclusion, Community/Process always kept,
       edge inclusion when at least one endpoint is writable, MEMBER_OF
       edges via graph-wide endpoints, empty subgraph case.

5. Medium (Claude Finding 6): AGENTS.md "Keeping the Index Fresh"
   said "only changed files are re-parsed." Imprecise — the pipeline
   parses every file every run; the cache skips tree-sitter for chunks
   whose contents haven't changed. Reworded to match the design doc.

Test plan still expects:
  [x] Typecheck clean
  [x] All 28 new unit tests pass
  [x] All previously-failing tests still pass on the rebased branch
  [x] Equivalence verified locally (incremental ≡ --force, byte-identical
      stats on this repo)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): round 3 review feedback — bounded BFS, atomic meta, integration test, docs

Addresses remaining findings on PR #1479 from Claude's re-review of
commit ad7bd31 + verifies the outstanding Bugbot HIGH severity.

1. F1 — Transitive importer expansion (Claude, was Medium-but-noted).
   Previous 1-hop importer expansion missed barrel re-export chains
   (A imports C, C re-exports B; when B changes, only C was pulled in
   — A was left with potentially-stale CALLS edges to refined targets).
   Replaced the single pass with a bounded BFS over the IMPORTS graph
   (depth ≤ 4). Catches nested barrel pyramids without ballooning into
   a near-full rebuild on monorepos with deep re-export trees. `--force`
   remains the escape hatch documented in GUARDRAILS.md for cases that
   exceed the bound.

2. F2 — Integration test for incremental orchestration (Claude, BLOCKER,
   DoD §2.7). The unit tests added in ad7bd31 covered `diffFileHashes`,
   `extractChangedSubgraph`, `computeChunkHash`, `pruneCache`, and the
   Map/Set JSON round-trip — but none of them exercised the real
   `runFullAnalysis` orchestration. Added gitnexus/test/unit/
   incremental-orchestration.test.ts with four end-to-end tests against
   a real git-initialized fixture repo + real LadybugDB:

     a. First run populates fileHashes + schemaVersion and clears
        incrementalInProgress on success.
     b. Second run on unchanged state takes the alreadyUpToDate fast
        path (early-return).
     c. Second run after a source edit takes the incremental path
        (not full rebuild) and rotates fileHashes for the touched file
        while keeping the dirty flag cleared.
     d. A pre-set incrementalInProgress flag forces a full rebuild
        that clears it (crash-recovery wire).

   These would catch any regression that wires `isIncremental` from a
   pre-pipeline prediction (the Bugbot finding from commit 5eb0597) or
   accidentally re-gates the embedding re-insert on `!isIncremental`
   (the Bugbot finding from commit 60c10f1).

3. F3 — GUARDRAILS.md docs accuracy (Claude, Low). Line 33 still said
   "only changed files are re-parsed" — AGENTS.md was already corrected
   in ad7bd31 but GUARDRAILS.md was missed. Reworded to match.

4. F5 — Atomic saveMeta (Claude, Medium; vvladescu-tb fork). The dirty
   flag (`incrementalInProgress`) travels through meta.json. A crash
   mid-write would leave a corrupt meta.json that `loadMeta` would
   silently treat as "no prior index", losing the flag and skipping
   recovery. Switched to tmp-file + rename matching saveParseCache.

5. Bugbot's "Subgraph edges reference nodes absent from subgraph"
   (HIGH severity). Verified as FALSE POSITIVE: `getNodeLabel` in
   lbug-adapter.ts derives labels from the node-ID string (parses
   the table prefix), not from the in-memory graph. The CSV
   generator writes (src_id, dst_id, type) rows without consulting
   node objects; `splitRelCsvByLabelPair` routes by ID-derived label;
   `COPY ... (from=X, to=Y)` resolves both endpoints against the live
   LadybugDB where unchanged-file nodes still exist. No fix needed.

All 213 tests pass locally (including the 4 new integration tests
and the previously-failing CI tests).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): address Bugbot round-4 findings (added-file shadow seed + dedupe)

Bugbot review on commit e23e4400 surfaced two new findings against the
incremental writeback in run-analyze.ts:

  HIGH — Incremental BFS misses importers of newly added files.
    queryImporters() reads the pre-pipeline DB. For a NEWLY ADDED
    file there are no IMPORTS rows pointing to it yet, so unchanged
    files whose pre-existing import statements now resolve to the
    newcomer keep stale CALLS edges pointing at the OLD resolution
    target.

  LOW — Deleted files double-counted in filesToDelete.
    hashDiff.deleted entries can reappear in writableFiles via the
    BFS expansion (queryImporters can return a now-deleted path),
    so deleteNodesForFile() ran twice for the same file.

Fixes:

  - Add gitnexus/src/core/incremental/shadow-candidates.ts: derive
    the pre-existing file paths whose JS/TS module-resolution claim
    an added file can steal. Pattern catalogue: same-basename/
    different-extension, bare-file-beats-directory-index, and
    directory-index-beats-bare-file. Emit both POSIX and Windows
    separators because the prior fileHashes map may have been
    written from either OS.

  - In run-analyze.ts, seed the BFS frontier with shadow candidates
    that exist in the prior meta.fileHashes. Their importers — found
    via queryImporters — get pulled into the writable set so their
    CALLS edges re-resolve against the new file.

  - Dedupe filesToDelete via Set to avoid the double-call.

Tests: gitnexus/test/unit/incremental-shadow-candidates.test.ts —
8 cases covering each shadow pattern, separator handling, .d.ts as
a single extension token, deduplication, and the no-self-shadow
invariant. All 40 incremental tests (file-hash, parse-cache,
subgraph-extract, shadow-candidates, orchestration) pass locally.

Note on the third Bugbot finding ("Subgraph edges reference nodes
absent from subgraph"): re-anchored from a prior review pass — the
code at subgraph-extract.ts:48 is unchanged. Already verified as a
false positive: getNodeLabel parses labels from ID strings, CSV
write is by ID, and COPY resolves against the live DB.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(incremental): exact-equality stats invariant + analyze ≡ analyze --force

Addresses the only remaining Claude production-readiness review finding
on PR #1479 (Low-Medium, test-quality only — Claude itself said it does
NOT block merge, but the central PR claim "incremental ≡ full rebuild"
deserves explicit CI coverage rather than implicit trust).

Changes to gitnexus/test/unit/incremental-orchestration.test.ts:

1) Tighten the existing "comment-only edit takes incremental path" test.
   - Replace toBeGreaterThan(0) bounds assertions on stats.files and
     stats.nodes with exact toBe(firstMeta) per-field equality across
     files / nodes / edges / communities / processes. DoD §2.7 calls
     out bounds-only assertions as masking regressions that drop half
     the graph; this swap closes that gap.
   - Rationale: a comment-only edit must change the file content hash
     (driving the incremental path) without changing any graph data.
     Therefore every stat MUST be identical to the first run. Anything
     else is a regression.

2) New test: incremental output is byte-equivalent to a full rebuild.
   - Run analyze → comment-only edit → analyze (incremental writeback)
     → analyze --force (full rebuild from same on-disk state).
   - Assert files / nodes / edges / communities / processes are exactly
     equal across the incremental and the --force passes.
   - This is the PR's central correctness contract, now proven by a
     test that exercises the real runtime path end-to-end against a
     real on-disk LadybugDB.

All 5 orchestration tests pass locally (52s), including the new
equivalence test — every stat field matches exactly between incremental
and --force on the mini-repo fixture.

tsc --noEmit clean.

* fix(incremental): F1 cross-file edge consistency + F4 stable chunk sort + unit coverage (#1511)

Patch addressing two of the still-open changes-requested findings on PR
#1479, rebased onto the current feat/incremental-indexing head. F3
(parser fingerprint in the cache key), F5 (atomic saveMeta), and F6
(AGENTS.md phrasing) were already handled on the branch, so the
corresponding parts of the original patch were dropped as redundant.

  F1 (Blocker) — Cross-file edges between unchanged files
    Adds `computeEffectiveWriteSet(graph, toWriteSet)` to
    subgraph-extract.ts: a single pass over the new graph's edges that
    pulls the unchanged-side file of every writable-boundary-crossing
    edge into the write set. run-analyze composes it ON TOP of the
    existing importer-BFS expansion and feeds the combined set to BOTH
    `deleteNodesForFile` and `extractChangedSubgraph`, so the delete
    cascade and the writeback subgraph cover identical files (asymmetry
    would leave stale rows or PK-conflict at COPY time). The BFS reads
    IMPORTS from the pre-pipeline DB (catches files that *stopped*
    importing a changed file); the edge walk reads the new graph
    (catches refined CALLS edges the pre-run DB couldn't predict, e.g.
    a barrel re-export shifting a symbol from B to D). `extractChangedSubgraph`
    stays a pure filter — all expansion is the orchestrator's job.

  F4 (Medium) — Restore alphabetical chunk sort
    `parseableScanned` is sorted before chunking. Filesystem-scan order
    isn't stable enough across runs/platforms (notably macOS APFS) to
    keep chunk hashes consistent, so the parse cache thrashes without
    it. The pre-existing Ruby cross-file resolution order-dependency the
    old comment cited is independent — the sort surfaces it but doesn't
    cause it; tracked separately rather than leaving the cache cold.

  Tests — incremental-subgraph-extract.test.ts
    Locks the F1 invariants: `extractChangedSubgraph` is a pure filter
    (includes only the set it's given, plus graph-wide nodes; edges
    fire on one writable endpoint), and `computeEffectiveWriteSet`
    covers the barrel-re-export scenario, the symmetric edge-into-
    changed-file case, the no-boundary-crossed no-op, graph-wide-node
    edges, and input-immutability. Supersedes the prior
    extractChangedSubgraph-only test file on the branch.

Co-authored-by: Val Vladescu <vvladescu-tb@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(call-processor): register properties in pre-pass to fix order-dependent field type disambiguation + regenerate golden snapshot

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2d66666f-861c-432e-a4b0-11f2aefca98a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(call-processor): port worker-path property enrichment into the sequential pre-pass

Copilot's pre-pass in 8184439 fixed the Ruby attr_accessor order-dependence,
but it copied the OLD in-loop registration logic, not the canonical worker
path in parse-worker.ts. That left the sequential and worker paths emitting
non-identical Property nodes/symbols for the same source — silently breaking
the `incremental ≡ --force` invariant the moment a repo crosses the worker
threshold between runs.

Two concrete divergences are closed here:

  * Node id: worker keys Property as `${file}:${className}.${propName}`
    (qualified). Pre-pass was using `${file}:${propName}` (unqualified).
    Same source produced different graph ids depending on which path ran.

  * Field metadata: worker enriches each routed property with
    `provider.fieldExtractor` + `getFieldInfo`, falling back to
    `routedFieldInfo.type` for `declaredType` when the routing payload
    lacks one (e.g. types discovered from `@address = Address.new`
    ctor assignments rather than YARD `@return [Type]`), and propagates
    `visibility` / `isStatic` / `isReadonly`. Pre-pass did none of this,
    so on the sequential path `resolveFieldAccessType` failed to walk
    chains where the type only came from the FieldExtractor.

The pre-pass now mirrors parse-worker.ts:1803-1898 verbatim, with one
deliberate difference: the FieldInfo cache is scoped to a single
`processCalls` invocation rather than module-level (the worker process
is short-lived; the main thread is not, and a module-level cache would
leak state between analyze runs).

Also drops the now-stale "Defer resolution: Ruby attr_accessor properties
are registered during this same loop" comment on `pendingWrites.push` —
the rationale is no longer accurate after Copilot's pre-pass, but the
deferral is still needed so write-access tracking sees inference that
completes during the main loop. Comment updated to reflect that.

Verification:
  * `tsc --noEmit`: 0 errors
  * test/unit (call-processor, call-routing, field-extraction, ruby-self-call): 224 passing
  * test/integration (ruby, ruby-sequential-mixin, pipeline-graph-golden): 137 passing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(call-processor): key fieldInfoCache by filePath:startIndex, not raw byte offset

Claude's review of 255bdf6 caught a real collision in the FieldInfoCache I
added: keying by `classNode.startIndex` alone is a per-file byte offset, so
two files that both begin with a class at byte 0 — extremely common in Ruby /
Python, where files frequently open with `class Foo`, `module Foo` — collide
on the same cache entry. The second file's `getFieldInfo` then returns the
first file's FieldInfo map, producing wrong `declaredType` / `visibility` /
`isReadonly` on its properties.

Same shape as the bug that already exists in parse-worker.ts:377 (also keyed
by `classNode.startIndex` in a module-level map, persistent across files
processed by the same worker). Fixing the symmetric pre-existing leak in
parse-worker.ts is a separate, scoped follow-up — left out of this commit to
keep the fix minimal and reviewable.

Cache map and key are now both string-typed. Composite key
`${context.filePath}:${classNode.startIndex}` keeps the within-file hit rate
(one FieldExtractor.extract() per class regardless of how many
`attr_accessor` lines it has) while eliminating cross-file aliasing.

Verification on the patched HEAD:
  * `tsc --noEmit`: 0 errors
  * test/unit (call-processor, call-routing, field-extraction, ruby-self-call): 224 passing
  * test/integration (ruby, ruby-sequential-mixin, pipeline-graph-golden): 137 passing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Val Vladescu <val.vladescu@thirdbridge.com>
Co-authored-by: Val Vladescu <vvladescu-tb@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-12 13:14:56 +01:00
Gergo Magyar
00d4cc08ce test(php): skip U1/U4 scope-resolver-only assertions in legacy DAG mode
Four tests added by U1 and U4 assert correctness wins that only the
scope-resolution pipeline delivers. The legacy DAG path has no
equivalent narrowing, so it emits the over-broad edges these tests
forbid:

- U1: MRO arity narrowing on class-name receivers (Case 2 in
  receiver-bound-calls.ts) — legacy has no arity check between MRO
  iterations, so 'Child::method(1)' falls through to Parent::method.
- U1: single-class arity gating — legacy resolves by name without
  arity, so an arity-incompatible call on a class with no parent
  still emits an edge.
- U4: 'phpEmitUnresolvedReceiverEdges' exact-required-arity gate —
  the legacy DAG has no equivalent unresolved-receiver fallback hook,
  so default-param over-arity and variadic-below-required shapes
  resolve via a different code path that over-emits.

Backporting these narrowing checks to the legacy DAG is out of scope,
matching the existing pattern in LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES
for record/pad zero-args (commit af9af4a9 U1) and the FQN cross-
namespace test (Codex PR #1497 finding 1).
2026-05-12 12:53:01 +01:00
Gergo Magyar
ea04c5ee52 docs(php): document '...' vs 'params' variadic-marker asymmetry (U5)
The shared narrowOverloadCandidates pass in scope-resolution/passes/
overload-narrowing.ts checks for 'params'/'params ' as the variadic
marker - the C# 'params' keyword convention. PHP uses '...' instead,
matching its source-language syntax. The discrepancy is harmless in
practice because PHP variadic methods set parameterCount=undefined,
which skips the 'max !== undefined && argCount > max' gate that hosts
the 'params' check entirely - so the C# branch is dead code for PHP.

Document the asymmetry at both sites with cross-references so a future
contributor reading either file lands on the full picture:

- overload-narrowing.ts:48 area: explains the 'params' branch is C#-
  specific and warns that adding new markers requires auditing each
  adapter's arity-metadata.ts.
- languages/php/arity-metadata.ts:46 area: explains PHP uses '...'
  intentionally and that the shared pass's 'params' branch is dead
  code for PHP because of the parameterCount=undefined setting.

No behavior change. Existing PHP variadic-resolution tests already
exercise the live '...' path.
2026-05-12 12:39:02 +01:00
Gergo Magyar
e90f77cb2a fix(php): tighten unresolved-receiver fallback to exact-required arity (U4)
phpEmitUnresolvedReceiverEdges emits 0.6-confidence CALLS edges for
member-call sites whose receiver has no type binding (PHP 'mixed'-typed
parameters, untyped variables) when the called method name has a single
workspace-wide candidate. The existing first-stage gate accepted any
argCount in 'min..max' (or '>= min' with variadic), which over-emits
when the lone candidate has optional / defaulted parameters - the
fallback fires for calls that happen to fit the wider range but
weren't meant for that target.

Add an EXACT-required-arity gate for the 0.6-confidence path only:
for fixed-arity candidates, require argCount === requiredParameterCount.
Variadic candidates ('...' marker in parameterTypes) keep the relaxed
'>= required' semantics already enforced by the first-stage check.

Adds php-unresolved-receiver-arity fixture with seven scenarios covering
happy path, default-param exact match, default-param over-arity (the
narrowed case), variadic at/above/below required, and the class-
detection sanity check. Pre-fix the over-arity scenario emits a stray
0.6 edge; post-fix it does not. All existing PHP integration tests
continue to pass.
2026-05-12 12:36:34 +01:00
Gergo Magyar
14a57b5226 test(php): lock dynamic-dispatch suppression in regression coverage (U3)
The PR #1497 adversarial review confirmed via grammar inspection that
dynamic PHP call shapes ($obj->$method(), $obj->{$method}(),
Class::$method(), $className::method(), call_user_func with variable
/ string / array callables, dynamic property reads) currently produce
zero captures and zero CALLS edges - because the member_call_expression
and scoped_call_expression query patterns constrain `name:` to
`(name)` rather than `(_)`, deliberately excluding variable_name
nodes.

That safety invariant was not regression-tested. A future query.ts
edit relaxing `name:` to `(_)` would silently emit false-positive
edges. Add a php-dynamic-calls fixture covering all 11 dynamic shapes
from Findings 1-7 plus a sanity-check static call that DOES emit an
edge - without the sanity check every zero-edge assertion would pass
even if the pipeline emitted no edges at all.

Adds inline SAFETY-INVARIANT comments above the three load-bearing
query patterns (member_call_expression, scoped_call_expression, and
the static-write property pattern) referencing this fixture as the
regression net.
2026-05-12 12:33:01 +01:00
Gergo Magyar
c82fa8369a fix(php): dedup typed-property catch-all double-match in captures (U2)
The untyped @declaration.variable catch-all pattern in query.ts (lines
101-103) has no `type:` constraint and tree-sitter therefore also
matches it against typed property declarations - emitting a second
capture for the same property_declaration anchor. Graph-level def-id
collision currently masks the duplicate at the node-emit layer, but
the catch-all capture still flows through scope-binding and name-keyed
registries with a `$`-prefixed name that the typed branch's `$`-strip
never normalizes - a known vector for receiver-binding lookup pollution.

The two tree-sitter patterns produce separate rawMatches entries with
separate `grouped` maps, so the dedup must be cross-match. Pre-scan
rawMatches once to collect anchor node IDs already covered by
@declaration.property, then skip @declaration.variable matches whose
anchor is in that set.

Adds php-typed-property-dedup fixture covering typed property,
constructor-promoted typed parameter, and a mixed-untyped declaration
to verify the catch-all path still fires for untyped properties.
2026-05-12 12:21:46 +01:00
Gergo Magyar
29e45f1971 fix(php): stop MRO walk on arity-incompatible most-derived (U1)
When the class-name receiver pass (Case 2 in receiver-bound-calls.ts)
found a most-derived definition that was arity-incompatible with the
call site, the previous code used `continue` to fall through to the
next ancestor in the MRO chain. If an ancestor happened to be arity-
compatible, the resolver emitted a false CALLS edge to it.

This is incorrect: PHP dispatches to the most-derived override at
runtime and throws `ArgumentCountError` when arities don't match -
it never silently redirects to an ancestor. Replace the inner
`continue` with `break` so the chain walk terminates and no edge
is emitted for the site.

Adds php-mro-arity-mismatch fixture and 5 regression tests covering:
- the bug scenario (Child::method/2 + Parent::method/1, call with 1 arg)
- arity-compatible happy path (Child::compat/1)
- no-parent case (Orphan with arity mismatch)
- happy path most-derived call
- class detection sanity check
2026-05-12 10:48:49 +01:00
Gergő Magyar
55b8790f75
Merge branch 'main' into lang/php-migration-v2 2026-05-12 09:38:45 +01:00
Copilot
d4f34905bc
feat: migrate Java to scope-based registry resolution (RFC #909 Ring 3) (#1482)
* Initial plan

* feat: implement Java scope-based resolution (RFC #909 Ring 3)

Add scope-resolution pipeline for Java, following the C# pattern:

- query.ts: tree-sitter query for scopes, declarations, imports,
  type bindings, and references against tree-sitter-java grammar
- captures.ts: orchestrator synthesizing import decomposition,
  receiver bindings (this/super), arity metadata, and reference arity
- import-decomposer.ts: decompose import_declaration nodes into
  kind/source/name markers (named, wildcard, static, static-wildcard)
- interpret.ts: convert captures to ParsedImport/ParsedTypeBinding
- receiver-binding.ts: synthesize this/super type-bindings on instance
  methods with superclass support
- arity-metadata.ts: extract parameter count/types using javaMethodConfig
- arity.ts: Java arity compatibility check with varargs support
- merge-bindings.ts: Java shadowing precedence (local > import > wildcard)
- simple-hooks.ts: bindingScopeFor, importOwningScope, receiverBinding
- import-target.ts: package path to file path resolution
- scope-resolver.ts: ScopeResolver implementation registered in registry

Wire scope hooks into javaProvider (java.ts) and register
javaScopeResolver in SCOPE_RESOLVERS registry. Add createResolverParityIt
wrapper to java.test.ts for parity testing.

All 172 existing Java tests pass. Java is NOT added to
MIGRATED_LANGUAGES — the resolver sits idle until the migration flag
is flipped.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix: address review findings 1-4 — varargs arity, static import resolution, importOwningScope, stripGeneric

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22308da3-59c9-47e6-8e52-738305b1b80a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: document registry-primary parity status and CI visibility gap in scope-resolver

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22308da3-59c9-47e6-8e52-738305b1b80a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: add generic type erasure fallback in stripGeneric + update scope-resolver docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/223f77ac-59a7-4487-9316-f2be05eac5d3

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: improve stripGeneric fallback regex — use valid Java identifier chars and handle nested generics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/223f77ac-59a7-4487-9316-f2be05eac5d3

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address adversarial review findings 1-6 — flaky test, wildcard import fixture, varargs fixed-prefix test, qualified generic stripping, JSDoc updates

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/172c8a1a-cdf3-4de8-9142-f2c12c14b0a6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: add inline comment explaining stripQualifier/stripGeneric call order

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/172c8a1a-cdf3-4de8-9142-f2c12c14b0a6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add varargs 0-arg fixture and strengthen wildcard import assertions

Finding 1: Added `badCall()` method with 0-arg `fmt.format()` call to the
varargs fixture. Test documents that legacy mode still resolves this call
(arity rejection is registry-primary only). The fixture now exercises both
the success path (2-arg, 3-arg) and the undersupplied path (0-arg).

Finding 2: Strengthened wildcard import test to assert `targetFilePath`
on the CALLS edge (`com/example/models/User.java`), confirming the call
resolved through the wildcard-imported type to the correct file.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2b4e5602-9833-485c-ab48-e1d54fdf8465

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-12 09:37:44 +01:00
Gergő Magyar
1d84164845
Merge branch 'main' into lang/php-migration-v2 2026-05-12 09:13:11 +01:00
Abhigyan Patwari
f33efa8714
Merge pull request #1523 from magyargergo/fix/claude-review-skip-permissions-ci
ci(claude): fix /review PR comments blocked by Bash approval in Actions
2026-05-12 09:12:46 +01:00
Gergo Magyar
2bf6d078aa ci(claude): allow Bash in code-review job without interactive approval
Claude Code defaults to prompting for Bash approval. In GitHub Actions there
is no human to approve, so gh pr comment and similar commands fail and the
PR receives no review comment. Pass --dangerously-skip-permissions for the
code-review step only (headless CI; token and checkout are already scoped).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-12 09:07:47 +01:00
Gergő Magyar
e7ed4df390
Merge branch 'main' into lang/php-migration-v2 2026-05-12 08:29:49 +01:00
Abhigyan Patwari
ba1ac3989d
Merge pull request #1522 from abhigyanpatwari/fix/claude-review-pr-comment
fix(ci): make /review reliably post PR comments
2026-05-12 08:29:39 +01:00
Gergo Magyar
c5c42a9649 fix(ci): make /review reliably post PR comments
The Run Claude Code Review step passed an invalid PR ref
(owner/repo/pull/N) which gh interprets as a branch name, causing
early gh pr view failures. More importantly, the prompt omitted
--comment, so the code-review plugin only displayed findings in
terminal output and never invoked gh pr comment to post to the PR.

Switch to a full PR URL and add --comment so the plugin posts the
review during the session, which also routes around upstream bugs
anthropics/claude-code-action#1061 and #1087 where the action's
post-step capture can silently drop output on issue_comment triggers.
2026-05-12 08:28:42 +01:00
Gergő Magyar
6ff0a4bd3b
Merge branch 'main' into lang/php-migration-v2 2026-05-12 08:04:11 +01:00
Abhigyan Patwari
a445d38e0b
Merge pull request #1521 from abhigyanpatwari/remove_flaky_test
chore(tests): remove flaky regression test for resource exhaustion
2026-05-12 08:00:04 +01:00
Gergo Magyar
0e2c0c77ec chore(tests): remove flaky regression test for resource exhaustion 2026-05-12 07:58:58 +01:00
Gergo Magyar
92e5fbe6ad Merge branch 'lang/php-migration-v2' of https://github.com/magyargergo/GitNexus into lang/php-migration-v2 2026-05-11 19:26:05 +01:00
Gergo Magyar
d776ca6a84 test(php): register FQN test in legacy-parity skip list (Codex #1497)
After U3 (FQN-keyed bindingAugmentations) lands, the FQN regression
passes under REGISTRY_PRIMARY_PHP=1 but still fails under =0. The
legacy DAG resolves receiver types via simple-name workspace lookup
and has no namespace-prefixed binding channel, so it cannot
distinguish `\App\Other\User` from a same-simple-name class
reachable via `use`. Per the established convention from commit
af9af4a9 ("parity with the legacy DAG is not a correctness
criterion when the legacy DAG itself has the same defect"), the
test registers in LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.php
rather than blocking the migration.

The other two assertions in the new describe block (class detection
and the saveLocal control) pass in both modes and stay green.

Verified:
- REGISTRY_PRIMARY_PHP=0: 174 passed, 4 skipped (3 existing + 1 new).
- REGISTRY_PRIMARY_PHP=1: 178 passed.
- C# + Python + TypeScript + Go + C resolver matrix: 794 passed.
- pickImplicitThisOverload unit tests: 5 passed.
- resolver-parity-expected-failures unit test: 3 passed.

Plan: docs/plans/2026-05-11-002-fix-php-fqn-and-overload-codex-findings-plan.md (U5)
2026-05-11 19:25:05 +01:00
Gergo Magyar
53d1b7ba2b fix(scope-resolution): require unique narrowing in pickImplicitThisOverload
Codex PR #1497 review, finding 2: pickImplicitThisOverload returned
candidates[0] after narrowOverloadCandidates without checking
uniqueness. When two same-name methods on the same class shared
identical arity and the call site lacked disambiguating argument-type
info, narrowing left both compatible and the resolver emitted a
high-confidence CALLS edge whose target depended on registration
order rather than a defensible resolution.

Tighten the picker:
- candidates.length === 1 -> return candidates[0] (unchanged for the
  unambiguous case)
- candidates.length !== 1 (zero or multiple) -> return undefined
  (call left unresolved; no edge emitted)

This mirrors pickUniqueGlobalCallable's existing pattern in the same
file.

Export pickImplicitThisOverload so a unit test can exercise it with
synthetic stubs — PHP cannot produce the multi-overload failure shape
(no method overloading in PHP), and C# integration coverage would
entangle the unit's contract with the broader C# resolver. The unit
test pins five cases: sole overload, narrowing-disambiguated, the
ambiguous multi-candidate case (the bug regression), no-match, and
no-enclosing-class.

Verified: 972/972 across PHP + C# + Python + TypeScript + Go + C
resolver suites; no regression in any language. tsc clean.

Plan: docs/plans/2026-05-11-002-fix-php-fqn-and-overload-codex-findings-plan.md (U4)
2026-05-11 18:58:59 +01:00
Gergo Magyar
4ffe5449ac fix(php): inject FQN-keyed module-scope bindings (Codex #1497 finding 1)
Extend populatePhpNamespaceSiblings with Step 3b: for every PHP file's
Module scope, inject a binding entry keyed by the fully-qualified
class name (`App\Models\User`) for every class-like def in the
workspace. This routes FQN-receivers like `\App\Other\User` to the
exact namespace-qualified class regardless of which simple-name `User`
the caller's `use` imports shadowed.

Why module-scope bindingAugmentations instead of mutating def.qualifiedName:
the shared QualifiedNameIndex consumes def.qualifiedName at finalize time,
but PHP class defs need to remain keyed by simple name throughout the rest
of the pipeline (MRO, method-dispatch-index, namespace-siblings step 3).
An earlier attempt to rewrite def.qualifiedName to namespace-prefixed form
cascaded into 32 unrelated test failures across receiver-binding, MRO, and
heritage. The bindingAugmentations channel is purpose-built for adding
post-finalize visibility without mutating shared semantic state, and
`findClassBindingInScope`'s scope-chain walk already consumes it via
`lookupBindingsAt` — wiring is zero-touch.

Cost: O(PHP files × class-like defs) augmentation entries. Typical PHP
project: hundreds × hundreds = bounded.

Verified locally:
- Registry-primary: 178/178 PHP tests pass (including the new FQN regression).
- Legacy DAG: 174 passed, 3 existing skips, 1 failure (the FQN test, expected
  — legacy DAG has no namespace augmentation channel; U5 registers it in
  LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.php).
- Cross-language: 794/794 (C#, Python, TypeScript, Go, C) pass. The PHP-only
  augmentation does not touch shared resolution code.
- tsc clean.

Plan: docs/plans/2026-05-11-002-fix-php-fqn-and-overload-codex-findings-plan.md (U3)
2026-05-11 18:12:35 +01:00
Gergo Magyar
1d69b3f5fa fix(php): preserve qualified form on TypeRef.rawName in normalizePhpType
Stop collapsing `\App\Models\User` to `User` in normalizePhpType (step 6).
Canonicalize the leading backslash off and preserve the qualified path
on TypeRef.rawName so downstream PHP receiver resolution can distinguish
the FQN target from a same-simple-name class reachable via `use`.

- Step 6 rewritten: `\App\Models\User` → `App\Models\User`,
  `App\Models\User` → unchanged, `User` → unchanged.
- Final validation regex relaxed from /^\w+$/ to /^\w+(?:\\w+)*$/ to
  accept qualified PHP identifiers while still rejecting empty segments,
  trailing backslashes, and non-identifier characters.
- Step 5 (single-arg generic strip) already passes qualified inner types
  through via its existing /^\w[\w\]*/ pattern — no change needed.

U3 will wire the qualified-name lookup path; this commit alone is a
no-op for resolution (the lookup chain still keys on simple names).
PHP suite remains green at 177 passed; the only failing test is the
intentional FQN regression from U1.

Plan: docs/plans/2026-05-11-002-fix-php-fqn-and-overload-codex-findings-plan.md (U2)
2026-05-11 17:05:22 +01:00
Gergo Magyar
0e5e54f2d3 test(php): land failing FQN cross-namespace regression (Codex #1497)
Fixture php-fqn-cross-namespace declares two `User` classes — `App\Models\User`
and `App\Other\User` — and a Service.php that imports the Models simple-name
via `use App\Models\User` but uses a fully-qualified `\App\Other\User` in a
parameter annotation. Adds three assertions:
- Both User classes are detected as Class nodes in distinct files.
- The FQN-parameter method's `$u->record()` resolves to app/Other/User.php
  (currently FAILS — Codex review finding 1).
- The simple-name-parameter method's `$u->record()` resolves to
  app/Models/User.php (control; still passes pre-fix).

This is the test-first gate Codex required: "Block merge until the PHP FQN
fixture fails before the fix and passes after it." Verified failing on
HEAD before U2-U3 land the fix.

Also sweeps every `toBeGreaterThanOrEqual` in php.test.ts to exact
`.toBe(N)` per DoD.md §2.7 ("avoid bounds-only assertions that mask
regressions"). Touched 10 assertions across the U2/U3 trait MRO blocks,
the namespace-aware free-call fallback block, and the new FQN block.
All previously-green tests stay green with exact counts — no latent
over-emission bugs were hidden by `>= 1`.

The ambiguous-overload finding (Codex finding 2) cannot be exercised by a
PHP integration fixture — PHP does not support method overloading, so
`model.methods.lookupAllByOwner` returns at most 1 entry per (class, name)
pair. That regression test ships as a unit test against
`pickImplicitThisOverload` in U4 once the function is exported.

Plan: docs/plans/2026-05-11-002-fix-php-fqn-and-overload-codex-findings-plan.md (U1)
2026-05-11 16:51:05 +01:00
Gergő Magyar
d89ec0a1df
Merge branch 'main' into lang/php-migration-v2 2026-05-11 16:50:23 +01:00
Antheurus
fcab1e2e82
fix(augment): add CONTAINS fallback when FTS indexes unavailable (#1476)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Release Candidate / Check if release candidate should run (push) Waiting to run
Release Candidate / ci (push) Blocked by required conditions
Release Candidate / Publish release candidate to npm (push) Blocked by required conditions
Release Candidate / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(augment): add CONTAINS fallback when FTS indexes unavailable

When the MCP server holds the KuzuDB write lock, the augment CLI opens
the DB read-only. FTS indexes cannot be created in read-only mode, so
searchFTSFromLbug returns ftsAvailable=false and an empty results array.
The existing early-return path silently produced no enrichment.

Add a Cypher name CONTAINS fallback that fires only when ftsAvailable is
false and BM25 produced no symbol matches. This covers the read-only DB
case (concurrent MCP server) and the first-run case (indexes not yet
built). The fallback is wrapped in .catch(() => []) and cannot throw.

When FTS indexes exist, this branch is never reached — behaviour is
unchanged for users without a concurrent MCP server.

* fix(augment): guard against CONTAINS '' and add no-FTS test coverage

Blocker 1 — CONTAINS '' on whitespace-leading patterns:
pattern.split(/\s+/)[0] returns "" when the input has leading whitespace
(e.g. "   ".split(/\s+/) → ["", ""]). In Kuzu, CONTAINS '' matches every
node with a name property, injecting arbitrary graph nodes into LLM context.

Fix: trim() before split, then guard on !firstWord || firstWord.length < 2.
No behaviour change for normal non-empty patterns.

Blocker 2 — zero test coverage on the FTS-unavailable code path:
The new CONTAINS fallback block (engine.ts lines 146-166) was exercised by
no existing test — all existing tests run with FTS indexes built. A second
withTestLbugDB fixture is added with no ftsIndexes, forcing searchFTSFromLbug
to return ftsAvailable: false, and asserts:
1. augment('login', ...) returns non-empty enrichment (fallback works)
2. augment('   ', ...) returns '' (CONTAINS '' guard holds)
3. augment('nxyz_notfound', ...) returns '' (no matching nodes)
4. executeQuery throwing returns '' (.catch(() => []) path)

* fix(augment): extend CONTAINS '' guard to FTS happy path and consolidate

The same split(/\s+/)[0] bug existed at line 125 (BM25 symbol filter,
FTS-available path) — a leading-whitespace pattern produced CONTAINS ''
there too, matching every node in BM25-matched files.

Fix: hoist patternFirstWord computation with trim() and the length guard
to the top of augment(), before any DB interaction. Both CONTAINS sites
(BM25 symbol filter and CONTAINS fallback) now use the single pre-validated
value. No behaviour change for normal patterns; the guard fires once for
all callers instead of being duplicated.

Also tighten the whitespace test in the no-FTS suite from 3 spaces to
4 spaces so it unambiguously exercises the patternFirstWord guard rather
than straddling the outer pattern.length < 3 boundary.

* test(augment): negative-safety test for ftsAvailable=true gate

Asserts the CONTAINS fallback does NOT fire when FTS is available but
BM25 returns zero results. Pins the safety property promised by the PR
description: behavior is unchanged for users without the read-only-DB
condition.

If anyone later loosens the gate to `symbolMatches.length === 0` alone,
this test fails.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 16:39:10 +01:00
Gergo Magyar
1326d7a3b2 Merge branch 'main' of https://github.com/abhigyanpatwari/GitNexus into lang/php-migration-v2 2026-05-11 15:06:09 +01:00
Gergő Magyar
88a269f1c1
Merge branch 'main' into lang/php-migration-v2 2026-05-11 15:05:43 +01:00
achianuri
fdf1effb2a
feat(cli): add --skip-skills and --index-only flags to analyze (resubmit of #742) (#1485)
* feat(cli): add --skip-skills and --index-only flags to analyze command

The `installSkills()` call in `generateAIContextFiles()` runs
unconditionally, injecting 6 skill files into `.claude/skills/gitnexus/`
even when `--skip-agents-md` is passed. This is problematic for bulk
indexing operations on read-only mirrors or third-party repos.

Add two new flags:
- `--skip-skills`: suppress standard GitNexus skill file injection
- `--index-only`: pure index mode that suppresses all file injection
  (AGENTS.md, CLAUDE.md, and skills), writing only to `.gitnexus/`

This gives users three levels of control:
- `--skip-agents-md` — suppress only root context files
- `--skip-skills` — suppress only skill injection
- `--index-only` — suppress everything (pure indexing)

Discovery context: while bulk-indexing 176 repos with
`--skip-agents-md`, all 144 indexed repos were contaminated with
`.claude/skills/gitnexus/` files requiring manual cleanup.

* fix(cli): address PR #742 review — gate community skills, drop dangling refs, add tests

Bot review (#742) flagged three issues with the original commit:

1. `--index-only --skills` still wrote community-derived skill files
   to `.claude/skills/generated/`. The `--skills` branch in analyze.ts
   was not gated by `skipAll`, so the "skip all file injection" contract
   was violated. Gate `generateSkillFiles()` with `!skipAll` so
   `--index-only` truly wins over `--skills`.

2. `--skip-skills` without `--skip-agents-md` produced AGENTS.md /
   CLAUDE.md that still referenced `.claude/skills/gitnexus/*/SKILL.md`
   files that were never installed — every agent load incurred 6
   failed reads. Pass `skipSkills` through to `generateGitNexusContent()`
   and omit the standard-skill rows (and the entire `## CLI` heading
   when the table is empty). Community skills, when present via
   `--skills`, are unaffected.

3. No filesystem tests for `skipSkills` / `indexOnly`. Add three
   regression guards to `test/unit/ai-context.test.ts`:
   - `.claude/skills/gitnexus/` is NOT created when skipSkills=true
   - Nothing is written when both skipAgentsMd and skipSkills are true
     (the resolved-flag state from --index-only)
   - AGENTS.md/CLAUDE.md routing table omits standard skill references
     when skipSkills=true, but preserves the load-bearing imperative
     sections (Always Do / Never Do / Resources)

* test(cli): PR 1485 review follow-ups (help text, gate test, --skip-skills docs)

- Assert --skip-skills and --index-only in analyze --help (skip-git-cli.test.ts).

- Export shouldGenerateCommunitySkillFiles; unit-test index-only+skills gate.

- Clarify --skip-skills does not suppress --skills community files; --index-only for full skip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): warn when --index-only silently overrides --skills

Address review findings on PR 1485 follow-ups:
- analyze.ts emits a one-line note when both --index-only and --skills
  are set, so users see why a pipeline re-index ran with no skill files
  written.
- index.ts --skills help text now flags the --index-only override.
- shouldGenerateCommunitySkillFiles JSDoc documents the dual role of
  the gate (community skills + AGENTS.md/CLAUDE.md re-generation).
- skip-git-cli.test.ts pins the override-warning surface end-to-end.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-11 15:00:58 +01:00
Gergo Magyar
f81c6753f3 Merge branch 'lang/php-migration-v2' of https://github.com/magyargergo/GitNexus into lang/php-migration-v2 2026-05-11 14:58:46 +01:00
Gergo Magyar
fd30b61c38 test(php): gate legacy-DAG-divergent assertions via parity helper
CI scope-parity / php parity was failing under REGISTRY_PRIMARY_PHP=0 on
three assertions added by commit af9af4a9 (U1 arity-narrowing, U3 trait
shadows parent). Per that commit's stance — "parity with the legacy DAG
is not a correctness criterion when the legacy DAG itself has the same
defect" — backporting these fixes to the legacy resolver is out of scope.

Adopt the existing sibling pattern (csharp/typescript/python use the
same helper):
- php.test.ts: switch to `const it = createResolverParityIt('php')` so
  expected-failure assertions skip under legacy mode.
- helpers.ts: register three test names in
  LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.php with the rationale.

Verified locally: 172 passed + 3 skipped under REGISTRY_PRIMARY_PHP=0,
175 passed under REGISTRY_PRIMARY_PHP=1, tsc clean.
2026-05-11 14:54:34 +01:00
Gergő Magyar
1796afb6f3
Merge branch 'main' into lang/php-migration-v2 2026-05-11 14:36:10 +01:00
henry201605
622f98ade5
feat(embeddings): forward dimensions param in HTTP embedding requests (#1498)
* feat(embeddings): forward GITNEXUS_EMBEDDING_DIMS as dimensions in HTTP request body

When GITNEXUS_EMBEDDING_DIMS is set, include it as the `dimensions` field
in the /v1/embeddings request body. This enables Matryoshka-capable models
(OpenAI text-embedding-3-*, Cohere embed-v3, Voyage) to return truncated
vectors at the requested size.

When the env var is unset, the request body remains `{ input, model }` —
no breaking change for backends that reject unknown fields.

Adds 4 unit tests covering both paths (with/without dimensions) on both
the batch embed and single-query embed code paths.

* fix(embeddings): address review findings — strict parseInt, multi-batch test, comment wording

1. Strict parseInt validation: reject non-numeric strings like '1024abc'
   by checking /^\d+$/ before parseInt (Finding 1).
2. Add multi-batch test asserting dimensions is forwarded in every fetch
   call when inputs exceed batch size (Finding 2).
3. Soften JSDoc comment: backends may ignore or reject the dimensions
   field rather than universally ignoring it (Finding 3).
4. Add test for invalid GITNEXUS_EMBEDDING_DIMS values.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 13:56:10 +01:00
Gergo Magyar
b1f0cd8cf1 Merge remote-tracking branch 'magyargergo/lang/php-migration-v2' into lang/php-migration-v2 2026-05-11 13:17:18 +01:00
Gergo Magyar
bc8b6a6fe2 fix(php): restore entryPointPatterns and astFrameworkPatterns
These were stripped by the PHP migration commit (69786b1) as
collateral damage during the rebase, but they have nothing to do
with scope resolution. They feed cluster/community detection and
Laravel route-attribute analysis at the higher analysis layers.

Restored verbatim from 69786b1~1:
- entryPointPatterns: 17 PHP idioms (Controller, handle/execute/
  boot/register, REST verbs, Service/Repository, find/save/delete)
- astFrameworkPatterns: Laravel routing block with 3.0x multiplier
  for Route::get / Route::post / #[Route(...)] attribute, etc.

PHP 175/175 still green; typecheck clean.
2026-05-11 13:16:50 +01:00
Gergő Magyar
f81e084686
Merge branch 'main' into lang/php-migration-v2 2026-05-11 13:14:38 +01:00
juyua9
55b7a79beb
fix(group): detect httpx async consumers (#1408)
* fix(group): detect httpx async consumers

* test(group): tighten httpx consumer coverage

* test(group): create extractor temp dirs safely

* fix(group): scope httpx async client tracking

* fix(group): tighten httpx module-scope tracking

Prevent module-scope httpx.AsyncClient tracking from matching same-name local variables inside functions.

Also documents the intentionally unsupported direct-import, alias, and typed-assignment forms, and extends the httpx extractor regression fixture to cover module-scope shadowing while keeping module-scope calls detected.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 13:13:49 +01:00
Gergő Magyar
ebccccf715
Merge branch 'main' into lang/php-migration-v2 2026-05-11 13:11:41 +01:00
Gergo Magyar
97093f9ce4 Merge remote-tracking branch 'magyargergo/lang/php-migration-v2' into lang/php-migration-v2 2026-05-11 13:10:54 +01:00
Gergo Magyar
af9af4a958 fix(php): four PHP semantic defects flagged by PR #1497 production review
Fixes four PHP-semantic defects identified by the production-readiness
review. The bar is "graph edges must reflect what PHP actually does",
not "matches legacy DAG" — parity with the legacy DAG is not a
correctness criterion when the legacy DAG itself has the same defect.

U1: Variadic requiredParameterCount
  arity-metadata.ts:49 now subtracts the variadic slot from required
  count: total - optionalCount - (hasVariadic ? 1 : 0). f(int $req,
  ...$rest) requires 1 arg, not undefined. Adds overload-narrowing and
  lookup-core arity-filter changes so resolvers actually drop candidates
  that are definitively arity-incompatible (was silently rescuing
  empty filter sets even when bounds were known).

U2: True transitive trait MRO
  buildPhpMro now uses a BFS worklist (collectTransitiveTraits) to
  flatten the trait-of-trait DAG to fixpoint instead of expanding one
  level. Adds (trait_declaration body (use_declaration ...)) heritage
  query so trait-uses-trait IMPLEMENTS edges are emitted at all. Fixes
  3+ level trait chains silently dropping methods.

U3: parent:: bypasses composed traits
  ScopeResolver gains optional buildExtendsOnlyMro hook; PHP returns
  the unaugmented EXTENDS-only chain via buildPhpExtendsOnlyMro.
  MethodDispatchIndex gains optional extendsOnlyMroFor accessor wired
  through buildPopulatedMethodDispatch. Super-branch dispatch in
  receiver-bound-calls now walks extendsOnlyMroFor when present, so
  parent::method() routes to the parent class even when a composed
  trait shadows the same name. Other languages leave the hook
  undefined and fall back to mroFor unchanged.

U4: Namespace-aware free-call fallback
  ScopeResolver gains optional isCallableVisibleFromCaller predicate;
  pickUniqueGlobalCallable applies it to filter cross-namespace
  candidates the caller can't reach without a use-function import.
  PHP impl checks same-namespace OR explicit use-function presence,
  using a side-channel namespace cache populated by
  populatePhpNamespaceSiblings. Fixes the legacy DAG's namespace-
  blind false-positive emissions. Existing php-calls fixture
  updated: write_audit now correctly imports its targets via
  use-function rather than relying on the false-positive name-only
  match.

Tests:
  - 175/175 PHP both flag states (REGISTRY_PRIMARY_PHP=0 and =1)
  - 794/794 C#/Python/TypeScript/Go/C (no cross-language regression)
  - typecheck clean
  - 4 new fixtures: php-variadic-arity-minimum, php-transitive-traits,
    php-parent-vs-trait, php-namespace-fallback-isolation

Plan: docs/plans/2026-05-11-001-fix-php-resolver-semantic-defects-plan.md
2026-05-11 13:10:04 +01:00
uwence
a3af0dcce2
fix(docker): include duckdb installer script in runtime image (#1502) 2026-05-11 11:38:59 +01:00
Gergő Magyar
8733809b0a
Merge branch 'main' into lang/php-migration-v2 2026-05-11 09:38:23 +01:00
Rin
6a8947217c
fix(server): sanitize repo name to prevent argument injection (#1305)
* fix(server): sanitize repo name to prevent argument injection

Sanitizes the extracted repository name to prevent argument injection during git clone operations and ensures compatibility with various file systems.

1. Strips leading dashes to prevent git command-line argument injection.

2. Replaces unsafe directory characters with underscores.

3. Blocks path traversal segments ('.' and '..') and Windows reserved names.

4. Fixes ReDoS vulnerability in parseRepoNameFromUrl regex.

5. Added unit tests for sanitization and path traversal edge cases.

* fix(server): expand Windows reserved name check to include extensions

- Updated sanitizeRepoName to block Windows reserved names (CON, NUL, etc.) even when they have extensions (e.g., CON.txt).
- Corrected regex and added unit tests for these edge cases to resolve CI failures on Windows.
- Ref: https://github.com/abhigyanpatwari/GitNexus/pull/1305#issuecomment-4407200914

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 09:38:07 +01:00
Gergo Magyar
a8201cd9c9 feat(scope-resolution): add emitUnresolvedReceiverEdges hook for dynamic languages
Adds an optional post-resolution pass on the ScopeResolver contract:
when a member-call receiver cannot be typed by the scope chain (no
TypeRef), the language may emit CALLS edges via a workspace-wide
unique-name lookup. Runs after emitReceiverBoundCalls and before
emitFreeCallFallback, gated per-language.

PHP wires the hook to recover member calls on mixed/untyped parameters
(e.g. save(mixed $entity) calling $entity->getId()), restoring parity
with the legacy DAG. Re-enables the previously skipped save → getId
test; PHP suite is now 160/160 with no skips in both flag states.
2026-05-11 09:23:50 +01:00