Merge branch 'main' into fix/issue-1518-docker-local-path

This commit is contained in:
Gergő Magyar 2026-05-29 20:50:22 +01:00 • committed by GitHub
commit 8974df0d25
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
102 changed files with 7640 additions and 302 deletions

View file

@ -0,0 +1,43 @@
# GitNexus PR Reviewer Swarm — Claude Code adapter
This is the **Claude Code** entrypoint for the cross-CLI GitNexus PR reviewer swarm. The
review logic itself is CLI-neutral and lives in **[`pr-swarm-review/`](../pr-swarm-review/README.md)**
— that README is the canonical guide and covers every CLI (Claude Code, Gemini, Copilot,
Cursor, Codex, and any AGENTS.md-aware agent).
## Invocation (Claude Code)
```
/gitnexus-pr-swarm-review <PR URL or PR number>
```
Runs in **Swarm mode**: the coordinator skill dispatches the seven `gitnexus-*` subagents in
parallel (lanes 1–2 first, 3–6 in parallel, lane 7 last as a hard gate).
## Files in this adapter
| File | Role |
|------|------|
| `.claude/skills/gitnexus-pr-swarm-review/SKILL.md` | Coordinator — runs Swarm mode per `pr-swarm-review/orchestration.md` |
| `.claude/agents/gitnexus-*.md` | Seven thin subagent wrappers; each reads its canonical persona in `pr-swarm-review/personas/` |
Each subagent keeps valid Claude Code frontmatter (model, tools, etc.); the mechanical
verifier lanes (`test-ci-verifier`, `branch-hygiene-reviewer`) run on Haiku, the analytical
lanes on Sonnet.
## Key properties
- **Read-only.** Tools limited to Read/Grep/Glob/Bash, and every persona enforces an
explicit permitted/prohibited Bash list. No agent edits files, commits, or posts.
- **Evidence-grounded**; **missing visibility becomes verification work**; **manually invoked.**
## Editing
Edit review behavior in the canonical files under `pr-swarm-review/` (orchestration +
personas), **not** in these wrappers. After adding or editing files in `.claude/agents/`,
restart Claude Code so it reloads the agent definitions.
## Relationship to `/gitnexus-pr-review`
Coexists with the single-agent `/gitnexus-pr-review` skill (a linear checklist using GitNexus
MCP tools). This swarm is the multi-persona deep production-readiness review.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-branch-hygiene-reviewer
description: "GitNexus branch hygiene and mergeability reviewer. Use to classify merge state, conflicts, stale branches, merge-from-main commits, unrelated churn, mixed domains, and whether rebase or split is required."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-haiku-4-5-20251001
maxTurns: 30
---
# GitNexus Branch Hygiene & Mergeability Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/02-branch-hygiene-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-docs-dod-reviewer
description: "GitNexus docs and Definition-of-Done reviewer. Use to translate repo guidance, linked issues, changed domains, docs requirements, release notes, and acceptance criteria into a PR-specific DoD."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 30
---
# GitNexus Docs & Definition-of-Done Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/06-docs-dod-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-pr-facts-historian
description: "GitNexus PR facts and repository-history investigator. Use to gather PR identity, visible GitHub state, changed files, commits, linked issues, related PRs, historical fixes, regressions, stale follow-ups, and missing visibility."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 40
---
# GitNexus PR Facts & Repository-History Investigator
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/01-pr-facts-historian.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-risk-architect
description: "GitNexus production-risk reviewer. Use for risk-model-first review of changed files, runtime behavior, multi-domain changes, user impact, failure modes, compatibility, and merge-blocking risk."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 40
---
# GitNexus Production-Risk Architect
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/03-risk-architect.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-security-boundary-reviewer
description: "GitNexus security and trust-boundary reviewer. Use for auth, permissions, secrets, injection, unsafe parsing, external input handling, hidden Unicode, YAML/Docker/workflow risks, and suspicious non-ASCII hygiene."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 35
---
# GitNexus Security & Trust-Boundary Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/05-security-boundary-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-synthesis-critic
description: "GitNexus final review synthesis critic. Use to check whether the final PR review is evidence-grounded, risk-prioritized, GitNexus-specific, non-generic, and follows required verdict rules."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 25
---
# GitNexus Final-Review Synthesis Critic
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/07-synthesis-critic.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,24 @@
---
name: gitnexus-test-ci-verifier
description: "GitNexus test and CI reviewer. Use to verify whether changed behavior is covered by targeted tests, whether CI actually runs those tests, and whether workflow changes weaken validation."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-haiku-4-5-20251001
maxTurns: 35
---
# GitNexus Test & CI Verifier
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/04-test-ci-verifier.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.

View file

@ -0,0 +1,31 @@
---
name: gitnexus-pr-swarm-review
description: "Run a GitNexus production-readiness pull request review using a coordinated reviewer swarm."
---
# GitNexus PR Swarm Review (Claude Code adapter)
Use this skill to review a GitNexus pull request and produce a production-readiness review.
```
/gitnexus-pr-swarm-review <PR URL or PR number>
```
You are the **swarm coordinator**. The full review contract — lanes, dependencies,
classifications, output structure, finding format, hidden-Unicode checks, and behavior
rules — is the canonical, CLI-neutral spec:
**`pr-swarm-review/orchestration.md`** — read it now and follow it.
This adapter only pins the Claude Code specifics:
- **Run in Swarm mode.** Dispatch each lane as its own subagent via the Agent tool. The
seven subagents are the project agents named `gitnexus-*` (one per persona); each reads
its canonical persona under `pr-swarm-review/personas/`. Run lanes 1–2 first, lanes 3–6
in parallel after, and lane 7 last on the draft.
- **Lane 7 is a hard gate.** Do not emit the final review while the synthesis critic's
"Required corrections before posting" section is non-empty — revise and re-run it.
- Stay **read-only**: investigate and report; never edit, commit, or post.
Do not flatten the review into a generic checklist; delegate to the subagents and
synthesize per `orchestration.md`.

View file

@ -0,0 +1,17 @@
# GitNexus PR Swarm Review
You are the GitNexus PR review coordinator. Review the pull request named after this command
(a PR URL or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given,
ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.

View file

@ -0,0 +1,19 @@
description = "GitNexus production-readiness PR swarm review (Solo mode)"
prompt = """
You are the GitNexus PR review coordinator. Review this pull request: {{args}}
(a PR URL or number for https://github.com/abhigyanpatwari/GitNexus). If no target was
given, ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly. It is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1-2 first, then 3-6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty — revise and re-run it otherwise.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
"""

View file

@ -0,0 +1,19 @@
---
description: 'GitNexus production-readiness PR swarm review (Solo mode)'
mode: 'agent'
---
You are the GitNexus PR review coordinator. Review the pull request the user names (a PR URL
or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given, ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.

6
.gitignore vendored
View file

@ -91,11 +91,13 @@ gitnexus/vendor/**/node_modules/
.claude-flow/
.claude/agents/
.claude/agents/*
!.claude/agents/gitnexus-*.md
.claude/commands/
.claude/helpers
.claude/skills/
.claude/skills/*
!.claude/skills/gitnexus/
!.claude/skills/gitnexus-pr-swarm-review/
.history/

View file

@ -44,6 +44,18 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
- **Cursor:** `.cursor/index.mdc` (always-on); `.cursor/rules/*.mdc` (glob-scoped). Legacy `.cursorrules` deprecated.
- **GitNexus:** skills in `.claude/skills/gitnexus/`; MCP rules in `gitnexus:start` block below.
## PR Swarm Review (cross-CLI)
To run a production-readiness review of a GitNexus pull request from **any** AI CLI, follow
the canonical, CLI-neutral spec **[`pr-swarm-review/orchestration.md`](pr-swarm-review/orchestration.md)**
(seven read-only review personas under `pr-swarm-review/personas/`). It defines two
execution modes with the same output contract: **Swarm mode** (parallel subagents, e.g.
Claude Code) and **Solo mode** (one agent runs all lanes sequentially — Codex, Gemini,
Cursor, Copilot, or any agent reading this file). Per-CLI entrypoints are thin wrappers
listed in [`pr-swarm-review/README.md`](pr-swarm-review/README.md); edit review logic only
in the canonical files, never in the wrappers. The review is read-only — it never edits,
commits, or posts.
## Changelog
| Date | Version | Change |

View file

@ -1096,6 +1096,17 @@ const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions):
);
console.log(` ${repoPath}`);
// Persistent (non-scrolling) warning when FTS indexing was skipped — the
// progress-bar log() that fired mid-run has already scrolled away, so the
// degraded-search state must also appear in the final summary (#1161).
if (result.ftsSkipped) {
console.log(
`\n Warning: full-text/BM25 search is disabled — the LadybugDB FTS extension was unavailable.\n` +
` Install it once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) then rerun, or\n` +
` run \`gitnexus analyze --repair-fts\` when connected. Run \`gitnexus doctor\` for details.`,
);
}
try {
await fs.access(getGlobalRegistryPath());
} catch {

View file

@ -2,6 +2,7 @@ import { getRuntimeCapabilities, getRuntimeFingerprint } from '../core/platform/
import { resolveEmbeddingConfig } from '../core/embeddings/config.js';
import { isHttpMode } from '../core/embeddings/http-client.js';
import { checkLbugNative } from '../core/lbug/native-check.js';
import { getExtensionInstallPolicy } from '../core/lbug/extension-loader.js';
import { t } from './i18n/index.js';
function isCombiningMark(codePoint: number): boolean {
@ -74,6 +75,17 @@ export const doctorCommand = async () => {
console.log(` ${label('doctor.labels.fullTextSearch', 18)}${capabilities.fts}`);
console.log(` ${label('doctor.labels.vectorIndex', 18)}${capabilities.vector}`);
console.log(` ${label('doctor.labels.semanticMode', 18)}${capabilities.semanticMode}`);
// Surface the optional-extension install policy so offline users can see
// whether analyze/query will reach the network (extension.ladybugdb.com).
// Literal label (like the 'native' line) to avoid adding i18n keys.
const installPolicy = getExtensionInstallPolicy();
const policyHint =
installPolicy === 'load-only'
? ' (offline; load only, no network install)'
: installPolicy === 'never'
? ' (optional extensions disabled)'
: ' (installs missing extensions over network)';
console.log(` ${padDisplayEnd('Ext install:', 18)}${installPolicy}${policyHint}`);
console.log(
` ${label('doctor.labels.exactScanLimit', 18)}${t('doctor.chunks', { count: capabilities.exactScanLimit })}`,
);

View file

@ -43,20 +43,38 @@ import {
STALE_HASH_SENTINEL,
} from '../lbug/schema.js';
import { loadVectorExtension } from '../lbug/lbug-adapter.js';
import type { ExtensionInstallPolicy } from '../lbug/extension-loader.js';
import { getExactScanLimit } from '../platform/capabilities.js';
import { logger } from '../logger.js';
const isDev = process.env.NODE_ENV === 'development';
const vectorUnavailableMessage =
'VECTOR extension is unavailable for this LadybugDB runtime; semantic search will use exact scan when embeddings exist.';
'VECTOR extension unavailable; semantic embeddings fall back to exact scan. ' +
'To enable vector search, install it once with network access ' +
'(GITNEXUS_LBUG_EXTENSION_INSTALL=auto), or pre-install it for offline use. ' +
'Set GITNEXUS_LBUG_EXTENSION_INSTALL=never to skip installs and silence this.';
/**
* Resolve the extension-install policy for the embedding WRITE path (analyze).
*
* Generating embeddings is an explicit opt-in to a feature that requires the
* VECTOR extension, so when the operator has NOT pinned a policy we default to
* `auto` (one bounded, out-of-process INSTALL) — matching the documented
* "auto = default for analyze" intent in extension-loader.ts. An explicit
* GITNEXUS_LBUG_EXTENSION_INSTALL=load-only|never|auto always wins, so an
* offline or locked-down operator is never silently forced onto the network
* (the #1153 regression caused by hard-coding `auto` here). Read on every call
* (not memoized) so test env stubbing works.
*/
export const resolveEmbeddingInstallPolicy = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'auto';
};
const ensureVectorExtensionAvailable = async (): Promise<boolean> => {
const vectorReady = await loadVectorExtension();
if (!vectorReady) {
return false;
}
return true;
return loadVectorExtension(undefined, { policy: resolveEmbeddingInstallPolicy() });
};
/**
* Bump this when the embedding text template changes in a way that should
@ -257,7 +275,7 @@ export const runEmbeddingPipeline = async (
try {
const vectorAvailable = await ensureVectorExtensionAvailable();
if (!vectorAvailable && isDev) {
if (!vectorAvailable) {
logger.warn(vectorUnavailableMessage);
}
@ -584,7 +602,11 @@ export const semanticSearch = async (
string,
{ distance: number; chunkIndex: number; startLine: number; endLine: number }
>();
if (await loadVectorExtension()) {
// Query/read path: NEVER spawn a network INSTALL on a user query. If the
// VECTOR extension was not pre-installed, fall back to exact scan rather than
// blocking the query on a download (offline-first; see extension-loader.ts
// "load-only" — used by all serve/MCP query paths).
if (await loadVectorExtension(undefined, { policy: 'load-only' })) {
try {
bestChunks = await collectBestChunks(k, async (fetchLimit) => {
const vectorQuery = `

View file

@ -188,6 +188,16 @@ function makeContract(
export interface ProtoServiceInfo {
package: string;
/**
* Optional. Value of `option java_package = "..."` declared in the
* same `.proto` file, when present and different from `package`.
* Empty string when the option is absent or equals `package`. Used by
* `detectionToContract()` to translate a Java import path back to the
* proto package whenever the proto explicitly publishes its generated
* Java code under a different namespace (a common pattern in
* Google-style protobuf projects).
*/
javaPackage: string;
serviceName: string;
methods: string[];
protoPath: string;
@ -207,6 +217,19 @@ function extractProtoImports(content: string): string[] {
return imports;
}
/**
* Extract `option java_package = "..."` from a `.proto` file, if any.
* The Java code generator places generated `XxxGrpc.java` classes under
* this package (instead of the proto `package` declaration) when the
* option is set. Real-world projects (Google Cloud Java APIs, internal
* shaded SDKs) routinely use this to publish their Java artifacts under
* a corporate namespace different from the wire-protocol package.
*/
function extractJavaPackageOption(content: string): string {
const m = content.match(/^\s*option\s+java_package\s*=\s*"([\w.]+)"\s*;/m);
return m?.[1] ?? '';
}
function longestSharedSegmentRun(aPath: string, bPath: string): number {
const a = aPath.split('/').filter(Boolean);
const b = bPath.split('/').filter(Boolean);
@ -228,8 +251,18 @@ function longestSharedSegmentRun(aPath: string, bPath: string): number {
async function buildProtoContext(repoPath: string): Promise<{
packagesByProto: Map<string, string>;
servicesByName: Map<string, ProtoServiceInfo[]>;
/**
* Reverse index: `option java_package` value → ProtoServiceInfo[]
* declared in `.proto` files that ship under that Java namespace.
* Only populated when `java_package` is set AND differs from
* `package`. Lets `detectionToContract()` translate an import-derived
* Java package back to its source proto package whenever the proto
* is in the same repository.
*/
servicesByJavaPackage: Map<string, ProtoServiceInfo[]>;
}> {
const servicesByName = new Map<string, ProtoServiceInfo[]>();
const servicesByJavaPackage = new Map<string, ProtoServiceInfo[]>();
// `.gitnexusignore` / `.gitignore` honoured via the shared IgnoreService —
// see `filesystem-walker.ts` for the canonical pattern. Replaces a
// hardcoded `[node_modules, .git, vendor]` array; those names plus the
@ -292,6 +325,13 @@ async function buildProtoContext(repoPath: string): Promise<{
const content = contents.get(normalizedRel);
if (!content) continue;
const pkg = resolvePackage(normalizedRel);
const javaPkgOption = extractJavaPackageOption(content);
// Only retain `javaPackage` when it actively diverges from `pkg`.
// When equal (or absent), the import-derived path produces the
// same FQN as the proto-derived path, so no translation is needed
// and we keep the field empty to avoid populating the reverse
// index with redundant entries.
const javaPackage = javaPkgOption && javaPkgOption !== pkg ? javaPkgOption : '';
const serviceBlocks = extractServiceBlocks(content);
for (const block of serviceBlocks) {
@ -303,6 +343,7 @@ async function buildProtoContext(repoPath: string): Promise<{
}
const info: ProtoServiceInfo = {
package: pkg,
javaPackage,
serviceName: block.name,
methods,
protoPath: normalizedRel,
@ -310,10 +351,16 @@ async function buildProtoContext(repoPath: string): Promise<{
const existing = servicesByName.get(block.name) ?? [];
existing.push(info);
servicesByName.set(block.name, existing);
if (javaPackage) {
const byJava = servicesByJavaPackage.get(javaPackage) ?? [];
byJava.push(info);
servicesByJavaPackage.set(javaPackage, byJava);
}
}
}
return { packagesByProto, servicesByName };
return { packagesByProto, servicesByName, servicesByJavaPackage };
}
export async function buildProtoMap(repoPath: string): Promise<Map<string, ProtoServiceInfo[]>> {
@ -377,6 +424,7 @@ export class GrpcExtractor implements ContractExtractor {
const out: ExtractedContract[] = [];
const protoContext = await buildProtoContext(repoPath);
const protoMap = protoContext.servicesByName;
const javaPackageMap = protoContext.servicesByJavaPackage;
// ─── Proto files — definitive provider source ─────────────────
// When tree-sitter-proto is available, .proto files are handled by
@ -435,7 +483,7 @@ export class GrpcExtractor implements ContractExtractor {
continue;
}
for (const d of detections) {
const contract = this.detectionToContract(d, rel, protoMap);
const contract = this.detectionToContract(d, rel, protoMap, javaPackageMap);
if (contract) out.push(contract);
}
}
@ -449,12 +497,163 @@ export class GrpcExtractor implements ContractExtractor {
* either a service-level (`grpc::pkg.Svc/*`) or method-level
* (`grpc::pkg.Svc/Method`) contract id, and selecting confidence
* based on whether the proto map had an entry.
*
* Resolution order for the package prefix:
*
* 1. **Java-package translation** (when detection
* supplied a `protoPackage` from a Java import).
* A `.proto` in the SAME repo may set `option
* java_package = "..."` to publish its generated
* Java classes under a namespace different from
* the proto `package`. Real-world projects (e.g.
* Google Cloud Java APIs) routinely do this.
* When the import-derived package matches that
* `java_package` value, translate back to the
* proto `package` so the resulting contract id
* is wire-correct rather than Java-namespace.
*
* 2. **Per-repo proto map check** (when the same
* service name has `.proto` candidates in this
* repo). The proto file is the authoritative
* source. If the proto's `package` agrees with
* the import's `protoPackage`, both paths produce
* the same FQN — emit it. If they DISAGREE (e.g.
* a typo'd Java import, or a mismatched
* java_package the reverse index didn't catch),
* trust the proto map and warn — the import
* MUST NOT silently overwrite an authoritative
* proto package.
*
* 3. **Import-derived FQN fallback** (when neither
* a `java_package` translation nor a proto map
* candidate exists in this repo). Typical for the
* "client-jar" pattern, where a consumer repo
* depends on a published stub jar and never
* carries the originating `.proto`. Use the
* import path verbatim as the proto package. Note
* the known limitation: when the published proto
* sets `option java_package` differing from
* `package`, the resulting FQN reflects the Java
* namespace rather than the proto namespace and
* will not match a provider repo's contract id —
* we cannot translate without sight of the proto.
*
* 4. **Per-repo proto map (no import)** — the legacy
* path. Used when the plugin didn't supply
* `protoPackage` (no import statement, wildcard
* import only, or non-Java languages that haven't
* been retrofitted yet).
*
* 5. **Short-name fallback** — when none of the
* above resolves a package, emit a service-only
* short-name contract id (`grpc::Svc/*`),
* preserving the pre-fix behaviour.
*/
private detectionToContract(
d: GrpcDetection,
filePath: string,
protoMap: Map<string, ProtoServiceInfo[]>,
javaPackageMap: Map<string, ProtoServiceInfo[]>,
): ExtractedContract | null {
if (d.protoPackage) {
// Step 1: java_package translation. The import-derived package
// may be the `option java_package` value of a `.proto` in the
// SAME repo. Look it up and, if found for the same service name,
// use the underlying proto `package` to build a wire-correct
// contract id.
const javaCandidates = javaPackageMap.get(d.protoPackage) ?? [];
const javaTranslated = javaCandidates.find((p) => p.serviceName === d.serviceName);
if (javaTranslated) {
const cid = d.methodName
? contractId(javaTranslated.package, d.serviceName, d.methodName)
: serviceContractId(javaTranslated.package, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: javaTranslated.package,
protoPackageSource: 'import-translated',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Step 2: proto map cross-check. When this repo also carries a
// `.proto` defining the same short service name, the proto is
// authoritative and decides the package. The import is only used
// to disambiguate among same-short-name candidates when the
// resolution heuristic can't pick a unique winner on path alone.
const candidates = protoMap.get(d.serviceName) ?? [];
if (candidates.length > 0) {
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
if (proto === null) {
// Ambiguous proto resolution; resolveProtoConflict already warned.
return null;
}
const protoPkg = proto.package;
if (protoPkg === d.protoPackage) {
// Both paths agree.
const cid = d.methodName
? contractId(protoPkg, d.serviceName, d.methodName)
: serviceContractId(protoPkg, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: protoPkg,
protoPackageSource: 'import',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Disagreement. Trust the proto file and emit a warning so
// operators can investigate the import. This protects against
// the symmetric Finding 2 case: a stale or typo'd Java import
// silently corrupting the contract id of a service whose
// `.proto` lives in the same repo.
logger.warn(
`[grpc-extractor] Java import package "${d.protoPackage}" for service ` +
`"${d.serviceName}" disagrees with local proto package "${protoPkg}" at ` +
`${filePath}; using proto package as authoritative source`,
);
const cid = d.methodName
? contractId(protoPkg, d.serviceName, d.methodName)
: serviceContractId(protoPkg, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: protoPkg,
protoPackageSource: 'proto-override',
importPackage: d.protoPackage,
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Step 3: import-derived fallback. No `.proto` in this repo
// names the service, and no `java_package` reverse-lookup
// matched. Emit the FQN with the import-derived package. This
// is the typical client-jar consumer path.
//
// Known limitation: when the published proto sets
// `option java_package` to a value that differs from
// `package`, this path produces a contract id that reflects
// the Java namespace, not the proto namespace, and will not
// match a provider repo. Resolving that case requires
// group-level proto knowledge, which is intentionally out of
// scope for this fix.
const cid = d.methodName
? contractId(d.protoPackage, d.serviceName, d.methodName)
: serviceContractId(d.protoPackage, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: d.protoPackage,
protoPackageSource: 'import',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Steps 4 + 5: legacy per-repo proto map resolution (no import).
const candidates = protoMap.get(d.serviceName) ?? [];
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
// If there were proto candidates but resolution was ambiguous, skip

View file

@ -78,6 +78,33 @@ const STUB_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
// `import <pkg>.<XxxGrpc>;` — captures the proto package of the
// imported gRPC class (e.g. `cn.unipus.ucf.admin.proto.client.service`
// for `import cn.unipus.ucf.admin.proto.client.service.ContentRpcServiceGrpc`).
// Used by `scan` to build a per-file `XxxGrpc → fullPackage` map so
// consumer-side detections can carry a fully-qualified contract id
// even when the consumer repo does not contain any `.proto` files.
//
// `import static …` is excluded by tree-sitter shape: the `name:`
// field is only present on the non-static form. `import w.x.*;` is
// also excluded for the same reason — wildcard imports have an
// `asterisk` child instead of a named identifier.
const GRPC_CLASS_IMPORT_PATTERNS = compilePatterns({
name: 'java-grpc-class-import',
language: Java,
patterns: [
{
meta: {},
query: `
(import_declaration
(scoped_identifier
scope: (_) @import_pkg
name: (identifier) @import_name (#match? @import_name "Grpc$")))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Check whether a `class_declaration` node has a `@GrpcService`
* annotation in its modifiers list. In tree-sitter-java, class-level
@ -118,6 +145,39 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
const out: GrpcDetection[] = [];
const emittedClassIds = new Set<number>();
// ─── Build per-file gRPC class import map ───────────────────────
// Maps `XxxGrpc` (short class name) → fully-qualified proto package
// (e.g. `cn.unipus.ucf.admin.proto.client.service`). Used below to
// tag both provider and consumer detections with a `protoPackage`
// so the orchestrator can build a fully-qualified contract id
// without depending on the current repo carrying any `.proto`
// files. This is the key fix for client-jar consumer repos.
//
// Same-short-name disambiguation: when two distinct `import` lines
// bring different `XxxGrpc` classes from different packages into
// the same file (rare for grpc — the second import would be a
// compile error in Java), the last one wins. Java's compiler
// forbids that case so we don't bother modelling it.
const grpcClassImports = new Map<string, string>();
for (const match of runCompiledPatterns(GRPC_CLASS_IMPORT_PATTERNS, tree)) {
const pkgNode = match.captures.import_pkg;
const nameNode = match.captures.import_name;
if (!pkgNode || !nameNode) continue;
grpcClassImports.set(nameNode.text, pkgNode.text);
}
/**
* Resolve the fully-qualified proto package for a short service
* name in this file. Looks up `<serviceName>Grpc` in the import
* map; returns `undefined` when the class is referenced via a
* fully-qualified name on every call site (no import line) or
* when only a wildcard import is present. The orchestrator falls
* back to the per-repo proto map in that case, preserving the
* pre-fix behaviour.
*/
const protoPackageFor = (serviceName: string): string | undefined =>
grpcClassImports.get(`${serviceName}Grpc`);
// ─── Providers: scoped form (`...Grpc.XxxImplBase`) ─────────────
for (const match of runCompiledPatterns(SCOPED_IMPL_BASE_PATTERNS, tree)) {
const classNode = match.captures.class;
@ -127,6 +187,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'provider',
serviceName,
@ -134,6 +195,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
...(protoPackage ? { protoPackage } : {}),
});
}
@ -147,6 +209,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'provider',
serviceName,
@ -154,6 +217,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
...(protoPackage ? { protoPackage } : {}),
});
}
@ -164,6 +228,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
const grpcMatch = GRPC_SUFFIX_RE.exec(grpcClsNode.text);
if (!grpcMatch) continue;
const serviceName = grpcMatch[1];
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'consumer',
serviceName,
@ -171,6 +236,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: 'java_stub',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
...(protoPackage ? { protoPackage } : {}),
});
}

View file

@ -36,6 +36,18 @@ export interface GrpcDetection {
confidenceWithProto: number;
/** Confidence when the proto map has no entry. */
confidenceWithoutProto: number;
/**
* Optional. Fully-qualified proto package the detection's service
* belongs to (e.g. `cn.unipus.ucf.admin.proto.client.service`),
* derived directly from the source file's import statements when
* available. When set, the orchestrator uses this package to build
* the contract id INSTEAD of consulting the per-repo proto map —
* letting consumer repos that don't carry `.proto` files (the
* client-jar architecture used by most Java gRPC microservices)
* still emit a fully-qualified contract id that matches the
* provider repo's contract id verbatim.
*/
protoPackage?: string;
}
/**

View file

@ -8,7 +8,13 @@ import { PYTHON_HTTP_PLUGIN } from './python.js';
import { PHP_HTTP_PLUGIN } from './php.js';
import { JAVASCRIPT_HTTP_PLUGIN, TYPESCRIPT_HTTP_PLUGIN, TSX_HTTP_PLUGIN } from './node.js';
export type { HttpDetection, HttpLanguagePlugin, HttpRole } from './types.js';
export type {
HttpDetection,
HttpFileDetections,
HttpLanguagePlugin,
HttpRole,
HttpScanInput,
} from './types.js';
/**
* File-extension → HTTP language plugin registry. The top-level

View file

@ -6,13 +6,21 @@ import {
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
import type {
HttpDetection,
HttpFileDetections,
HttpLanguagePlugin,
HttpScanInput,
} from './types.js';
/**
* Java HTTP plugin. Handles:
* - Spring `@RequestMapping` class prefixes + `@(Get|Post|...)Mapping` method annotations
* - Spring `RestTemplate.getForObject/...`, `WebClient.method(HttpMethod.X, ...)`
* - Spring `RestTemplate.getForObject/...`, `exchange(...)`
* - Spring `WebClient.method(HttpMethod.X, ...)`, `WebClient.get().uri(...)`
* - OkHttp `new Request.Builder().url("...")`
* - OpenFeign interfaces with Spring MVC method annotations
* - Java / Apache HttpClient literal request construction
*
* The plugin runs two pattern bundles: one to collect class-level
* `@RequestMapping` prefixes keyed by the enclosing class node, and a
@ -43,31 +51,132 @@ const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
// route prefixes — e.g. `produces = "application/json"` would corrupt
// every method route under that controller). The sibling
// `topic-patterns/java.ts` uses the same `key:` constraint approach.
const SPRING_CLASS_PREFIX_PATTERNS = compilePatterns({
name: 'java-spring-class-prefix',
interface SpringRouteBinding {
method: string;
path: string;
}
interface SpringMethodInfo {
name: string;
routes: SpringRouteBinding[];
}
interface SpringTypeInfo {
filePath: string;
kind: 'class' | 'interface';
name: string;
classPrefix: string;
implementedInterfaces: string[];
isController: boolean;
methods: SpringMethodInfo[];
}
// ─── Provider: Spring class/interface-level @RequestMapping prefix ───
const SPRING_TYPE_PREFIX_PATTERNS = compilePatterns({
name: 'java-spring-type-prefix',
language: Java,
patterns: [
{
meta: {},
query: `
(class_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @class
[
(class_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @type
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @type
]
`,
},
{
meta: {},
query: `
(class_declaration
[
(class_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#match? @key "^(path|value)$")
value: (string_literal) @prefix))))) @type
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#match? @key "^(path|value)$")
value: (string_literal) @prefix))))) @type
]
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const SPRING_TYPE_DECLARATION_PATTERNS = compilePatterns({
name: 'java-spring-type-declaration',
language: Java,
patterns: [
{
meta: {},
query: `
[
(class_declaration name: (identifier) @type_name) @type
(interface_declaration name: (identifier) @type_name) @type
]
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: OpenFeign interface-level prefixes ───────────────────
// Feign's `name`/`value` attributes identify a service, not an HTTP path,
// so only `path` is used as a URL prefix. `@RequestMapping` on a Feign
// interface is also common and does carry a path prefix.
const FEIGN_INTERFACE_PREFIX_PATTERNS = compilePatterns({
name: 'java-feign-interface-prefix',
language: Java,
patterns: [
{
meta: {},
query: `
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "FeignClient")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#eq? @key "path")
value: (string_literal) @prefix))))) @interface
`,
},
{
meta: {},
query: `
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @interface
`,
},
{
meta: {},
query: `
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#match? @key "^(path|value)$")
value: (string_literal) @prefix))))) @class
value: (string_literal) @prefix))))) @interface
`,
},
],
@ -116,6 +225,8 @@ const SPRING_METHOD_ROUTE_PATTERNS = compilePatterns({
// RestTemplate.put → PUT
// RestTemplate.delete → DELETE
// RestTemplate.patchForObject → PATCH
// Source-scan only: receiver must be named exactly `restTemplate`.
// Fields, `this.restTemplate`, aliases, and other injection names are deferred.
const REST_TEMPLATE_TO_HTTP: Record<string, string> = {
getForObject: 'GET',
getForEntity: 'GET',
@ -146,22 +257,48 @@ const REST_TEMPLATE_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<RestTemplateMeta>);
// ─── Consumer: Spring WebClient — webClient.method(HttpMethod.X, "path") ─
const WEB_CLIENT_PATTERNS = compilePatterns({
name: 'java-web-client',
const REST_TEMPLATE_EXCHANGE_PATTERNS = compilePatterns({
name: 'java-rest-template-exchange',
language: Java,
patterns: [
{
meta: { framework: 'spring-rest-template' },
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "restTemplate")
name: (identifier) @method (#eq? @method "exchange")
arguments: (argument_list
. (string_literal) @path
(field_access
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
field: (identifier) @http_method)))
`,
},
],
} satisfies LanguagePatterns<RestTemplateMeta>);
const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
get: 'GET',
post: 'POST',
put: 'PUT',
delete: 'DELETE',
patch: 'PATCH',
};
const WEB_CLIENT_SHORT_FORM_PATTERNS = compilePatterns({
name: 'java-web-client-short-form',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "webClient")
name: (identifier) @method (#eq? @method "method")
arguments: (argument_list
(field_access
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
field: (identifier) @http_method)
(string_literal) @path))
object: (method_invocation
object: (identifier) @obj (#eq? @obj "webClient")
name: (identifier) @verb (#match? @verb "^(get|post|put|delete|patch)$")
arguments: (argument_list))
name: (identifier) @uri_method (#eq? @uri_method "uri")
arguments: (argument_list . (string_literal) @path))
`,
},
],
@ -188,10 +325,58 @@ const OK_HTTP_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
const JAVA_HTTP_CLIENT_PATTERNS = compilePatterns({
name: 'java-http-client',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (method_invocation
object: (method_invocation
object: (identifier) @builderCls (#eq? @builderCls "HttpRequest")
name: (identifier) @newBuilder (#eq? @newBuilder "newBuilder")
arguments: (argument_list))
name: (identifier) @uri_method (#eq? @uri_method "uri")
arguments: (argument_list
(method_invocation
object: (identifier) @uriCls (#eq? @uriCls "URI")
name: (identifier) @create (#eq? @create "create")
arguments: (argument_list . (string_literal) @path))))
name: (identifier) @http_method (#match? @http_method "^(GET|POST|PUT|DELETE)$"))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const APACHE_HTTP_CLIENT_TO_HTTP: Record<string, string> = {
HttpGet: 'GET',
HttpPost: 'POST',
HttpPut: 'PUT',
HttpDelete: 'DELETE',
HttpPatch: 'PATCH',
};
const APACHE_HTTP_CLIENT_PATTERNS = compilePatterns({
name: 'java-apache-http-client',
language: Java,
patterns: [
{
meta: {},
query: `
(object_creation_expression
type: (type_identifier) @type (#match? @type "^Http(Get|Post|Put|Delete|Patch)$")
arguments: (argument_list . (string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Find the nearest enclosing class_declaration ancestor for a node, or
* null if the node is top-level. Tree-sitter's SyntaxNode.parent walks
* one level at a time.
* Find the nearest enclosing class/interface declaration ancestor for
* a node, or null if the node is top-level. Tree-sitter's
* SyntaxNode.parent walks one level at a time.
*/
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
@ -202,6 +387,15 @@ function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
return null;
}
function findEnclosingInterface(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'interface_declaration') return cur;
cur = cur.parent;
}
return null;
}
/**
* Join a class-level prefix and a method-level path into a single URL
* path. Mirrors the semantics of the original regex implementation:
@ -215,6 +409,184 @@ function joinPath(prefix: string, methodPath: string): string {
return `/${cleanPrefix}/${cleanSub}`;
}
function getNodeName(node: Parser.SyntaxNode): string | null {
return node.childForFieldName('name')?.text ?? null;
}
function hasAnnotation(node: Parser.SyntaxNode, names: string | readonly string[]): boolean {
const modifiers = node.namedChildren.find((child) => child.type === 'modifiers');
if (!modifiers) return false;
const allowed = new Set(typeof names === 'string' ? [names] : names);
const stack = [...modifiers.namedChildren];
while (stack.length > 0) {
const cur = stack.pop()!;
const annotationName = cur.childForFieldName('name')?.text ?? '';
const simpleName = annotationName.split('.').pop() ?? annotationName;
if (
(cur.type === 'annotation' || cur.type === 'marker_annotation') &&
(allowed.has(annotationName) || allowed.has(simpleName))
) {
return true;
}
stack.push(...cur.namedChildren);
}
return false;
}
function collectTypePrefixes(tree: Parser.Tree): Map<number, string> {
const prefixByTypeId = new Map<number, string>();
for (const match of runCompiledPatterns(SPRING_TYPE_PREFIX_PATTERNS, tree)) {
const prefixNode = match.captures.prefix;
const typeNode = match.captures.type;
if (!prefixNode || !typeNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix !== null) prefixByTypeId.set(typeNode.id, prefix);
}
return prefixByTypeId;
}
function collectMethodRoutes(tree: Parser.Tree): Map<number, SpringRouteBinding[]> {
const routesByMethodId = new Map<number, SpringRouteBinding[]>();
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
const annNode = match.captures.ann;
const pathNode = match.captures.path;
const methodNode = match.captures.method;
if (!annNode || !pathNode || !methodNode) continue;
const httpMethod = METHOD_ANNOTATION_TO_HTTP[annNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
const routes = routesByMethodId.get(methodNode.id) ?? [];
routes.push({ method: httpMethod, path: rawPath });
routesByMethodId.set(methodNode.id, routes);
}
return routesByMethodId;
}
function collectDirectMethods(typeNode: Parser.SyntaxNode): Parser.SyntaxNode[] {
const out: Parser.SyntaxNode[] = [];
const visit = (node: Parser.SyntaxNode): void => {
for (const child of node.namedChildren) {
if (child.type === 'method_declaration') {
out.push(child);
continue;
}
if (
child !== typeNode &&
(child.type === 'class_declaration' || child.type === 'interface_declaration')
) {
continue;
}
visit(child);
}
};
visit(typeNode);
return out;
}
function collectImplementedInterfaces(typeNode: Parser.SyntaxNode): string[] {
const interfacesNode = typeNode.childForFieldName('interfaces');
if (!interfacesNode) return [];
const out: string[] = [];
const visit = (node: Parser.SyntaxNode): void => {
if (node.type === 'type_identifier' || node.type === 'scoped_type_identifier') {
out.push(node.text.split('.').pop() ?? node.text);
return;
}
for (const child of node.namedChildren) visit(child);
};
visit(interfacesNode);
return out;
}
function collectSpringTypes(filePath: string, tree: Parser.Tree): SpringTypeInfo[] {
const prefixByTypeId = collectTypePrefixes(tree);
const routesByMethodId = collectMethodRoutes(tree);
const out: SpringTypeInfo[] = [];
for (const match of runCompiledPatterns(SPRING_TYPE_DECLARATION_PATTERNS, tree)) {
const typeNode = match.captures.type;
const typeNameNode = match.captures.type_name;
if (!typeNode || !typeNameNode) continue;
const kind = typeNode.type === 'interface_declaration' ? 'interface' : 'class';
const methods = collectDirectMethods(typeNode)
.map((methodNode) => ({
name: getNodeName(methodNode),
routes: routesByMethodId.get(methodNode.id) ?? [],
}))
.filter((method): method is SpringMethodInfo => method.name !== null);
out.push({
filePath,
kind,
name: typeNameNode.text,
classPrefix: prefixByTypeId.get(typeNode.id) ?? '',
implementedInterfaces: kind === 'class' ? collectImplementedInterfaces(typeNode) : [],
isController: kind === 'class' && hasAnnotation(typeNode, ['RestController', 'Controller']),
methods,
});
}
return out;
}
function scanSpringProject(files: readonly HttpScanInput[]): HttpFileDetections[] {
const types = files.flatMap((file) => collectSpringTypes(file.filePath, file.tree));
const interfaceRoutes = new Map<string, Map<string, SpringRouteBinding[]> | null>();
for (const type of types) {
if (type.kind !== 'interface') continue;
if (interfaceRoutes.has(type.name)) {
interfaceRoutes.set(type.name, null);
continue;
}
const methodMap = new Map<string, SpringRouteBinding[]>();
for (const method of type.methods) {
const routes = method.routes.map((route) => ({
method: route.method,
path: type.classPrefix ? joinPath(type.classPrefix, route.path) : route.path,
}));
if (routes.length > 0) methodMap.set(method.name, routes);
}
interfaceRoutes.set(type.name, methodMap);
}
const detectionsByFile = new Map<string, HttpDetection[]>();
for (const type of types) {
if (type.kind !== 'class' || !type.isController) continue;
for (const method of type.methods) {
if (method.routes.length > 0) continue;
const inheritedRoutes = type.implementedInterfaces.flatMap((interfaceName) => {
const routeMap = interfaceRoutes.get(interfaceName);
if (!routeMap) return [];
const routes = routeMap.get(method.name) ?? [];
return routes.map((route) => ({
method: route.method,
path: joinPath(type.classPrefix, route.path),
}));
});
for (const route of inheritedRoutes) {
const detections = detectionsByFile.get(type.filePath) ?? [];
detections.push({
role: 'provider',
framework: 'spring',
method: route.method,
path: route.path,
name: method.name,
confidence: 0.8,
});
detectionsByFile.set(type.filePath, detections);
}
}
}
return [...detectionsByFile.entries()].map(([filePath, detections]) => ({
filePath,
detections,
}));
}
export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'java-http',
language: Java,
@ -222,13 +594,16 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
const out: HttpDetection[] = [];
// ─── Providers: Spring class prefix + method annotations ────────
const prefixByClassId = new Map<number, string>();
for (const match of runCompiledPatterns(SPRING_CLASS_PREFIX_PATTERNS, tree)) {
const prefixByTypeId = collectTypePrefixes(tree);
const feignPrefixByInterfaceId = new Map<number, string>();
for (const match of runCompiledPatterns(FEIGN_INTERFACE_PREFIX_PATTERNS, tree)) {
const prefixNode = match.captures.prefix;
const classNode = match.captures.class;
if (!prefixNode || !classNode) continue;
const interfaceNode = match.captures.interface;
if (!prefixNode || !interfaceNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix !== null) prefixByClassId.set(classNode.id, prefix);
if (prefix !== null && !feignPrefixByInterfaceId.has(interfaceNode.id))
feignPrefixByInterfaceId.set(interfaceNode.id, prefix);
}
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
@ -241,8 +616,23 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
const enclosingInterface = findEnclosingInterface(methodNode);
if (enclosingInterface && hasAnnotation(enclosingInterface, 'FeignClient')) {
const prefix = feignPrefixByInterfaceId.get(enclosingInterface.id) ?? '';
const fullPath = joinPath(prefix, rawPath);
out.push({
role: 'consumer',
framework: 'openfeign',
method: httpMethod,
path: fullPath,
name: nameNode?.text ?? null,
confidence: 0.7,
});
continue;
}
const enclosingClass = findEnclosingClass(methodNode);
const prefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
if (!enclosingClass) continue;
const prefix = prefixByTypeId.get(enclosingClass.id) ?? '';
const fullPath = joinPath(prefix, rawPath);
out.push({
role: 'provider',
@ -273,8 +663,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: WebClient.method(HttpMethod.X, "path") ──────────
for (const match of runCompiledPatterns(WEB_CLIENT_PATTERNS, tree)) {
for (const match of runCompiledPatterns(REST_TEMPLATE_EXCHANGE_PATTERNS, tree)) {
const httpMethodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!httpMethodNode || !pathNode) continue;
@ -282,7 +671,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
framework: 'spring-rest-template',
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
@ -290,6 +679,28 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: WebClient.get().uri("path") short form ─────────
// Source-scan only: receiver must be named exactly `webClient`.
// The real long-form chain `webClient.method(HttpMethod.X).uri("/x")`
// needs multi-hop chain analysis and is intentionally deferred.
for (const match of runCompiledPatterns(WEB_CLIENT_SHORT_FORM_PATTERNS, tree)) {
const verbNode = match.captures.verb;
const pathNode = match.captures.path;
if (!verbNode || !pathNode) continue;
const httpMethod = WEB_CLIENT_SHORT_TO_HTTP[verbNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
const pathNode = match.captures.path;
@ -306,6 +717,45 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: Java HttpClient request builder ─────────────────
// Java's builder exposes GET/POST/PUT/DELETE helpers. PATCH uses
// `.method("PATCH", body)`, which is intentionally deferred.
for (const match of runCompiledPatterns(JAVA_HTTP_CLIENT_PATTERNS, tree)) {
const httpMethodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!httpMethodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'java-http-client',
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.65,
});
}
// ─── Consumers: Apache HttpClient request constructors ──────────
for (const match of runCompiledPatterns(APACHE_HTTP_CLIENT_PATTERNS, tree)) {
const typeNode = match.captures.type;
const pathNode = match.captures.path;
if (!typeNode || !pathNode) continue;
const httpMethod = APACHE_HTTP_CLIENT_TO_HTTP[typeNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'apache-http-client',
method: httpMethod,
path,
name: null,
confidence: 0.65,
});
}
return out;
},
scanProject: scanSpringProject,
};

View file

@ -17,18 +17,22 @@ import type { HttpDetection, HttpLanguagePlugin } from './types.js';
* named annotation arguments (`@GetMapping(value = "/x")` and
* `@GetMapping(path = "/x")`) are supported.
*
* **Consumers** (this PR) — three call-site patterns common in Kotlin
* **Consumers** — four call-site patterns common in Kotlin
* Spring projects:
*
* 1. `restTemplate.getForObject("/x", ...)` and friends
* 2. `webClient.get().uri("/x")` (short form, 1 verb hop + 1 uri hop)
* 3. `Request.Builder().url("/x")` (OkHttp)
* 1. `restTemplate.getForObject("/x", ...)` and friends (#1855)
* 2. `webClient.get().uri("/x")` — short form (#1855)
* 3. `Request.Builder().url("/x")` — OkHttp (#1855)
* 4. `webClient.method(HttpMethod.X).uri("/y")` — long form (this PR)
*
* The long-form `webClient.method(HttpMethod.X).uri("/y")` chain is
* intentionally deferred to a follow-up: it requires walk-up logic
* to recover the verb from a sibling `call_expression`, and we can
* land 80% of real-world Kotlin Spring consumer coverage with the
* three simpler patterns above.
* The long form puts the verb on a sibling `call_expression` two hops
* away from the path. Rather than introducing imperative walk-up logic,
* we use a single deeper tree-sitter query that matches the full chain
* structurally — see `WEB_CLIENT_LONG_PATTERNS` below. The verb is
* captured directly as the `simple_identifier` of `HttpMethod.X`, so
* variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
* are intentionally NOT picked up — those need a graph-aware resolver
* and are out of scope for source-scan.
*
* tree-sitter-kotlin (fwcd) AST shapes used here:
* class_declaration
@ -109,6 +113,16 @@ const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
patch: 'PATCH',
};
/**
* Allowed HTTP verbs for the WebClient long-form path
* `webClient.method(HttpMethod.X).uri("/y")`. Compiled once at module
* load (instead of inside the scan loop) per maintainer feedback on
* PR #1884. Mirrors the keys of `WEB_CLIENT_SHORT_TO_HTTP` above —
* keeping HEAD/OPTIONS/TRACE intentionally excluded for symmetry
* with the short form and the Java plugin.
*/
const WEB_CLIENT_LONG_VERB_RE = /^(GET|POST|PUT|DELETE|PATCH)$/;
/**
* Build the plugin only if the Kotlin grammar is available. Compiling
* the queries against a null grammar would throw at module load time
@ -265,8 +279,9 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
// - outer call's first value_argument is a string literal
//
// The long-form `webClient.method(HttpMethod.GET).uri("/x")` chain
// uses an extra navigation hop and an enum field access — it's
// intentionally out of scope here (see file header).
// uses an extra navigation hop and an enum field access — handled
// by `WEB_CLIENT_LONG_PATTERNS` below, separately so each query is
// straightforward to reason about.
const WEB_CLIENT_SHORT_PATTERNS = compilePatterns({
name: 'kotlin-web-client-short',
language,
@ -290,6 +305,59 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: Spring WebClient (long form) ───────────────────────────
// The fluent long form passes the verb as a `HttpMethod.X` enum field
// access through `.method(...)`, then carries the path on a separate
// `.uri(...)` hop further down the chain:
//
// webClient.method(HttpMethod.GET).uri("/x").retrieve().awaitBody<T>()
//
// Compared to the short form there are two extra structural hops:
// - the inner `.method(...)` `call_expression` has a `value_argument`
// whose payload is itself a `navigation_expression` (HttpMethod → .GET)
// - the outer `.uri(...)` is reached via one more
// `navigation_expression` wrapping that inner call
//
// We capture the verb at the `simple_identifier` under `HttpMethod`'s
// `navigation_suffix`. That `simple_identifier` is the literal field
// name (`GET`, `POST`, ...) used in source — Kotlin enum fields by
// convention are upper-case, matching `HttpMethod` from
// `org.springframework.http`. We forward the captured text as-is.
//
// Variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
// do NOT match — they fail the `(navigation_expression ...)` shape
// because the value_argument carries a bare `simple_identifier` instead
// of a `HttpMethod.X` field access. This is intentional: source-scan
// can't follow the binding without graph context. Pinned by an
// anti-overreach test in the consumer suite.
const WEB_CLIENT_LONG_PATTERNS = compilePatterns({
name: 'kotlin-web-client-long',
language,
patterns: [
{
meta: {},
query: `
(call_expression
(navigation_expression
(call_expression
(navigation_expression
(simple_identifier) @obj (#eq? @obj "webClient")
(navigation_suffix
(simple_identifier) @method_call (#eq? @method_call "method")))
(call_suffix
(value_arguments
. (value_argument
(navigation_expression
(simple_identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
(navigation_suffix (simple_identifier) @verb))))))
(navigation_suffix (simple_identifier) @uri (#eq? @uri "uri")))
(call_suffix
(value_arguments . (value_argument . (string_literal) @path))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: OkHttp Request.Builder().url("/x") ─────────────────────
// Kotlin parses `Request.Builder()` as a `call_expression` whose
// callee is a `navigation_expression` (Request → .Builder), NOT as
@ -437,6 +505,33 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
});
}
// ─── Consumers: WebClient long form (.method(HttpMethod.X) → .uri) ─
for (const match of runCompiledPatterns(WEB_CLIENT_LONG_PATTERNS, tree)) {
const verbNode = match.captures.verb;
const pathNode = match.captures.path;
if (!verbNode || !pathNode) continue;
// The captured text is the literal `HttpMethod.X` field name.
// Spring's `org.springframework.http.HttpMethod` defines GET,
// POST, PUT, DELETE, PATCH, HEAD, OPTIONS, TRACE — we only
// emit for the five verbs we already handle elsewhere, so
// exotic ones are silently skipped (consistent with the
// short form's WEB_CLIENT_SHORT_TO_HTTP guard). The accepted
// verb regex is hoisted to module scope (see
// `WEB_CLIENT_LONG_VERB_RE` near the top of this file).
const verbText = verbNode.text;
if (!WEB_CLIENT_LONG_VERB_RE.test(verbText)) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: verbText,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
const pathNode = match.captures.path;

View file

@ -6,7 +6,7 @@ import {
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
import type { HttpDetection, HttpLanguagePlugin, RepoContext } from './types.js';
/**
* Python HTTP plugin. Handles:
@ -29,9 +29,13 @@ const FASTAPI_VERBS: Record<string, string> = {
patch: 'PATCH',
};
// ─── Provider: FastAPI @app.get/... ──────────────────────────────────
const FASTAPI_PATTERNS = compilePatterns({
name: 'python-fastapi',
// ─── Provider: FastAPI @app.<verb> / @router.<verb> ──────────────────
// Two separate patterns so we can tag detections by decorator object.
// Only `@router.*` detections participate in `include_router(prefix=)`
// path-prefix joining (see `PythonRepoContext` + `joinPrefix`); `@app.*`
// routes already carry their final path verbatim.
const FASTAPI_APP_PATTERNS = compilePatterns({
name: 'python-fastapi-app',
language: Python,
patterns: [
{
@ -48,6 +52,138 @@ const FASTAPI_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
const FASTAPI_ROUTER_PATTERNS = compilePatterns({
name: 'python-fastapi-router',
language: Python,
patterns: [
{
meta: {},
query: `
(decorator
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "router")
attribute: (identifier) @method (#match? @method "^(get|post|put|delete|patch)$"))
arguments: (argument_list . (string) @path)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── include_router(<router_obj>, prefix='/x') across the repo ────────
// Two shapes are common:
// app.include_router(assistant.router, prefix='/ai')
// app.include_router(my_router, prefix='/ai')
// The first names the originating module via `<module>.router`; the second
// references a name imported into the host file. We capture both.
const INCLUDE_ROUTER_ATTR_PATTERNS = compilePatterns({
name: 'python-fastapi-include-router-attr',
language: Python,
patterns: [
{
meta: {},
// Match any `<host>.include_router(<module>.router, ..., prefix='/x')`
// call. We deliberately do NOT pin `<host>` to the literal name `app`
// — production code routinely uses `api`, `application`, `asgi_app`,
// etc. The shape (`include_router` invoked with a router argument and
// a `prefix=` keyword) is specific enough on its own; restricting the
// host produces false negatives without removing meaningful false
// positives.
query: `
(call
function: (attribute
attribute: (identifier) @incl (#eq? @incl "include_router"))
arguments: (argument_list
(attribute
object: (identifier) @router_module
attribute: (identifier) @router_attr (#eq? @router_attr "router"))
(keyword_argument
name: (identifier) @kw (#eq? @kw "prefix")
value: (string) @prefix)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const INCLUDE_ROUTER_NAME_PATTERNS = compilePatterns({
name: 'python-fastapi-include-router-name',
language: Python,
patterns: [
{
meta: {},
// Same `<host>` rationale as INCLUDE_ROUTER_ATTR_PATTERNS — see above.
query: `
(call
function: (attribute
attribute: (identifier) @incl (#eq? @incl "include_router"))
arguments: (argument_list
(identifier) @router_name
(keyword_argument
name: (identifier) @kw (#eq? @kw "prefix")
value: (string) @prefix)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// `from .api.assistant import router` style — used together with
// INCLUDE_ROUTER_NAME so we can map a local name back to its module
// path, then back to the file the router was declared in.
const FROM_IMPORT_ROUTER_PATTERNS = compilePatterns({
name: 'python-fastapi-from-import-router',
language: Python,
patterns: [
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (dotted_name (identifier) @imported (#eq? @imported "router")))
`,
},
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (aliased_import
name: (dotted_name (identifier) @imported (#eq? @imported "router"))
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// `from api import users` / `from api import users as u` — module-level
// imports where the imported name is itself the module that owns
// `<name>.router`. Lets Shape A (`<host>.include_router(<name>.router, …)`)
// look up the full package path of `<name>` and pin the prefix onto the
// exact file (`api/users.py`) rather than every file basenamed `users.py`.
const FROM_IMPORT_MODULE_PATTERNS = compilePatterns({
name: 'python-fastapi-from-import-module',
language: Python,
patterns: [
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (dotted_name (identifier) @imported))
`,
},
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (aliased_import
name: (dotted_name (identifier) @imported)
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: requests.get/post/... ──────────────────────────────────
const REQUESTS_VERB_PATTERNS = compilePatterns({
name: 'python-requests-verb',
@ -447,15 +583,226 @@ const HTTPX_ASYNC_CLIENT_GENERIC_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── prepareRepo: build router-module → prefix list map ─────────────
//
// FastAPI splits route declarations across files: handler decorators
// live in `api/<feature>.py` while `app.include_router(<x>.router,
// prefix='/ai')` lives in `main.py`. A per-file plugin scan therefore
// can't see the prefix that ought to be applied. We resolve this by
// running a one-shot pre-pass over the repo: for every file that
// hosts an `app.include_router(...)` we record the module the router
// came from (either via `module.router` attribute access, or via a
// local name resolved through a `from <module> import router` import)
// together with the prefix string. At scan time the python plugin
// looks up the current file's module key in this map and joins each
// prefix with each `@router.<verb>` decorator's path.
//
// Multiple prefixes for the same module are kept and emitted as
// separate detections — this matches FastAPI's behaviour when one
// router is mounted under several prefixes.
//
// Module keying is two-tiered to avoid prefix bleed between same-named
// files in different packages (e.g. `api/users.py` vs `admin/users.py`):
// • short key — file basename without `.py` (`users`)
// • long key — `<parent-dir>/<basename>` (`api/users`)
// The pre-pass records prefixes against the long key whenever the import
// site supplies enough context (`from api.users import router as ...` →
// long key `api/users`); otherwise it falls back to the short key.
// At scan time the file's own long key is consulted first; only when no
// long-key entry targets this file do we look up the short key. This
// preserves the previous coarse-grained behaviour where context is
// missing while delivering precision wherever the import statement
// gives us a multi-segment module path.
interface PythonRepoContext {
/** `<parent>/<stem>` → set of prefixes (precise, package-aware) */
prefixesByLongKey: Map<string, Set<string>>;
/** stem only → set of prefixes (basename fallback, may collide) */
prefixesByShortKey: Map<string, Set<string>>;
}
/** Strip `.py` and return the bare basename (e.g. `api/users.py` → `users`). */
function fileShortKey(rel: string): string {
const slash = rel.lastIndexOf('/');
const file = slash >= 0 ? rel.slice(slash + 1) : rel;
return file.endsWith('.py') ? file.slice(0, -3) : file;
}
/**
* Long key for a `.py` file: parent directory + stem, joined with `/`.
* Files at the repo root return the empty string (no parent), in which
* case callers should fall back to the short key.
*/
function fileLongKey(rel: string): string {
const noExt = rel.endsWith('.py') ? rel.slice(0, -3) : rel;
const lastSlash = noExt.lastIndexOf('/');
if (lastSlash < 0) return '';
const beforeLast = noExt.slice(0, lastSlash);
const stem = noExt.slice(lastSlash + 1);
const prevSlash = beforeLast.lastIndexOf('/');
const parent = prevSlash >= 0 ? beforeLast.slice(prevSlash + 1) : beforeLast;
return `${parent}/${stem}`;
}
/** Last `.`-separated segment of a (possibly relative) module path. */
function lastSegmentOfDotted(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const dot = stripped.lastIndexOf('.');
return dot >= 0 ? stripped.slice(dot + 1) : stripped;
}
/**
* Last two `.`-separated segments of a (possibly relative) module path
* joined with `/`, e.g. `api.users` → `api/users`. Single-segment paths
* and pure-dot inputs return the empty string; callers should fall back
* to the short key in that case.
*/
function lastTwoSegmentsAsLongKey(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const last = stripped.lastIndexOf('.');
if (last <= 0) return '';
const beforeLast = stripped.slice(0, last);
const stem = stripped.slice(last + 1);
const prev = beforeLast.lastIndexOf('.');
const parent = prev >= 0 ? beforeLast.slice(prev + 1) : beforeLast;
return `${parent}/${stem}`;
}
function recordPrefix(target: Map<string, Set<string>>, key: string, prefix: string): void {
const set = target.get(key) ?? new Set<string>();
set.add(prefix);
target.set(key, set);
}
function buildPythonRepoContext(
files: string[],
parser: Parser,
readFile: (rel: string) => string | null,
parseSource: (parser: Parser, src: string) => Parser.Tree | null,
): PythonRepoContext {
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
// Pre-pass over .py files. We deliberately run this even on files
// that don't contain `include_router` — the cost of an extra parse
// is bounded by the file count, and detecting `include_router`
// beforehand would require its own grep/scan.
for (const rel of files) {
if (!rel.endsWith('.py')) continue;
const src = readFile(rel);
if (!src) continue;
if (!src.includes('include_router')) continue;
parser.setLanguage(Python);
const tree = parseSource(parser, src);
if (!tree) continue;
// Local name → (short, long) map for the current file, populated
// from `from <module> import router [as <alias>]` statements. The
// alias (or 'router' when there is no alias) is the local name
// we'll later see passed to `<host>.include_router`.
interface LocalImport {
moduleShort: string;
moduleLong: string;
}
const localNameToModule = new Map<string, LocalImport>();
for (const m of runCompiledPatterns(FROM_IMPORT_ROUTER_PATTERNS, tree)) {
const moduleNode = m.captures.module;
const aliasNode = m.captures.alias;
const importedNode = m.captures.imported;
if (!moduleNode || !importedNode) continue;
const localName = aliasNode?.text ?? importedNode.text;
const moduleShort = lastSegmentOfDotted(moduleNode.text);
if (!moduleShort) continue;
const moduleLong = lastTwoSegmentsAsLongKey(moduleNode.text);
localNameToModule.set(localName, { moduleShort, moduleLong });
}
// Module-alias map: name imported from a multi-segment package →
// long key. Lets Shape A look up the precise file for `<name>.router`
// even when `<name>` collides with another package's basename.
const localNameToModuleAlias = new Map<string, string>();
for (const m of runCompiledPatterns(FROM_IMPORT_MODULE_PATTERNS, tree)) {
const moduleNode = m.captures.module;
const importedNode = m.captures.imported;
const aliasNode = m.captures.alias;
if (!moduleNode || !importedNode) continue;
// Skip the `router` shape — already handled by FROM_IMPORT_ROUTER_PATTERNS
// above and stored under its router-aware semantics.
if (importedNode.text === 'router') continue;
const moduleLong = lastTwoSegmentsAsLongKey(`${moduleNode.text}.${importedNode.text}`);
if (!moduleLong) continue;
const localName = aliasNode?.text ?? importedNode.text;
localNameToModuleAlias.set(localName, moduleLong);
}
// Shape A: `<host>.include_router(<module>.router, prefix='/x')`.
// The call site gives us only a short module name. We promote to a
// long key when the same file imports `<module>` via either
// `from <pkg> import <module>` (recorded in `localNameToModuleAlias`
// — the typical pattern) or, less commonly, a router-aware import
// statement. Only fall back to the basename short key when neither
// alias is available.
for (const m of runCompiledPatterns(INCLUDE_ROUTER_ATTR_PATTERNS, tree)) {
const modNode = m.captures.router_module;
const prefixNode = m.captures.prefix;
if (!modNode || !prefixNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix === null) continue;
const moduleShort = modNode.text;
const aliasLong = localNameToModuleAlias.get(moduleShort);
const sameFileImport = localNameToModule.get(moduleShort);
const longKey = aliasLong ?? sameFileImport?.moduleLong;
if (longKey) {
recordPrefix(prefixesByLongKey, longKey, prefix);
} else {
recordPrefix(prefixesByShortKey, moduleShort, prefix);
}
}
// Shape B: `<host>.include_router(my_router, prefix='/x')` — resolve
// `my_router` via the import map built above. Whenever the import
// statement supplied a multi-segment module path the long key is
// recorded, eliminating cross-package collisions.
for (const m of runCompiledPatterns(INCLUDE_ROUTER_NAME_PATTERNS, tree)) {
const nameNode = m.captures.router_name;
const prefixNode = m.captures.prefix;
if (!nameNode || !prefixNode) continue;
const localImp = localNameToModule.get(nameNode.text);
if (!localImp) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix === null) continue;
if (localImp.moduleLong) {
recordPrefix(prefixesByLongKey, localImp.moduleLong, prefix);
} else {
recordPrefix(prefixesByShortKey, localImp.moduleShort, prefix);
}
}
}
return { prefixesByLongKey, prefixesByShortKey };
}
function joinPrefix(prefix: string, route: string): string {
// Mirror FastAPI's path joining: trim trailing slash off prefix,
// ensure exactly one leading slash on the result.
const p = prefix.replace(/\/+$/, '');
const r = route.startsWith('/') ? route : `/${route}`;
return `${p}${r}`;
}
export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'python-http',
language: Python,
scan(tree) {
prepareRepo({ files, parser, readFile, parseSource }): RepoContext {
return buildPythonRepoContext(files, parser, readFile, parseSource);
},
scan(tree, repoContext, fileRel) {
const out: HttpDetection[] = [];
const httpxAsyncClients = collectHttpxAsyncClients(tree);
const ctx = repoContext as PythonRepoContext | undefined;
// Providers: FastAPI
for (const match of runCompiledPatterns(FASTAPI_PATTERNS, tree)) {
// Providers: FastAPI @app.<verb>("/path") — already absolute path.
for (const match of runCompiledPatterns(FASTAPI_APP_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
@ -473,6 +820,47 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// Providers: FastAPI @router.<verb>("/path") — must be joined
// with the prefix(es) declared at the include_router site. When
// no prefix is found we still emit the unprefixed path so this
// change is strictly additive vs. the prior @app-only behaviour;
// when the same router is mounted under multiple prefixes we emit
// one detection per prefix.
for (const match of runCompiledPatterns(FASTAPI_ROUTER_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const httpMethod = FASTAPI_VERBS[methodNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
// Long key first (precise, package-aware), short key as fallback.
// Mirrors the ingestion-side resolution in parse-impl.ts so the
// graph nodes and group contracts agree on which prefix applies.
const longKey = fileRel ? fileLongKey(fileRel) : '';
const longPrefixes = longKey ? ctx?.prefixesByLongKey.get(longKey) : undefined;
const shortKey = fileRel ? fileShortKey(fileRel) : '';
const shortPrefixes =
longPrefixes || !shortKey ? undefined : ctx?.prefixesByShortKey.get(shortKey);
const prefixSet = longPrefixes ?? shortPrefixes;
const paths =
prefixSet && prefixSet.size > 0
? [...prefixSet].map((p) => joinPrefix(p, rawPath))
: [rawPath];
for (const p of paths) {
out.push({
role: 'provider',
framework: 'fastapi',
method: httpMethod,
path: p,
name: null,
confidence: 0.8,
});
}
}
// Consumers: requests.<verb>
for (const match of runCompiledPatterns(REQUESTS_VERB_PATTERNS, tree)) {
const methodNode = match.captures.method;

View file

@ -40,6 +40,16 @@ export interface HttpDetection {
confidence: number;
}
export interface HttpScanInput {
filePath: string;
tree: Parser.Tree;
}
export interface HttpFileDetections {
filePath: string;
detections: HttpDetection[];
}
/**
* One language-scoped HTTP plugin. The plugin owns the tree-sitter
* grammar and the `scan` function that translates a parsed tree into
@ -51,15 +61,54 @@ export interface HttpDetection {
* `LanguagePatterns.language` in `tree-sitter-scanner.ts` — the
* grammar modules export different shapes.
*/
/**
* Per-repo state a plugin can build during a `prepareRepo` pass before
* any per-file `scan` is invoked. The orchestrator threads this opaque
* value back into each `scan` call so plugins can resolve cross-file
* facts (e.g. FastAPI `app.include_router(prefix=...)` mappings live
* in `main.py` but apply to handlers declared in `api/*.py`).
*
* Plugins that have no cross-file state can omit `prepareRepo` and
* receive `undefined`.
*/
export type RepoContext = unknown;
export interface HttpLanguagePlugin {
/** Human-readable plugin name for diagnostics. */
name: string;
/** tree-sitter grammar object (passed to the shared parser). */
language: unknown;
/**
* Optional pre-pass: walk the relevant files in the repo and produce
* an opaque context that `scan` can use to resolve cross-file facts.
* Implementations must not throw — return undefined on any error so
* the orchestrator falls back to context-less scanning.
*/
prepareRepo?(args: {
repoPath: string;
files: string[];
parser: Parser;
readFile: (rel: string) => string | null;
parseSource: (parser: Parser, src: string) => Parser.Tree | null;
}): RepoContext | undefined;
/**
* Scan a parsed tree and return zero or more HTTP detections. Plugins
* must not throw — they should swallow per-match errors so a single
* malformed construct does not abort the whole file.
*
* `repoContext` is whatever the plugin's `prepareRepo` produced (or
* `undefined` if there is no `prepareRepo`).
*
* `fileRel` is the repo-relative path of the file being scanned;
* plugins that resolve cross-file facts (e.g. FastAPI router prefix
* joining) need it to key into `repoContext`. Optional so existing
* single-file plugins can keep their unary `scan(tree)` shape.
*/
scan(tree: Parser.Tree): HttpDetection[];
scan(tree: Parser.Tree, repoContext?: RepoContext, fileRel?: string): HttpDetection[];
/**
* Optional project-level scan hook for language rules that require
* multiple files, such as Java controllers inheriting Spring mappings
* from annotated interfaces.
*/
scanProject?(files: readonly HttpScanInput[]): HttpFileDetections[];
}

View file

@ -6,7 +6,13 @@ import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-patterns/index.js';
import {
getPluginForFile,
HTTP_SCAN_GLOB,
type HttpDetection,
type HttpLanguagePlugin,
type HttpScanInput,
} from './http-patterns/index.js';
/**
* Language-agnostic orchestrator for HTTP route (provider + consumer)
@ -160,31 +166,85 @@ export class HttpRouteExtractor implements ContractExtractor {
// both graph-assisted enrichment and source-scan emission.
const parser = new Parser();
const cachedDetections = new Map<string, HttpDetection[]>();
const getDetections = (rel: string): HttpDetection[] => {
const cached = cachedDetections.get(rel);
if (cached) return cached;
const cachedInputs = new Map<
string,
{ plugin: HttpLanguagePlugin; input: HttpScanInput; repoContext: unknown } | null
>();
const projectDetections = new Map<string, HttpDetection[]>();
let projectScanComplete = false;
// Per-plugin cross-file context (e.g. Python's FastAPI router →
// include_router(prefix=...) map). Built lazily on first
// `getDetections` call for a file the plugin handles, scoped to the
// file list returned by `getScannedFiles`. Stored by plugin name so
// a repo with multiple languages keeps each plugin's context
// independent.
const repoContextByPlugin = new Map<string, unknown>();
const ensureRepoContext = async (
plugin: ReturnType<typeof getPluginForFile>,
): Promise<unknown> => {
if (!plugin || typeof plugin.prepareRepo !== 'function') return undefined;
if (repoContextByPlugin.has(plugin.name)) return repoContextByPlugin.get(plugin.name);
try {
const ctx = plugin.prepareRepo({
repoPath,
files: await getScannedFiles(),
parser,
readFile: (rel) => readSafe(repoPath, rel),
parseSource: (p, src) => parseSourceSafe(p, src),
});
repoContextByPlugin.set(plugin.name, ctx);
return ctx;
} catch {
repoContextByPlugin.set(plugin.name, undefined);
return undefined;
}
};
const getScanInput = async (
rel: string,
): Promise<{
plugin: HttpLanguagePlugin;
input: HttpScanInput;
repoContext: unknown;
} | null> => {
if (cachedInputs.has(rel)) return cachedInputs.get(rel) ?? null;
const plugin = getPluginForFile(rel);
if (!plugin) {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
const repoContext = await ensureRepoContext(plugin);
const content = readSafe(repoPath, rel);
if (!content) {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
try {
parser.setLanguage(plugin.language);
const tree = parseSourceSafe(parser, content);
const detections = plugin.scan(tree);
cachedDetections.set(rel, detections);
return detections;
const input = { filePath: rel, tree };
const item = { plugin, input, repoContext };
cachedInputs.set(rel, item);
return item;
} catch {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
};
const getDetections = async (rel: string): Promise<HttpDetection[]> => {
const cached = cachedDetections.get(rel);
if (cached) return cached;
const scanInput = await getScanInput(rel);
const ownDetections = scanInput
? scanInput.plugin.scan(scanInput.input.tree, scanInput.repoContext, rel)
: [];
const detections = [...ownDetections, ...(projectDetections.get(rel) ?? [])];
cachedDetections.set(rel, detections);
return detections;
};
// Glob the source-scan file list at most once per extract() —
// both provider and consumer fallback paths share the same list.
let scannedFiles: string[] | null = null;
@ -194,20 +254,46 @@ export class HttpRouteExtractor implements ContractExtractor {
return scannedFiles;
};
const collectProjectDetections = async (files: string[]): Promise<void> => {
if (projectScanComplete) return;
projectScanComplete = true;
const byPlugin = new Map<HttpLanguagePlugin, HttpScanInput[]>();
for (const rel of files) {
const scanInput = await getScanInput(rel);
if (!scanInput?.plugin.scanProject) continue;
const items = byPlugin.get(scanInput.plugin) ?? [];
items.push(scanInput.input);
byPlugin.set(scanInput.plugin, items);
}
for (const [plugin, inputs] of byPlugin) {
const results = plugin.scanProject?.(inputs) ?? [];
for (const result of results) {
const existing = projectDetections.get(result.filePath) ?? [];
projectDetections.set(result.filePath, [...existing, ...result.detections]);
}
}
cachedDetections.clear();
};
const files = await getScannedFiles();
await collectProjectDetections(files);
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
// Source scan always runs to capture routes in languages/files not covered
// by graph edges; the glob and per-file parse results are cached above.
const providers = this.mergeGraphAndSourceContracts(
graphProviders,
this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
await this.extractProvidersSourceScan(files, getDetections),
);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
const consumers = this.mergeGraphAndSourceContracts(
graphConsumers,
this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
await this.extractConsumersSourceScan(files, getDetections),
);
return [...providers, ...consumers];
@ -232,7 +318,7 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractProvidersGraph(
db: CypherExecutor,
getDetections: (rel: string) => HttpDetection[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@ -254,7 +340,7 @@ export class HttpRouteExtractor implements ContractExtractor {
// helpers — tree-sitter gives both pieces of information
// structurally. Always run the lookup: even when method is set by
// `methodFromRouteReason`, we still need the handler name.
const detections = filePath ? getDetections(filePath) : [];
const detections = filePath ? await getDetections(filePath) : [];
const providerDetections = detections.filter((d) => d.role === 'provider');
let handlerName: string | null = null;
const normalizedRoute = normalizeHttpPath(routePath);
@ -331,13 +417,13 @@ export class HttpRouteExtractor implements ContractExtractor {
// ─── Source-scan providers ─────────────────────────────────────────
private extractProvidersSourceScan(
private async extractProvidersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = getDetections(rel);
const detections = await getDetections(rel);
for (const d of detections) {
if (d.role !== 'provider') continue;
const pathNorm = normalizeHttpPath(d.path);
@ -366,7 +452,7 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractConsumersGraph(
db: CypherExecutor,
getDetections: (rel: string) => HttpDetection[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@ -382,7 +468,7 @@ export class HttpRouteExtractor implements ContractExtractor {
let method = 'GET';
// Prefer the plugin's detected method if we can find a matching
// fetch/axios call in the same file.
const detections = filePath ? getDetections(filePath) : [];
const detections = filePath ? await getDetections(filePath) : [];
// Symmetric to the provider path: if multiple consumer calls in
// the same file share the same normalized path (e.g. a GET
// fetch AND a POST fetch to `/api/orders`), `.find()` silently
@ -436,13 +522,13 @@ export class HttpRouteExtractor implements ContractExtractor {
// ─── Source-scan consumers ─────────────────────────────────────────
private extractConsumersSourceScan(
private async extractConsumersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = getDetections(rel);
const detections = await getDetections(rel);
for (const d of detections) {
if (d.role !== 'consumer') continue;
const pathNorm = normalizeConsumerPath(d.path);

View file

@ -150,9 +150,24 @@ export const processCobol = (
const entry = copybookMap.get(name.toUpperCase());
return entry ? entry.path : null;
};
// Memoize preprocessed copybook content for the duration of this
// processCobol call. A single copybook is COPYed by many programs (and at
// many COPY sites within a program); without this cache
// preprocessCobolSource would re-run once per COPY site —
// O(programs × copybooks) preprocessing passes over the same content.
// Keyed by the resolved copybook path. REPLACING is applied later by the
// expander on the returned (pre-REPLACING) content (see
// cobol-copy-expander.ts readFile→applyReplacing), so caching the
// pre-REPLACING preprocessed text here is safe and per-call-scoped.
const preprocessedCopyCache = new Map<string, string>();
const readCopy = (copyPath: string): string | null => {
const cached = preprocessedCopyCache.get(copyPath);
if (cached !== undefined) return cached;
const content = copybookByPath.get(copyPath);
return content ? preprocessCobolSource(content) : null;
if (!content) return null; // preserves original falsy→null (missing/empty)
const preprocessed = preprocessCobolSource(content);
preprocessedCopyCache.set(copyPath, preprocessed);
return preprocessed;
};
// Track module names for cross-program CALL resolution

View file

@ -80,7 +80,11 @@ export function emitCobolScopeCaptures(
: rangeOf(startLine, startCol, endLine, endCol);
const grouped: Record<string, Capture> = {
'@scope.module': capture('@scope.module', nameRange, name),
'@scope.module': capture(
'@scope.module',
rangeOf(startLine, startCol, endLine, endCol),
name,
),
'@declaration.program': capture(
'@declaration.program',
rangeOf(startLine, startCol, endLine, endCol),
@ -118,7 +122,11 @@ export function emitCobolScopeCaptures(
: rangeOf(startLine, startCol, endLine, endCol);
const grouped: Record<string, Capture> = {
'@scope.module': capture('@scope.module', nameRange, prog.name),
'@scope.module': capture(
'@scope.module',
rangeOf(startLine, startCol, endLine, endCol),
prog.name,
),
'@declaration.program': capture(
'@declaration.program',
rangeOf(startLine, startCol, endLine, endCol),

View file

@ -24,22 +24,18 @@
* V2 additionally walks class ancestors (via MRO), so base-class enclosing
* namespaces also contribute associated namespaces.
*
* **GitNexus approximation (not strict ISO C++ ADL):** passing a qualified
* function reference like `utils::worker` contributes `utils` to the associated
* set, enabling resolution of unqualified calls like `with_callback(utils::worker)`
* to `utils::with_callback`. Under ISO C++ `[basic.lookup.argdep]`, associated
* entities for function-type arguments come from the **parameter types and return
* type** of each function in the overload set — NOT the function's enclosing
* namespace. For `void worker()`, the standard-compliant associated set is empty.
* GitNexus instead contributes the enclosing namespace of any Function/Method
* def whose simple name matches, because it enables the dominant real-world ADL
* pattern at reasonable precision cost.
* Function-reference arguments follow ISO C++ `[basic.lookup.argdep]`:
* associated entities come from the parameter types and return type of each
* referenced function in the overload set, not from the function's enclosing
* namespace. For `void worker()`, the associated set is empty. For
* `void worker(api::Token)` or `api::Token make_token()`, `api` is associated
* through `Token`.
*
* For qualified refs (e.g. `utils::worker`) the namespace is confirmed via a
* workspace lookup (only contributed when a Function/Method named `worker` exists
* in `utils`). For unqualified refs the workspace is searched for any Function
* def with that simple name. Locally-declared function-pointer variables
* (e.g. `void (*g)()`) and function parameters are excluded from this path.
* For qualified refs (e.g. `utils::worker`) the workspace lookup is restricted
* to functions/methods named `worker` in `utils`; for unqualified refs the
* workspace is searched for matching functions/methods by simple name. Locally
* declared function-pointer variables and function parameters are excluded
* from this path.
*
* ADL candidates are merged with ordinary unqualified-lookup candidates
* in the free-call fallback before overload narrowing.
@ -70,6 +66,7 @@
import type { ParsedFile, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { normalizeCppParamType } from './arity-metadata.js';
import { isCppInlineNamespaceScope } from './inline-namespaces.js';
/**
@ -97,11 +94,8 @@ export interface CppAdlArgInfo {
/** When set, the arg is a potential free-function reference (not a locally-
* declared function-pointer variable or function parameter). Contains the
* identifier text as written in source (e.g. `"utils::worker"` or
* `"worker"`). GitNexus approximation: the function's enclosing namespace
* is contributed to the ADL associated set. For qualified refs a workspace
* lookup confirms a Function/Method with that simple name exists in the
* namespace before contributing; for unqualified refs every namespace
* containing a matching Function/Method def is contributed. */
* `"worker"`). Resolution contributes associated namespaces from each
* referenced Function/Method def's parameter and return types. */
readonly functionRefText?: string;
}
@ -207,7 +201,12 @@ export function pickCppAdlCandidates(
for (const arg of args) {
collectAssociatedNamespacesForAdlArg(arg, scopes, associatedNamespaces);
if (arg.functionRefText !== undefined) {
collectFunctionRefNamespaces(arg.functionRefText, parsedFiles, associatedNamespaces);
collectFunctionTypeAssociatedNamespaces(
arg.functionRefText,
scopes,
parsedFiles,
associatedNamespaces,
);
}
}
if (associatedNamespaces.size === 0) return undefined;
@ -472,23 +471,12 @@ function findCppClassDefBySimpleName(
}
/**
* Contribute associated namespaces for a function-reference argument.
*
* - **Qualified refs** (`utils::worker`, `outer::inner::fn`): the namespace
* is extracted from the qualifier text (converting `::` to `.` for dot-joined
* QName matching). A workspace lookup then **verifies** that a Function or
* Method def named `worker` (the simple name after the last `::`) actually
* exists in the extracted namespace. This prevents false positives from
* namespace-qualified variables, enum values, and static data members, which
* also produce `qualified_identifier` AST nodes in tree-sitter-cpp (the
* AST node type alone does not distinguish functions from non-function names).
* - **Unqualified refs** (`worker`): the workspace is searched for any
* Function/Method def whose simple name matches. Every distinct enclosing
* namespace found is added — overloads across the same namespace produce
* a single entry; GitNexus does not select a specific overload at this stage.
* Contribute associated namespaces for a function-reference argument by walking
* the referenced overload set's parameter and return types.
*/
function collectFunctionRefNamespaces(
function collectFunctionTypeAssociatedNamespaces(
refText: string,
scopes: ScopeResolutionIndexes,
parsedFiles: readonly ParsedFile[],
out: Set<string>,
): void {
@ -511,30 +499,130 @@ function collectFunctionRefNamespaces(
for (const def of scope.ownedDefs) {
if (def.type !== 'Function' && def.type !== 'Method') continue;
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
if (simple === simpleName) {
out.add(nsText);
return; // Namespace confirmed; no need to scan further files.
}
if (simple === simpleName) collectAssociatedNamespacesForFunctionDef(def, scopes, out);
}
}
}
return;
}
// Unqualified: search all namespace scopes for a Function def with this
// simple name and contribute its enclosing namespace.
// Unqualified function references are approximated workspace-wide, matching
// the previous V1 lookup scope. The stricter part of this PR is what each
// overload contributes: only namespaces from parameter/return types, never
// the function's own enclosing namespace.
for (const parsed of parsedFiles) {
const scopesById = new Map<ScopeId, (typeof parsed.scopes)[number]>();
for (const sc of parsed.scopes) scopesById.set(sc.id, sc);
for (const scope of parsed.scopes) {
if (scope.kind !== 'Namespace') continue;
for (const def of scope.ownedDefs) {
if (def.type !== 'Function' && def.type !== 'Method') continue;
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
if (simple !== refText) continue;
const nsQName = computeNamespaceQName(scope, scopesById);
if (nsQName !== '') out.add(nsQName);
collectAssociatedNamespacesForFunctionDef(def, scopes, out);
}
}
}
}
function collectAssociatedNamespacesForFunctionDef(
def: SymbolDefinition,
scopes: ScopeResolutionIndexes,
out: Set<string>,
): void {
const parameterTypes = def.parameterTypeClasses?.map((typeClass) => typeClass.base);
for (const paramType of parameterTypes ?? def.parameterTypes ?? []) {
collectAssociatedNamespacesForFunctionTypeText(paramType, scopes, out);
}
if (def.returnType !== undefined) {
collectAssociatedNamespacesForFunctionTypeText(def.returnType, scopes, out);
}
}
function collectAssociatedNamespacesForFunctionTypeText(
typeText: string,
scopes: ScopeResolutionIndexes,
out: Set<string>,
): void {
for (const token of extractCppTypeNameTokens(typeText)) {
if (isIgnoredCppAdlNamespace(token.namespaceName)) continue;
addAssociatedNamespaceForClassName(token.simpleName, scopes, out);
if (token.namespaceName !== '') out.add(token.namespaceName);
}
}
function extractCppTypeNameTokens(typeText: string): readonly {
readonly simpleName: string;
readonly namespaceName: string;
}[] {
const cleaned = normalizeCppParamType(typeText);
if (cleaned === '' || isPrimitiveCppAdlType(cleaned)) return [];
const out: { simpleName: string; namespaceName: string }[] = [];
const seen = new Set<string>();
const tokenSource = typeText.includes('<') ? `${cleaned} ${typeText}` : cleaned;
for (const rawToken of tokenSource.match(/[A-Za-z_]\w*(?:::[A-Za-z_]\w*)*/g) ?? []) {
if (isPrimitiveCppAdlType(rawToken)) continue;
const segments = rawToken.split('::').filter((part) => part.length > 0);
const simpleName = segments.at(-1) ?? '';
if (simpleName === '' || isPrimitiveCppAdlType(simpleName)) continue;
const namespaceName = segments.length > 1 ? segments.slice(0, -1).join('.') : '';
const key = `${namespaceName}\0${simpleName}`;
if (seen.has(key)) continue;
seen.add(key);
out.push({
simpleName,
namespaceName,
});
}
return out;
}
const CPP_ADL_PRIMITIVE_OR_KEYWORD_TYPES = new Set<string>([
'alignas',
'alignof',
'auto',
'bool',
'char',
'char8_t',
'char16_t',
'char32_t',
'class',
'const',
'consteval',
'constexpr',
'constinit',
'decltype',
'double',
'enum',
'explicit',
'extern',
'float',
'inline',
'int',
'long',
'mutable',
'noexcept',
'null',
'register',
'short',
'signed',
'static',
'string',
'struct',
'template',
'thread_local',
'typename',
'union',
'unknown',
'unsigned',
'void',
'volatile',
'wchar_t',
'...',
]);
function isPrimitiveCppAdlType(typeText: string): boolean {
return CPP_ADL_PRIMITIVE_OR_KEYWORD_TYPES.has(typeText);
}
function isIgnoredCppAdlNamespace(namespaceName: string): boolean {
return namespaceName === 'std' || namespaceName.startsWith('std.');
}

View file

@ -126,6 +126,14 @@ export function emitCppScopeCaptures(
JSON.stringify(arity.parameterTypeClasses),
);
}
const returnType = extractCppDeclarationReturnType(fnNode);
if (returnType !== undefined) {
grouped['@declaration.return-type'] = syntheticCapture(
'@declaration.return-type',
fnNode,
returnType,
);
}
if (hasExplicitSpecifier(fnNode)) {
grouped['@declaration.is-explicit'] = syntheticCapture(
'@declaration.is-explicit',
@ -417,6 +425,30 @@ export function emitCppScopeCaptures(
return out;
}
function extractCppDeclarationReturnType(fnNode: SyntaxNode): string | undefined {
const typeNode = fnNode.childForFieldName('type');
if (typeNode === null) return undefined;
const funcDeclarator = findFunctionDeclarator(fnNode);
if (funcDeclarator !== null && isCppUnsupportedReturnTypeDeclarator(funcDeclarator)) {
return undefined;
}
const typeText = typeNode.text.trim();
if (typeText !== 'auto') return typeText.length > 0 ? typeText : undefined;
if (funcDeclarator === null) return typeText;
for (let i = 0; i < funcDeclarator.namedChildCount; i++) {
const child = funcDeclarator.namedChild(i);
if (child?.type !== 'trailing_return_type') continue;
const typeDesc = child.firstNamedChild;
return typeDesc?.text.trim() || typeText;
}
return typeText;
}
function isCppUnsupportedReturnTypeDeclarator(funcDeclarator: SyntaxNode): boolean {
const text = funcDeclarator.text;
return /\boperator\b/.test(text) || /(^|[(:\s])~\s*[A-Za-z_]\w*/.test(text);
}
/**
* Walk every C++ class/struct base clause and emit `@reference.inherits`
* captures for each base so scope resolution can resolve them into EXTENDS

View file

@ -28,17 +28,19 @@
* aliased `using static X = Y.Z;`, attributed namespace declarations,
* and preprocessor-guarded declarations correctly because the
* tree-sitter grammar parses them as real nodes (not textual
* coincidences).
* coincidences). When the orchestrator's `treeCache` has no Tree for a
* file — the worker path, where native Trees can't cross MessageChannels
* — `extractFileStructure` falls back to a line scanner rather than
* re-parsing every file from scratch (that re-parse dominated worker-mode
* scope-resolution time). See `extractCsharpStructureViaScanner`.
*/
import type { SyntaxNode } from 'tree-sitter';
import type { BindingRef, ParsedFile, Scope, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { getCsharpParser } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
interface CsharpFileStructure {
export interface CsharpFileStructure {
/** Declared namespace names in file source order. Empty array means
* the file has no `namespace X;` / `namespace X { }` declaration
* and sits in the default (global) namespace. */
@ -48,18 +50,174 @@ interface CsharpFileStructure {
readonly usingStaticPaths: readonly string[];
}
/** Build a structural view of a C# file by walking the tree-sitter
* AST. Prefers `cachedTree` (handed in via `treeCache`) so we don't
* re-parse files the orchestrator already parsed for `extractParsedFile`;
* falls back to a fresh parse on cache miss. Parser singleton is
* shared across calls. */
// Line-anchored matchers for the worker-path fallback (see
// `extractCsharpStructureViaScanner`). Anchored at line start (after
// indentation); the scanner additionally tracks block-comment / string
// state across lines so a keyword at the start of a line inside one of
// those regions is skipped.
const CS_NAMESPACE_RE = /^[ \t]*namespace[ \t]+([A-Za-z_@][A-Za-z0-9_.]*)/;
// `global using static`, plain `using static`, and the aliased
// `using static Alias = NS.Type;` form (the AST keeps the RHS path, so
// the optional `Alias =` is skipped and only the dotted path captured).
const CS_USING_STATIC_RE =
/^[ \t]*(?:global[ \t]+)?using[ \t]+static[ \t]+(?:[A-Za-z_@][A-Za-z0-9_]*[ \t]*=[ \t]*)?([A-Za-z_@][A-Za-z0-9_.]*)/;
/** Multi-line lexical state carried line-to-line by the scanner. */
type CsScanState = 'code' | 'block' | 'verbatim' | 'raw';
/** Advance the scanner's lexical state across one line, consuming block
* comments (slash-star), line comments (`//`), single-line regular /
* interpolated strings, verbatim strings (`@"…"`), and raw string literals
* (`"""…"""`, fence length tracked in `rawFence`). Returns the state and
* raw-fence length in effect at the START of the next line. Single-line
* strings and `//` comments resolve back to `code` before end of line; only
* block comments and multi-line strings carry state forward. */
function advanceCsScanState(
line: string,
state: CsScanState,
rawFence: number,
): [CsScanState, number] {
const n = line.length;
let i = 0;
while (i < n) {
if (state === 'block') {
const end = line.indexOf('*/', i);
if (end === -1) return ['block', rawFence];
i = end + 2;
state = 'code';
} else if (state === 'verbatim') {
// Ends at a `"` that is not doubled (`""` is an escaped quote).
while (i < n) {
if (line[i] === '"') {
if (line[i + 1] === '"') {
i += 2;
continue;
}
break;
}
i++;
}
if (i >= n) return ['verbatim', rawFence];
i += 1;
state = 'code';
} else if (state === 'raw') {
// Ends at a run of `"` at least `rawFence` long.
let closed = false;
while (i < n) {
if (line[i] === '"') {
let k = i;
while (k < n && line[k] === '"') k++;
if (k - i >= rawFence) {
i = k;
state = 'code';
rawFence = 0;
closed = true;
break;
}
i = k;
} else {
i++;
}
}
if (!closed) return ['raw', rawFence];
} else {
const c = line[i];
const next = line[i + 1];
if (c === '/' && next === '/') return ['code', rawFence]; // line comment to EOL
if (c === '/' && next === '*') {
state = 'block';
i += 2;
} else if (c === '@' && next === '"') {
state = 'verbatim';
i += 2;
} else if ((c === '$' && next === '@') || (c === '@' && next === '$')) {
if (line[i + 2] === '"') {
state = 'verbatim'; // interpolated verbatim ($@"…" / @$"…")
i += 3;
} else {
i++;
}
} else if (c === '"') {
let k = i;
while (k < n && line[k] === '"') k++;
const run = k - i;
if (run >= 3) {
state = 'raw';
rawFence = run;
i = k;
} else if (run === 2) {
i = k; // "" — empty string
} else {
// single-line regular / interpolated string; consume to closer
let j = i + 1;
while (j < n) {
if (line[j] === '\\') {
j += 2;
continue;
}
if (line[j] === '"') break;
j++;
}
i = j >= n ? n : j + 1;
}
} else {
i++;
}
}
}
return [state, rawFence];
}
/** Line-scanner used when no cached tree is available (worker-parsed files
* can't transfer native tree-sitter Trees across MessageChannels, so
* `treeCache` is empty for them). Re-parsing every C# file here with
* tree-sitter was the dominant scope-resolution cost on large worker-mode
* runs — for a multi-thousand-file solution this loop alone re-parsed the
* whole repo a second time. The scanner extracts the same `namespaces` /
* `usingStaticPaths` the AST walk produces for line-anchored declarations,
* while tracking block-comment and string state across lines (via
* `advanceCsScanState`) so a `namespace` / `using static` keyword at the
* start of a line inside a block comment, verbatim string, or raw string
* literal is NOT mistaken for a declaration. The remaining trade-off vs the
* AST is a declaration whose keyword is not at the start of a code line
* (split across lines, or sharing a line with a comment/string closer).
* Mirrors PHP's `extractNamespaceViaScanner` (issue #1741). */
export function extractCsharpStructureViaScanner(content: string): CsharpFileStructure {
const namespaces: string[] = [];
const usingStaticPaths: string[] = [];
let state: CsScanState = 'code';
let rawFence = 0;
for (const line of content.split('\n')) {
// Only match when the line START is real code — keywords reached while
// inside a block comment / multi-line string are skipped.
if (state === 'code') {
const ns = CS_NAMESPACE_RE.exec(line);
if (ns !== null) {
namespaces.push(ns[1]!);
} else {
const us = CS_USING_STATIC_RE.exec(line);
if (us !== null) usingStaticPaths.push(us[1]!);
}
}
[state, rawFence] = advanceCsScanState(line, state, rawFence);
}
return { namespaces, usingStaticPaths };
}
/** Build a structural view of a C# file. Prefers `cachedTree` (handed in
* via `treeCache`) and walks the tree-sitter AST — the authoritative
* path that sees `global using static`, aliased `using static X = Y.Z;`,
* attributed namespace declarations, and preprocessor-guarded nodes
* correctly. On cache miss (worker-parsed files, whose native Trees
* can't cross MessageChannels) it falls back to the line scanner instead
* of a fresh tree-sitter parse — the parse here dominated worker-mode
* scope-resolution time. Parser singleton is shared across calls. */
function extractFileStructure(content: string, cachedTree: unknown): CsharpFileStructure {
if (!cachedTree) {
return extractCsharpStructureViaScanner(content);
}
type CsharpTree = ReturnType<ReturnType<typeof getCsharpParser>['parse']>;
const tree =
(cachedTree as CsharpTree | undefined) ??
parseSourceSafe(getCsharpParser(), content, undefined, {
bufferSize: getTreeSitterBufferSize(content),
});
const tree = cachedTree as CsharpTree;
const namespaces: string[] = [];
const usingStaticPaths: string[] = [];
@ -277,11 +435,17 @@ export function populateCsharpNamespaceSiblings(
// scope, so `Record(...)` (without `Logger.` qualifier) resolves
// to `Logger.Record`. AST walk above captured these (including
// `global using static` and aliased forms).
// Pre-index files by path once: the member-injection lookup below would
// otherwise be an O(files) scan per `using static` import.
const fileByPath = new Map<string, ParsedFile>(parsedFiles.map((p) => [p.filePath, p]));
for (const parsed of parsedFiles) {
const struct = structureByFile.get(parsed.filePath);
if (struct === undefined) continue;
const moduleScope = parsed.scopes.find((s) => s.kind === 'Module');
if (moduleScope === undefined) continue;
// Per-file de-dup sets keyed by simple name, seeded lazily from the
// augmentation bucket — replaces the per-member O(A) `.some` scan below.
const seenByName = new Map<string, Set<string>>();
for (const fullPath of struct.usingStaticPaths) {
const lastDot = fullPath.lastIndexOf('.');
@ -302,7 +466,7 @@ export function populateCsharpNamespaceSiblings(
// Inject the class's member methods into the importer's module
// scope. `memberByOwner` wasn't built yet here, so we walk the
// file's localDefs to find members with `ownerId === targetDef.nodeId`.
const targetFile = parsedFiles.find((p) => p.filePath === targetDef.filePath);
const targetFile = fileByPath.get(targetDef.filePath);
if (targetFile === undefined) continue;
for (const memberDef of targetFile.localDefs) {
if ((memberDef as { ownerId?: string }).ownerId !== targetDef.nodeId) continue;
@ -316,7 +480,14 @@ export function populateCsharpNamespaceSiblings(
// `lookupBindingsAt`, which fans out across `bindings` +
// `bindingAugmentations`.
const bucketArr = getAugmentationBucket(augmentations, moduleScope.id, simpleName);
if (bucketArr.some((b) => b.def.nodeId === memberDef.nodeId)) continue;
let seen = seenByName.get(simpleName);
if (seen === undefined) {
seen = new Set<string>();
for (const b of bucketArr) seen.add(b.def.nodeId);
seenByName.set(simpleName, seen);
}
if (seen.has(memberDef.nodeId)) continue;
seen.add(memberDef.nodeId);
bucketArr.push({ def: memberDef, origin: 'import' });
}
}
@ -332,6 +503,9 @@ export function populateCsharpNamespaceSiblings(
for (const parsed of parsedFiles) {
const moduleScope = parsed.scopes.find((s) => s.kind === 'Module');
if (moduleScope === undefined) continue;
// Per-file de-dup sets keyed by simple name, seeded lazily from the
// augmentation bucket — replaces the per-def O(A) `.some` scan below.
const seenByName = new Map<string, Set<string>>();
for (const imp of parsed.parsedImports) {
if (imp.kind !== 'namespace') continue;
const targetNs = imp.targetRaw;
@ -344,41 +518,113 @@ export function populateCsharpNamespaceSiblings(
const simpleName = q.includes('.') ? q.slice(q.lastIndexOf('.') + 1) : q;
if (simpleName === '') continue;
const bucketArr = getAugmentationBucket(augmentations, moduleScope.id, simpleName);
if (bucketArr.some((b) => b.def.nodeId === def.nodeId)) continue;
let seen = seenByName.get(simpleName);
if (seen === undefined) {
seen = new Set<string>();
for (const b of bucketArr) seen.add(b.def.nodeId);
seenByName.set(simpleName, seen);
}
if (seen.has(def.nodeId)) continue;
seen.add(def.nodeId);
bucketArr.push({ def, origin: 'namespace' });
}
}
}
for (const [, bucket] of buckets) {
// De-dup by (nodeId, filePath) across multiple declarations (e.g.
// partial classes declaring the same name in two files — we take
// both and leave de-dup to downstream consumers of bindings).
// Workspace-level binding channel for global-namespace types (see the
// global fast-path below). `lookupBindingsAt` consults this as a third
// source after finalized + per-scope augmented bindings. Its inner arrays
// are mutable by contract (append-only, like `bindingAugmentations` — see
// the ScopeResolutionIndexes doc + validateBindingsImmutability), so the
// ReadonlyMap→Map cast is localized to this one line and all writes go
// through `getWorkspaceBucket`.
const workspace = indexes.workspaceFqnBindings as Map<string, BindingRef[]>;
for (const [nsName, bucket] of buckets) {
// Group sibling defs by simple name. Append in place — the previous
// `[...prev, def]` copy made this O(D²) per bucket, which on the
// global (`''`) namespace bucket of a large Unity solution (tens of
// thousands of type defs) was a primary slowness/OOM source. We keep
// every declaration (e.g. partial classes across files) and leave
// de-dup to downstream consumers.
const defsByName = new Map<string, SymbolDefinition[]>();
for (const def of bucket.classDefs) {
// Simple name = last segment of qualifiedName (e.g. `App.User` → `User`).
const q = def.qualifiedName ?? '';
const key = q.includes('.') ? q.slice(q.lastIndexOf('.') + 1) : q;
if (key === '') continue;
const arr = [...(defsByName.get(key) ?? [])];
let arr = defsByName.get(key);
if (arr === undefined) {
arr = [];
defsByName.set(key, arr);
}
arr.push(def);
defsByName.set(key, arr);
}
// Global-namespace fast path (Unity OOM guard). Types declared in the
// default (global) namespace are visible from EVERY file in C# — the
// global namespace is always implicitly in scope — so one workspace-
// level entry per simple name is both semantically correct and O(D)
// instead of the O(S·D) per-scope augmentation that materialized
// billions of BindingRefs on large Unity solutions (tens of thousands
// of global types × tens of thousands of scopes). `walkScopeChain`
// checks local `scope.bindings` first, so local declarations still
// shadow these workspace entries; a file resolving its own global type
// hits the local binding before this map. Dedup by `def.nodeId` keeps
// partial-class / duplicate declarations from double-emitting.
if (nsName === '') {
for (const [name, defs] of defsByName) {
const bucket = getWorkspaceBucket(workspace, name);
const seen = new Set<string>();
for (const b of bucket) seen.add(b.def.nodeId);
for (const def of defs) {
if (seen.has(def.nodeId)) continue; // dedup by nodeId (keeps partials, drops re-emits)
seen.add(def.nodeId);
bucket.push({ def, origin: 'namespace' });
}
}
continue;
}
// Pre-index the first scope per file once (O(S)) instead of an
// O(S) `.find` re-run for every (scope, name) pair, which made the
// injection loop O(S²·D) and was the dominant cost on large buckets.
// Multiple scopes share a filePath (Module + Namespace); the local
// shadow check only needs that file's lexical `Scope.bindings`, which
// is identical regardless of which of those scopes we read.
const firstScopeByFile = new Map<string, Scope>();
for (const s of bucket.scopes) {
if (!firstScopeByFile.has(s.filePath)) firstScopeByFile.set(s.filePath, s.scope);
}
for (const { scopeId, filePath } of bucket.scopes) {
const localScope = firstScopeByFile.get(filePath);
for (const [name, defs] of defsByName) {
// Skip names already present locally — `origin: 'local'` in
// scope.bindings would naturally shadow the cross-file
// namespace entry, but we also keep this index lean.
const local = bucket.scopes.find((s) => s.filePath === filePath)?.scope.bindings.get(name);
const local = localScope?.bindings.get(name);
if (local !== undefined && local.some((b) => b.origin === 'local')) continue;
let bucketArr: BindingRef[] | null = null;
// Bind the augmentation bucket and its seeded de-dup set together
// under one nullable lifecycle, so neither needs a non-null
// assertion (they are always set or unset as a pair). Stays lazy:
// nothing is allocated for a name with no cross-file defs.
let inject: { bucket: BindingRef[]; seen: Set<string> } | null = null;
for (const def of defs) {
if (def.filePath === filePath) continue; // don't self-reference
if (bucketArr === null) bucketArr = getAugmentationBucket(augmentations, scopeId, name);
if (bucketArr.some((b) => b.def.nodeId === def.nodeId)) continue;
bucketArr.push({ def, origin: 'namespace' });
if (inject === null) {
const bucket = getAugmentationBucket(augmentations, scopeId, name);
// Seed the de-dup set from any entries an earlier pass
// (using-static / cross-namespace imports) already added,
// replacing the per-def O(A) `.some` scan.
const seen = new Set<string>();
for (const b of bucket) seen.add(b.def.nodeId);
inject = { bucket, seen };
}
if (inject.seen.has(def.nodeId)) continue;
inject.seen.add(def.nodeId);
inject.bucket.push({ def, origin: 'namespace' });
}
}
}
@ -409,6 +655,22 @@ function getAugmentationBucket(
return bucketArr;
}
/** Get-or-create a mutable inner bucket inside the `workspaceFqnBindings`
* channel (the scope-independent third channel; see
* `ScopeResolutionIndexes.workspaceFqnBindings`). Like
* `getAugmentationBucket`, the inner arrays are mutable by contract —
* callers `push` directly. Keeping the get-or-create here means the one
* ReadonlyMap→Map cast at the call site is the only place the mutable
* view is taken. */
function getWorkspaceBucket(workspace: Map<string, BindingRef[]>, name: string): BindingRef[] {
let bucketArr = workspace.get(name);
if (bucketArr === undefined) {
bucketArr = [];
workspace.set(name, bucketArr);
}
return bucketArr;
}
function isTypeDef(def: SymbolDefinition): boolean {
return (
def.type === 'Class' ||

View file

@ -39,6 +39,56 @@ import {
interpretGoTypeBinding,
} from './go/index.js';
const GO_BUILT_INS: ReadonlySet<string> = new Set([
// built-in functions
'make',
'new',
'len',
'cap',
'append',
'copy',
'delete',
'close',
'panic',
'recover',
'print',
'println',
'complex',
'real',
'imag',
'clear',
'min',
'max',
// built-in types
'error',
'bool',
'string',
'int',
'int8',
'int16',
'int32',
'int64',
'uint',
'uint8',
'uint16',
'uint32',
'uint64',
'uintptr',
'float32',
'float64',
'complex64',
'complex128',
'byte',
'rune',
'any',
'comparable',
// built-in values
'true',
'false',
'nil',
'iota',
]);
export const goProvider = defineLanguage({
id: SupportedLanguages.Go,
extensions: ['.go'],
@ -92,6 +142,7 @@ export const goProvider = defineLanguage({
variableExtractor: createVariableExtractor(goVariableConfig),
classExtractor: createClassExtractor(goClassConfig),
heritageExtractor: createHeritageExtractor(goHeritageConfig),
builtInNames: GO_BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks ──────────
emitScopeCaptures: emitGoScopeCaptures,

View file

@ -38,6 +38,7 @@ import { splitImportStatement } from '../typescript/import-decomposer.js';
import { getJsParser, getJsScopeQuery, jsCachedTreeMatchesGrammar } from './query.js';
import { computeTsArityMetadata } from '../typescript/arity-metadata.js';
import { synthesizeTsReceiverBinding } from '../typescript/receiver-binding.js';
import { isArrayMethodCallbackArrow } from '../typescript/array-callback.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
@ -640,6 +641,21 @@ export function emitJsScopeCaptures(
}
}
// #1876: drop @declaration.function for array higher-order-method
// callbacks (`const x = arr.map(a => …)`). The HOC-wrapped-arrow
// pattern matches them, but the binding holds a value, not a callable.
// The binding keeps its separate @declaration.const / .variable match,
// and the arrow's own @scope.function match (a different pattern) is
// untouched, so inner-call attribution falls through to the enclosing
// scope instead of a phantom Function.
const fnDeclAnchor = grouped['@declaration.function'];
if (fnDeclAnchor !== undefined) {
const arrowNode = findFunctionNode(tree.rootNode, fnDeclAnchor.range);
if (arrowNode !== null && isArrayMethodCallbackArrow(arrowNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declarations.
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
if (declAnchor !== undefined) {

View file

@ -148,6 +148,12 @@ const JAVASCRIPT_SCOPE_QUERY = `
;; HOC-wrapped variable declarations: const X = HOC((args) => { ... }).
;; Covers React.forwardRef, memo, useCallback, useMemo, observer,
;; debounce, and any user-defined HOC factory.
;;
;; #1876: this shape also matches array higher-order-method callbacks
;; (const x = arr.map(a => ...)), where x is a value, not a function.
;; Those are filtered out emit-side in captures.ts via
;; isArrayMethodCallbackArrow (member-expression callee whose property
;; is a known Array method), so only the @declaration.const survives.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name

View file

@ -0,0 +1,98 @@
/**
* Array higher-order-method callback detection (issue #1876).
*
* The HOC-wrapped-arrow declaration pattern in the JS/TS scope queries
* (`const X = call((args) => …)`) was added for React idioms
* (`forwardRef` / `memo` / `useCallback`). It has the same AST shape as
* an array higher-order-method call (`const x = arr.map(a => …)`), so
* those callbacks also match and produce a spurious `@declaration.function`
* named after the binding — duplicating the `@declaration.const` /
* `@declaration.variable` def that the same binding already gets.
*
* For an array-method callback the binding holds a *value* (the method's
* result), not a callable, so the `Function` def is semantically wrong.
* `isArrayMethodCallbackArrow` lets the emitter (`captures.ts`) drop that
* `@declaration.function` match, leaving only the value def.
*
* Shared by both the JavaScript and TypeScript capture emitters — the
* relevant grammar nodes (`arrow_function`, `function_expression`,
* `arguments`, `call_expression`, `member_expression`,
* `property_identifier`) are identical across `tree-sitter-javascript`
* and `tree-sitter-typescript`.
*
* Pure given the input node. No I/O, no globals.
*/
import type { SyntaxNode } from '../../utils/ast-helpers.js';
/**
* Array prototype higher-order methods whose result is a value, not a
* function. A callback passed to one of these is an anonymous callback,
* never a top-level function definition. Identifier-callee HOCs
* (`forwardRef(...)`, `useCallback(...)`, custom factories) are
* deliberately NOT listed — they keep their `Function` classification.
*
* Trade-off (unchanged from before #1876): a custom *fluent-API* member
* call with a callback whose method name is not in this set
* (`qb.where(x => …)`) still classifies as `Function`. There is no clean
* syntactic line beyond the well-known Array surface, so the set is
* intentionally closed and easy to extend.
*
* Receiver-blind, by design: the match keys on the method NAME only, never
* the receiver type (tree-sitter has no type information here). So an in-set
* name on a NON-array receiver — `Map`/`Set` `.forEach`, an RxJS
* `observable.map(…)`, a query builder `.sort(…)`, a lodash chain
* `.filter(…)` — is ALSO treated as a callback and has its
* `@declaration.function` dropped. This is an accepted limitation, not a
* regression: those bindings hold the call's *result value*, not a callable,
* so a value def is the correct classification anyway. The only genuine loss
* is a bespoke DSL whose in-set-named method returns something callable —
* rare enough to accept rather than guard with type inference. Pinned by the
* "in-set method on a non-array receiver" case in `*-captures.test.ts`.
*/
export const ARRAY_CALLBACK_METHODS: ReadonlySet<string> = new Set([
'map',
'filter',
'find',
'findIndex',
'findLast',
'findLastIndex',
'forEach',
'reduce',
'reduceRight',
'some',
'every',
'flatMap',
'sort',
]);
/**
* True when `node` (an `arrow_function` / `function_expression`) is the
* callback argument of an array higher-order-method call, i.e. the
* enclosing call's callee is a `member_expression` whose property is one
* of {@link ARRAY_CALLBACK_METHODS}.
*
* Returns false for direct assignments (`const fn = () => {}` — parent is
* `variable_declarator`, not `arguments`) and for identifier-callee HOCs
* (`forwardRef(() => …)` — callee is an `identifier`, not a
* `member_expression`), so neither is ever suppressed.
*
* Intentional non-suppressing gaps (preserve current behavior, no
* regression): parenthesized callee `(arr.map)(cb)` (`parenthesized_expression`)
* and computed callee `arr['map'](cb)` (`subscript_expression`).
*/
export function isArrayMethodCallbackArrow(node: SyntaxNode): boolean {
const args = node.parent;
if (args === null || args.type !== 'arguments') return false;
const call = args.parent;
if (call === null || call.type !== 'call_expression') return false;
const callee = call.childForFieldName('function');
if (callee === null || callee.type !== 'member_expression') return false;
const property = callee.childForFieldName('property');
if (property === null || property.type !== 'property_identifier') return false;
return ARRAY_CALLBACK_METHODS.has(property.text);
}

View file

@ -37,6 +37,7 @@ import { getTsParser, getTsScopeQuery, tsCachedTreeMatchesGrammar } from './quer
import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { synthesizeTsReceiverBinding } from './receiver-binding.js';
import { computeTsArityMetadata } from './arity-metadata.js';
import { isArrayMethodCallbackArrow } from './array-callback.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
@ -252,6 +253,25 @@ export function emitTsScopeCaptures(
}
}
// #1876: drop @declaration.function for array higher-order-method
// callbacks (`const x = arr.map(a => …)`). The HOC-wrapped-arrow
// pattern matches them, but the binding holds a value, not a callable.
// The binding keeps its separate @declaration.const / .variable match,
// and the arrow's own @scope.function match (a different pattern) is
// untouched, so inner-call attribution falls through to the enclosing
// scope instead of a phantom Function.
const fnDeclAnchor = grouped['@declaration.function'];
if (fnDeclAnchor !== undefined) {
const arrowNode = findFunctionNode(
tree.rootNode,
fnDeclAnchor.range,
groupedNodes['@declaration.function'],
);
if (arrowNode !== null && isArrayMethodCallbackArrow(arrowNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declaration anchors
// before pushing the match. The registry uses these to narrow
// overloads — TypeScript supports overload signatures via

View file

@ -250,20 +250,22 @@ const TYPESCRIPT_SCOPE_QUERY = `
;; that promotes the binding to the parent scope (where \`const X\`
;; lives).
;;
;; Trade-off — chained array-method form: \`const x = arr.find((y) => p(y))\`
;; has the same syntactic shape and would also match, naming the
;; \`.find\` callback as \`x\`. The resulting \`Function:x\` is mostly
;; harmless: \`x\` is consumed as a value (\`if (x) { ... }\`), never
;; invoked as a function, so it gets zero incoming \`CALLS\` edges. The
;; one outgoing edge \`Function:x → p\` is a minor mis-attribution that
;; could in principle be fixed by adding a \`function: [(identifier)
;; (member_expression)]\` predicate that excludes property-identifiers
;; matching a known array-method blocklist (\`map\` / \`filter\` / \`find\`
;; / \`reduce\` / \`forEach\` / \`some\` / \`every\`). We don't do that here
;; because (a) the false-positive cost is negligible, (b) the blocklist
;; would need maintenance, and (c) any user-defined fluent-API method
;; with a callback argument would still false-positive — there's no
;; clean syntactic line.
;; #1876 — chained array-method form: \`const x = arr.find((y) => p(y))\`
;; has the same syntactic shape and matches here too, naming the
;; \`.find\` callback as \`x\`. Because \`x\` holds a value (the method
;; result), not a callable, the spurious \`Function:x\` def is dropped
;; emit-side in captures.ts: \`isArrayMethodCallbackArrow\` skips any
;; \`@declaration.function\` whose enclosing call has a member-expression
;; callee with a known Array-method property (\`ARRAY_CALLBACK_METHODS\`:
;; \`map\` / \`filter\` / \`find\` / \`reduce\` / \`forEach\` / \`some\` /
;; \`every\` / …). Only the \`@declaration.variable\` survives, so the
;; binding is a single value def and calls inside the callback attribute
;; to the enclosing scope rather than \`Function:x\`.
;;
;; Residual (intentional): a user-defined fluent-API method with a
;; callback (\`qb.where(x => …)\`) is NOT in the blocklist and still
;; classifies as \`Function\` — there's no clean syntactic line beyond
;; the well-known Array surface, so the set is closed and easy to extend.
;;
;; Trade-off — multi-arrow arguments: \`const x = call(arrow1, arrow2)\`
;; would emit TWO matches with the same name \`x\`. tree-sitter-query

View file

@ -77,11 +77,15 @@ export interface ScopeResolutionIndexes {
* are returned first and win duplicate `def.nodeId` metadata, with
* unique augmentations appended after. See I8. */
readonly bindingAugmentations: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
/** Workspace-level FQN binding lookup. Populated by PHP namespace-
* siblings Step 3b as a shared map instead of per-scope duplication.
* Consulted by `lookupBindingsAt` as a third source after finalized
* and per-scope augmented bindings. Keys are backslash-separated FQNs
* (e.g. `App\Models\User`). */
/** Workspace-level binding lookup, shared instead of per-scope
* duplication. Consulted by `lookupBindingsAt` as a third source after
* finalized and per-scope augmented bindings. Language-specific
* namespace-sibling hooks populate it with disjoint key formats that
* never collide — e.g. backslash-separated FQNs (`App\Models\User`) for
* backslash-namespace languages, and bare simple names (`User`) for
* global-/default-namespace types that are visible from every file. The
* shared map gives those workspace-wide names one entry each instead of
* O(scopes × defs) per-scope augmentation. */
readonly workspaceFqnBindings: ReadonlyMap<string, readonly BindingRef[]>;
/** Pre-resolution usage facts; consumed by the resolution phase. */
readonly referenceSites: readonly ReferenceSite[];

View file

@ -55,6 +55,11 @@ import type {
ExtractedORMQuery,
FetchWrapperDef,
} from './workers/parse-worker.js';
import type {
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from './route-extractors/fastapi-router-bindings.js';
import {
getTreeSitterBufferSize,
getTreeSitterContentByteLength,
@ -72,6 +77,9 @@ export interface WorkerExtractedData {
fetchCalls: ExtractedFetchCall[];
fetchWrapperDefs: FetchWrapperDef[];
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
routerModuleAliases: ExtractedRouterModuleAlias[];
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
@ -114,6 +122,9 @@ export const mergeChunkResults = (
const allFetchCalls: ExtractedFetchCall[] = [];
const allFetchWrapperDefs: FetchWrapperDef[] = [];
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
@ -152,6 +163,9 @@ export const mergeChunkResults = (
for (const item of result.fetchCalls) allFetchCalls.push(item);
for (const item of result.fetchWrapperDefs ?? []) allFetchWrapperDefs.push(item);
for (const item of result.decoratorRoutes) allDecoratorRoutes.push(item);
for (const item of result.routerIncludes ?? []) allRouterIncludes.push(item);
for (const item of result.routerImports ?? []) allRouterImports.push(item);
for (const item of result.routerModuleAliases ?? []) allRouterModuleAliases.push(item);
for (const item of result.toolDefs) allToolDefs.push(item);
if (result.ormQueries) for (const item of result.ormQueries) allORMQueries.push(item);
for (const item of result.constructorBindings) allConstructorBindings.push(item);
@ -169,6 +183,9 @@ export const mergeChunkResults = (
fetchCalls: allFetchCalls,
fetchWrapperDefs: allFetchWrapperDefs,
decoratorRoutes: allDecoratorRoutes,
routerIncludes: allRouterIncludes,
routerImports: allRouterImports,
routerModuleAliases: allRouterModuleAliases,
toolDefs: allToolDefs,
ormQueries: allORMQueries,
constructorBindings: allConstructorBindings,
@ -210,6 +227,9 @@ const processParsingWithWorkers = async (
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],

View file

@ -63,6 +63,11 @@ import type {
FileConstructorBindings,
FetchWrapperDef,
} from '../workers/parse-worker.js';
import type {
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from '../route-extractors/fastapi-router-bindings.js';
import type { ExtractedHeritage } from '../model/heritage-map.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import type { PipelineOptions } from '../pipeline.js';
@ -357,6 +362,9 @@ export async function runChunkedParseAndResolve(
const allFetchWrapperDefs: FetchWrapperDef[] = [];
const allExtractedRoutes: ExtractedRoute[] = [];
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const deferredWorkerCalls: ExtractedCall[] = [];
@ -675,6 +683,15 @@ export async function runChunkedParseAndResolve(
if (chunkWorkerData.decoratorRoutes?.length) {
for (const item of chunkWorkerData.decoratorRoutes) allDecoratorRoutes.push(item);
}
if (chunkWorkerData.routerIncludes?.length) {
for (const item of chunkWorkerData.routerIncludes) allRouterIncludes.push(item);
}
if (chunkWorkerData.routerImports?.length) {
for (const item of chunkWorkerData.routerImports) allRouterImports.push(item);
}
if (chunkWorkerData.routerModuleAliases?.length) {
for (const item of chunkWorkerData.routerModuleAliases) allRouterModuleAliases.push(item);
}
if (chunkWorkerData.toolDefs?.length) {
for (const item of chunkWorkerData.toolDefs) allToolDefs.push(item);
}
@ -1085,6 +1102,157 @@ export async function runChunkedParseAndResolve(
importCtx.index = EMPTY_INDEX;
importCtx.normalizedFileList = [];
// FastAPI router-prefix resolution (cross-file).
//
// Workers emit two kinds of records per Python file:
// • `routerIncludes` — every `app.include_router(<routerExpr>, prefix='/x')`
// site, where `routerExpr` is either `<module>.router` (Shape A) or a
// bare local name (Shape B).
// • `routerImports` — every `from <module> import router [as <alias>]`,
// mapping a local name to a module key (the basename of the source
// module). These let us resolve Shape-B router includes back to the
// module that defines the router.
//
// We build `module-basename → Set<prefix>` and then walk
// `allDecoratorRoutes`: any decorator route emitted from a `router.<verb>`
// decorator inherits its file-basename's prefix. When a router is mounted
// under multiple prefixes we duplicate the route entry, mirroring FastAPI's
// runtime behaviour.
if (allRouterIncludes.length > 0 && allDecoratorRoutes.length > 0) {
// Group `routerImports` by file so we can resolve Shape-B locals against
// imports declared in the SAME file as the include_router call. We carry
// both the short module key (file basename) and, when available, the long
// key (`<dir>/<basename>`) so cross-package same-name modules don't blur
// their prefixes together. `routerModuleAliases` lifts the same long-key
// information for Shape-A includes whose receiving module was imported
// via `from <pkg> import <module>`.
interface LocalImport {
moduleKey: string;
moduleKeyLong: string | undefined;
}
const importsByFile = new Map<string, Map<string, LocalImport>>();
for (const imp of allRouterImports) {
let m = importsByFile.get(imp.filePath);
if (!m) {
m = new Map();
importsByFile.set(imp.filePath, m);
}
m.set(imp.localName, {
moduleKey: imp.moduleKey,
moduleKeyLong: imp.moduleKeyLong,
});
}
// Module-alias map keyed by file: `localName` (the imported module
// identifier in this file) → long key. Shape-A receivers like
// `users.router` are matched against this map; the long key, when
// present, scopes the prefix to the precise source file.
const moduleAliasesByFile = new Map<string, Map<string, string>>();
for (const alias of allRouterModuleAliases) {
let m = moduleAliasesByFile.get(alias.filePath);
if (!m) {
m = new Map();
moduleAliasesByFile.set(alias.filePath, m);
}
m.set(alias.localName, alias.moduleKeyLong);
}
// Two parallel maps: long-key (precise) and short-key (basename
// fallback). Long-key entries are preferred when the file's own long
// key matches; short-key entries match any file with that basename and
// remain the fallback when no long key is known (e.g. Shape A includes
// without a corresponding import statement).
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
const recordPrefix = (target: Map<string, Set<string>>, key: string, prefix: string): void => {
let set = target.get(key);
if (!set) {
set = new Set();
target.set(key, set);
}
set.add(prefix);
};
for (const inc of allRouterIncludes) {
// Shape A: `<module>.router`. The worker emits `routerExpr` already
// including `.router`, so split it back. We only know a short module
// key here — the call site doesn't carry the dotted package path. If
// the same file imports `<module>` via `from <pkg> import <module>`
// (recorded in `allRouterModuleAliases`) we promote to a long key.
const dotIdx = inc.routerExpr.indexOf('.router');
if (dotIdx > 0) {
const moduleShort = inc.routerExpr.slice(0, dotIdx);
const aliasLong = moduleAliasesByFile.get(inc.filePath)?.get(moduleShort);
if (aliasLong) {
recordPrefix(prefixesByLongKey, aliasLong, inc.prefix);
} else {
recordPrefix(prefixesByShortKey, moduleShort, inc.prefix);
}
continue;
}
// Shape B: bare local name. Resolve through this file's imports. The
// import line gives us a long key whenever the module path was multi-
// segment, so cross-package collisions are eliminated for Shape B.
const localImp = importsByFile.get(inc.filePath)?.get(inc.routerExpr);
if (!localImp) continue;
if (localImp.moduleKeyLong) {
recordPrefix(prefixesByLongKey, localImp.moduleKeyLong, inc.prefix);
} else {
recordPrefix(prefixesByShortKey, localImp.moduleKey, inc.prefix);
}
}
if (prefixesByLongKey.size > 0 || prefixesByShortKey.size > 0) {
const fileLongKey = (rel: string): string => {
// Strip `.py`, then take the last two path segments. `api/users.py`
// → `api/users`. Files at the repo root return the empty string,
// which can never match a long-key entry (those always include a
// parent directory) and so fall through to the short-key lookup.
const noExt = rel.endsWith('.py') ? rel.slice(0, -3) : rel;
const lastSlash = noExt.lastIndexOf('/');
if (lastSlash < 0) return '';
const beforeLast = noExt.slice(0, lastSlash);
const stem = noExt.slice(lastSlash + 1);
const prevSlash = beforeLast.lastIndexOf('/');
const parent = prevSlash >= 0 ? beforeLast.slice(prevSlash + 1) : beforeLast;
return `${parent}/${stem}`;
};
const fileShortKey = (rel: string): string => {
const slash = rel.lastIndexOf('/');
const file = slash >= 0 ? rel.slice(slash + 1) : rel;
return file.endsWith('.py') ? file.slice(0, -3) : file;
};
const expanded: ExtractedDecoratorRoute[] = [];
for (const dr of allDecoratorRoutes) {
if (dr.decoratorReceiver !== 'router' || !dr.filePath.endsWith('.py')) {
expanded.push(dr);
continue;
}
// Long-key lookup first; only fall back to the short key when no
// long-key prefix targets this file. This avoids prefix leakage
// between e.g. `api/users.py` and `admin/users.py`.
const longKey = fileLongKey(dr.filePath);
const longPrefixes = longKey ? prefixesByLongKey.get(longKey) : undefined;
const shortPrefixes = longPrefixes
? undefined
: prefixesByShortKey.get(fileShortKey(dr.filePath));
const prefixes = longPrefixes ?? shortPrefixes;
if (!prefixes || prefixes.size === 0) {
expanded.push(dr);
continue;
}
for (const prefix of prefixes) {
expanded.push({ ...dr, prefix });
}
}
allDecoratorRoutes.length = 0;
for (const dr of expanded) allDecoratorRoutes.push(dr);
}
}
return {
exportedTypeMap,
allFetchCalls,

View file

@ -198,7 +198,6 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
}
const ensureSlash = (path: string) => (path.startsWith('/') ? path : '/' + path);
let duplicateRoutes = 0;
const namedRouteRegistry = new Map<string, string>();
const addRoute = (url: string, entry: RouteEntry) => {
@ -220,7 +219,8 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
}
for (const dr of allDecoratorRoutes) {
addRoute(ensureSlash(dr.routePath), {
const url = normalizeExtractedRoutePath(dr.routePath, dr.prefix ?? null);
addRoute(url, {
filePath: dr.filePath,
source: `decorator-${dr.decoratorName}`,
});

View file

@ -81,6 +81,7 @@ export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> = new Set<Suppo
SupportedLanguages.Java,
SupportedLanguages.Rust,
SupportedLanguages.Ruby,
SupportedLanguages.Cobol,
]);
/**

View file

@ -0,0 +1,275 @@
/**
* FastAPI router-prefix detection — pure functions, no worker thread.
*
* NOT A WORKER. This module exports plain synchronous functions; it
* does not import `worker_threads`, does not call `parentPort`, and
* is not a new worker entry point. It lives next to the other route
* extractors (expo, nextjs, php, laravel) for that reason.
*
* The implementation was historically inlined in `workers/parse-worker.ts`,
* but parse-worker.ts is itself the worker entry point and cannot be
* loaded from the main thread (see the same constraint used by
* `test/unit/call-attribution-issue-1166.test.ts`). Splitting the pure
* extraction here lets unit tests import the function directly without
* booting a worker, satisfying DoD §2.7.
*
* Worker phase is per-file, so the heavy cross-file resolution lives in
* `pipeline-phases/parse-impl.ts`. Here we only extract two raw record
* kinds and let the pipeline aggregate them across files:
*
* • {@link ExtractedRouterInclude} — every
* `<host>.include_router(<routerExpr>, prefix='/x')` site, where
* `<routerExpr>` is either `<module>.router` (Shape A) or a bare
* local name (Shape B). `<host>` is intentionally unconstrained:
* production code uses `app`, `api`, `application`, `asgi_app`,
* etc., and the call shape (`include_router` invoked with a
* `prefix=` keyword) is specific enough on its own.
*
* • {@link ExtractedRouterImport} — every
* `from <module> import router [as <alias>]`, captured for both
* absolute and relative module paths (`from .calls import …`).
* parse-impl uses the imports to resolve Shape-B local names back
* to the file that declares the router.
*
* Module keying is two-tiered to avoid prefix bleed between same-named
* files in different packages (e.g. `api/users.py` vs `admin/users.py`):
*
* • short key — basename without `.py` (`users`)
* • long key — `<parent-dir>/<basename>` (`api/users`)
*
* Imports always carry the short key and, when the module path was
* multi-segment, also the long key. parse-impl matches against the
* long key first and falls back to the short key, so cross-package
* collisions are eliminated for Shape B and minimised for Shape A.
*
* The functions in this module are pure (no Worker / parentPort
* dependency) so they can be unit-tested directly without booting a
* worker thread.
*/
/**
* One `<host>.include_router(<routerExpr>, prefix='/x')` site.
*
* `routerExpr` is the raw text of the first argument — either
* `<module>.router` (Shape A) or a bare local name (Shape B).
* parse-impl resolves Shape B against {@link ExtractedRouterImport}
* records emitted by the same file.
*/
export interface ExtractedRouterInclude {
filePath: string;
routerExpr: string;
prefix: string;
lineNumber: number;
}
/**
* One `from <module> import router [as <alias>]` discovered in a
* Python file.
*
* `moduleKey` is the short key (last `.`-segment of the module path,
* e.g. `api.users` → `users`). `moduleKeyLong` is the long key (last
* two segments joined with `/`, e.g. `api/users`); it is the empty
* string / undefined when the import is single-segment (e.g.
* `from users import router`) or pure-dots (e.g. `from . import
* router`). The long key, when present, gives parse-impl a precise
* way to bind a Shape-B `include_router` call to exactly one Python
* file even when other packages contain a same-named module.
*/
export interface ExtractedRouterImport {
filePath: string;
localName: string;
moduleKey: string;
moduleKeyLong?: string;
}
/**
* One `from <package> import <module>` discovered in a Python file
* where `<module>` is later used as a Shape-A include receiver
* (`<host>.include_router(<module>.router, prefix='/x')`). Without
* this record parse-impl would have to fall back to the short key
* `<module>`, which collides between e.g. `api/users.py` and
* `admin/users.py`. The record carries the long key
* (`<package>/<module>`) so parse-impl can pin the prefix onto the
* exact source file.
*
* Only emitted when the import path was multi-segment (a single
* `from users import users` would yield no long key). All fields
* carry the same module-key semantics as
* {@link ExtractedRouterImport}.
*/
export interface ExtractedRouterModuleAlias {
filePath: string;
/** Local name in the importing file (== imported name or its alias). */
localName: string;
/** Long key (`<parent>/<stem>`) — non-empty for every emitted record. */
moduleKeyLong: string;
}
// `<host>.include_router(<module>.router, ..., prefix='/x')` (Shape A).
// `<host>` is left unrestricted — common production names include
// `app`, `api`, `application`, `asgi_app`. Pinning to the literal
// `app` would silently drop these.
const INCLUDE_ROUTER_ATTR_RE =
/\b(?:[A-Za-z_][\w.]*)\.include_router\s*\(\s*([A-Za-z_][\w]*)\.router\b[^)]*?\bprefix\s*=\s*(['"])([^'"]*)\2/g;
// `<host>.include_router(<local_name>, ..., prefix='/x')` (Shape B).
const INCLUDE_ROUTER_NAME_RE =
/\b(?:[A-Za-z_][\w.]*)\.include_router\s*\(\s*([A-Za-z_][\w]*)\b[^)]*?\bprefix\s*=\s*(['"])([^'"]*)\2/g;
// Module path: a sequence of dots (`.`, `..`, `...`) for "current
// package" imports, OR an optional leading-dot prefix followed by a
// dotted identifier (`api.users`, `.api.users`, `..siblings.users`).
// The latter is the common case and the only one we can map back to
// a module stem.
const FROM_IMPORT_ROUTER_RE = /^\s*from\s+(\.+|\.*[A-Za-z_][\w.]*)\s+import\s+([^#\n]+)/gm;
/**
* Last `.`-separated segment of a (possibly relative) Python module
* path. Strips any leading dots first so `from .api.assistant import
* …` and `from api.assistant import …` both yield `assistant`.
* Pure-dot inputs (`.`, `..`) have no segment and return the empty
* string; callers should skip empty results.
*/
export function lastDottedSegment(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const dot = stripped.lastIndexOf('.');
return dot >= 0 ? stripped.slice(dot + 1) : stripped;
}
/**
* Last two `.`-separated segments of a (possibly relative) module
* path joined with `/`, e.g. `api.users` → `api/users`. Mirrors the
* long-key shape used for files (`api/users.py` → `api/users`).
* Returns the empty string when no parent segment is available
* (single-segment imports or pure dots); callers should fall back
* to the short key in that case.
*/
export function lastTwoSegmentsAsPath(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const last = stripped.lastIndexOf('.');
if (last <= 0) return '';
const beforeLast = stripped.slice(0, last);
const stem = stripped.slice(last + 1);
const prev = beforeLast.lastIndexOf('.');
const parent = prev >= 0 ? beforeLast.slice(prev + 1) : beforeLast;
return `${parent}/${stem}`;
}
/**
* Scan a single Python file's source text for FastAPI router
* `include_router` sites and `from <module> import router` imports,
* appending raw records to the supplied collectors.
*
* `outModuleAliases` is optional: when supplied, every multi-segment
* `from <pkg> import <name>` (other than `router` itself) is recorded
* as a module alias so parse-impl can pin Shape-A
* `<name>.include_router(...)` calls onto the exact module file. When
* omitted, the function preserves the pre-existing behaviour and
* skips the alias collection — this keeps the function signature
* back-compat with older callers (and the parse-cache replay path).
*/
export function extractFastAPIRouterBindings(
filePath: string,
content: string,
outIncludes: ExtractedRouterInclude[],
outImports: ExtractedRouterImport[],
outModuleAliases?: ExtractedRouterModuleAlias[],
): void {
if (!content.includes('include_router') && !content.includes('router')) return;
// `from <module> import router [as <alias>]`. We capture every name
// in the import list. `router` (with or without an `as` alias) maps
// to outImports; every other name lands in outModuleAliases when a
// long key is available, so Shape-A `<name>.router` includes can be
// pinned to the exact module file.
if (content.includes(' import ')) {
FROM_IMPORT_ROUTER_RE.lastIndex = 0;
let m: RegExpExecArray | null;
while ((m = FROM_IMPORT_ROUTER_RE.exec(content)) !== null) {
const moduleText = m[1];
const importList = m[2];
const moduleShort = lastDottedSegment(moduleText);
if (!moduleShort) continue;
// Long key for the imported MODULE itself (used by router
// imports — `from api.users import router` sets
// `moduleKeyLong = api/users`).
const moduleLong = lastTwoSegmentsAsPath(moduleText);
// Strip surrounding parens / trailing whitespace; split on
// commas. (Multiline import groups already have their newlines
// present in the captured list.)
const cleaned = importList.replace(/[()]/g, '').trim();
for (const rawPart of cleaned.split(',')) {
const part = rawPart.trim();
if (!part) continue;
// `router` or `router as foo` → ExtractedRouterImport.
const routerAlias = /^router(?:\s+as\s+([A-Za-z_]\w*))?$/.exec(part);
if (routerAlias) {
const localName = routerAlias[1] ?? 'router';
outImports.push({
filePath,
localName,
moduleKey: moduleShort,
...(moduleLong ? { moduleKeyLong: moduleLong } : {}),
});
continue;
}
// Any other `<name>` or `<name> as <alias>` — recorded as a
// module alias so parse-impl can pin Shape-A includes. The
// long key here is computed against the IMPORTED MODULE PATH
// (`<moduleText>.<name>`), not the package path that `<name>`
// was imported FROM. `from api import users` therefore yields
// `api/users`, the same long key as the file it points at.
if (!outModuleAliases) continue;
const otherAlias = /^([A-Za-z_]\w*)(?:\s+as\s+([A-Za-z_]\w*))?$/.exec(part);
if (!otherAlias) continue;
const importedName = otherAlias[1];
const localName = otherAlias[2] ?? importedName;
const aliasLong = lastTwoSegmentsAsPath(`${moduleText}.${importedName}`);
if (!aliasLong) continue;
outModuleAliases.push({
filePath,
localName,
moduleKeyLong: aliasLong,
});
}
}
}
if (!content.includes('include_router')) return;
// Shape A: `<host>.include_router(<module>.router, prefix='/x')`.
INCLUDE_ROUTER_ATTR_RE.lastIndex = 0;
let m: RegExpExecArray | null;
while ((m = INCLUDE_ROUTER_ATTR_RE.exec(content)) !== null) {
outIncludes.push({
filePath,
routerExpr: `${m[1]}.router`,
prefix: m[3],
lineNumber: content.substring(0, m.index).split('\n').length,
});
}
// Shape B: `<host>.include_router(my_router, prefix='/x')`.
// Resolution to a module key happens in parse-impl using
// outImports from the same file.
INCLUDE_ROUTER_NAME_RE.lastIndex = 0;
while ((m = INCLUDE_ROUTER_NAME_RE.exec(content)) !== null) {
// Skip cases that already matched Shape A — INCLUDE_ROUTER_NAME_RE
// is intentionally permissive and would re-capture `<mod>.router`
// as the bare name `mod`. Discriminate by re-checking the
// immediate source around the captured argument position.
const argStart = m.index + m[0].indexOf(m[1]);
const dotProbe = content.slice(argStart + m[1].length, argStart + m[1].length + 8);
if (/^\s*\.\s*router/.test(dotProbe)) continue;
outIncludes.push({
filePath,
routerExpr: m[1],
prefix: m[3],
lineNumber: content.substring(0, m.index).split('\n').length,
});
}
}

View file

@ -750,6 +750,61 @@ function normalizeNodeLabel(kindStr: string): SymbolDefinition['type'] | undefin
}
}
/** Function-like labels: callable defs that must keep incoming CALLS edges. */
const NODE_BEARING_FUNCTION_LABELS: ReadonlySet<SymbolDefinition['type']> = new Set([
'Function',
'Method',
'Constructor',
]);
/** Value labels: non-callable bindings (a `const`/`let`/`var` holds a value). */
const NODE_BEARING_VALUE_LABELS: ReadonlySet<SymbolDefinition['type']> = new Set([
'Const',
'Variable',
]);
/**
* Collapse rule for the deferred node-creation migration (#1876).
*
* When graph-node creation moves from the legacy DAG onto the
* registry-primary path, a single source binding can carry more than one
* `SymbolDefinition` for the same name in the same scope — e.g. a direct
* arrow `const fn = () => {}` is classified BOTH as a `Function` (the
* arrow) and a `Variable` (the binding). Emitting one graph node per def
* would reproduce exactly the duplicate-node bug this issue tracks.
*
* `selectNodeBearingDef` picks the ONE def that should bear the graph node
* for such a binding group:
*
* 1. a function-like def (`Function` / `Method` / `Constructor`) if any —
* the binding is callable and must keep incoming `CALLS` edges;
* 2. otherwise a value def (`Const` / `Variable`) — the binding holds a
* value (e.g. an array-method result after the U1/U2 narrowing);
* 3. otherwise the first def — deterministic fallback for label sets this
* rule does not rank.
*
* INPUT CONTRACT: `group` must be the defs bound to ONE name within ONE
* scope (a binding group). It deliberately does NOT dedup by range —
* `SymbolDefinition` carries no range and `makeDefId` encodes only the
* start position, so containment is uncomputable here; the caller forms the
* group (e.g. from a scope's `ownedDefs` keyed by name) before calling.
*
* Pure. No production call site yet — this dead export is intentional and
* tracked by #1876 (the deferred node-creation migration); it is the
* executable contract that follow-up will consume, pinned today by the
* scope-extractor unit test.
*/
export function selectNodeBearingDef(
group: readonly SymbolDefinition[],
): SymbolDefinition | undefined {
if (group.length === 0) return undefined;
const functionLike = group.find((def) => NODE_BEARING_FUNCTION_LABELS.has(def.type));
if (functionLike !== undefined) return functionLike;
const value = group.find((def) => NODE_BEARING_VALUE_LABELS.has(def.type));
if (value !== undefined) return value;
return group[0];
}
function makeDefId(
filePath: string,
range: Range,
@ -1087,6 +1142,7 @@ const KNOWN_SUB_TAGS: ReadonlySet<string> = new Set<string>([
'@declaration.required-parameter-count',
'@declaration.parameter-types',
'@declaration.parameter-type-classes',
'@declaration.return-type',
'@declaration.template-constraints',
'@declaration.is-explicit',
]);

View file

@ -116,6 +116,20 @@ export const scopeResolutionPhase: PipelinePhase<ScopeResolutionOutput> = {
preExtractedByPath.set(pf.filePath, pf);
}
// Drop pre-extracted entries for standalone providers — these
// languages are skipped by the canonical guard below (line 164)
// and never consume preExtractedByPath, so holding onto their
// entries leaks memory until the cleanup loop at 262-264 which
// also never runs for skipped providers.
for (const [path] of preExtractedByPath) {
const lang = getLanguageFromFilename(path);
if (lang === null) continue;
const provider = SCOPE_RESOLVERS.get(lang);
if (provider?.languageProvider.parseStrategy === 'standalone') {
preExtractedByPath.delete(path);
}
}
let totalFiles = 0;
let totalImports = 0;
let totalRefs = 0;
@ -158,6 +172,14 @@ export const scopeResolutionPhase: PipelinePhase<ScopeResolutionOutput> = {
for (const [lang, provider] of SCOPE_RESOLVERS) {
if (!isRegistryPrimary(lang)) continue;
// Standalone providers (COBOL, JCL) don't emit graph edges yet
// through the scope-resolution path. This is the canonical guard:
// runScopeResolution is never called for standalone providers, which
// keeps cobolPhase as the sole IMPORTS edge producer. Keep this guard
// in sync with any additional standalone providers added to
// SCOPE_RESOLVERS.
if (provider.languageProvider.parseStrategy === 'standalone') continue;
const langFiles = scannedFiles.filter((f) => getLanguageFromFilename(f.path) === lang);
if (langFiles.length === 0) continue;

View file

@ -1,6 +1,6 @@
/**
* Dev-mode runtime validator for the two-channel binding lifecycle
* (Contract Invariant I8 in `contract/scope-resolver.ts`).
* Dev-mode runtime validator for the post-finalize binding-channel
* lifecycle (Contract Invariant I8 in `contract/scope-resolver.ts`).
*
* The two channels:
* - `indexes.bindings` — finalize-output channel. After
@ -74,5 +74,21 @@ export function validateBindingsImmutability(
}
}
// Third channel: `workspaceFqnBindings` (scope-independent, shared map
// populated by language namespace-sibling hooks — PHP FQN keys, C#
// global-namespace simple names). Like bindingAugmentations its inner
// arrays are mutable by contract (hooks `push()` directly), so freezing
// one is the same defect as freezing an augmentation bucket.
for (const [name, bucket] of indexes.workspaceFqnBindings) {
if (Object.isFrozen(bucket)) {
onWarn(
`binding-immutability: indexes.workspaceFqnBindings[${name}] is FROZEN — ` +
`the workspace channel is mutable by contract; freezing it defeats the ` +
`append-only purpose. See ScopeResolver Invariant I8.`,
);
violations++;
}
}
return violations;
}

View file

@ -99,6 +99,14 @@ const EMPTY_NAMES: Iterable<string> = Object.freeze([]) as readonly string[];
* Fast paths (zero allocation) when at most one channel is populated:
* returns the underlying `Map.keys()` iterator directly. Only when both
* channels carry names do we materialize a `Set` for deduplication.
*
* Scope: enumerates only the per-scope `bindings` and `bindingAugmentations`
* channels. It deliberately EXCLUDES the scope-independent
* `workspaceFqnBindings` channel (PHP FQN keys, C# global-namespace simple
* names). `lookupBindingsAt` consults that third channel when resolving a
* specific name, but name *enumeration* here does not — those names apply at
* every scope and would flood per-scope callers. Callers that need
* workspace-level names must read `workspaceFqnBindings` directly.
*/
export function namesAtScope(scopeId: ScopeId, scopes: ScopeResolutionIndexes): Iterable<string> {
const finalized = scopes.bindings.get(scopeId);

View file

@ -23,6 +23,11 @@ import {
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import type { SymbolTableReader } from '../model/symbol-table.js';
import type { ExtractedHeritage } from '../model/heritage-map.js';
import type {
ExtractedRouterInclude,
ExtractedRouterImport,
ExtractedRouterModuleAlias,
} from '../route-extractors/fastapi-router-bindings.js';
/** Language grammar type accepted by Parser.setLanguage(). */
type TreeSitterLanguage = Parameters<typeof Parser.prototype.setLanguage>[0];
@ -209,6 +214,19 @@ export interface ExtractedDecoratorRoute {
httpMethod: string;
decoratorName: string;
lineNumber: number;
/**
* Decorator receiver identifier (e.g. `router` for `@router.get(...)`,
* `app` for `@app.get(...)`). Used by parse-impl to decide which routes
* participate in `include_router(prefix=...)` joining.
*/
decoratorReceiver?: string;
/**
* FastAPI `app.include_router(prefix='/x')` prefix that applies to
* this route. Filled by parse-impl after cross-file aggregation; the
* routes phase joins it via `normalizeExtractedRoutePath`. `null` /
* absent ⇒ no prefix applies.
*/
prefix?: string | null;
}
export interface ExtractedToolDef {
@ -275,6 +293,18 @@ export interface ParseWorkerResult {
fetchCalls: ExtractedFetchCall[];
fetchWrapperDefs: FetchWrapperDef[];
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
/**
* Optional. `from <pkg> import <module>` records from Python files
* where `<module>` is later used as a Shape-A include receiver
* (`<host>.include_router(<module>.router, prefix='/x')`). parse-impl
* uses these to promote Shape-A short-key entries to long keys, so
* same-named modules in different packages don't share prefixes.
* Optional for cache backward compatibility (older cache entries
* predate the field; consumers must guard with `if (… ?? [])`).
*/
routerModuleAliases?: ExtractedRouterModuleAlias[];
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
@ -740,6 +770,9 @@ const processBatch = (
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@ -779,9 +812,34 @@ const processBatch = (
for (const [language, langFiles] of byLanguage) {
const provider = getProvider(language);
const queryString = provider.treeSitterQueries;
if (!queryString) continue;
// Track if we need to handle tsx separately
if (!queryString) {
// Standalone providers (regex-based, no tree-sitter) that implement
// emitScopeCaptures feed into the scope-resolution pipeline via
// extractParsedFile directly — no tree-sitter involved.
if (provider.emitScopeCaptures) {
for (const file of langFiles) {
const parsedFile = extractParsedFile(
provider,
file.content,
file.path,
(message) => {
if (parentPort) {
parentPort.postMessage({ type: 'warning', message });
} else {
logger.warn(message);
}
},
undefined, // no cachedTree for standalone providers
);
if (parsedFile !== undefined) {
result.parsedFiles.push(parsedFile);
result.fileCount++;
onFileProcessed?.();
}
}
}
continue;
}
const tsxFiles: ParseWorkerInput[] = [];
const regularFiles: ParseWorkerInput[] = [];
@ -968,6 +1026,18 @@ export function extractORMQueries(
}
}
// ============================================================================
// FastAPI router prefix detection (Python)
// ============================================================================
//
// The extraction lives in `../route-extractors/fastapi-router-bindings`
// (a pure-function module — NOT a worker, no `worker_threads`, no
// `parentPort`). It's imported here only so the worker entry can call it
// per file; this module does not re-export it. Downstream consumers
// import the function and its types directly from `route-extractors/`.
import { extractFastAPIRouterBindings } from '../route-extractors/fastapi-router-bindings.js';
const processFileGroup = (
files: ParseWorkerInput[],
language: SupportedLanguages,
@ -1200,6 +1270,7 @@ const processFileGroup = (
if (captureMap['decorator'] && captureMap['decorator.name']) {
const decoratorName = captureMap['decorator.name'].text;
const decoratorArg = captureMap['decorator.arg']?.text;
const decoratorReceiver = captureMap['decorator.receiver']?.text;
const decoratorNode = captureMap['decorator'];
// Store by the decorator's end line — the definition follows immediately after
fileDecorators.set(decoratorNode.endPosition.row, {
@ -1219,6 +1290,7 @@ const processFileGroup = (
httpMethod,
decoratorName,
lineNumber: decoratorNode.startPosition.row + lineOffset,
...(decoratorReceiver ? { decoratorReceiver } : {}),
});
}
// MCP/RPC tool detection: @mcp.tool(), @app.tool(), @server.tool()
@ -1994,6 +2066,20 @@ const processFileGroup = (
// Extract ORM queries (Prisma, Supabase)
extractORMQueries(file.path, parseContent, result.ormQueries);
// Extract FastAPI include_router(prefix=...) and `from <mod> import router`
// sites. parse-impl aggregates these into a per-module prefix map and
// injects the resolved prefix onto each ExtractedDecoratorRoute that
// came from a `@router.<verb>` decorator. Python-only.
if (language === SupportedLanguages.Python) {
extractFastAPIRouterBindings(
file.path,
parseContent,
result.routerIncludes,
result.routerImports,
(result.routerModuleAliases ??= []),
);
}
// Vue: emit CALLS edges for components used in <template>
if (language === SupportedLanguages.Vue) {
const templateComponents = extractTemplateComponents(file.content);
@ -2026,6 +2112,9 @@ let accumulated: ParseWorkerResult = {
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@ -2055,6 +2144,12 @@ const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult) => {
appendAll(target.fetchCalls, src.fetchCalls);
appendAll(target.fetchWrapperDefs, src.fetchWrapperDefs);
appendAll(target.decoratorRoutes, src.decoratorRoutes);
if (src.routerIncludes) appendAll(target.routerIncludes, src.routerIncludes);
if (src.routerImports) appendAll(target.routerImports, src.routerImports);
if (src.routerModuleAliases) {
target.routerModuleAliases ??= [];
appendAll(target.routerModuleAliases, src.routerModuleAliases);
}
appendAll(target.toolDefs, src.toolDefs);
appendAll(target.ormQueries, src.ormQueries);
appendAll(target.constructorBindings, src.constructorBindings);
@ -2147,6 +2242,9 @@ parentPort!.on('message', (msg: WorkerIncomingMessage) => {
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],

View file

@ -51,6 +51,28 @@ const alreadyAvailable = (message: string): boolean =>
message.includes('already exists');
const resolvePolicyFromEnv = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'load-only';
};
export const getExtensionInstallPolicy = (): ExtensionInstallPolicy => resolvePolicyFromEnv();
/**
* Install policy for the **analyze (write) path**.
*
* The global default (`resolvePolicyFromEnv`) is `load-only` so serve/query
* read paths never require outbound network access (PR #1161, offline-first).
* The analyze path is different: it owns building the search indexes, so it
* defaults to `auto` — LOAD the extension if present, otherwise attempt one
* bounded out-of-process INSTALL. This keeps FTS symmetric with the
* VECTOR/embeddings path (which already defaults to `auto`) and matches the
* #726 contract. An explicit `GITNEXUS_LBUG_EXTENSION_INSTALL` value still
* wins, so operators can force `load-only`/`never` for fully offline analyze;
* `auto` LOADs-first, so offline machines still degrade gracefully when the
* INSTALL cannot reach the network.
*/
export const resolveAnalyzeInstallPolicy = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'auto';
@ -148,7 +170,7 @@ export const installDuckDbExtensionOutOfProcess = async (
* subsequent analyze or query calls.
*
* Policy precedence (most specific wins):
* per-call `opts.policy` → constructor `options.policy` → env → `auto`
* per-call `opts.policy` → constructor `options.policy` → env → `load-only`
*/
export class ExtensionManager {
private readonly capabilities = new Map<string, ExtensionCapability>();

View file

@ -24,8 +24,10 @@ import {
deleteNodesForFile,
deleteAllCommunitiesAndProcesses,
queryImporters,
loadFTSExtension,
} from './lbug/lbug-adapter.js';
import { createSearchFTSIndexes, verifySearchFTSIndexes } from './search/fts-indexes.js';
import { resolveAnalyzeInstallPolicy } from './lbug/extension-loader.js';
import {
startWalCheckpointDriver,
type WalCheckpointDriver,
@ -144,8 +146,26 @@ export interface AnalyzeResult {
pipelineResult?: any;
/** True when analyze only repaired FTS indexes and skipped pipeline re-analysis. */
ftsRepairedOnly?: boolean;
/**
* True when the FTS extension was unavailable so search-index creation was
* skipped (offline-first degradation). The graph is fully queryable; only
* full-text/BM25 search is disabled. Lets callers (CLI summary, server) and
* the persisted meta surface the degraded state instead of reporting healthy.
*/
ftsSkipped?: boolean;
}
/**
* Logged when the optional FTS extension cannot be loaded or installed during
* a full analyze. Kept as a named constant so the env-var/command guidance
* stays in one place (mirrors the VECTOR message in embedding-pipeline.ts).
*/
const FTS_UNAVAILABLE_MESSAGE =
'FTS extension unavailable; skipping search-index creation. ' +
'Full-text/BM25 search will be disabled until the LadybugDB FTS extension is ' +
'installed once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) or ' +
'pre-installed for offline use. Run `gitnexus doctor` for details.';
// Re-export the pure flag-derivation helper so external callers (and tests)
// keep importing from this module's stable surface.
export { deriveEmbeddingMode, DEFAULT_EMBEDDING_NODE_LIMIT } from './embedding-mode.js';
@ -684,23 +704,41 @@ export async function runFullAnalysis(
}
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
// The analyze (write) path owns building the search indexes, so it uses
// the `auto` install policy (LOAD-first, then one bounded INSTALL) —
// symmetric with the VECTOR/embeddings path below and consistent with the
// #726 contract. The global `load-only` default (PR #1161) governs the
// serve/query read paths, not this one. When the extension still cannot be
// loaded (genuinely offline + not pre-installed, or policy forced to
// load-only/never), degrade gracefully — exactly like the VECTOR path — so
// analyze still produces a fully queryable graph; only full-text/BM25
// search falls back. `--repair-fts` (whose sole job is FTS) still fails
// loudly on its own path above.
progress('fts', 85, 'Creating search indexes...');
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
const ftsAvailable = await loadFTSExtension(undefined, {
policy: resolveAnalyzeInstallPolicy(),
});
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
if (missingIndexNames.length > 0) {
throw new Error(
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
);
if (ftsAvailable) {
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
});
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
if (missingIndexNames.length > 0) {
throw new Error(
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
);
}
progress('fts', 90, 'Search indexes ready');
} else {
log(FTS_UNAVAILABLE_MESSAGE);
progress('fts', 90, 'Search indexes skipped (FTS unavailable)');
}
progress('fts', 90, 'Search indexes ready');
// ── Phase 3.5: Re-insert cached embeddings ────────────────────────
// Runs on BOTH the full-rebuild path and the incremental path:
@ -889,7 +927,14 @@ export async function runFullAnalysis(
},
capabilities: {
graph: { provider: 'ladybugdb', status: runtimeCapabilities.graph },
fts: { provider: 'ladybugdb-fts', status: runtimeCapabilities.fts },
// Reflect what this analyze run actually produced: when the FTS
// extension was unavailable the indexes were skipped, so record
// 'unavailable' rather than the static runtime default. Keeps
// meta.json / `gitnexus doctor` honest about degraded search.
fts: {
provider: 'ladybugdb-fts',
status: ftsAvailable ? runtimeCapabilities.fts : 'unavailable',
},
vectorSearch: {
provider: effectiveSemanticMode === 'vector-index' ? 'ladybugdb-vector' : 'exact-scan',
status: embeddingCount > 0 ? effectiveSemanticMode : 'unavailable',
@ -989,6 +1034,7 @@ export async function runFullAnalysis(
repoPath,
stats: meta.stats,
pipelineResult,
ftsSkipped: !ftsAvailable,
};
} catch (err) {
// Ensure LadybugDB is closed even on error. Stop the driver first

View file

@ -2907,9 +2907,11 @@ export class LocalBackend {
limit?: number;
offset?: number;
summaryOnly?: boolean;
skipPerSymbolEnrichment?: boolean;
},
): Promise<any> {
const { maxDepth, relationTypes, includeTests, minConfidence } = opts;
const skipPerSymbolEnrichment = opts.skipPerSymbolEnrichment ?? false;
const hasExplicitLimit = typeof opts.limit === 'number' && Number.isFinite(opts.limit);
const paginationLimit = hasExplicitLimit
? Math.max(1, Math.min(Math.trunc(opts.limit!), 10000))
@ -3066,13 +3068,25 @@ export class LocalBackend {
const directCount = (grouped[1] || []).length;
let affectedProcesses: any[] = [];
let affectedModules: any[] = [];
// Per-symbol process membership: maps impacted symbol id -> list of processes
// it participates in. Populated by a second chunked Cypher pass below when
// any process is affected at all. Surfaced as `processes: [...]` on each
// byDepth item so consumers can tell which caller belongs to which cron/
// webhook/route without a follow-up query.
const perSymbolProcesses = new Map<
string,
Array<{ id: string; label: string; processType: string; step: number }>
>();
// Chunking bounds for batched DB round-trips. Declared at function scope so
// both the in-block enrichment passes and the post-pagination per-symbol
// process enrichment can reference them.
const CHUNK_SIZE = 100;
// Max number of chunks to process to avoid unbounded DB round-trips.
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
if (impacted.length > 0) {
const CHUNK_SIZE = 100;
// Max number of chunks to process to avoid unbounded DB round-trips.
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
// ── Process enrichment: batched chunking (bounded by MAX_CHUNKS) ─
// Uses merged Cypher query (WITH + OPTIONAL MATCH) to fetch
// process + entry point info in 1 round-trip per chunk. Converted to
@ -3218,6 +3232,10 @@ export class LocalBackend {
}))
.sort((a, b) => b.total_hits - a.total_hits);
// Per-symbol process membership is populated post-pagination (see below)
// so it covers exactly the symbols returned in byDepth, not a pre-capped
// flat slice that could miss depth-2+ symbols when depth-1 is large.
// ── Module enrichment: use same cap as process enrichment and parameterized queries
const maxItems = Math.min(impacted.length, MAX_CHUNKS * CHUNK_SIZE);
const cappedImpacted = impacted.slice(0, maxItems);
@ -3360,7 +3378,7 @@ export class LocalBackend {
return base;
}
// Apply limit/offset pagination per depth level
// Apply limit/offset pagination per depth level.
const paginatedGrouped: Record<number, any[]> = {};
let anyTruncated = false;
for (const [depth, items] of Object.entries(grouped)) {
@ -3372,8 +3390,82 @@ export class LocalBackend {
}
}
// ── Per-symbol process membership enrichment (post-pagination) ───────
// Runs after paginatedGrouped is built so we enrich only the IDs that
// actually appear in the response. This eliminates the false-empty
// processes:[] case where a depth-2+ symbol's flat position in `impacted`
// exceeded MAX_CHUNKS*CHUNK_SIZE even though it is returned by byDepth.
// Also uses DISTINCT + MIN(r.step) per (symbol, process) pair to avoid
// duplicate entries when a symbol has multiple STEP_IN_PROCESS edges.
// Skipped entirely when `skipPerSymbolEnrichment` is set (group cross-repo
// fan-out, which consumes byDepth but not byDepth[].processes); the
// attach-loop below still stamps an empty processes:[] for shape stability.
let perSymbolEnrichmentCapped = false;
if (affectedProcesses.length > 0 && !skipPerSymbolEnrichment) {
// Collect unique IDs from the paginated result in one pass.
const pageIds = new Set<string>();
for (const items of Object.values(paginatedGrouped)) {
for (const it of items) {
const id = String(it.id ?? '');
if (id) pageIds.add(id);
}
}
// Bound the enrichment to the same ceiling as the aggregation pass
// (MAX_CHUNKS * CHUNK_SIZE) so a large paginated page cannot trigger
// unbounded DB round-trips (DoD 2.6). When capped, mark the result
// partial so callers know some returned symbols may carry an empty
// processes:[] that is a cap artifact, not a true absence.
const maxPageIds = MAX_CHUNKS * CHUNK_SIZE;
let pageIdArr = Array.from(pageIds);
if (pageIdArr.length > maxPageIds) {
pageIdArr = pageIdArr.slice(0, maxPageIds);
perSymbolEnrichmentCapped = true;
}
for (let i = 0; i < pageIdArr.length; i += CHUNK_SIZE) {
const chunkIds = pageIdArr.slice(i, i + CHUNK_SIZE);
try {
const rows = await executeParameterized(
repo.id,
`
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE s.id IN $ids
RETURN s.id AS sid, p.id AS pid, p.heuristicLabel AS pName,
p.processType AS pType, MIN(r.step) AS step
`,
{ ids: chunkIds },
).catch(() => []);
for (const row of rows) {
const sid = row.sid ?? row[0];
if (!sid) continue;
const procEntry = {
id: String(row.pid ?? row[1] ?? ''),
label: String(row.pName ?? row[2] ?? ''),
processType: String(row.pType ?? row[3] ?? ''),
step: Number(row.step ?? row[4] ?? -1),
};
const list = perSymbolProcesses.get(String(sid));
if (list) list.push(procEntry);
else perSymbolProcesses.set(String(sid), [procEntry]);
}
} catch (e) {
logQueryError('impact:per-symbol-process-chunk', e);
}
}
}
// Attach processes field to each paginated item.
for (const items of Object.values(paginatedGrouped)) {
for (const it of items) {
it.processes = perSymbolProcesses.get(String(it.id)) ?? [];
}
}
return {
...base,
// Surface partial if the per-symbol enrichment was capped, even when the
// BFS traversal itself completed — some returned symbols may carry an
// empty processes:[] that is a cap artifact rather than a true absence.
...(perSymbolEnrichmentCapped && { partial: true }),
...(anyTruncated && {
pagination: {
...(Number.isFinite(paginationLimit) && { limit: paginationLimit }),
@ -3467,11 +3559,20 @@ export class LocalBackend {
];
try {
// skipPerSymbolEnrichment suppresses ONLY the per-symbol STEP_IN_PROCESS
// enrichment pass while preserving byDepth. Group-mode cross-repo fan-out
// may fan across many repos; the per-symbol pass adds up to MAX_CHUNKS
// extra round-trips per repo, which is unacceptable at group scale. But
// cross-impact fan-out DOES consume byDepth (cross-impact.ts reads
// fan.byDepth to populate group by_depth), so summaryOnly would wrongly
// drop it. Group callers do not consume byDepth[].processes, so skipping
// only that enrichment is the correct, targeted suppression.
return await this._runImpactBFS(repo, sym, symType, dir, {
maxDepth: opts.maxDepth,
relationTypes,
includeTests: opts.includeTests,
minConfidence: opts.minConfidence,
skipPerSymbolEnrichment: true,
});
} catch {
return null;

View file

@ -284,6 +284,37 @@ Follow these steps:
/**
* Start the MCP server on stdio transport (for CLI use).
*/
/** Force-exit fallback budget if graceful shutdown cleanup hangs. */
const SHUTDOWN_FORCE_EXIT_MS = 5_000;
/** Conventional 128 + signal-number exit codes for graceful termination. */
export const SHUTDOWN_EXIT_CODES = { SIGINT: 130, SIGTERM: 143 } as const;
type SignalRegistrar = (
event: 'SIGINT' | 'SIGTERM',
listener: (...args: unknown[]) => void,
) => void;
/**
* Wire SIGINT/SIGTERM to a graceful shutdown using NUMERIC exit codes.
*
* Node invokes signal listeners with the signal NAME string as the first
* argument, so registering an `(exitCode = 0) => process.exit(exitCode)`
* shutdown directly passes `'SIGTERM'` into `process.exit()` and crashes with
* `ERR_INVALID_ARG_TYPE` (#1132). These wrappers discard the signal argument
* and pass the conventional 128+signal code instead. `on` is injectable so the
* mapping can be unit-tested without touching the real process.
*/
export function installSignalShutdown(
shutdown: (exitCode?: number) => unknown,
on: SignalRegistrar = (event, listener) => {
process.on(event, listener);
},
): void {
on('SIGINT', () => void shutdown(SHUTDOWN_EXIT_CODES.SIGINT));
on('SIGTERM', () => void shutdown(SHUTDOWN_EXIT_CODES.SIGTERM));
}
export async function startMCPServer(backend: LocalBackend): Promise<void> {
const server = createMCPServer(backend);
@ -321,6 +352,11 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
const shutdown = async (exitCode = 0) => {
if (shuttingDown) return;
shuttingDown = true;
// Safety net: if backend.disconnect()/server.close() hangs, still exit so a
// SIGINT/SIGTERM reliably terminates the process. Unref'd so the timer alone
// never keeps the event loop alive.
const forceExit = setTimeout(() => process.exit(exitCode), SHUTDOWN_FORCE_EXIT_MS);
forceExit.unref();
try {
await backend.disconnect();
} catch {}
@ -329,12 +365,16 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
} catch {}
const { flushLoggerSync } = await import('../core/logger.js');
flushLoggerSync();
clearTimeout(forceExit);
process.exit(exitCode);
};
// Handle graceful shutdown
process.on('SIGINT', shutdown);
process.on('SIGTERM', shutdown);
// Handle graceful shutdown. Node invokes signal listeners with the signal
// NAME (e.g. 'SIGTERM') as the first argument; registering `shutdown`
// directly passed that string to process.exit() and crashed with
// ERR_INVALID_ARG_TYPE (#1132). Map each signal to its conventional
// 128+signal exit code instead.
installSignalShutdown(shutdown);
// Log crashes to stderr so they aren't silently lost.
// uncaughtException is fatal — shut down.
@ -342,14 +382,16 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
// killing the server for one missed catch would be worse than logging it.
process.on('uncaughtException', (err) => {
process.stderr.write(`GitNexus MCP uncaughtException: ${err?.stack || err}\n`);
shutdown(1);
void shutdown(1);
});
process.on('unhandledRejection', (reason: any) => {
process.stderr.write(`GitNexus MCP unhandledRejection: ${reason?.stack || reason}\n`);
});
// Handle stdio errors — stdin close means the parent process is gone
process.stdin.on('end', shutdown);
process.stdin.on('error', () => shutdown());
process.stdout.on('error', () => shutdown());
// Handle stdio errors — stdin close means the parent process is gone.
// Wrap so the event payload (e.g. an Error for 'error') can never reach
// process.exit() as a non-numeric exit code, and void the returned promise.
process.stdin.on('end', () => void shutdown(0));
process.stdin.on('error', () => void shutdown(0));
process.stdout.on('error', () => void shutdown(0));
}

View file

@ -336,7 +336,7 @@ Output includes:
- summary: direct callers, processes affected, modules affected
- affected_processes: which execution flows break and at which step
- affected_modules: which functional areas are hit (direct vs indirect)
- byDepth: affected symbols grouped by traversal depth (paginated by limit/offset; omitted when summaryOnly:true — use byDepthCounts for totals per depth, pagination object when truncated)
- byDepth: affected symbols grouped by traversal depth (paginated by limit/offset; omitted when summaryOnly:true — use byDepthCounts for totals per depth, pagination object when truncated). Each item includes a processes:[{id,label,processType,step}] field listing the execution flows that symbol participates in. Empty when the symbol has no process membership. Can ALSO be empty when partial:true is set — either the process-aggregation pass hit its cap before detecting affected processes, or per-symbol enrichment was capped on a very large page. When partial:true, do NOT treat processes:[] as proof of no participation; cross-check the top-level affected_processes list.
Depth groups:
- d=1: WILL BREAK (direct callers/importers)

View file

@ -0,0 +1,14 @@
from fastapi import APIRouter
# Same module name as `api/users.py`. Before the long-key fix the
# basename `users` collided across packages, leaking `/users` (the
# prefix mounted on `api/users.py`) onto these admin routes.
# parse-impl now keys prefixes by `<dir>/<stem>` whenever the import
# statement carried enough context, so this file's `@router.get`
# routes must NOT be prefixed with `/users`.
router = APIRouter()
@router.get("/audit")
def audit():
return []

View file

@ -0,0 +1,8 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/list")
def list_calls():
return []

View file

@ -0,0 +1,13 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/list")
def list_users():
return []
@router.post("/create")
def create_user(payload):
return {"ok": True}

View file

@ -0,0 +1,13 @@
from fastapi import FastAPI
from api import users
from api.calls import router as calls_router
from .relative import router as rel_router
# Hostname is `application`, NOT `app` — exercises the unrestricted-host
# path through both the parse-worker regex and the group-layer
# tree-sitter pattern. Pinning to literal `app` would silently drop
# every prefix here.
application = FastAPI()
application.include_router(users.router, prefix="/users", tags=["users"])
application.include_router(calls_router, prefix="/calls")
application.include_router(rel_router, prefix="/rel")

View file

@ -0,0 +1,8 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/info")
def info():
return {}

View file

@ -0,0 +1,7 @@
#include "lib.h"
namespace caller {
void run() {
run_callback(utils::make_token);
}
}

View file

@ -0,0 +1,11 @@
#pragma once
namespace api {
struct Token {
friend void run_callback(Token t) {}
};
}
namespace utils {
api::Token make_token();
}

View file

@ -0,0 +1,7 @@
#include "lib.h"
namespace caller {
void run() {
run_callback(utils::worker);
}
}

View file

@ -0,0 +1,11 @@
#pragma once
namespace api {
struct Token {
friend void run_callback(Token t) {}
};
}
namespace utils {
void worker(api::Token token);
}

View file

@ -0,0 +1,27 @@
function transform(account) {
return account.id;
}
function predicate(account) {
return account.active;
}
// Control: a normal named function whose body calls `transform` directly.
// Proves the registry-primary resolver wires same-file free calls for this
// fixture, so the callback assertions below are not vacuous.
function run(account) {
return transform(account);
}
const accountsList = [];
// #1876: array higher-order-method callbacks at module scope. Pre-fix the JS
// scope model emitted a phantom `Function:exportData` / `Function:firstActive`
// for these callbacks (they match the HOC-wrapped-arrow declaration pattern),
// and calls INSIDE the callbacks (`transform`, `predicate`) attributed to that
// phantom Function. Post-fix the callback is no longer a `Function` def, so the
// inner calls fall through to the enclosing File scope.
const exportData = accountsList.map((account) => transform(account));
const firstActive = accountsList.find((account) => predicate(account));
module.exports = { run, exportData, firstActive };

View file

@ -9,7 +9,8 @@
* Seed data is NOT included — each test provides its own via options.seed.
*/
import path from 'path';
import { describe, beforeAll, afterAll } from 'vitest';
import { describe, beforeAll, beforeEach, afterAll } from 'vitest';
import { resolveAnalyzeInstallPolicy } from '../../src/core/lbug/extension-loader.js';
import { createTempDir, type TestDBHandle } from './test-db.js';
import { NODE_TABLES, EMBEDDING_TABLE_NAME } from '../../src/core/lbug/schema.js';
@ -73,6 +74,15 @@ export function withTestLbugDB(
// init on Windows CI regularly exceeds 30s due to native resource setup.
const timeout = options?.timeout ?? 120_000;
// Suites that seed FTS indexes need the optional FTS extension. It is not
// guaranteed on every machine (e.g. the macOS platform-sensitive CI runner,
// where it is neither pre-installed nor installable). Track availability so
// setup can skip FTS seeding instead of throwing, and so every test in the
// suite is skipped rather than failing against a missing index. (PR #1161.)
const ftsRequired = !!options?.ftsIndexes?.length;
let ftsAvailable = true;
let ftsSkipWarned = false;
const setup = async () => {
const tmpHandle = await createTempDir('gitnexus-lbug-');
const dbPath = path.join(tmpHandle.dbPath, 'lbug');
@ -84,6 +94,16 @@ export function withTestLbugDB(
// already open for this dbPath (no new native objects created).
await adapter.initLbug(dbPath);
// 1b. Probe the FTS extension for suites that need it, mirroring the
// analyze write path (`auto`: LOAD-first, then one bounded INSTALL).
// When it still cannot load, the suite is skipped (see beforeEach)
// and FTS seeding below is bypassed so setup never throws.
if (ftsRequired) {
ftsAvailable = await adapter.loadFTSExtension(undefined, {
policy: resolveAnalyzeInstallPolicy(),
});
}
// 2. Drop stale FTS indexes from previous test file
if (options?.ftsIndexes?.length) {
for (const idx of options.ftsIndexes) {
@ -108,8 +128,9 @@ export function withTestLbugDB(
}
}
// 5. Create FTS indexes on fresh data
if (options?.ftsIndexes?.length) {
// 5. Create FTS indexes on fresh data (only when the extension loaded;
// otherwise the suite is skipped via beforeEach below).
if (options?.ftsIndexes?.length && ftsAvailable) {
for (const idx of options.ftsIndexes) {
await adapter.createFTSIndex(idx.table, idx.indexName, idx.columns);
}
@ -166,6 +187,21 @@ export function withTestLbugDB(
// collisions when multiple withTestLbugDB calls share the same file.
describe(`withTestLbugDB(${prefix})`, () => {
beforeAll(setup, timeout);
// Skip FTS-dependent suites when the extension could not be loaded or
// installed on this machine. Without this, tests would assert against a
// missing index and fail. Warn once so the skip is visible, not silent.
beforeEach((ctx) => {
if (ftsRequired && !ftsAvailable) {
if (!ftsSkipWarned) {
ftsSkipWarned = true;
console.warn(
`[withTestLbugDB(${prefix})] Skipping FTS-dependent tests — the LadybugDB ` +
`FTS extension is unavailable (not pre-installed and could not be installed).`,
);
}
ctx.skip();
}
});
// Explicit timeout: KuzuDB's C++ destructor can hang on Windows during
// native resource cleanup. The vitest hookTimeout (120s) should apply
// automatically, but some vitest versions fall back to testTimeout (30s)

View file

@ -0,0 +1,252 @@
/**
* COBOL ingestion pipeline benchmark.
*
* Generates synthetic COBOL codebases at increasing scales and measures
* wall-clock time and peak heap through the full pipeline — scanning,
* preprocessing, COPY expansion, CALL resolution, and scope extraction.
*
* Run: GITNEXUS_BENCH=1 npx vitest run test/integration/cobol-pipeline-benchmark.test.ts
*
* Results are identical under both REGISTRY_PRIMARY_COBOL modes because
* cobolPhase runs in both modes. Under =1, scope-resolution is skipped for
* COBOL (standalone guard at phase.ts:164), so node/edge counts come entirely
* from the legacy cobolPhase.
*
* IMPORTANT — this benchmark measures scaling in FILE COUNT, so per-file work
* must stay constant as fileCount grows. Each program therefore COPYs a fixed
* number of shared copybooks (COPYBOOKS_PER_PROGRAM), independent of fileCount.
* Do NOT make every program COPY all copybooks: copybookCount grows as
* floor(fileCount/5), so copy-all makes emitted data-item nodes — and thus
* total work — O(fileCount²), which measures copybook fan-out rather than
* file-count scaling. The pipeline itself is O(fileCount) (verified: with
* constant fan-out, node count and wall-clock scale exactly linearly); the
* node-ratio assertion below guards against reintroducing the O(n²) pattern.
*/
import { describe, it, expect } from 'vitest';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { runPipelineFromRepo } from '../../src/core/ingestion/pipeline.js';
const BENCH_ENABLED = process.env.GITNEXUS_BENCH === '1';
interface BenchResult {
fileCount: number;
programCount: number;
paragraphCount: number;
copybookCount: number;
elapsedMs: number;
peakHeapMB: number;
nodeCount: number;
edgeCount: number;
}
function generateCobolFixture(
fileCount: number,
paragraphsPerProgram: number,
): { dir: string; programCount: number; paragraphCount: number; copybookCount: number } {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), `cobol-bench-${fileCount}-`));
const copybookDir = path.join(dir, 'copybooks');
fs.mkdirSync(copybookDir, { recursive: true });
const programCount = fileCount;
const paragraphCount = fileCount * paragraphsPerProgram;
// Generate shared copybooks (1 per 5 programs, at least 2)
const copybookCount = Math.max(2, Math.floor(fileCount / 5));
const copybookNames: string[] = [];
for (let c = 0; c < copybookCount; c++) {
const name = `BENCH${String(c + 1).padStart(4, '0')}`;
copybookNames.push(name);
const copyContent = [
` 01 ${name}-RECORD.`,
` 05 ${name}-KEY PIC X(10).`,
` 05 ${name}-VALUE PIC 9(08).`,
` 05 ${name}-FLAG PIC X(01).`,
'',
].join('\n');
fs.writeFileSync(path.join(copybookDir, `${name}.cpy`), copyContent);
}
for (let f = 0; f < fileCount; f++) {
const programName = `PGM${String(f + 1).padStart(4, '0')}`;
const paragraphs: string[] = [];
for (let p = 0; p < paragraphsPerProgram; p++) {
const paraName = `${String(p + 1).padStart(4, '0')}-PARA`;
// Every paragraph has a PERFORM to the next paragraph (or wraps around)
const nextParaIdx = (p + 1) % paragraphsPerProgram;
const nextParaName = `${String(nextParaIdx + 1).padStart(4, '0')}-PARA`;
const performLine = ` PERFORM ${nextParaName}.`;
// Cross-file CALL: every 3rd paragraph calls another program
const crossFileIdx = (f + p + 1) % fileCount;
const crossProgram = `PGM${String(crossFileIdx + 1).padStart(4, '0')}`;
const callLine =
p % 3 === 0
? ` CALL '${crossProgram}' USING ${copybookNames[p % copybookCount]}-KEY.`
: '';
// COPY in paragraphs adds preprocessing stress — non-idiomatic but
// exercises the preprocessor's expansion path per-paragraph.
const copyLine = ` COPY ${copybookNames[f % copybookCount]}.`;
paragraphs.push(
` ${paraName}.`,
copyLine,
performLine,
callLine,
` DISPLAY '${programName} ${paraName}'.`,
'',
);
}
// Each program COPYs a CONSTANT number of shared copybooks (independent of
// fileCount) so per-file work stays O(1) and the benchmark measures true
// file-count scaling. Copybooks are chosen by program index so they remain
// shared across programs (fan-in), still exercising cross-program copybook
// reuse and multi-COPY-per-program expansion. (Copying ALL copybooks here
// would make per-file work — and emitted data-item nodes — grow with
// fileCount, i.e. O(fileCount²); see the file header.)
const COPYBOOKS_PER_PROGRAM = 3;
const wsCopybooks = [
...new Set(
Array.from(
{ length: COPYBOOKS_PER_PROGRAM },
(_, k) => copybookNames[(f + k) % copybookCount],
),
),
];
const content = [
` IDENTIFICATION DIVISION.`,
` PROGRAM-ID. ${programName}.`,
` ENVIRONMENT DIVISION.`,
` DATA DIVISION.`,
` WORKING-STORAGE SECTION.`,
...wsCopybooks.map((n) => ` COPY ${n}.`),
` PROCEDURE DIVISION.`,
...paragraphs,
` STOP RUN.`,
` END PROGRAM ${programName}.`,
'',
].join('\n');
fs.writeFileSync(path.join(dir, `${programName}.cbl`), content);
}
return { dir, programCount, paragraphCount, copybookCount };
}
async function runBenchmark(
fileCount: number,
paragraphsPerProgram: number,
budgetMs: number,
): Promise<BenchResult> {
const { dir, programCount, paragraphCount, copybookCount } = generateCobolFixture(
fileCount,
paragraphsPerProgram,
);
let peakHeapMB = 0;
const heapSampler = setInterval(() => {
const heap = process.memoryUsage().heapUsed / 1024 / 1024;
if (heap > peakHeapMB) peakHeapMB = heap;
}, 50);
try {
const start = Date.now();
const result = await Promise.race([
runPipelineFromRepo(dir, () => {}, { skipGraphPhases: true }),
new Promise<never>((_, reject) =>
setTimeout(
() => reject(new Error(`Pipeline exceeded ${budgetMs}ms at ${fileCount} files`)),
budgetMs,
),
),
]);
const elapsedMs = Date.now() - start;
return {
fileCount,
programCount,
paragraphCount,
copybookCount,
elapsedMs,
peakHeapMB: Math.round(peakHeapMB),
nodeCount: result.graph.nodeCount,
edgeCount: result.graph.relationshipCount,
};
} finally {
clearInterval(heapSampler);
fs.rmSync(dir, { recursive: true, force: true });
}
}
function printResults(label: string, results: BenchResult[]) {
console.log(`\n${label}`);
console.log(
'┌──────────┬──────────┬────────────┬──────────┬───────────┬──────────┬───────┬───────┐',
);
console.log(
'│ Files │ Programs │ Paragraphs │ Copybooks│ Time (ms) │ Heap MB │ Nodes │ Edges │',
);
console.log(
'├──────────┼──────────┼────────────┼──────────┼───────────┼──────────┼───────┼───────┤',
);
for (const r of results) {
console.log(
`│ ${String(r.fileCount).padStart(8)} │ ${String(r.programCount).padStart(8)} │ ${String(r.paragraphCount).padStart(10)} │ ${String(r.copybookCount).padStart(8)} │ ${String(r.elapsedMs).padStart(9)} │ ${String(r.peakHeapMB).padStart(8)} │ ${String(r.nodeCount).padStart(5)} │ ${String(r.edgeCount).padStart(5)} │`,
);
}
console.log(
'└──────────┴──────────┴────────────┴──────────┴───────────┴──────────┴───────┴───────┘',
);
if (results.length >= 2) {
console.log('\nScaling ratios (time_ratio / file_ratio):');
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
const scaling = timeRatio / fileRatio;
console.log(
` ${results[i - 1].fileCount} \u2192 ${results[i].fileCount}: ${scaling.toFixed(2)}x (${scaling < 1.5 ? 'linear' : scaling < 3 ? 'superlinear' : 'WARNING: quadratic'})`,
);
}
}
}
describe.skipIf(!BENCH_ENABLED)('COBOL pipeline benchmark', () => {
it('scales with file count', async () => {
const scales = [100, 250, 500, 1000];
const results: BenchResult[] = [];
for (const fileCount of scales) {
const paragraphsPerProgram = 3;
const result = await runBenchmark(fileCount, paragraphsPerProgram, 300_000);
results.push(result);
console.log(
` ${fileCount} files: ${result.elapsedMs}ms, ${result.peakHeapMB}MB heap, ${result.nodeCount} nodes, ${result.edgeCount} edges`,
);
}
printResults('COBOL Pipeline', results);
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
// Wall-clock is noisy (GC/CI load); keep a coarse upper bound here.
expect(timeRatio / fileRatio).toBeLessThan(4);
// Deterministic regression guard: with constant per-program copybook
// fan-out the emitted node count is exactly linear in fileCount
// (ratio ≈ 1.0). If someone reintroduces O(fileCount²) work — e.g. by
// making every program COPY all copybooks — node growth jumps to ~2x
// per file-doubling and this fails. Node count is deterministic, so
// this is a non-flaky guard unlike the wall-clock check above.
const nodeRatio = results[i].nodeCount / results[i - 1].nodeCount;
expect(nodeRatio / fileRatio).toBeLessThan(1.3);
}
}, 600_000);
});

View file

@ -0,0 +1,253 @@
/**
* C# ingestion pipeline benchmark.
*
* Generates synthetic C# codebases at increasing scales and measures
* wall-clock time and peak heap through the full pipeline — parsing,
* scope extraction, C# namespace-siblings (same-namespace cross-file
* visibility, using-static, cross-namespace imports), and call
* resolution.
*
* Mirrors test/integration/php-pipeline-benchmark.test.ts. Two shapes:
* 1. "spread" — files distributed across many namespaces (the common
* case; each namespace bucket stays small).
* 2. "concentrated" — every file in the SAME (or global/no) namespace,
* so a single namespace bucket holds all type defs. This is the
* shape that drove the Unity-solution OOM: `populateCsharpNamespaceSiblings`
* materialises O(scopes × defs) BindingRefs into that one bucket.
* The concentrated test is the regression guard for that path.
*
* Run: GITNEXUS_BENCH=1 npx vitest run test/integration/csharp-pipeline-benchmark.test.ts
*/
import { describe, it, expect } from 'vitest';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { runPipelineFromRepo } from '../../src/core/ingestion/pipeline.js';
const BENCH_ENABLED = process.env.GITNEXUS_BENCH === '1';
interface BenchResult {
fileCount: number;
classCount: number;
namespaceCount: number;
elapsedMs: number;
peakHeapMB: number;
nodeCount: number;
edgeCount: number;
}
type FixtureShape = 'spread' | 'concentrated';
function generateCsharpFixture(
fileCount: number,
namespacesPerLevel: number,
shape: FixtureShape,
): { dir: string; classCount: number; namespaceCount: number } {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), `csharp-bench-${shape}-${fileCount}-`));
// "spread": square grid of namespaces. "concentrated": a single
// global (no-namespace) bucket so every type lands in the `''` bucket
// — the OOM-prone path.
const namespaces: string[] = [];
if (shape === 'spread') {
for (let i = 0; i < namespacesPerLevel; i++) {
for (let j = 0; j < namespacesPerLevel; j++) {
namespaces.push(`App.Module${i}.Sub${j}`);
}
}
} else {
namespaces.push(''); // global / no namespace declaration
}
const classCount = fileCount;
const namespaceCount = namespaces.length;
for (let f = 0; f < fileCount; f++) {
const ns = namespaces[f % namespaces.length]!;
const className = `Class${f}`;
// Concentrated files share a flat directory; spread files mirror the
// namespace as a directory tree (matches typical C# project layout).
const targetDir = ns === '' ? dir : path.join(dir, ns.replace(/\./g, '/'));
fs.mkdirSync(targetDir, { recursive: true });
const siblingIdx = (f + 1) % fileCount;
const siblingClass = `Class${siblingIdx}`;
const crossNsIdx = (f + Math.floor(fileCount / 3)) % fileCount;
const crossNs = namespaces[crossNsIdx % namespaces.length]!;
const crossClass = `Class${crossNsIdx}`;
const usesCross = ns !== '' && ns !== crossNs;
const body = [
ns !== '' ? `namespace ${ns};` : '',
usesCross ? `using ${crossNs};` : '',
'',
`public class ${className}`,
'{',
' private int id;',
' private string name;',
'',
' public int GetId()',
' {',
' return this.id;',
' }',
'',
` public ${siblingClass} Process()`,
' {',
` var sibling = new ${siblingClass}();`,
' return sibling;',
' }',
usesCross
? [
'',
` public ${crossClass} CrossCall()`,
' {',
` var cross = new ${crossClass}();`,
' cross.GetId();',
' return cross;',
' }',
].join('\n')
: '',
'}',
'',
]
.filter(Boolean)
.join('\n');
fs.writeFileSync(path.join(targetDir, `${className}.cs`), body);
}
// Minimal SDK-style csproj so the C# project-loading phase engages
// (matches the real-world Unity/.NET solution path).
const csproj = [
'<Project Sdk="Microsoft.NET.Sdk">',
' <PropertyGroup>',
' <TargetFramework>net8.0</TargetFramework>',
' <Nullable>enable</Nullable>',
' </PropertyGroup>',
'</Project>',
'',
].join('\n');
fs.writeFileSync(path.join(dir, 'Bench.csproj'), csproj);
return { dir, classCount, namespaceCount };
}
async function runBenchmark(
fileCount: number,
nsLevels: number,
shape: FixtureShape,
budgetMs: number,
): Promise<BenchResult> {
const { dir, classCount, namespaceCount } = generateCsharpFixture(fileCount, nsLevels, shape);
let peakHeapMB = 0;
const heapSampler = setInterval(() => {
const heap = process.memoryUsage().heapUsed / 1024 / 1024;
if (heap > peakHeapMB) peakHeapMB = heap;
}, 50);
let budgetTimer: ReturnType<typeof setTimeout> | undefined;
try {
const start = Date.now();
const result = await Promise.race([
runPipelineFromRepo(dir, () => {}, { skipGraphPhases: true }),
new Promise<never>((_, reject) => {
budgetTimer = setTimeout(
() =>
reject(new Error(`Pipeline exceeded ${budgetMs}ms at ${fileCount} files (${shape})`)),
budgetMs,
);
}),
]);
const elapsedMs = Date.now() - start;
return {
fileCount,
classCount,
namespaceCount,
elapsedMs,
peakHeapMB: Math.round(peakHeapMB),
nodeCount: result.graph.nodeCount,
edgeCount: result.graph.relationshipCount,
};
} finally {
clearInterval(heapSampler);
clearTimeout(budgetTimer);
fs.rmSync(dir, { recursive: true, force: true });
}
}
function printResults(label: string, results: BenchResult[]) {
console.log(`\n${label}`);
console.log('┌──────────┬─────────┬──────────┬───────────┬──────────┬───────┬───────┐');
console.log('│ Files │ Classes │ NS Count │ Time (ms) │ Heap MB │ Nodes │ Edges │');
console.log('├──────────┼─────────┼──────────┼───────────┼──────────┼───────┼───────┤');
for (const r of results) {
console.log(
`│ ${String(r.fileCount).padStart(8)} │ ${String(r.classCount).padStart(7)} │ ${String(r.namespaceCount).padStart(8)} │ ${String(r.elapsedMs).padStart(9)} │ ${String(r.peakHeapMB).padStart(8)} │ ${String(r.nodeCount).padStart(5)} │ ${String(r.edgeCount).padStart(5)} │`,
);
}
console.log('└──────────┴─────────┴──────────┴───────────┴──────────┴───────┴───────┘');
if (results.length >= 2) {
console.log('\nScaling ratios (time_ratio / file_ratio):');
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
const scaling = timeRatio / fileRatio;
console.log(
` ${results[i - 1].fileCount} → ${results[i].fileCount}: ${scaling.toFixed(2)}x (${scaling < 1.5 ? 'linear' : scaling < 3 ? 'superlinear' : 'WARNING: quadratic'})`,
);
}
}
}
describe.skipIf(!BENCH_ENABLED)('C# pipeline benchmark', () => {
it('scales with file count — namespaces spread across the solution', async () => {
const scales = [100, 250, 500];
const results: BenchResult[] = [];
for (const fileCount of scales) {
const nsLevels = Math.max(2, Math.ceil(Math.sqrt(fileCount / 4)));
const result = await runBenchmark(fileCount, nsLevels, 'spread', 180_000);
results.push(result);
console.log(
` ${fileCount} files: ${result.elapsedMs}ms, ${result.peakHeapMB}MB heap, ${result.nodeCount} nodes, ${result.edgeCount} edges`,
);
}
printResults('C# Pipeline — Namespaces Spread', results);
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
expect(timeRatio / fileRatio).toBeLessThan(3);
}
}, 600_000);
it('scales with file count — all types in one (global) namespace bucket', async () => {
// Regression guard for the Unity-solution OOM: a single namespace
// bucket holds every type def, so naive per-scope binding
// materialisation is O(files²). Time must stay sub-quadratic and the
// run must not OOM.
const scales = [100, 250, 500];
const results: BenchResult[] = [];
for (const fileCount of scales) {
const result = await runBenchmark(fileCount, 1, 'concentrated', 180_000);
results.push(result);
console.log(
` ${fileCount} files: ${result.elapsedMs}ms, ${result.peakHeapMB}MB heap, ${result.nodeCount} nodes, ${result.edgeCount} edges`,
);
}
printResults('C# Pipeline — Concentrated Global Namespace', results);
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
expect(timeRatio / fileRatio).toBeLessThan(3);
}
}, 600_000);
});

View file

@ -0,0 +1,125 @@
/**
* End-to-end coverage of the FastAPI `include_router(prefix=…)` fix.
*
* The PR claims to update both layers — ingestion (graph `Route`
* nodes) and group (HTTP contracts). The group side is exercised by
* `test/unit/group/http-route-extractor.test.ts`; this file pins the
* **ingestion** side by running the full pipeline against a realistic
* fixture and inspecting the resulting `Route` graph nodes.
*
* What this test pins:
*
* 1. **Shape A** (`from api import users` +
* `application.include_router(users.router, prefix='/users')`)
* produces `Route` nodes whose `name` is the prefixed full path
* (`/users/list`, `/users/create`) — not the bare decorator path.
*
* 2. **Shape B with relative import**
* (`from .calls import router as calls_router` +
* `application.include_router(calls_router, prefix='/calls')`)
* works end-to-end. Before the regex fix, the leading-dot module
* path was rejected and the prefix was silently dropped.
*
* 3. **Same-name modules in different packages** do not bleed
* prefixes. Before the long-key fix, both `api/users.py` and
* `admin/users.py` shared the basename `users`, so `admin/users`
* routes inherited the `/users` prefix that was only meant for
* `api/users.py`.
*
* 4. **Non-`app` host names** (`application = FastAPI()`) work in the
* ingestion regex. The group-layer counterpart is pinned by the
* `non-app host` cases in `http-route-extractor.test.ts`.
*
* The fixture lives at `test/fixtures/fastapi-prefix-app/` so the
* pipeline can scan a real on-disk repo (mirroring how `gitnexus
* analyze` is used in production) and so reviewers can inspect the
* inputs without reading test source.
*/
import { describe, it, expect, beforeAll } from 'vitest';
import path from 'node:path';
import { runPipelineFromRepo } from '../../src/core/ingestion/pipeline.js';
import type { PipelineResult } from '../../types/pipeline.js';
const FIXTURE = path.resolve(__dirname, '..', 'fixtures', 'fastapi-prefix-app');
describe('FastAPI include_router(prefix=…) — ingestion pipeline', () => {
let result: PipelineResult;
beforeAll(async () => {
// Force the worker-pool code path on this small fixture (~5
// files). Without this the pipeline takes the sequential
// fallback, which historically does NOT run the FastAPI router
// bindings extractor — the very behaviour we want to pin lives
// exclusively inside the worker entry point.
result = await runPipelineFromRepo(FIXTURE, () => {}, {
workerThresholdsForTest: { minFiles: 1, minBytes: 1 },
});
}, 60_000);
function routeNames(): string[] {
const out: string[] = [];
result.graph.forEachNode((n) => {
if (n.label === 'Route') out.push(String(n.properties.name));
});
return out.sort();
}
it('joins Shape-A `<mod>.router` prefixes with sub-router decorator paths', () => {
// `application.include_router(users.router, prefix='/users')` in
// main.py + `@router.get('/list')` / `@router.post('/create')` in
// api/users.py → `/users/list`, `/users/create`.
const names = routeNames();
expect(names).toContain('/users/list');
expect(names).toContain('/users/create');
// The bare decorator paths must NOT survive when a prefix
// mapping exists — one router yields exactly one Route node per
// prefix, not the prefixed AND the unprefixed copy.
expect(names.filter((n) => n === '/list')).toHaveLength(0);
expect(names.filter((n) => n === '/create')).toHaveLength(0);
});
it('joins Shape-B with absolute named import (`from api.calls import router as …`)', () => {
// main.py mounts `api.calls` under `/calls`; api/calls.py has
// `@router.get('/list')`. Long-key resolution is required here:
// `users` (under `/users`) and `calls` are distinct module
// basenames, but the long key (`api/calls`) is what makes the
// binding deterministic.
const names = routeNames();
expect(names).toContain('/calls/list');
});
it('joins Shape-B with relative import (`from .relative import router as …`)', () => {
// FINDING 2: the worker regex `[A-Za-z_][\w.]*` used to reject
// module paths starting with `.`, silently dropping every
// leading-dot relative import. relative.py declares `/info`; the
// expected joined route is `/rel/info`.
const names = routeNames();
expect(names).toContain('/rel/info');
});
it('does NOT bleed `/users` prefix onto the same-name `admin/users.py`', () => {
// FINDING 3: `api/users.py` and `admin/users.py` collide on the
// short module key `users`. main.py only mounts the `api/users`
// router under `/users`, so the admin file's `@router.get('/audit')`
// must surface as the bare `/audit` — never as `/users/audit`.
const names = routeNames();
expect(names).toContain('/audit');
expect(names.filter((n) => n === '/users/audit')).toHaveLength(0);
});
it('emits exactly one Route node per (router method, prefix) pair', () => {
// Defence-in-depth: counts the unique route nodes for the
// prefixed routes to make sure the duplication path in
// parse-impl (`for prefix of prefixes`) didn't accidentally
// double-emit when only a single prefix was registered.
const names = routeNames();
const counts = new Map<string, number>();
for (const n of names) counts.set(n, (counts.get(n) ?? 0) + 1);
expect(counts.get('/users/list')).toBe(1);
expect(counts.get('/users/create')).toBe(1);
expect(counts.get('/calls/list')).toBe(1);
expect(counts.get('/rel/info')).toBe(1);
expect(counts.get('/audit')).toBe(1);
});
});

View file

@ -0,0 +1,81 @@
/**
* JavaScript: CALLS-edge attribution for calls inside array higher-order-
* method callbacks (issue #1876).
*
* `const exportData = accountsList.map(account => transform(account))` matches
* the HOC-wrapped-arrow declaration pattern, so before this fix the JS scope
* model emitted a phantom `Function:exportData` for the `.map` callback (on
* top of the value binding). Calls nested in the callback (`transform`) then
* attributed to that phantom `Function` instead of the enclosing scope.
*
* U1 drops the `@declaration.function` for array-method callbacks, so the
* binding is value-only and the inner call falls through to the File scope —
* exactly the Zustand module-level-call behavior already pinned for TS.
*
* SCOPE: this asserts the registry-primary CALLS-edge ATTRIBUTION change only.
* The duplicate *graph node* (`Function:exportData`) is created by the legacy
* parse-worker node path, which this change does not touch; collapsing it is
* the deferred node-creation migration. Accordingly this file makes NO node-
* count assertion.
*
* Registry-primary-only correctness win: under the forced-legacy parity flag
* (`REGISTRY_PRIMARY_JAVASCRIPT=0`) the legacy DAG still emits the phantom
* attribution, so the suite is skipped there (mirrors the per-language
* expected-failure handling in `resolvers/helpers.ts`).
*/
import { describe, it, expect, beforeAll } from 'vitest';
import path from 'path';
import {
FIXTURES,
getRelationships,
isLegacyResolverParityRun,
runPipelineFromRepo,
type PipelineResult,
} from './resolvers/helpers.js';
describe.skipIf(isLegacyResolverParityRun('javascript'))(
'JavaScript array-method-callback CALLS attribution (#1876)',
() => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'javascript-array-method-callback'),
() => {},
);
}, 60000);
it('control: run() body calls transform directly (resolver is wired)', () => {
const calls = getRelationships(result, 'CALLS').filter((c) => c.target === 'transform');
expect(calls.map((c) => `${c.source} → ${c.target}`)).toContain('run → transform');
});
it('call inside .map callback attributes to File, not a phantom Function:exportData', () => {
const calls = getRelationships(result, 'CALLS').filter((c) => c.target === 'transform');
const fromExportData = calls.filter((c) => c.source === 'exportData');
expect(
fromExportData,
'transform must NOT be attributed to exportData (phantom Function)',
).toEqual([]);
const fromFile = calls.filter((c) => c.sourceLabel === 'File');
expect(
fromFile,
'the .map callback call to transform must source from the File node (exactly once)',
).toHaveLength(1);
});
it('call inside .find callback attributes to File, not a phantom Function:firstActive', () => {
const calls = getRelationships(result, 'CALLS').filter((c) => c.target === 'predicate');
const fromFirstActive = calls.filter((c) => c.source === 'firstActive');
expect(
fromFirstActive,
'predicate must NOT be attributed to firstActive (phantom Function)',
).toEqual([]);
const fromFile = calls.filter((c) => c.sourceLabel === 'File');
expect(
fromFile,
'the .find callback call to predicate must source from the File node (exactly once)',
).toHaveLength(1);
});
},
);

View file

@ -23,6 +23,26 @@ import { withTestLbugDB } from '../helpers/test-indexed-db.js';
*/
const itLbugReopen = process.platform === 'win32' ? it.skip : it;
/**
* The FTS extension is optional and defaults to a `load-only` install policy
* (PR #1161 — offline-first), so on a machine where it was never pre-installed
* it cannot load. The tests below exercise the FTS *primitives* directly and
* have nothing to assert without the extension — skip them rather than fail.
* Graceful degradation when FTS is unavailable is covered at the analyze /
* query layer (see run-analyze.ts and the BM25 fallback tests).
*/
const FTS_UNAVAILABLE_NOTE =
'FTS extension unavailable (load-only policy; not pre-installed on this machine)';
/**
* Dynamically skip an FTS-primitive test when the extension cannot load.
* `ctx.skip()` aborts the test, so callers should `await` this first thing.
*/
const skipUnlessFtsAvailable = async (ctx: { skip: (note?: string) => void }): Promise<void> => {
const { loadFTSExtension } = await import('../../src/core/lbug/lbug-adapter.js');
if (!(await loadFTSExtension())) ctx.skip(FTS_UNAVAILABLE_NOTE);
};
// ─── Core LadybugDB Adapter ─────────────────────────────────────────────
withTestLbugDB(
@ -47,7 +67,8 @@ withTestLbugDB(
expect(folderRows).toHaveLength(1);
});
it('createFTSIndex: creates FTS index on Function table without error', async () => {
it('createFTSIndex: creates FTS index on Function table without error', async (ctx) => {
await skipUnlessFtsAvailable(ctx);
const { createFTSIndex } = await import('../../src/core/lbug/lbug-adapter.js');
await expect(
@ -55,7 +76,8 @@ withTestLbugDB(
).resolves.toBeUndefined();
});
it('loadFTSExtension(conn): loads on an explicit connection and returns true', async () => {
it('loadFTSExtension(conn): loads on an explicit connection and returns true', async (ctx) => {
await skipUnlessFtsAvailable(ctx);
const lbug = (await import('@ladybugdb/core')).default;
const { loadFTSExtension, getDatabase } =
await import('../../src/core/lbug/lbug-adapter.js');
@ -119,7 +141,8 @@ withTestLbugDB(
});
describe('error handling', () => {
it('createFTSIndex handles already-existing index gracefully', async () => {
it('createFTSIndex handles already-existing index gracefully', async (ctx) => {
await skipUnlessFtsAvailable(ctx);
const { createFTSIndex } = await import('../../src/core/lbug/lbug-adapter.js');
// First call creates the index (may already exist from earlier test)
@ -131,7 +154,8 @@ withTestLbugDB(
).resolves.toBeUndefined();
});
it('ensureFTSIndex is idempotent and caches across writable calls (#1224)', async () => {
it('ensureFTSIndex is idempotent and caches across writable calls (#1224)', async (ctx) => {
await skipUnlessFtsAvailable(ctx);
const { ensureFTSIndex } = await import('../../src/core/lbug/lbug-adapter.js');
// First call creates the index. Second call must short-circuit on the
@ -174,7 +198,8 @@ withTestLbugDB(
itLbugReopen(
'initLbug loads FTS so reopened HTTP-style sessions can query existing indexes',
async () => {
async (ctx) => {
await skipUnlessFtsAvailable(ctx);
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const indexName = 'function_fts_init_probe';

View file

@ -131,6 +131,8 @@ const accumulated = {
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],

View file

@ -14,7 +14,7 @@ import path from 'path';
import fs from 'fs';
import { emitCobolScopeCaptures } from '../../../src/core/ingestion/languages/cobol/captures.js';
const FIXTURES = path.resolve(process.cwd(), 'test/fixtures/cobol');
const FIXTURES = path.resolve(__dirname, '..', '..', 'fixtures', 'cobol');
// ---------------------------------------------------------------------------
// Helpers

View file

@ -18,6 +18,12 @@ import {
runPipelineFromRepo,
type PipelineResult,
} from './helpers.js';
import { isRegistryPrimary } from '../../../src/core/ingestion/registry-primary-flag.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { extractParsedFile } from '../../../src/core/ingestion/scope-extractor-bridge.js';
import { cobolProvider } from '../../../src/core/ingestion/languages/cobol.js';
const isPrimary = isRegistryPrimary(SupportedLanguages.Cobol);
describe('COBOL full system extraction', () => {
let result: PipelineResult;
@ -715,4 +721,48 @@ describe('COBOL full system extraction', () => {
expect(getRelationships(result, 'ACCESSES').length).toBe(25);
});
});
// =====================================================================
// SCOPE-RESOLUTION MODE: when REGISTRY_PRIMARY_COBOL=1, the scope-
// resolution pipeline produces captures from standalone providers.
// These tests verify that the scope-resolution output matches expected
// capture counts for the cobol-app fixture.
// =====================================================================
describe('scope-resolution mode', () => {
// Scope-resolution captures are only produced when registry-primary
// flips COBOL into the scope-resolution pipeline (REGISTRY_PRIMARY_COBOL=1).
// Under legacy mode (=0), the legacy cobolPhase produces graph edges
// tested above — scope-resolution captures are not expected.
it('scope-resolution pipeline produces capture output when REGISTRY_PRIMARY_COBOL=1', () => {
if (!isPrimary) {
// Legacy mode (REGISTRY_PRIMARY_COBOL=0): scope-resolution phases
// are skipped (skipGraphPhases=true), so parsedFiles is not populated.
return;
}
// Registry-primary mode: standalone provider wiring in parse-worker
// produces scope captures via emitCobolScopeCaptures
expect(result.graph).not.toBeNull();
expect(Object.keys(result.graph.nodes ?? {}).length).toBeGreaterThan(0);
});
it('extractParsedFile works for standalone COBOL provider', () => {
const source = `
IDENTIFICATION DIVISION.
PROGRAM-ID. TESTPROG.
PROCEDURE DIVISION.
DISPLAY 'hello'.
STOP RUN.
END PROGRAM TESTPROG.
`;
const parsedFile = extractParsedFile(cobolProvider, source, 'TESTPROG.cbl', () => {});
expect(parsedFile).not.toBeNull();
// Use toBe for strict equality — not.toBeNull() per DoD
expect(parsedFile!.scopes.length).toBeGreaterThan(0);
expect(typeof parsedFile!.moduleScope).toBe('string');
expect(parsedFile!.moduleScope.length).toBeGreaterThan(0);
});
});
});

View file

@ -2963,41 +2963,66 @@ describe('C++ ADL — block-scope function declaration suppresses ADL', () => {
});
// ---------------------------------------------------------------------------
// ADL V2 — free-function reference args contribute their namespace.
// ADL V2 - strict function-type associated entities.
//
// GitNexus approximation (not strict ISO C++ ADL): when a qualified_identifier
// like `utils::worker` is passed as an argument, GitNexus contributes the
// enclosing namespace (`utils`) to the associated set, provided a Function or
// Method named `worker` is found in the `utils` namespace at resolution time.
// Under ISO C++ [basic.lookup.argdep] the associated entities for a function-type
// argument come from the parameter types and return type of the overload set —
// NOT the function's enclosing namespace. For `void worker()`, the standard-
// compliant associated set is empty. The approximation captures the dominant
// real-world pattern (pass a utility function → find its sibling) at the cost
// of potential false positives when an unrelated function with the same simple
// name exists in the same namespace (bounded by the workspace-function lookup).
// Function-reference arguments follow strict ISO C++ ADL: GitNexus walks the
// referenced overload set's parameter and return types instead of contributing
// the referenced function's enclosing namespace.
// For `void worker()`, the associated set is empty; for `void worker(api::Token)`
// or `api::Token make_token()`, `api` is associated through `Token`.
// ---------------------------------------------------------------------------
describe('C++ ADL — qualified free-function reference contributes its namespace', () => {
describe('C++ ADL - free-function reference does not contribute its namespace', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(path.join(FIXTURES, 'cpp-adl-free-func-ref'), () => {});
}, 60000);
it('with_callback(utils::worker) resolves to utils::with_callback via ADL', () => {
it('with_callback(utils::worker) emits zero CALLS edges when worker has no class parameter or return type', () => {
const calls = getRelationships(result, 'CALLS');
const cbCalls = calls.filter((c) => c.source === 'run' && c.target === 'with_callback');
// Ordinary lookup inside caller::run finds nothing (no `using`, no local
// declaration). utils::worker is a qualified_identifier argument, so ADL
// contributes `utils` to the associated-namespace set. utils::with_callback
// is then discovered as the sole candidate.
expect(cbCalls.length).toBe(1);
expect(cbCalls[0].targetFilePath).toContain('utils.h');
expect(cbCalls.length).toBe(0);
});
});
describe('C++ ADL — overloaded free-function reference does not crash', () => {
describe('C++ ADL - free-function reference contributes parameter-type associated namespace', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'cpp-adl-free-func-ref-strict'),
() => {},
);
}, 60000);
it('run_callback(utils::worker) resolves hidden friend through worker(api::Token)', () => {
const calls = getRelationships(result, 'CALLS');
const cbCalls = calls.filter((c) => c.source === 'run' && c.target === 'run_callback');
expect(cbCalls.length).toBe(1);
expect(cbCalls[0].targetFilePath).toContain('lib.h');
});
});
describe('C++ ADL - free-function reference contributes return-type associated namespace', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'cpp-adl-free-func-ref-return-strict'),
() => {},
);
}, 60000);
it('run_callback(utils::make_token) resolves hidden friend through api::Token return type', () => {
const calls = getRelationships(result, 'CALLS');
const cbCalls = calls.filter((c) => c.source === 'run' && c.target === 'run_callback');
expect(cbCalls.length).toBe(1);
expect(cbCalls[0].targetFilePath).toContain('lib.h');
});
});
describe('C++ ADL - overloaded free-function reference stays strict', () => {
let result: PipelineResult;
beforeAll(async () => {
@ -3007,15 +3032,10 @@ describe('C++ ADL — overloaded free-function reference does not crash', () =>
);
}, 60000);
it('with_callback(utils::worker) with overloaded utils::worker still resolves utils::with_callback via ADL', () => {
it('with_callback(utils::worker) with overloaded utils::worker still emits zero CALLS edges', () => {
const calls = getRelationships(result, 'CALLS');
const cbCalls = calls.filter((c) => c.source === 'run' && c.target === 'with_callback');
// utils::worker has two overloads (worker() and worker(int)). V1
// simplification: contribute the namespace if ANY overload exists in the
// workspace, regardless of which one would be selected. The namespace
// `utils` is still added, and utils::with_callback is discovered.
expect(cbCalls.length).toBe(1);
expect(cbCalls[0].targetFilePath).toContain('utils.h');
expect(cbCalls.length).toBe(0);
});
});
@ -3035,10 +3055,10 @@ describe('C++ ADL — namespace-qualified variable arg does NOT contribute names
// data::value is a namespace-qualified integer variable. tree-sitter-cpp
// produces a qualified_identifier AST node regardless of whether `value`
// denotes a function, variable, enum, or static member. The GitNexus guard
// in collectFunctionRefNamespaces verifies that a Function/Method named
// `value` exists in the `data` namespace before contributing it. Since
// `data::value` is an int variable, `data` is never added to the associated
// set, so data::process is never found as an ADL candidate.
// in collectFunctionTypeAssociatedNamespaces verifies that a Function/Method
// named `value` exists in the `data` namespace before walking any function
// type. Since `data::value` is an int variable, no function type is walked,
// so data::process is never found as an ADL candidate.
expect(processCalls.length).toBe(0);
});
});

View file

@ -2447,9 +2447,12 @@ describe('C# class-name receiver write ACCESSES (merged Case 2 kind-aware branch
// cross-namespace `using` and a colliding local class. Pins both fixes in
// the resolver dataset:
//
// 1. emitCsharpScopeCaptures + extractFileStructure must use the adaptive
// `getTreeSitterBufferSize` on cache miss, otherwise UserService.cs
// fails to reparse with "Invalid argument" and CreateUser is dropped.
// 1. emitCsharpScopeCaptures must use the adaptive `getTreeSitterBufferSize`
// on cache miss, otherwise UserService.cs fails to reparse with "Invalid
// argument" and CreateUser is dropped. (extractFileStructure no longer
// re-parses on cache miss — it uses the line scanner,
// extractCsharpStructureViaScanner — so this fixture's line-anchored
// namespaces are read identically by either branch.)
// 2. populateCsharpNamespaceSiblings must append to bindingAugmentations
// instead of mutating frozen finalize-produced BindingRef[] arrays;
// otherwise the cross-namespace inject loop throws "Cannot add property

View file

@ -363,6 +363,11 @@ const LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES: Readonly<Record<string, Readonly
// Scope-resolver-only correctness wins; backporting is out of scope.
'process(data::value) emits zero CALLS edges \u2014 data::value is a variable, not a function',
'run_with(callback) emits zero CALLS edges when callback is a parameter, not a function reference',
// PR #1633: strict function-type ADL no longer contributes the referenced
// function's enclosing namespace. The legacy DAG still resolves these via
// simple-name global fallback.
'with_callback(utils::worker) emits zero CALLS edges when worker has no class parameter or return type',
'with_callback(utils::worker) with overloaded utils::worker still emits zero CALLS edges',
// PR #1599 adversarial review findings: nearest-scope ADL blocker
// semantics and block-scope function declaration ADL suppression are
// scope-resolver-only. The legacy DAG has no scope-aware ADL blocker

View file

@ -647,6 +647,229 @@ describe('LocalBackend.callTool', () => {
expect(result.target).toBeDefined();
});
it('impact byDepth items include a processes field (default empty when no processes)', async () => {
// Resolver returns target; BFS returns one frontier caller; no STEP_IN_PROCESS rows.
(executeParameterized as any).mockResolvedValue([
{ id: 'func:main', name: 'main', type: 'Function', filePath: 'src/index.ts' },
]);
(executeQuery as any).mockResolvedValue([
{
id: 'func:caller',
name: 'caller',
type: 'Function',
filePath: 'src/uses-main.ts',
relType: 'CALLS',
confidence: 0.9,
},
]);
const result = await backend.callTool('impact', { target: 'main', direction: 'upstream' });
const d1 = result.byDepth?.[1] || result.byDepth?.['1'] || [];
expect(d1.length).toBeGreaterThan(0);
for (const item of d1) {
expect(item).toHaveProperty('processes');
expect(Array.isArray(item.processes)).toBe(true);
}
});
it('impact populates byDepth processes when STEP_IN_PROCESS rows exist', async () => {
(executeParameterized as any).mockImplementation((_repoId: string, cypher: string) => {
// Symbol resolver name-lookup
if (cypher.includes('WHERE n.name =')) {
return Promise.resolve([
{ id: 'func:main', name: 'main', type: 'Function', filePath: 'src/index.ts' },
]);
}
// Aggregation pass (must return at least one row so per-symbol pass is gated open)
if (cypher.includes('COUNT(DISTINCT s.id)')) {
return Promise.resolve([
{
pId: 'proc:cron_daily',
name: 'Daily cron',
heuristicLabel: 'Daily cron',
processType: 'cron',
entryPointId: 'func:cron_entry',
hits: 1,
minStep: 1,
stepCount: 5,
epName: 'cron_entry',
epType: 'Function',
epFilePath: 'src/cron.ts',
},
]);
}
// New per-symbol pass added by this change
if (cypher.includes('RETURN s.id AS sid')) {
return Promise.resolve([
{
sid: 'func:caller',
pid: 'proc:cron_daily',
pName: 'Daily cron',
pType: 'cron',
step: 2,
},
]);
}
return Promise.resolve([]);
});
(executeQuery as any).mockResolvedValue([
{
id: 'func:caller',
name: 'caller',
type: 'Function',
filePath: 'src/uses-main.ts',
relType: 'CALLS',
confidence: 0.9,
},
]);
const result = await backend.callTool('impact', { target: 'main', direction: 'upstream' });
const d1 = result.byDepth?.[1] || result.byDepth?.['1'] || [];
const caller = d1.find((it: any) => it.id === 'func:caller');
expect(caller).toBeDefined();
expect(caller.processes).toHaveLength(1);
expect(caller.processes[0]).toMatchObject({
id: 'proc:cron_daily',
label: 'Daily cron',
processType: 'cron',
step: 2,
});
});
it('impact summaryOnly:true skips the per-symbol STEP_IN_PROCESS enrichment pass', async () => {
// Resolver returns target; BFS returns one caller; aggregation returns one process row.
(executeParameterized as any).mockImplementation((_repoId: string, cypher: string) => {
if (cypher.includes('WHERE n.name =')) {
return Promise.resolve([
{ id: 'func:main', name: 'main', type: 'Function', filePath: 'src/index.ts' },
]);
}
if (cypher.includes('COUNT(DISTINCT s.id)')) {
return Promise.resolve([
{
pId: 'proc:daily',
name: 'Daily cron',
heuristicLabel: 'Daily cron',
processType: 'cron',
entryPointId: 'func:cron_entry',
hits: 1,
minStep: 1,
stepCount: 5,
epName: 'cron_entry',
epType: 'Function',
epFilePath: 'src/cron.ts',
},
]);
}
return Promise.resolve([]);
});
(executeQuery as any).mockResolvedValue([
{
id: 'func:caller',
name: 'caller',
type: 'Function',
filePath: 'src/a.ts',
relType: 'CALLS',
confidence: 0.9,
},
]);
const result = await backend.callTool('impact', {
target: 'main',
direction: 'upstream',
summaryOnly: true,
});
// summaryOnly should return base fields only, no byDepth
expect(result.summary).toBeDefined();
expect(result.byDepth).toBeUndefined();
// The per-symbol enrichment query contains 'RETURN s.id AS sid'; verify it
// was never called (the gate should have suppressed it).
const perSymbolCalls = (executeParameterized as any).mock.calls.filter(
([, cypher]: [string, string]) =>
typeof cypher === 'string' && cypher.includes('RETURN s.id AS sid'),
);
expect(perSymbolCalls).toHaveLength(0);
});
it('impactByUid preserves byDepth while skipping per-symbol enrichment (group fan-out)', async () => {
// Regression guard for the cross-repo by_depth contract: impactByUid must
// suppress only the per-symbol STEP_IN_PROCESS pass, NOT the whole byDepth
// field. cross-impact.ts reads fan.byDepth to populate group `by_depth`;
// using summaryOnly here would silently empty it.
//
// impactByUid takes an explicit repoId and calls refreshRepos() internally.
// Use a fresh backend whose repo path is already absolute/resolved so the
// derived repoId stays stable across that refresh (an unresolved POSIX
// fixture path triggers the path-collision rehash and drops the key).
const resolvedRepoPath = path.resolve('/tmp/test-project');
(listRegisteredRepos as any).mockResolvedValue([
{ ...MOCK_REPO_ENTRY, path: resolvedRepoPath },
]);
backend = new LocalBackend();
await backend.init();
(executeParameterized as any).mockImplementation((_repoId: string, cypher: string) => {
// UID resolver
if (cypher.includes('WHERE n.id = $uid')) {
return Promise.resolve([
{ id: 'func:main', name: 'main', filePath: 'src/index.ts', type: 'Function' },
]);
}
// Aggregation pass (returns a process row so affectedProcesses > 0; if the
// per-symbol pass were not skipped, this would open its gate)
if (cypher.includes('COUNT(DISTINCT s.id)')) {
return Promise.resolve([
{
pId: 'proc:daily',
name: 'Daily cron',
heuristicLabel: 'Daily cron',
processType: 'cron',
entryPointId: 'func:cron_entry',
hits: 1,
minStep: 1,
stepCount: 5,
epName: 'cron_entry',
epType: 'Function',
epFilePath: 'src/cron.ts',
},
]);
}
return Promise.resolve([]);
});
(executeQuery as any).mockResolvedValue([
{
id: 'func:caller',
name: 'caller',
type: 'Function',
filePath: 'src/uses-main.ts',
relType: 'CALLS',
confidence: 0.9,
},
]);
const result = await backend.impactByUid('test-project', 'uid:main', 'upstream', {
maxDepth: 5,
relationTypes: ['CALLS'],
minConfidence: 0,
includeTests: true,
});
// byDepth must survive (Finding A regression guard)
expect(result).not.toBeNull();
expect(result.byDepth).toBeDefined();
const d1 = result.byDepth?.[1] || result.byDepth?.['1'] || [];
expect(d1.find((it: any) => it.id === 'func:caller')).toBeDefined();
// The per-symbol enrichment query must never fire under skipPerSymbolEnrichment
const perSymbolCalls = (executeParameterized as any).mock.calls.filter(
([, cypher]: [string, string]) =>
typeof cypher === 'string' && cypher.includes('RETURN s.id AS sid'),
);
expect(perSymbolCalls).toHaveLength(0);
});
it('dispatches detect_changes tool', async () => {
// detect_changes calls execFileSync which we haven't mocked at module level,
// so it will throw a git error — that's fine, we test the error path

View file

@ -0,0 +1,99 @@
import { describe, it, expect } from 'vitest';
import { extractCsharpStructureViaScanner } from '../../src/core/ingestion/languages/csharp/namespace-siblings.js';
// Scanner fallback used on the worker path, where native tree-sitter Trees
// can't cross MessageChannels so `treeCache` is empty. It must reproduce
// the AST walk's `namespaces` / `usingStaticPaths` for the common
// line-anchored declaration forms (see namespace-siblings.ts).
describe('extractCsharpStructureViaScanner', () => {
it('extracts a file-scoped namespace declaration', () => {
const src = `namespace App.Models;\n\npublic class User {}`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Models']);
});
it('extracts a block namespace declaration', () => {
const src = `namespace App.Services\n{\n public class Svc {}\n}`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Services']);
});
it('extracts multiple namespaces in source order', () => {
const src = `namespace A.One\n{\n}\nnamespace A.Two\n{\n}`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['A.One', 'A.Two']);
});
it('returns empty namespaces for a global (no-namespace) file', () => {
const src = `public class Global {}\n`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual([]);
});
it('captures a plain `using static` path', () => {
const src = `using static System.Math;\nnamespace App;`;
const out = extractCsharpStructureViaScanner(src);
expect(out.usingStaticPaths).toEqual(['System.Math']);
expect(out.namespaces).toEqual(['App']);
});
it('captures a `global using static` path', () => {
const src = `global using static App.Utils.Logger;\n`;
expect(extractCsharpStructureViaScanner(src).usingStaticPaths).toEqual(['App.Utils.Logger']);
});
it('captures the RHS path of an aliased `using static`', () => {
const src = `using static M = App.Utils.MathUtils;\n`;
expect(extractCsharpStructureViaScanner(src).usingStaticPaths).toEqual(['App.Utils.MathUtils']);
});
it('does not treat a plain `using` directive as using-static', () => {
const src = `using System.Collections.Generic;\nusing App.Models;\n`;
expect(extractCsharpStructureViaScanner(src).usingStaticPaths).toEqual([]);
});
it('does not treat a `using var`/`using (...)` statement as using-static', () => {
const src = `using var stream = File.Open(p);\nusing (var x = Get()) { }\n`;
expect(extractCsharpStructureViaScanner(src).usingStaticPaths).toEqual([]);
});
it('ignores a `// namespace X` line comment', () => {
const src = `// namespace Fake.Comment;\nnamespace App.Real;`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Real']);
});
it('handles indentation before declarations', () => {
const src = `\t\tnamespace App.Indented;\n`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Indented']);
});
it('handles an empty file', () => {
const out = extractCsharpStructureViaScanner('');
expect(out.namespaces).toEqual([]);
expect(out.usingStaticPaths).toEqual([]);
});
// Cross-line comment/string state: a keyword at the start of a line inside
// a block comment or multi-line string must NOT be read as a declaration
// (the worker path would otherwise mis-bucket the file vs the AST).
it('skips a `namespace` line inside a block comment', () => {
const src = `/*\nnamespace Fake.InComment;\n*/\nnamespace App.Real;`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Real']);
});
it('skips a `using static` line inside a block comment', () => {
const src = `/*\nusing static Fake.Helpers;\n*/\nusing static App.Real.Helpers;`;
expect(extractCsharpStructureViaScanner(src).usingStaticPaths).toEqual(['App.Real.Helpers']);
});
it('skips a `namespace` line inside a raw string literal', () => {
const src = `var sql = """\nnamespace Fake.InRaw;\n""";\nnamespace App.Real;`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Real']);
});
it('skips a `namespace` line inside a verbatim string literal', () => {
const src = `var s = @"\nnamespace Fake.InVerbatim;\n";\nnamespace App.Real;`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Real']);
});
it('still reads a real declaration after a closed same-line block comment', () => {
const src = `/* header */ class C {}\nnamespace App.Real;`;
expect(extractCsharpStructureViaScanner(src).namespaces).toEqual(['App.Real']);
});
});

View file

@ -3,6 +3,7 @@ import { createHash } from 'crypto';
import {
contentHashForNode,
EMBEDDING_TEXT_VERSION,
resolveEmbeddingInstallPolicy,
} from '../../src/core/embeddings/embedding-pipeline.js';
import { generateEmbeddingText } from '../../src/core/embeddings/text-generator.js';
import type { EmbeddableNode, EmbeddingProgress } from '../../src/core/embeddings/types.js';
@ -12,6 +13,55 @@ import { STALE_HASH_SENTINEL } from '../../src/core/lbug/schema.js';
const CLASS_CHUNK_SIZE = 90;
const CLASS_OVERLAP = 10;
// ────────────────────────────────────────────────────────────────────────────
// resolveEmbeddingInstallPolicy (offline-first, #1153)
// ────────────────────────────────────────────────────────────────────────────
describe('resolveEmbeddingInstallPolicy (#1153)', () => {
const ENV = 'GITNEXUS_LBUG_EXTENSION_INSTALL';
const original = process.env[ENV];
const restore = () => {
if (original === undefined) delete process.env[ENV];
else process.env[ENV] = original;
};
it('defaults to auto when unset (embeddings are an explicit network-capable opt-in)', () => {
delete process.env[ENV];
try {
expect(resolveEmbeddingInstallPolicy()).toBe('auto');
} finally {
restore();
}
});
it('honors an explicit load-only override (offline operator is not forced onto the network)', () => {
process.env[ENV] = 'load-only';
try {
expect(resolveEmbeddingInstallPolicy()).toBe('load-only');
} finally {
restore();
}
});
it('honors an explicit never override', () => {
process.env[ENV] = 'never';
try {
expect(resolveEmbeddingInstallPolicy()).toBe('never');
} finally {
restore();
}
});
it('falls back to auto for invalid values', () => {
process.env[ENV] = 'bogus';
try {
expect(resolveEmbeddingInstallPolicy()).toBe('auto');
} finally {
restore();
}
});
});
// ────────────────────────────────────────────────────────────────────────────
// contentHashForNode
// ────────────────────────────────────────────────────────────────────────────

View file

@ -0,0 +1,287 @@
/**
* Unit tests for {@link extractFastAPIRouterBindings} — the per-file
* regex extractor that the parse worker calls on every Python file.
* The cross-file aggregation that turns these raw records into prefix
* maps lives in parse-impl and is covered by
* `fastapi-prefix-pipeline.test.ts` (integration) plus
* `http-route-extractor.test.ts` (group layer). This file pins the
* shape the worker emits, so a regression in either regex or in the
* import-list parsing fails here first.
*
* What this file is responsible for:
* • Shape A `app.include_router(<mod>.router, prefix=…)` and
* Shape B `app.include_router(<local>, prefix=…)` are both
* captured.
* • `<host>.include_router` matches any host name, not just `app`.
* • Module path keying is two-tiered: short basename (always) and
* long `<parent>/<stem>` key (whenever the import path was
* multi-segment).
* • Relative imports (`from .calls import …`,
* `from ..siblings.calls import …`) are captured.
* • `as`-aliased imports route the prefix to the alias, not to
* `router`.
* • Nothing is emitted when `include_router` is absent or has no
* `prefix=` keyword.
*/
import { describe, it, expect } from 'vitest';
import {
extractFastAPIRouterBindings,
lastDottedSegment,
lastTwoSegmentsAsPath,
type ExtractedRouterInclude,
type ExtractedRouterImport,
} from '../../src/core/ingestion/route-extractors/fastapi-router-bindings.js';
function run(filePath: string, content: string) {
const includes: ExtractedRouterInclude[] = [];
const imports: ExtractedRouterImport[] = [];
extractFastAPIRouterBindings(filePath, content, includes, imports);
return { includes, imports };
}
describe('lastDottedSegment', () => {
it('returns the last segment of an absolute dotted path', () => {
expect(lastDottedSegment('api.users')).toBe('users');
expect(lastDottedSegment('api.v2.users')).toBe('users');
});
it('strips leading dots from a relative path', () => {
expect(lastDottedSegment('.users')).toBe('users');
expect(lastDottedSegment('..api.users')).toBe('users');
expect(lastDottedSegment('...users')).toBe('users');
});
it('returns the input when there is no dot after stripping', () => {
expect(lastDottedSegment('users')).toBe('users');
});
it('returns the empty string for pure-dot inputs', () => {
expect(lastDottedSegment('.')).toBe('');
expect(lastDottedSegment('..')).toBe('');
expect(lastDottedSegment('...')).toBe('');
});
});
describe('lastTwoSegmentsAsPath', () => {
it('joins the last two segments with `/`', () => {
expect(lastTwoSegmentsAsPath('api.users')).toBe('api/users');
expect(lastTwoSegmentsAsPath('app.api.users')).toBe('api/users');
});
it('strips leading dots before joining', () => {
expect(lastTwoSegmentsAsPath('..api.users')).toBe('api/users');
});
it('returns the empty string when the path has only one segment', () => {
// Single-segment imports cannot be promoted to a long key.
expect(lastTwoSegmentsAsPath('users')).toBe('');
expect(lastTwoSegmentsAsPath('.users')).toBe('');
});
it('returns the empty string for pure-dot inputs', () => {
expect(lastTwoSegmentsAsPath('.')).toBe('');
expect(lastTwoSegmentsAsPath('..')).toBe('');
});
});
describe('extractFastAPIRouterBindings — Shape A (`<mod>.router`)', () => {
it('captures app.include_router(<mod>.router, prefix=…)', () => {
const { includes } = run(
'main.py',
[
'from fastapi import FastAPI',
'from api import users',
'app = FastAPI()',
"app.include_router(users.router, prefix='/users', tags=['users'])",
'',
].join('\n'),
);
expect(includes).toHaveLength(1);
expect(includes[0]).toMatchObject({
filePath: 'main.py',
routerExpr: 'users.router',
prefix: '/users',
});
// Line number is 1-indexed and points to the include_router call.
expect(includes[0].lineNumber).toBe(4);
});
it('captures non-`app` host variables', () => {
// FINDING 4: production code commonly uses `api`, `application`,
// `asgi_app` etc. Pinning the regex to `app.` would silently drop
// these, which used to leave the ingestion and group layers
// disagreeing on whether a prefix was applied.
const { includes } = run(
'main.py',
[
'from fastapi import FastAPI',
'from api import users',
'api = FastAPI()',
"api.include_router(users.router, prefix='/users')",
'',
].join('\n'),
);
expect(includes).toHaveLength(1);
expect(includes[0].routerExpr).toBe('users.router');
expect(includes[0].prefix).toBe('/users');
});
it('captures multiple Shape-A includes in the same file', () => {
const { includes } = run(
'main.py',
[
'from api import users, calls',
'app = FastAPI()',
"app.include_router(users.router, prefix='/users')",
"app.include_router(calls.router, prefix='/calls')",
'',
].join('\n'),
);
expect(includes).toHaveLength(2);
expect(includes.map((i) => i.routerExpr).sort()).toEqual(['calls.router', 'users.router']);
});
});
describe('extractFastAPIRouterBindings — Shape B (bare local name)', () => {
it('captures app.include_router(<local>, prefix=…) and the import', () => {
const { includes, imports } = run(
'main.py',
[
'from fastapi import FastAPI',
'from api.users import router as users_router',
'app = FastAPI()',
"app.include_router(users_router, prefix='/users')",
'',
].join('\n'),
);
expect(imports).toHaveLength(1);
expect(imports[0]).toMatchObject({
filePath: 'main.py',
localName: 'users_router',
moduleKey: 'users',
moduleKeyLong: 'api/users',
});
expect(includes).toHaveLength(1);
expect(includes[0]).toMatchObject({
filePath: 'main.py',
routerExpr: 'users_router',
prefix: '/users',
});
});
it('captures the unaliased shape `from <mod> import router`', () => {
const { imports } = run('main.py', ['from api.users import router', ''].join('\n'));
expect(imports).toHaveLength(1);
expect(imports[0]).toMatchObject({
localName: 'router',
moduleKey: 'users',
moduleKeyLong: 'api/users',
});
});
it('does NOT re-capture Shape A as Shape B (`<mod>.router` is not bare)', () => {
// Anti-regression: INCLUDE_ROUTER_NAME_RE is intentionally
// permissive (`(identifier)`). Without the lookahead in
// extractFastAPIRouterBindings it would re-capture the bare
// module name `users` from `users.router` and add a phantom
// include with `routerExpr: "users"`.
const { includes } = run(
'main.py',
["app.include_router(users.router, prefix='/users')", ''].join('\n'),
);
const shapes = includes.map((i) => i.routerExpr).sort();
expect(shapes).toEqual(['users.router']);
});
});
describe('extractFastAPIRouterBindings — relative imports', () => {
it('captures single-dot relative imports (`from .calls import router as …`)', () => {
// FINDING 2: the previous regex `[A-Za-z_][\w.]*` rejected
// module paths starting with `.`, silently dropping every
// relative-import Shape-B include. The PR description's own
// motivating example used this shape — now pinned.
const { imports } = run(
'main.py',
['from .calls import router as calls_router', ''].join('\n'),
);
expect(imports).toHaveLength(1);
expect(imports[0]).toMatchObject({
localName: 'calls_router',
moduleKey: 'calls',
});
// Single-segment relative paths cannot be promoted to a long key.
expect(imports[0].moduleKeyLong).toBeUndefined();
});
it('captures multi-segment relative imports and emits a long key', () => {
const { imports } = run(
'main.py',
['from ..api.users import router as users_router', ''].join('\n'),
);
expect(imports).toHaveLength(1);
expect(imports[0]).toMatchObject({
localName: 'users_router',
moduleKey: 'users',
moduleKeyLong: 'api/users',
});
});
});
describe('extractFastAPIRouterBindings — long-key precision', () => {
it('emits long key `api/users` for a multi-segment absolute import', () => {
// FINDING 3: short-key-only collides for `api/users.py` vs
// `admin/users.py`. The long key gives parse-impl the precision
// it needs to bind a Shape-B include to the right file.
const { imports } = run('main.py', ['from api.users import router', ''].join('\n'));
expect(imports[0].moduleKeyLong).toBe('api/users');
});
it('omits the long key for a single-segment top-level import', () => {
const { imports } = run('main.py', ['from users import router', ''].join('\n'));
expect(imports[0].moduleKey).toBe('users');
expect(imports[0].moduleKeyLong).toBeUndefined();
});
});
describe('extractFastAPIRouterBindings — negative cases', () => {
it('emits nothing for files without any include_router or import', () => {
const { includes, imports } = run('helpers.py', 'def add(a, b):\n return a + b\n');
expect(includes).toEqual([]);
expect(imports).toEqual([]);
});
it('does not capture include_router calls without a prefix= keyword', () => {
const { includes } = run(
'main.py',
['app.include_router(users.router, tags=["users"])', ''].join('\n'),
);
expect(includes).toEqual([]);
});
it('does not capture include_router calls with a non-string prefix', () => {
// The current regex requires a string literal for the prefix
// value. Variables / f-strings / concatenations are not
// resolvable at parse time.
const { includes } = run(
'main.py',
['app.include_router(users.router, prefix=PREFIX_USERS)', ''].join('\n'),
);
expect(includes).toEqual([]);
});
it('ignores non-router names in `from … import` lists', () => {
const { imports } = run('main.py', ['from api.users import schemas, helpers', ''].join('\n'));
expect(imports).toEqual([]);
});
it('correctly handles a mixed import list (router + others)', () => {
const { imports } = run(
'main.py',
['from api.users import router, schemas, helpers', ''].join('\n'),
);
expect(imports).toHaveLength(1);
expect(imports[0].localName).toBe('router');
expect(imports[0].moduleKey).toBe('users');
});
});

View file

@ -17,6 +17,7 @@ import {
serviceContractId,
} from '../../../src/core/group/extractors/grpc-extractor.js';
import type { ProtoServiceInfo } from '../../../src/core/group/extractors/grpc-extractor.js';
import { buildProviderIndex, runWildcardMatch } from '../../../src/core/group/matching.js';
import type { RepoHandle } from '../../../src/core/group/types.js';
import { _captureLogger } from '../../../src/core/logger.js';
@ -384,6 +385,566 @@ public class AuthGrpcService extends AuthServiceGrpc.AuthServiceImplBase {
});
});
// ─── Java client-jar / import-derived FQN ─────────────────────────
// The "client-jar" architecture is the dominant pattern for Java
// gRPC microservices: the service owner publishes a pre-compiled
// stub jar to a Maven repository, and consumer repos depend on the
// jar instead of carrying the originating `.proto` files. Examples:
// gRPC official quickstart, Alibaba HSF, ByteDance KiteX-Java,
// google-cloud-java SDK.
//
// Before this fix, the extractor only resolved a fully-qualified
// contract id (`grpc::<package>.<Service>/*`) when the consumer
// repo also carried a matching `.proto` file. Client-jar consumers
// had no proto, so they fell back to a short-name contract id
// (`grpc::<Service>/*`) that never matched the provider repo's
// package-qualified contract id — cross-repo grpc cross-link count
// dropped to zero on every realistic Java micro-service group.
//
// The fix derives the FQN directly from the consumer file's `import
// <pkg>.<XxxGrpc>;` statement, which is always present (without it
// the Java code wouldn't even compile). The package from the import
// is exactly the proto package, so the contract id matches the
// provider's verbatim — no `.proto` lookup needed.
describe('Java client-jar consumer (import-derived FQN)', () => {
it('test_consumer_with_import_emits_fqn_contract_id_without_local_proto', async () => {
// No .proto file in this repo — the consumer ONLY has the import.
writeFile(
'src/main/java/AuthClient.java',
`package my.app;
import io.grpc.ManagedChannel;
import com.acme.auth.proto.AuthServiceGrpc;
public class AuthClient {
private final AuthServiceGrpc.AuthServiceBlockingStub stub;
public AuthClient(ManagedChannel ch) {
this.stub = AuthServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(1);
expect(consumers[0].contractId).toBe('grpc::com.acme.auth.proto.AuthService/*');
// Confidence stays at the "with proto" tier: the import
// statement is at least as authoritative as a per-repo proto
// map, so consumers shouldn't be penalised for not carrying
// a redundant `.proto` file.
expect(consumers[0].confidence).toBe(0.75);
expect(consumers[0].meta.protoPackageSource).toBe('import');
expect(consumers[0].meta.package).toBe('com.acme.auth.proto');
});
it('test_provider_with_import_emits_fqn_contract_id_without_local_proto', async () => {
// Same idea on the provider side: a server impl class lives in
// a repo that does NOT carry the originating `.proto`. The
// import on `AuthServiceGrpc` is enough to derive the FQN.
writeFile(
'src/main/java/AuthServerImpl.java',
`package my.server;
import com.acme.auth.proto.AuthServiceGrpc;
import io.grpc.stub.StreamObserver;
public class AuthServerImpl extends AuthServiceGrpc.AuthServiceImplBase {
@Override
public void login(LoginRequest req, StreamObserver<LoginResponse> obs) {}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers).toHaveLength(1);
expect(providers[0].contractId).toBe('grpc::com.acme.auth.proto.AuthService/*');
expect(providers[0].confidence).toBe(0.8);
expect(providers[0].meta.protoPackageSource).toBe('import');
});
it('test_same_short_name_different_packages_resolves_to_distinct_fqns', async () => {
// The motivating real-world case (unipus_cloud_framework):
// `ContentRpcService` is defined in TWO different proto packages
// by two different client modules.
//
// ucf-api-client/Service.proto → cn.unipus.ucf.api.proto.client.service.ContentRpcService
// ucf-admin-client/Service.proto → cn.unipus.ucf.admin.proto.client.service.ContentRpcService
//
// A short-name fallback would silently merge consumers of the
// two services into one bogus contract id; the import-derived
// FQN keeps them distinct.
writeFile(
'src/main/java/ApiContentClient.java',
`package my.app.api;
import io.grpc.ManagedChannel;
import cn.unipus.ucf.api.proto.client.service.ContentRpcServiceGrpc;
public class ApiContentClient {
private final ContentRpcServiceGrpc.ContentRpcServiceBlockingStub stub;
public ApiContentClient(ManagedChannel ch) {
this.stub = ContentRpcServiceGrpc.newBlockingStub(ch);
}
}`,
);
writeFile(
'src/main/java/AdminContentClient.java',
`package my.app.admin;
import io.grpc.ManagedChannel;
import cn.unipus.ucf.admin.proto.client.service.ContentRpcServiceGrpc;
public class AdminContentClient {
private final ContentRpcServiceGrpc.ContentRpcServiceBlockingStub stub;
public AdminContentClient(ManagedChannel ch) {
this.stub = ContentRpcServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(2);
const ids = consumers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::cn.unipus.ucf.admin.proto.client.service.ContentRpcService/*',
'grpc::cn.unipus.ucf.api.proto.client.service.ContentRpcService/*',
]);
});
it('test_local_proto_overrides_unrelated_import_with_same_short_name', async () => {
// Symmetric to Finding 2: when the consumer repo carries its
// OWN `.proto` defining the same short service name, the proto
// is authoritative and wins over a Java import that points at a
// different package. Without this Step-2 cross-check, a typo'd
// or stale Java import (or genuinely unrelated same-name
// service in the same repo) would silently corrupt the
// contract id of the locally-defined service.
writeFile(
'protos/local-other.proto',
`syntax = "proto3";
package local.unrelated;
service AuthService {
rpc Ping (PingRequest) returns (PingResponse);
}`,
);
writeFile(
'src/main/java/AuthClient.java',
`package my.app;
import io.grpc.ManagedChannel;
import com.acme.auth.proto.AuthServiceGrpc;
public class AuthClient {
private final AuthServiceGrpc.AuthServiceBlockingStub stub;
public AuthClient(ManagedChannel ch) {
this.stub = AuthServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(1);
// Local proto wins. The disagreement is recorded so operators
// can investigate the divergent import.
expect(consumers[0].contractId).toBe('grpc::local.unrelated.AuthService/*');
expect(consumers[0].meta.protoPackageSource).toBe('proto-override');
expect(consumers[0].meta.importPackage).toBe('com.acme.auth.proto');
});
it('test_consumer_without_import_falls_back_to_proto_map', async () => {
// No import line — perhaps a fully-qualified call site like
// `com.acme.auth.proto.AuthServiceGrpc.newBlockingStub(...)`,
// or a refactor that broke the import. The current STUB_PATTERNS
// captures only `(identifier) @grpc_cls`, so it skips the
// fully-qualified form. With no detection there's also nothing
// for the proto-map fallback to anchor onto. We assert the
// benign no-op (no false-positive emitted) — the proto-map
// fallback path is exercised by the dedicated test below.
writeFile(
'src/main/java/AuthClient.java',
`package my.app;
import io.grpc.ManagedChannel;
public class AuthClient {
private final com.acme.auth.proto.AuthServiceGrpc.AuthServiceBlockingStub stub;
public AuthClient(ManagedChannel ch) {
this.stub = com.acme.auth.proto.AuthServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
// STUB_PATTERNS only captures bare-identifier `XxxGrpc`, so the
// fully-qualified `com.acme.auth.proto.AuthServiceGrpc.newStub(...)`
// form is intentionally not matched. Pinning behaviour so the
// import-driven path doesn't accidentally introduce a regression.
expect(consumers).toHaveLength(0);
});
it('test_short_import_consumer_with_local_proto_still_uses_proto_map', async () => {
// Backward-compat: when the consumer repo HAS a matching
// `.proto` (the legacy path) AND the import is present, both
// paths agree — but we want to confirm the import-driven path
// takes precedence and emits the same FQN with the
// `protoPackageSource: 'import'` marker.
writeFile(
'protos/auth.proto',
`syntax = "proto3";
package com.acme.auth.proto;
service AuthService {
rpc Login (LoginRequest) returns (LoginResponse);
}`,
);
writeFile(
'src/main/java/AuthClient.java',
`package my.app;
import io.grpc.ManagedChannel;
import com.acme.auth.proto.AuthServiceGrpc;
public class AuthClient {
private final AuthServiceGrpc.AuthServiceBlockingStub stub;
public AuthClient(ManagedChannel ch) {
this.stub = AuthServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(1);
expect(consumers[0].contractId).toBe('grpc::com.acme.auth.proto.AuthService/*');
// Marker confirms import path won, not the proto map. Both
// would have produced the same FQN, but only the import path
// is robust against client-jar consumers and same-short-name
// collisions.
expect(consumers[0].meta.protoPackageSource).toBe('import');
});
it('test_static_and_wildcard_imports_are_ignored', async () => {
// `import static …` and `import w.x.*;` shouldn't pollute the
// import map. Pinned via the tree-sitter query shape (the
// `name:` field is only present on the non-static, non-wildcard
// form). When the only `XxxGrpc` reference comes through one
// of these unsupported import styles, the consumer detection
// emits nothing-import-derived and the legacy short-name
// fallback applies.
writeFile(
'src/main/java/AuthClient.java',
`package my.app;
import static com.acme.auth.proto.Constants.SOMETHING;
import com.acme.unrelated.*;
import io.grpc.ManagedChannel;
public class AuthClient {
private final com.acme.auth.proto.AuthServiceGrpc.AuthServiceBlockingStub stub;
public AuthClient(ManagedChannel ch) {
this.stub = com.acme.auth.proto.AuthServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
// STUB_PATTERNS doesn't match fully-qualified call forms; this
// pins that adding GRPC_CLASS_IMPORT_PATTERNS doesn't accidentally
// lift the static / wildcard imports into the FQN map (which
// would have created a phantom detection).
expect(consumers).toHaveLength(0);
});
it('test_provider_in_client_jar_consumer_repo_emits_provider_too', async () => {
// Same repo holds a SERVER impl whose only knowledge of the
// proto package is the import — no `.proto` is present. The
// provider detection should also use the import-derived FQN.
writeFile(
'src/main/java/AuthServer.java',
`package my.server;
import com.acme.auth.proto.AuthServiceGrpc;
import io.grpc.stub.StreamObserver;
@GrpcService
public class AuthServer extends AuthServiceGrpc.AuthServiceImplBase {
@Override
public void login(LoginRequest req, StreamObserver<LoginResponse> obs) {}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers).toHaveLength(1);
expect(providers[0].contractId).toBe('grpc::com.acme.auth.proto.AuthService/*');
expect(providers[0].confidence).toBe(0.8);
expect(providers[0].meta.protoPackageSource).toBe('import');
});
it('test_unipus_admin_and_api_consumers_in_one_repo_do_not_collide', async () => {
// End-to-end version of the same-short-name case: a single
// consumer repo imports BOTH `ContentRpcService` flavours from
// unipus_cloud_framework. Ensures the per-file import map is
// file-local (each file's import wins for that file's call sites)
// rather than blurring across the whole repo.
writeFile(
'src/main/java/api/ApiContentClient.java',
`package my.app.api;
import io.grpc.ManagedChannel;
import cn.unipus.ucf.api.proto.client.service.ContentRpcServiceGrpc;
public class ApiContentClient {
public ApiContentClient(ManagedChannel ch) {
ContentRpcServiceGrpc.newBlockingStub(ch);
}
}`,
);
writeFile(
'src/main/java/admin/AdminContentClient.java',
`package my.app.admin;
import io.grpc.ManagedChannel;
import cn.unipus.ucf.admin.proto.client.service.ContentRpcServiceGrpc;
public class AdminContentClient {
public AdminContentClient(ManagedChannel ch) {
ContentRpcServiceGrpc.newBlockingStub(ch);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(2);
const ids = new Set(consumers.map((c) => c.contractId));
expect(ids.has('grpc::cn.unipus.ucf.api.proto.client.service.ContentRpcService/*')).toBe(
true,
);
expect(ids.has('grpc::cn.unipus.ucf.admin.proto.client.service.ContentRpcService/*')).toBe(
true,
);
});
});
// ─── Java `option java_package` divergence ────────────────────
// Java protobuf projects frequently set
// `option java_package = "..."` to publish their generated Java
// classes under a namespace different from the proto `package`
// declaration. Google Cloud Java SDKs are the canonical example:
// proto `package google.cloud.speech.v1` + `option java_package =
// "com.google.cloud.speech.v1"`. Without specific handling, the
// import-derived FQN would reflect the Java namespace instead of
// the wire-protocol namespace and never match a provider's
// contract id.
//
// The cases below pin the four resolution branches in
// `detectionToContract`:
//
// 1. java_package translation (same-repo provider with the
// option set; consumer in the same repo imports via the
// java_package — the reverse index translates back to the
// proto package);
// 2. proto-map cross-check (local proto exists for the same
// service short name and AGREES with the import — both paths
// produce the same FQN, marker confirms import path took
// precedence);
// 2b. proto-map cross-check (local proto DISAGREES with the
// import — the proto wins authoritatively, the import package
// is recorded as `meta.importPackage` for diagnostics);
// 3. import-derived fallback known limitation (consumer repo
// carries no proto AND the published proto sets a divergent
// java_package — we cannot translate without the proto in
// reach, so the FQN reflects the Java namespace and will not
// match a provider repo. This is documented as a scope
// limitation; the test pins the limitation to catch any
// accidental change in behaviour).
describe('Java option java_package divergence', () => {
it('test_provider_proto_with_diverging_java_package_emits_proto_package_FQN', async () => {
// Provider side: proto declares both `package` and a
// different `option java_package`. The provider contract id
// must use the proto `package` — that's the wire identity any
// consumer (regardless of its language) will see at runtime.
writeFile(
'proto/speech.proto',
`syntax = "proto3";
package google.cloud.speech.v1;
option java_package = "com.google.cloud.speech.v1";
service Speech {
rpc Recognize (RecognizeRequest) returns (RecognizeResponse);
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter((c) => c.role === 'provider');
const recognize = providers.find((c) => c.contractId.endsWith('Speech/Recognize'));
expect(recognize).toBeDefined();
// Wire-protocol package, NOT the java_package value.
expect(recognize!.contractId).toBe('grpc::google.cloud.speech.v1.Speech/Recognize');
});
it('test_consumer_with_java_package_translation_uses_proto_package', async () => {
// Same repo carries the proto with a divergent java_package
// AND a Java consumer that imports via the java_package. The
// reverse index built by `buildProtoContext` should translate
// the import back to the proto package so the consumer's
// contract id matches the provider's.
writeFile(
'proto/speech.proto',
`syntax = "proto3";
package google.cloud.speech.v1;
option java_package = "com.google.cloud.speech.v1";
service Speech {
rpc Recognize (RecognizeRequest) returns (RecognizeResponse);
}`,
);
writeFile(
'src/main/java/SpeechClient.java',
`package my.app;
import io.grpc.ManagedChannel;
import com.google.cloud.speech.v1.SpeechGrpc;
public class SpeechClient {
public SpeechClient(ManagedChannel ch) {
SpeechGrpc.newBlockingStub(ch).recognize(null);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(1);
// The reverse-index translation kicked in:
// import "com.google.cloud.speech.v1"
// ↓ (javaPackageMap lookup)
// proto pkg "google.cloud.speech.v1" ← used in contract id
expect(consumers[0].contractId).toBe('grpc::google.cloud.speech.v1.Speech/*');
expect(consumers[0].meta.protoPackageSource).toBe('import-translated');
expect(consumers[0].meta.package).toBe('google.cloud.speech.v1');
});
it('test_consumer_without_local_proto_and_diverging_java_package_is_known_limitation', async () => {
// Client-jar consumer: zero `.proto` in this repo, and the
// published proto (somewhere else) uses a divergent
// java_package. We have no way to translate from
// java_package back to proto package without sight of the
// source proto. The current behaviour is to use the
// import-derived java_package literally; the resulting
// contract id will not match a provider's. This is a
// documented scope limitation — resolving it requires
// group-level proto knowledge that's out of scope for this
// change. The test pins the limitation so it cannot
// regress silently.
writeFile(
'src/main/java/SpeechClient.java',
`package my.app;
import io.grpc.ManagedChannel;
import com.google.cloud.speech.v1.SpeechGrpc;
public class SpeechClient {
public SpeechClient(ManagedChannel ch) {
SpeechGrpc.newBlockingStub(ch).recognize(null);
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers).toHaveLength(1);
// Pinned limitation: the FQN reflects the Java namespace.
expect(consumers[0].contractId).toBe('grpc::com.google.cloud.speech.v1.Speech/*');
expect(consumers[0].meta.protoPackageSource).toBe('import');
});
});
// ─── End-to-end wildcard match (Finding 3) ────────────────────
// The 9 unit tests above pin contract-id shape; this block pins
// the next stage of the pipeline — `runWildcardMatch` against a
// provider index — so a regression in either contract-id format
// OR in the matcher's wildcard logic would fail here. Per DoD §2.7
// ("tests cover the real changed path"), exercising the pipeline
// end to end is the production-readiness signal we need.
describe('Java client-jar consumer — end-to-end wildcard match', () => {
it('test_e2e_client_jar_consumer_FQN_creates_wildcard_cross_link', async () => {
// Two-repo group fixture, written into separate subdirectories
// of tmpDir so the per-repo `extract()` can run isolated.
const providerDir = path.join(tmpDir, 'provider-repo');
const consumerDir = path.join(tmpDir, 'consumer-repo');
fs.mkdirSync(path.join(providerDir, 'proto'), { recursive: true });
fs.mkdirSync(path.join(consumerDir, 'src/main/java'), { recursive: true });
fs.writeFileSync(
path.join(providerDir, 'proto/auth.proto'),
`syntax = "proto3";
package com.acme.auth.proto;
service AuthService {
rpc Login (LoginRequest) returns (LoginResponse);
}`,
);
// Consumer repo carries NO `.proto` — typical client-jar pattern.
fs.writeFileSync(
path.join(consumerDir, 'src/main/java/AuthClient.java'),
`package my.app;
import io.grpc.ManagedChannel;
import com.acme.auth.proto.AuthServiceGrpc;
public class AuthClient {
public AuthClient(ManagedChannel ch) {
AuthServiceGrpc.newBlockingStub(ch).login(null);
}
}`,
);
const providerExtracted = await extractor.extract(null, providerDir, makeRepo(providerDir));
const consumerExtracted = await extractor.extract(null, consumerDir, makeRepo(consumerDir));
// Stamp `repo` on the contracts so they look like StoredContract;
// matching.ts skips same-repo cross-links by comparing this field.
const stored = [
...providerExtracted.map((c) => ({ ...c, repo: 'provider' })),
...consumerExtracted.map((c) => ({ ...c, repo: 'consumer' })),
];
const providerIndex = buildProviderIndex(stored);
const consumerWildcards = stored.filter(
(c) => c.role === 'consumer' && c.contractId.endsWith('/*'),
);
const result = runWildcardMatch(consumerWildcards, providerIndex);
// The consumer's contract id is the package-qualified service
// wildcard (`grpc::com.acme.auth.proto.AuthService/*`); the
// provider emits a method-level id (`grpc::com.acme.auth.proto.
// AuthService/Login`). The wildcard matcher pairs them and
// produces exactly one cross-link.
expect(result.matched).toHaveLength(1);
const cross = result.matched[0];
expect(cross.contractId).toBe('grpc::com.acme.auth.proto.AuthService/*');
expect(cross.matchType).toBe('wildcard');
expect(cross.from.repo).toBe('consumer');
expect(cross.to.repo).toBe('provider');
});
});
describe('Python detection', () => {
it('test_extract_python_add_servicer_returns_provider', async () => {
writeFile(

View file

@ -41,6 +41,8 @@ describe('HttpRouteExtractor', () => {
});
});
const toPosixPath = (filePath: string): string => filePath.replace(/\\/g, '/');
describe('provider extraction — graph-first (Strategy A)', () => {
it('extracts routes from Route/HANDLES_ROUTE graph + source scan for method', async () => {
const dir = path.join(tmpDir, 'graph-first');
@ -832,6 +834,181 @@ class UserController {
},
);
it('does not emit annotated Java interfaces as concrete Spring provider routes', async () => {
const dir = path.join(tmpDir, 'spring-interface-only');
fs.mkdirSync(path.join(dir, 'src/rest'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src/rest/DepartmentApi.java'),
`
package com.example.rest;
import org.springframework.web.bind.annotation.*;
@RequestMapping("/departments")
public interface DepartmentApi {
@GetMapping("")
Object list();
@GetMapping("/{name}")
Object getByName(@PathVariable String name);
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers).toHaveLength(0);
});
it('inherits Spring interface route mappings when controller methods omit annotations', async () => {
const dir = path.join(tmpDir, 'spring-interface-inherited-methods');
fs.mkdirSync(path.join(dir, 'src/rest'), { recursive: true });
fs.mkdirSync(path.join(dir, 'src/controller'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src/rest/StatusApi.java'),
`
package com.example.rest;
import org.springframework.web.bind.annotation.*;
@RequestMapping("/status")
public interface StatusApi {
@GetMapping("")
Object getStatus();
}
`,
);
fs.writeFileSync(
path.join(dir, 'src/controller/StatusController.java'),
`
package com.example.controller;
import com.example.rest.StatusApi;
import org.springframework.web.bind.annotation.*;
@RestController
public class StatusController implements StatusApi {
@Override
public Object getStatus() { return null; }
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
const statusRoute = providers.find((c) => c.contractId === 'http::GET::/status');
expect(statusRoute).toBeDefined();
expect(toPosixPath(statusRoute!.symbolRef.filePath)).toBe(
'src/controller/StatusController.java',
);
expect(statusRoute!.symbolName).toBe('getStatus');
expect(providers.filter((c) => c.symbolRef.filePath.includes('StatusApi.java'))).toHaveLength(
0,
);
});
it('combines controller class mapping with inherited interface method mapping', async () => {
const dir = path.join(tmpDir, 'spring-interface-controller-prefix');
fs.mkdirSync(path.join(dir, 'src/rest'), { recursive: true });
fs.mkdirSync(path.join(dir, 'src/controller'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src/rest/UserApi.java'),
`
package com.example.rest;
import org.springframework.web.bind.annotation.*;
public interface UserApi {
@GetMapping("/users")
Object listUsers();
}
`,
);
fs.writeFileSync(
path.join(dir, 'src/controller/UserController.java'),
`
package com.example.controller;
import com.example.rest.UserApi;
import org.springframework.web.bind.annotation.*;
@RestController
@RequestMapping("/api")
public class UserController implements UserApi {
@Override
public Object listUsers() { return null; }
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
const usersRoute = providers.find((c) => c.contractId === 'http::GET::/api/users');
expect(usersRoute).toBeDefined();
expect(toPosixPath(usersRoute!.symbolRef.filePath)).toBe(
'src/controller/UserController.java',
);
});
it('skips ambiguous inherited routes when interfaces share a simple name', async () => {
const dir = path.join(tmpDir, 'spring-interface-simple-name-collision');
fs.mkdirSync(path.join(dir, 'src/a'), { recursive: true });
fs.mkdirSync(path.join(dir, 'src/b'), { recursive: true });
fs.mkdirSync(path.join(dir, 'src/controller'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src/a/StatusApi.java'),
`
package com.example.a;
import org.springframework.web.bind.annotation.*;
public interface StatusApi {
@GetMapping("/a/status")
Object getStatus();
}
`,
);
fs.writeFileSync(
path.join(dir, 'src/b/StatusApi.java'),
`
package com.example.b;
import org.springframework.web.bind.annotation.*;
public interface StatusApi {
@GetMapping("/b/status")
Object getStatus();
}
`,
);
fs.writeFileSync(
path.join(dir, 'src/controller/StatusController.java'),
`
package com.example.controller;
import com.example.a.StatusApi;
import org.springframework.web.bind.annotation.*;
@RestController
public class StatusController implements StatusApi {
@Override
public Object getStatus() { return null; }
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers.find((c) => c.contractId === 'http::GET::/b/status')).toBeUndefined();
expect(providers.find((c) => c.contractId === 'http::GET::/a/status')).toBeUndefined();
expect(
providers.filter((c) => c.symbolRef.filePath.includes('StatusController.java')),
).toHaveLength(0);
});
it('extracts Express router.get patterns', async () => {
const dir = path.join(tmpDir, 'express');
fs.mkdirSync(path.join(dir, 'src/routes'), { recursive: true });
@ -1364,7 +1541,7 @@ shadowed_module_client.get("/module-level-rebind-fp")
).toBeUndefined();
});
it('extracts Java RestTemplate, WebClient and OkHttp calls', async () => {
it('extracts Java Spring RestTemplate, WebClient and OkHttp literal calls', async () => {
const dir = path.join(tmpDir, 'java-consumer');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
@ -1378,7 +1555,8 @@ import okhttp3.Request;
class ApiClient {
void run(RestTemplate restTemplate, WebClient webClient) {
restTemplate.getForObject("/api/users/{id}", String.class, 42);
webClient.method(HttpMethod.PATCH, "/api/users/42");
restTemplate.exchange("/api/users/{id}/details", HttpMethod.GET, null, String.class);
webClient.post().uri("/api/users");
new Request.Builder().url("/api/orders/42").build();
}
}
@ -1390,19 +1568,289 @@ class ApiClient {
expect(consumers.find((c) => c.contractId === 'http::GET::/api/users/{param}')).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::PATCH::/api/users/{param}'),
consumers.find((c) => c.contractId === 'http::GET::/api/users/{param}/details'),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::GET::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find(
(c) =>
c.contractId === 'http::GET::/api/users/{param}/details' &&
c.meta.framework === 'spring-rest-template' &&
c.confidence === 0.7,
),
).toBeDefined();
expect(
consumers.find(
(c) =>
c.contractId === 'http::POST::/api/users' &&
c.meta.framework === 'spring-web-client' &&
c.confidence === 0.7,
),
).toBeDefined();
});
// ─── Kotlin consumers (RestTemplate / WebClient short / OkHttp) ──
it('does NOT match Java WebClient long-form method(HttpMethod).uri(...) yet', async () => {
const dir = path.join(tmpDir, 'java-web-client-long-form');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'LongFormClient.java'),
`
import org.springframework.http.HttpMethod;
import org.springframework.web.reactive.function.client.WebClient;
class LongFormClient {
void run(WebClient webClient) {
webClient.method(HttpMethod.PATCH).uri("/api/users/42").retrieve();
}
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(
consumers.find((c) => c.contractId === 'http::PATCH::/api/users/{param}'),
).toBeUndefined();
});
it('extracts OpenFeign clients as consumers, not providers', async () => {
const dir = path.join(tmpDir, 'java-openfeign-consumer');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'OrderClient.java'),
`
import org.springframework.cloud.openfeign.FeignClient;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.PostMapping;
import org.springframework.web.bind.annotation.PathVariable;
@FeignClient(name = "order-service", url = "\${order.service.url}", path = "/api")
interface OrderClient {
@GetMapping("/orders/{id}")
OrderDto getOrder(@PathVariable("id") String id);
@PostMapping(path = "/orders")
OrderDto createOrder(OrderDto body);
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
const providers = contracts.filter((c) => c.role === 'provider');
expect(
consumers.find((c) => c.contractId === 'http::GET::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find(
(c) =>
c.contractId === 'http::POST::/api/orders' &&
c.meta.framework === 'openfeign' &&
c.confidence === 0.7,
),
).toBeDefined();
expect(
providers.find((c) => c.symbolRef.filePath.endsWith('OrderClient.java')),
).toBeUndefined();
});
it('extracts OpenFeign clients without an interface path prefix', async () => {
const dir = path.join(tmpDir, 'java-openfeign-no-prefix');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'HealthClient.java'),
`
import org.springframework.cloud.openfeign.FeignClient;
import org.springframework.web.bind.annotation.GetMapping;
@FeignClient(name = "health-service")
interface HealthClient {
@GetMapping("/health")
String health();
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
const providers = contracts.filter((c) => c.role === 'provider');
expect(
consumers.find(
(c) =>
c.contractId === 'http::GET::/health' &&
c.meta.framework === 'openfeign' &&
c.confidence === 0.7,
),
).toBeDefined();
expect(
providers.find((c) => c.symbolRef.filePath.endsWith('HealthClient.java')),
).toBeUndefined();
});
it('does not treat @FeignClient text in an interface body as a Feign annotation', async () => {
const dir = path.join(tmpDir, 'java-non-feign-interface-text');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'NotFeignClient.java'),
`
import org.springframework.web.bind.annotation.GetMapping;
interface NotFeignClient {
String MARKER = "@FeignClient";
@GetMapping("/not-feign")
String call();
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
const providers = contracts.filter((c) => c.role === 'provider');
expect(consumers.find((c) => c.contractId === 'http::GET::/not-feign')).toBeUndefined();
expect(providers.find((c) => c.contractId === 'http::GET::/not-feign')).toBeUndefined();
});
it('extracts OpenFeign clients with @RequestMapping interface prefixes', async () => {
const dir = path.join(tmpDir, 'java-openfeign-request-mapping-prefix');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'InventoryClient.java'),
`
import org.springframework.cloud.openfeign.FeignClient;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestMapping;
@FeignClient(name = "inventory-service")
@RequestMapping(path = "/api")
interface InventoryClient {
@GetMapping("/inventory/{id}")
InventoryDto getInventory(String id);
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(
consumers.find(
(c) =>
c.contractId === 'http::GET::/api/inventory/{param}' &&
c.meta.framework === 'openfeign',
),
).toBeDefined();
});
it('prefers @FeignClient(path=...) over @RequestMapping prefixes on OpenFeign clients', async () => {
const dir = path.join(tmpDir, 'java-openfeign-prefix-precedence');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'PrecedenceClient.java'),
`
import org.springframework.cloud.openfeign.FeignClient;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestMapping;
@FeignClient(name = "order-service", path = "/feign-path")
@RequestMapping("/rm-path")
interface PrecedenceClient {
@GetMapping("/orders")
OrderDto getOrders();
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers.find((c) => c.contractId === 'http::GET::/feign-path/orders')).toBeDefined();
expect(consumers.find((c) => c.contractId === 'http::GET::/rm-path/orders')).toBeUndefined();
});
it('extracts Java and Apache HttpClient literal request construction', async () => {
const dir = path.join(tmpDir, 'java-http-client-consumer');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'HttpClients.java'),
`
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import org.apache.http.client.methods.HttpGet;
import org.apache.http.client.methods.HttpPost;
import org.apache.http.client.methods.HttpPut;
import org.apache.http.client.methods.HttpDelete;
import org.apache.http.client.methods.HttpPatch;
class HttpClients {
void run(HttpClient client) throws Exception {
HttpRequest get = HttpRequest.newBuilder()
.uri(URI.create("/api/users/1"))
.GET()
.build();
HttpRequest post = HttpRequest.newBuilder()
.uri(URI.create("/api/users"))
.POST(HttpRequest.BodyPublishers.ofString("{}"))
.build();
new HttpGet("/api/orders/2");
new HttpPost("/api/orders");
new HttpPut("/api/orders/3");
new HttpDelete("/api/orders/4");
new HttpPatch("/api/orders/5");
}
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers.find((c) => c.contractId === 'http::GET::/api/users/{param}')).toBeDefined();
expect(
consumers.find(
(c) =>
c.contractId === 'http::POST::/api/users' &&
c.meta.framework === 'java-http-client' &&
c.confidence === 0.65,
),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::GET::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find(
(c) =>
c.contractId === 'http::POST::/api/orders' &&
c.meta.framework === 'apache-http-client' &&
c.confidence === 0.65,
),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::PUT::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::DELETE::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::PATCH::/api/orders/{param}'),
).toBeDefined();
});
// ─── Kotlin consumers (RestTemplate / WebClient short+long / OkHttp) ──
// Same shape as the Java consumer test above, but parsed by the
// tree-sitter-kotlin grammar via `KOTLIN_HTTP_PLUGIN`. Three
// consumer flavors covered here (long-form WebClient
// `webClient.method(HttpMethod.X).uri(...)` is intentionally
// deferred to a follow-up — see kotlin.ts file header).
// tree-sitter-kotlin grammar via `KOTLIN_HTTP_PLUGIN`. Four
// consumer flavors covered here: RestTemplate (#1855), WebClient
// short form (#1855), OkHttp (#1855), and WebClient long form
// (`webClient.method(HttpMethod.X).uri(...)`, this PR / #1884) —
// see kotlin.ts file header for the full list.
//
// tree-sitter-kotlin is an optionalDependency. If the binding is
// unavailable, `getPluginForFile` returns undefined for `.kt` and
@ -1581,28 +2029,104 @@ class OkPostClient(private val client: OkHttpClient, private val body: RequestBo
},
);
itKotlinConsumer('extracts Kotlin WebClient long form GET', async () => {
const dir = path.join(tmpDir, 'kotlin-web-client-long-get');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'LongGetClient.kt'),
`package com.example
import org.springframework.http.HttpMethod
import org.springframework.web.reactive.function.client.WebClient
import org.springframework.web.reactive.function.client.awaitBody
class LongGetClient(private val webClient: WebClient) {
suspend fun run() {
val r = webClient.method(HttpMethod.GET).uri("/api/users").retrieve().awaitBody<User>()
}
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
const route = consumers.find((c) => c.contractId === 'http::GET::/api/users');
expect(route).toBeDefined();
expect(route!.meta.framework).toBe('spring-web-client');
});
itKotlinConsumer('extracts Kotlin WebClient long form POST/PUT/DELETE/PATCH', async () => {
const dir = path.join(tmpDir, 'kotlin-web-client-long-verbs');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'LongVerbClient.kt'),
`package com.example
import org.springframework.http.HttpMethod
import org.springframework.web.reactive.function.client.WebClient
import org.springframework.web.reactive.function.client.awaitBody
import org.springframework.web.reactive.function.client.awaitBodilessEntity
class LongVerbClient(private val webClient: WebClient) {
suspend fun run() {
webClient.method(HttpMethod.POST).uri("/api/orders").retrieve().awaitBody<Order>()
webClient.method(HttpMethod.PUT).uri("/api/orders/1").retrieve().awaitBody<Order>()
webClient.method(HttpMethod.DELETE).uri("/api/orders/2").retrieve().awaitBodilessEntity()
webClient.method(HttpMethod.PATCH).uri("/api/orders/3").retrieve().awaitBody<Order>()
}
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
expect(consumers.find((c) => c.contractId === 'http::POST::/api/orders')).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::PUT::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::DELETE::/api/orders/{param}'),
).toBeDefined();
expect(
consumers.find((c) => c.contractId === 'http::PATCH::/api/orders/{param}'),
).toBeDefined();
// All four should be tagged as `spring-web-client` so polyglot
// repos coalesce on the same framework key as the short form.
// The fixture is fully deterministic — exactly 4 long-form calls,
// no short-form / RestTemplate / OkHttp calls mixed in — so an
// exact count is meaningful (DoD §2.7). If a future change
// accidentally emits a 5th consumer (e.g. duplicate query firing,
// or a regressed receiver constraint matching unrelated calls),
// this assertion catches it.
const wcConsumers = consumers.filter((c) => c.meta.framework === 'spring-web-client');
expect(wcConsumers).toHaveLength(4);
});
itKotlinConsumer(
'does NOT match Kotlin WebClient long form (deferred to follow-up)',
'short-form query does NOT also fire on Kotlin WebClient long form (no double-emit)',
async () => {
// Anti-overreach: confirm the short-form query does NOT
// accidentally fire on the long-form chain
// `webClient.method(HttpMethod.GET).uri(...)`. The long form
// is intentionally unsupported in this PR; if a future change
// to the short-form query starts capturing it we want a loud
// signal here. Long-form support will arrive in a follow-up
// with a dedicated query + verb walk-up helper.
const dir = path.join(tmpDir, 'kotlin-web-client-long');
// The long-form query handles `webClient.method(HttpMethod.X).uri(...)`,
// and the short-form query handles `webClient.get().uri(...)`. Both
// queries carry sibling `(navigation_suffix (simple_identifier) @verb)`
// constraints — short form requires the verb name itself
// (`get`/`post`/...), long form requires the literal name
// `method`. The two are disjoint.
//
// This test pins that disjointness: a single `.method(HttpMethod.GET)`
// call must emit ONE consumer, not two (one from each query).
const dir = path.join(tmpDir, 'kotlin-web-client-long-no-double');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'LegacyClient.kt'),
path.join(dir, 'src', 'NoDoubleClient.kt'),
`package com.example
import org.springframework.http.HttpMethod
import org.springframework.web.reactive.function.client.WebClient
import org.springframework.web.reactive.function.client.awaitBody
class LegacyClient(private val webClient: WebClient) {
class NoDoubleClient(private val webClient: WebClient) {
suspend fun run() {
val r = webClient.method(HttpMethod.GET).uri("/api/legacy").retrieve().awaitBody<String>()
webClient.method(HttpMethod.GET).uri("/api/single").retrieve().awaitBody<String>()
}
}
`,
@ -1611,12 +2135,49 @@ class LegacyClient(private val webClient: WebClient) {
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
// No consumer should be emitted from this file by the
// current short-form query. Documented as a known limitation.
const fromLegacy = consumers.filter((c) =>
c.symbolRef.filePath.endsWith('LegacyClient.kt'),
const fromThisFile = consumers.filter((c) =>
c.symbolRef.filePath.endsWith('NoDoubleClient.kt'),
);
expect(fromLegacy).toHaveLength(0);
expect(fromThisFile).toHaveLength(1);
expect(fromThisFile[0].contractId).toBe('http::GET::/api/single');
},
);
itKotlinConsumer(
'does NOT match Kotlin WebClient long form with variable-bound verb',
async () => {
// Anti-overreach: source-scan can't follow `val verb = HttpMethod.X`
// back to the literal — that's a graph-aware concern. The long-form
// query requires `(navigation_expression HttpMethod . verb)` as the
// `value_argument` shape, so a bare `simple_identifier` (the
// variable name) fails to match. Pin this so a future relaxation
// of the value_argument shape cannot silently start guessing the
// verb from arbitrary identifiers.
const dir = path.join(tmpDir, 'kotlin-web-client-long-var-verb');
fs.mkdirSync(path.join(dir, 'src'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'src', 'VariableVerbClient.kt'),
`package com.example
import org.springframework.http.HttpMethod
import org.springframework.web.reactive.function.client.WebClient
import org.springframework.web.reactive.function.client.awaitBody
class VariableVerbClient(private val webClient: WebClient) {
suspend fun run() {
val verb = HttpMethod.PATCH
val r = webClient.method(verb).uri("/api/dynamic").retrieve().awaitBody<String>()
}
}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const consumers = contracts.filter((c) => c.role === 'consumer');
const fromThisFile = consumers.filter((c) =>
c.symbolRef.filePath.endsWith('VariableVerbClient.kt'),
);
expect(fromThisFile).toHaveLength(0);
},
);
@ -1811,6 +2372,84 @@ async def create_user(user: UserCreate):
expect(providers.find((c) => c.contractId === 'http::GET::/users')).toBeDefined();
expect(providers.find((c) => c.contractId === 'http::POST::/users')).toBeDefined();
});
it('joins FastAPI @router.<verb> path with include_router(prefix=...) from main.py (attribute shape)', async () => {
const dir = path.join(tmpDir, 'fastapi-router-attr');
fs.mkdirSync(path.join(dir, 'api'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'main.py'),
`from fastapi import FastAPI
from api import assistant
app = FastAPI()
app.include_router(assistant.router, prefix='/ai', tags=['ai'])
`,
);
fs.writeFileSync(
path.join(dir, 'api/assistant.py'),
`from fastapi import APIRouter
router = APIRouter()
@router.post("/assistant")
async def assistant(req):
return {}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers.find((c) => c.contractId === 'http::POST::/ai/assistant')).toBeDefined();
// bare unprefixed form should not be emitted when a prefix mapping exists
expect(providers.find((c) => c.contractId === 'http::POST::/assistant')).toBeUndefined();
});
it('joins FastAPI @router.<verb> path with include_router(prefix=...) (named-import shape)', async () => {
const dir = path.join(tmpDir, 'fastapi-router-named');
fs.mkdirSync(path.join(dir, 'api'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'main.py'),
`from fastapi import FastAPI
from api.predict import router as predict_router
app = FastAPI()
app.include_router(predict_router, prefix='/ai')
`,
);
fs.writeFileSync(
path.join(dir, 'api/predict.py'),
`from fastapi import APIRouter
router = APIRouter()
@router.get("/concurrent")
async def concurrent():
return {}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers.find((c) => c.contractId === 'http::GET::/ai/concurrent')).toBeDefined();
});
it('emits @router.<verb> path unmodified when no include_router prefix is configured', async () => {
const dir = path.join(tmpDir, 'fastapi-router-no-prefix');
fs.mkdirSync(path.join(dir, 'api'), { recursive: true });
fs.writeFileSync(path.join(dir, 'main.py'), `app = None\n`);
fs.writeFileSync(
path.join(dir, 'api/loose.py'),
`from fastapi import APIRouter
router = APIRouter()
@router.get("/standalone")
async def standalone():
return {}
`,
);
const contracts = await extractor.extract(null, dir, makeRepo(dir));
const providers = contracts.filter((c) => c.role === 'provider');
expect(providers.find((c) => c.contractId === 'http::GET::/standalone')).toBeDefined();
});
});
describe('consumer extraction — graph-first (Strategy A)', () => {

View file

@ -97,7 +97,10 @@ describe('impact: batching and grouping', () => {
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
const params = args[2] || {};
if (query.includes('STEP_IN_PROCESS')) {
// Match only the aggregation chunk (which uses COUNT(DISTINCT s.id)),
// not the per-symbol enrichment pass added by impact byDepth processes
// (which also matches STEP_IN_PROCESS but has a different RETURN shape).
if (query.includes('STEP_IN_PROCESS') && query.includes('COUNT(DISTINCT s.id)')) {
// Count ids passed in as params.ids
const ids = Array.isArray(params.ids) ? params.ids : [];
const cnt = ids.length;
@ -263,7 +266,10 @@ describe('impact: batching and grouping', () => {
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
const params = args[2] || {};
if (query.includes('STEP_IN_PROCESS')) {
// Match only the aggregation chunk (which uses COUNT(DISTINCT s.id)),
// not the per-symbol enrichment pass added by impact byDepth processes
// (which also matches STEP_IN_PROCESS but has a different RETURN shape).
if (query.includes('STEP_IN_PROCESS') && query.includes('COUNT(DISTINCT s.id)')) {
const ids = Array.isArray(params.ids) ? params.ids : [];
chunkSizes.push(ids.length);
return [

View file

@ -25,6 +25,8 @@ const minimalResult = (overrides: Partial<ParseWorkerResult> = {}): ParseWorkerR
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],

View file

@ -2,6 +2,7 @@ import { describe, expect, it, vi } from 'vitest';
import {
ExtensionManager,
getExtensionInstallChildProcessArgs,
getExtensionInstallPolicy,
getExtensionInstallTimeoutMs,
type ExtensionInstallResult,
} from '../../src/core/lbug/extension-loader.js';
@ -222,6 +223,64 @@ describe('installDuckDbExtensionOutOfProcess child process', () => {
});
});
describe('getExtensionInstallPolicy', () => {
it('defaults to load-only when env var is unset', () => {
const original = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
delete process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
try {
expect(getExtensionInstallPolicy()).toBe('load-only');
} finally {
if (original === undefined) {
delete process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
} else {
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = original;
}
}
});
it('returns auto when env var is set to auto', () => {
const original = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = 'auto';
try {
expect(getExtensionInstallPolicy()).toBe('auto');
} finally {
if (original === undefined) {
delete process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
} else {
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = original;
}
}
});
it('returns never when env var is set to never', () => {
const original = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = 'never';
try {
expect(getExtensionInstallPolicy()).toBe('never');
} finally {
if (original === undefined) {
delete process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
} else {
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = original;
}
}
});
it('falls back to load-only for invalid env var values', () => {
const original = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = 'bogus';
try {
expect(getExtensionInstallPolicy()).toBe('load-only');
} finally {
if (original === undefined) {
delete process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
} else {
process.env.GITNEXUS_LBUG_EXTENSION_INSTALL = original;
}
}
});
});
describe('getExtensionInstallTimeoutMs', () => {
it('reads a positive override from the environment', () => {
const original = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS;

View file

@ -41,6 +41,8 @@ const emptyWorkerResult = (filePath: string, name: string): ParseWorkerResult =>
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@ -74,7 +76,7 @@ fs.writeFileSync(${JSON.stringify(markerPath)}, 'spawned');
parentPort.postMessage({ type: 'ready' });
const accumulated = {
nodes: [], relationships: [], symbols: [], imports: [], calls: [], assignments: [], heritage: [],
routes: [], fetchCalls: [], fetchWrapperDefs: [], decoratorRoutes: [], toolDefs: [], ormQueries: [], constructorBindings: [],
routes: [], fetchCalls: [], fetchWrapperDefs: [], decoratorRoutes: [], routerIncludes: [], routerImports: [], toolDefs: [], ormQueries: [], constructorBindings: [],
fileScopeBindings: [], parsedFiles: [], skippedLanguages: {}, fileCount: 0,
};
parentPort.on('message', (msg) => {

View file

@ -19,6 +19,7 @@ describe('runFullAnalysis FTS repair and verification failure paths', () => {
vi.doUnmock('../../src/core/lbug/lbug-adapter.js');
vi.doUnmock('../../src/core/search/fts-indexes.js');
vi.doUnmock('../../src/core/ingestion/pipeline.js');
vi.doUnmock('../../src/storage/repo-manager.js');
vi.resetModules();
vi.clearAllMocks();
});
@ -211,6 +212,8 @@ describe('runFullAnalysis FTS repair and verification failure paths', () => {
deleteNodesForFile: vi.fn(async () => undefined),
deleteAllCommunitiesAndProcesses: vi.fn(async () => undefined),
queryImporters: vi.fn(async () => []),
// FTS extension loads → analyze proceeds to create + verify indexes.
loadFTSExtension: vi.fn(async () => true),
}));
vi.doMock('../../src/core/search/fts-indexes.js', () => ({
createSearchFTSIndexes: vi.fn(async () => undefined),
@ -240,4 +243,66 @@ describe('runFullAnalysis FTS repair and verification failure paths', () => {
await tmpRepo.cleanup();
}
});
it('full analyze degrades gracefully (no throw, warns, skips index creation) when FTS extension is unavailable', async () => {
// Offline-first degradation: when loadFTSExtension() returns false, the
// analyze path must NOT call createSearchFTSIndexes / verifySearchFTSIndexes
// and must NOT throw — it logs a warning and completes (#1161).
const createSearchFTSIndexes = vi.fn(async () => undefined);
const verifySearchFTSIndexes = vi.fn(async () => []);
vi.doMock('../../src/core/lbug/lbug-adapter.js', () => ({
initLbug: vi.fn(async () => undefined),
loadGraphToLbug: vi.fn(async () => undefined),
getLbugStats: vi.fn(async () => ({ nodes: 1, edges: 0, communities: 0, processes: 0 })),
executeQuery: vi.fn(async () => []),
executeWithReusedStatement: vi.fn(async () => []),
closeLbug: vi.fn(async () => undefined),
loadCachedEmbeddings: vi.fn(async () => ({ embeddingNodeIds: new Set(), embeddings: [] })),
deleteNodesForFile: vi.fn(async () => undefined),
deleteAllCommunitiesAndProcesses: vi.fn(async () => undefined),
queryImporters: vi.fn(async () => []),
// FTS extension cannot load (offline + not pre-installed, or policy forced).
loadFTSExtension: vi.fn(async () => false),
}));
vi.doMock('../../src/core/search/fts-indexes.js', () => ({
createSearchFTSIndexes,
verifySearchFTSIndexes,
}));
vi.doMock('../../src/core/ingestion/pipeline.js', () => ({
runPipelineFromRepo: vi.fn(async (repoPath: string) => ({
repoPath,
totalFileCount: 1,
graph: { forEachNode: () => undefined },
})),
}));
// Avoid touching the global registry / repo .gitnexusignore from a unit test.
vi.doMock('../../src/storage/repo-manager.js', async (importActual) => ({
...(await importActual<typeof import('../../src/storage/repo-manager.js')>()),
registerRepo: vi.fn(async () => 'degraded-repo'),
ensureGitNexusIgnored: vi.fn(async () => undefined),
}));
const tmpRepo = await createTempDir('gitnexus-run-analyze-fts-degrade-');
try {
const logs: string[] = [];
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
const result = await runFullAnalysis(
tmpRepo.dbPath,
{ force: true },
{ onProgress: () => {}, onLog: (msg: string) => logs.push(msg) },
);
expect(result.ftsSkipped).toBe(true);
expect(createSearchFTSIndexes).not.toHaveBeenCalled();
expect(verifySearchFTSIndexes).not.toHaveBeenCalled();
expect(logs.join('\n')).toMatch(/FTS extension unavailable; skipping search-index creation/i);
// The degraded state is persisted so meta.json / doctor stay honest.
const { storagePath } = getStoragePaths(tmpRepo.dbPath);
const meta = JSON.parse(await fs.readFile(`${storagePath}/meta.json`, 'utf-8'));
expect(meta.capabilities.fts.status).toBe('unavailable');
} finally {
await tmpRepo.cleanup();
}
});
});

View file

@ -283,7 +283,7 @@ describe('populateCsharpNamespaceSiblings', () => {
expect(Object.isFrozen(augmented)).toBe(false);
});
it('parses UTF-8-heavy cache-miss files before namespace sibling injection', () => {
it('scans (no re-parse) UTF-8-heavy cache-miss files before namespace sibling injection', () => {
const sibling = classDef('def:b.B', 'b.cs', 'Demo.B');
const moduleA = scope('scope:a:module', 'Module', 'a.cs');
const moduleB = scope('scope:b:module', 'Module', 'b.cs');
@ -322,6 +322,128 @@ describe('populateCsharpNamespaceSiblings', () => {
expect(bindingAugmentations.get(moduleA.id)?.get('B')?.[0]?.def.nodeId).toBe('def:b.B');
});
it('routes global-namespace types to workspaceFqnBindings, not per-scope augmentations (#1871 OOM guard)', () => {
// Types declared with NO `namespace` (the global/default namespace) are
// visible from every C# file, so the hook writes ONE workspace-level entry
// per simple name instead of O(scopes x defs) per-scope augmentations —
// the fix for the #1871 Unity-scale OOM. This pins both halves of that
// contract: global types are reachable via `workspaceFqnBindings`, and the
// per-scope augmentation channel stays empty for them.
//
// Note: the mock MUST supply `workspaceFqnBindings` — the global fast path
// reads `indexes.workspaceFqnBindings` directly, so omitting it (as the
// other tests in this suite do) would throw.
const defA = classDef('def:a.A', 'a.cs', 'A'); // simple name => global namespace
const defB = classDef('def:b.B', 'b.cs', 'B');
const moduleA = scope('scope:a:module', 'Module', 'a.cs');
const classA = scope('scope:a:class', 'Class', 'a.cs', moduleA.id, [defA]);
const moduleB = scope('scope:b:module', 'Module', 'b.cs');
const classB = scope('scope:b:class', 'Class', 'b.cs', moduleB.id, [defB]);
const parsedFiles: ParsedFile[] = [
{
filePath: 'a.cs',
moduleScope: moduleA.id,
scopes: Object.freeze([moduleA, classA]),
parsedImports: Object.freeze([]),
localDefs: Object.freeze([defA]),
referenceSites: Object.freeze([]),
} as ParsedFile,
{
filePath: 'b.cs',
moduleScope: moduleB.id,
scopes: Object.freeze([moduleB, classB]),
parsedImports: Object.freeze([]),
localDefs: Object.freeze([defB]),
referenceSites: Object.freeze([]),
} as ParsedFile,
];
const bindingAugmentations = new Map<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>();
const workspaceFqnBindings = new Map<string, readonly BindingRef[]>();
populateCsharpNamespaceSiblings(
parsedFiles,
{
bindings: new Map(),
bindingAugmentations,
workspaceFqnBindings,
} as unknown as ScopeResolutionIndexes,
{
fileContents: new Map([
['a.cs', 'class A { }\n'], // no `namespace` => global
['b.cs', 'class B { }\n'],
]),
},
);
// Global types are reachable workspace-wide via simple-name keys.
expect(workspaceFqnBindings.get('A')?.map((b) => b.def.nodeId)).toEqual(['def:a.A']);
expect(workspaceFqnBindings.get('B')?.map((b) => b.def.nodeId)).toEqual(['def:b.B']);
// O(D) invariant: one entry per unique simple name, never scopes x defs.
expect(workspaceFqnBindings.size).toBe(2);
// The whole point of the fast path: no per-scope augmentation explosion.
expect(bindingAugmentations.size).toBe(0);
// Workspace entries carry the cross-file `namespace` origin (so shadowing
// precedence in lookupBindingsAt orders them after local/finalized).
expect(workspaceFqnBindings.get('A')?.[0]?.origin).toBe('namespace');
});
it('keeps every declaration of a repeated global simple name (partial classes across files)', () => {
// Two global-namespace files each declare `Foo` (a partial class split
// across files => distinct nodeIds, same simple name). Both must survive
// in the workspace channel — the fast path de-dups by nodeId, not by name,
// so partial-class members from both files stay resolvable.
const foo1 = classDef('def:foo1.Foo', 'foo1.cs', 'Foo');
const foo2 = classDef('def:foo2.Foo', 'foo2.cs', 'Foo');
const moduleA = scope('scope:foo1:module', 'Module', 'foo1.cs');
const classA = scope('scope:foo1:class', 'Class', 'foo1.cs', moduleA.id, [foo1]);
const moduleB = scope('scope:foo2:module', 'Module', 'foo2.cs');
const classB = scope('scope:foo2:class', 'Class', 'foo2.cs', moduleB.id, [foo2]);
const parsedFiles: ParsedFile[] = [
{
filePath: 'foo1.cs',
moduleScope: moduleA.id,
scopes: Object.freeze([moduleA, classA]),
parsedImports: Object.freeze([]),
localDefs: Object.freeze([foo1]),
referenceSites: Object.freeze([]),
} as ParsedFile,
{
filePath: 'foo2.cs',
moduleScope: moduleB.id,
scopes: Object.freeze([moduleB, classB]),
parsedImports: Object.freeze([]),
localDefs: Object.freeze([foo2]),
referenceSites: Object.freeze([]),
} as ParsedFile,
];
const bindingAugmentations = new Map<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>();
const workspaceFqnBindings = new Map<string, readonly BindingRef[]>();
populateCsharpNamespaceSiblings(
parsedFiles,
{
bindings: new Map(),
bindingAugmentations,
workspaceFqnBindings,
} as unknown as ScopeResolutionIndexes,
{
fileContents: new Map([
['foo1.cs', 'class Foo { }\n'],
['foo2.cs', 'class Foo { }\n'],
]),
},
);
// Both partial declarations are kept (de-dup is by nodeId, not name).
expect(
workspaceFqnBindings
.get('Foo')
?.map((b) => b.def.nodeId)
.sort(),
).toEqual(['def:foo1.Foo', 'def:foo2.Foo']);
expect(bindingAugmentations.size).toBe(0);
});
});
describe('csharpReceiverBinding', () => {

View file

@ -0,0 +1,124 @@
/**
* Coverage for the JavaScript scope-captures orchestrator, focused on the
* #1876 array-method-callback narrowing.
*
* `const x = arr.map(a => …)` must NOT produce a `@declaration.function`
* named `x` (the binding holds a value, not a callable) — only the
* `@declaration.const`. Identifier-callee HOCs (`forwardRef`, `useMemo`)
* and direct arrow assignments keep their `@declaration.function`.
*
* Runs against tree-sitter-javascript so it catches grammar drift before
* the integration parity gate.
*/
import { describe, it, expect } from 'vitest';
import { emitJsScopeCaptures } from '../../../../src/core/ingestion/languages/javascript/captures.js';
function matchesFor(src: string) {
return emitJsScopeCaptures(src, 'test.js');
}
/** True when some match carries `tag` and its @declaration.name is `name`. */
function hasDecl(src: string, tag: string, name: string): boolean {
return matchesFor(src).some((m) => m[tag] !== undefined && m['@declaration.name']?.text === name);
}
/** Count matches carrying `tag` (any name). */
function countTag(src: string, tag: string): number {
return matchesFor(src).filter((m) => m[tag] !== undefined).length;
}
describe('emitJsScopeCaptures — #1876 array-method-callback narrowing', () => {
it('does not emit @declaration.function for `const x = arr.map(a => …)`', () => {
const src = 'const exportData = accountsList.map(account => ({ id: account.id }));';
expect(hasDecl(src, '@declaration.const', 'exportData')).toBe(true);
expect(hasDecl(src, '@declaration.function', 'exportData')).toBe(false);
// Exactly one binding-bearing declaration for the name.
expect(countTag(src, '@declaration.function')).toBe(0);
});
// Every method in ARRAY_CALLBACK_METHODS except `map` (covered above).
it.each([
'filter',
'find',
'findIndex',
'findLast',
'findLastIndex',
'reduce',
'reduceRight',
'forEach',
'some',
'every',
'flatMap',
'sort',
])('suppresses the Function def for array method .%s()', (method) => {
const src = `const x = arr.${method}((a) => a);`;
expect(hasDecl(src, '@declaration.function', 'x')).toBe(false);
expect(hasDecl(src, '@declaration.const', 'x')).toBe(true);
});
it('keeps @declaration.function for an identifier-callee HOC (forwardRef)', () => {
const src = 'const Button = forwardRef((props, ref) => null);';
expect(hasDecl(src, '@declaration.function', 'Button')).toBe(true);
});
it('keeps @declaration.function for useMemo (identifier callee, unchanged this round)', () => {
const src = 'const value = useMemo(() => compute(), []);';
expect(hasDecl(src, '@declaration.function', 'value')).toBe(true);
});
it('keeps dual classification for a direct arrow `const fn = () => {}`', () => {
const src = 'const fn = () => { doThing(); };';
expect(hasDecl(src, '@declaration.function', 'fn')).toBe(true);
expect(hasDecl(src, '@declaration.const', 'fn')).toBe(true);
});
it('keeps @declaration.function for a non-array fluent-API member call (accepted limitation)', () => {
const src = 'const q = qb.where((row) => row.ok);';
expect(hasDecl(src, '@declaration.function', 'q')).toBe(true);
});
it('suppresses an in-set method name on a NON-array receiver (accepted receiver-blind limitation)', () => {
// The predicate keys on the method NAME only, never the receiver type —
// tree-sitter has no type info. So `.map` on an RxJS observable (or
// Map/Set `.forEach`, a query builder `.sort`, a lodash chain `.filter`)
// is also treated as a callback and loses its Function def. Accepted: the
// binding holds the call's result value, so a value def is correct anyway.
const src = 'const stream = source$.map((event) => handle(event));';
expect(hasDecl(src, '@declaration.function', 'stream')).toBe(false);
expect(hasDecl(src, '@declaration.const', 'stream')).toBe(true);
});
it('suppresses the outer .map() callback in a chained array call', () => {
const src = 'const x = arr.filter((a) => a).map((b) => b);';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(false);
expect(hasDecl(src, '@declaration.const', 'x')).toBe(true);
});
it('suppresses through an export_statement wrapper', () => {
const src = 'export const x = arr.map((a) => a);';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(false);
expect(hasDecl(src, '@declaration.const', 'x')).toBe(true);
});
it('suppresses a function_expression callback', () => {
const src = 'const x = arr.map(function (a) { return a; });';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(false);
expect(hasDecl(src, '@declaration.const', 'x')).toBe(true);
});
it('suppresses an optional-chained array call `arr?.map(...)`', () => {
const src = 'const x = arr?.map((a) => a);';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(false);
});
it('does NOT suppress a parenthesized callee `(arr.map)(cb)` (intentional gap)', () => {
const src = 'const x = (arr.map)((a) => a);';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(true);
});
it('does NOT suppress a computed callee `arr["map"](cb)` (intentional gap)', () => {
const src = 'const x = arr["map"]((a) => a);';
expect(hasDecl(src, '@declaration.function', 'x')).toBe(true);
});
});

View file

@ -16,8 +16,13 @@ import type {
ReferenceKind,
Scope,
ScopeKind,
SymbolDefinition,
} from 'gitnexus-shared';
import { extract, type ScopeExtractorHooks } from '../../../src/core/ingestion/scope-extractor.js';
import {
extract,
selectNodeBearingDef,
type ScopeExtractorHooks,
} from '../../../src/core/ingestion/scope-extractor.js';
// ─── Synthetic-capture helpers ──────────────────────────────────────────────
@ -550,3 +555,54 @@ describe('end-to-end fixture (all 5 passes together)', () => {
expect(result.moduleScope).toBe(mod.id);
});
});
describe('selectNodeBearingDef — #1876 one-node-per-binding collapse rule', () => {
const def = (type: SymbolDefinition['type'], name = 'x'): SymbolDefinition => ({
nodeId: `def:test.ts#1:0:${type}:${name}`,
filePath: 'test.ts',
type,
qualifiedName: name,
});
it('returns undefined for an empty group', () => {
expect(selectNodeBearingDef([])).toBeUndefined();
});
it('returns the only def for a single-element group', () => {
const only = def('Variable');
expect(selectNodeBearingDef([only])).toBe(only);
});
it('prefers a Function over a co-bound Variable (direct arrow / HOC)', () => {
const fn = def('Function');
const variable = def('Variable');
// Order-independent: function-like wins regardless of position.
expect(selectNodeBearingDef([variable, fn])).toBe(fn);
expect(selectNodeBearingDef([fn, variable])).toBe(fn);
});
it('prefers a Method over a co-bound value def', () => {
const method = def('Method');
const variable = def('Variable');
expect(selectNodeBearingDef([variable, method])).toBe(method);
});
it('returns the value def when no function-like def is present (array-method result)', () => {
const constDef = def('Const');
expect(selectNodeBearingDef([constDef])).toBe(constDef);
const variable = def('Variable');
expect(selectNodeBearingDef([variable])).toBe(variable);
});
it('prefers a value def even when an unranked label appears first', () => {
const cls = def('Class');
const variable = def('Variable');
expect(selectNodeBearingDef([cls, variable])).toBe(variable);
});
it('falls back to the first def for label sets the rule does not rank', () => {
const cls = def('Class');
const iface = def('Interface');
expect(selectNodeBearingDef([cls, iface])).toBe(cls);
});
});

View file

@ -565,3 +565,100 @@ describe('emitTsScopeCaptures — edge cases', () => {
expect(() => emitTsScopeCaptures('', 'test.ts')).not.toThrow();
});
});
describe('emitTsScopeCaptures — #1876 array-method-callback narrowing', () => {
// True when some match carries `tag` and its @declaration.name is `name`.
const declWithName = (src: string, tag: string, name: string): boolean =>
emitTsScopeCaptures(src, 'test.ts').some(
(m) => m[tag] !== undefined && m['@declaration.name']?.text === name,
);
it('does not emit @declaration.function for `const x = arr.map(a => …)`', () => {
const src = 'const exportData = accountsList.map((account) => ({ id: account.id }));';
expect(declWithName(src, '@declaration.variable', 'exportData')).toBe(true);
expect(declWithName(src, '@declaration.function', 'exportData')).toBe(false);
});
// Every method in ARRAY_CALLBACK_METHODS except `map` (covered above).
it.each([
'filter',
'find',
'findIndex',
'findLast',
'findLastIndex',
'reduce',
'reduceRight',
'forEach',
'some',
'every',
'flatMap',
'sort',
])('suppresses the Function def for array method .%s()', (method) => {
const src = `const x = arr.${method}((a) => a);`;
expect(declWithName(src, '@declaration.function', 'x')).toBe(false);
expect(declWithName(src, '@declaration.variable', 'x')).toBe(true);
});
it('keeps @declaration.function for an identifier-callee HOC (forwardRef)', () => {
const src = 'const Button = forwardRef((props, ref) => null);';
expect(declWithName(src, '@declaration.function', 'Button')).toBe(true);
});
it('keeps @declaration.function for useCallback (identifier callee, unchanged this round)', () => {
const src = 'const cb = useCallback(() => doThing(), []);';
expect(declWithName(src, '@declaration.function', 'cb')).toBe(true);
});
it('keeps dual classification for a direct arrow `const fn = () => {}`', () => {
const src = 'const fn = () => { doThing(); };';
expect(declWithName(src, '@declaration.function', 'fn')).toBe(true);
expect(declWithName(src, '@declaration.variable', 'fn')).toBe(true);
});
it('keeps @declaration.function for a non-array fluent-API member call (accepted limitation)', () => {
const src = 'const q = qb.where((row) => row.ok);';
expect(declWithName(src, '@declaration.function', 'q')).toBe(true);
});
it('suppresses an in-set method name on a NON-array receiver (accepted receiver-blind limitation)', () => {
// Receiver-blind by design — see array-callback.ts. An in-set method name
// on a non-array receiver (RxJS observable, Map/Set, query builder) also
// loses its Function def. Accepted: the binding holds a value, not a callable.
const src = 'const stream = source$.map((event) => handle(event));';
expect(declWithName(src, '@declaration.function', 'stream')).toBe(false);
expect(declWithName(src, '@declaration.variable', 'stream')).toBe(true);
});
it('suppresses the outer .map() callback in a chained array call', () => {
const src = 'const x = arr.filter((a) => a).map((b) => b);';
expect(declWithName(src, '@declaration.function', 'x')).toBe(false);
expect(declWithName(src, '@declaration.variable', 'x')).toBe(true);
});
it('suppresses through an export_statement wrapper', () => {
const src = 'export const x = arr.map((a) => a);';
expect(declWithName(src, '@declaration.function', 'x')).toBe(false);
expect(declWithName(src, '@declaration.variable', 'x')).toBe(true);
});
it('suppresses a function_expression callback', () => {
const src = 'const x = arr.map(function (a) { return a; });';
expect(declWithName(src, '@declaration.function', 'x')).toBe(false);
expect(declWithName(src, '@declaration.variable', 'x')).toBe(true);
});
it('suppresses an optional-chained array call `arr?.map(...)`', () => {
const src = 'const x = arr?.map((a) => a);';
expect(declWithName(src, '@declaration.function', 'x')).toBe(false);
});
it('does NOT suppress a parenthesized callee `(arr.map)(cb)` (intentional gap)', () => {
const src = 'const x = (arr.map)((a) => a);';
expect(declWithName(src, '@declaration.function', 'x')).toBe(true);
});
it('does NOT suppress a computed callee `arr["map"](cb)` (intentional gap)', () => {
const src = 'const x = arr["map"]((a) => a);';
expect(declWithName(src, '@declaration.function', 'x')).toBe(true);
});
});

View file

@ -22,10 +22,12 @@ const mkRef = (nodeId: string): BindingRef =>
const mkIndexes = (
bindings: Map<ScopeId, Map<string, readonly BindingRef[]>>,
augmentations: Map<ScopeId, Map<string, BindingRef[]>>,
workspace: Map<string, readonly BindingRef[]> = new Map(),
): ScopeResolutionIndexes =>
({
bindings,
bindingAugmentations: augmentations,
workspaceFqnBindings: workspace,
}) as unknown as ScopeResolutionIndexes;
describe('validateBindingsImmutability', () => {
@ -88,6 +90,24 @@ describe('validateBindingsImmutability', () => {
expect(onWarn.mock.calls[0][0]).toMatch(/I8/);
});
it('warns when a bucket in indexes.workspaceFqnBindings IS frozen', () => {
vi.stubEnv('NODE_ENV', 'development');
const workspace = new Map<string, readonly BindingRef[]>([
['User', Object.freeze([mkRef('def:User')]) as BindingRef[]],
]);
const onWarn = vi.fn();
const violations = validateBindingsImmutability(
mkIndexes(new Map(), new Map(), workspace),
onWarn,
);
expect(violations).toBe(1);
expect(onWarn).toHaveBeenCalledTimes(1);
expect(onWarn.mock.calls[0][0]).toMatch(/indexes\.workspaceFqnBindings/);
expect(onWarn.mock.calls[0][0]).toMatch(/I8/);
});
it('does not detect semantically wrong frozen replacements in indexes.bindings', () => {
vi.stubEnv('NODE_ENV', 'development');
const bindings = new Map<ScopeId, Map<string, readonly BindingRef[]>>([

View file

@ -32,9 +32,11 @@ const ref = (nodeId: string, origin: BindingRef['origin'] = 'local'): BindingRef
function indexesWith({
finalized,
augmented,
workspace,
}: {
finalized?: readonly BindingRef[];
augmented?: readonly BindingRef[];
workspace?: readonly BindingRef[];
}): ScopeResolutionIndexes {
const bindings = new Map<ScopeId, Map<string, readonly BindingRef[]>>();
if (finalized !== undefined) {
@ -43,7 +45,13 @@ function indexesWith({
}
const bindingAugmentations = new Map<ScopeId, Map<string, readonly BindingRef[]>>();
if (augmented !== undefined) bindingAugmentations.set(SCOPE, new Map([['name', augmented]]));
return { bindings, bindingAugmentations } as unknown as ScopeResolutionIndexes;
const workspaceFqnBindings = new Map<string, readonly BindingRef[]>();
if (workspace !== undefined) workspaceFqnBindings.set('name', workspace);
return {
bindings,
bindingAugmentations,
workspaceFqnBindings,
} as unknown as ScopeResolutionIndexes;
}
function scope(id: ScopeId, bindings = new Map<string, readonly BindingRef[]>()): Scope {
@ -105,6 +113,33 @@ describe('lookupBindingsAt', () => {
expect(out.find((b) => b.def.nodeId === 'A')!.origin).toBe('import');
});
// Third channel: workspaceFqnBindings (scope-independent — global-namespace
// C# types / PHP FQNs). Consulted LAST, after finalized + augmented.
it('returns the workspace bucket when it is the only channel', () => {
const workspace = [ref('W', 'namespace')];
const out = lookupBindingsAt(SCOPE, 'name', indexesWith({ workspace }));
expect(out).toEqual(workspace);
expect(out).toBe(workspace); // identity preserved when only one channel populates
});
it('appends workspace entries after finalized and augmented', () => {
const finalized = [ref('A', 'import')];
const augmented = [ref('B', 'namespace')];
const workspace = [ref('C', 'namespace')];
const out = lookupBindingsAt(SCOPE, 'name', indexesWith({ finalized, augmented, workspace }));
expect(out.map((b) => b.def.nodeId)).toEqual(['A', 'B', 'C']);
});
it('dedupes workspace entries already present in finalized/augmented (workspace loses)', () => {
const finalized = [ref('A', 'import')];
const augmented = [ref('B', 'namespace')];
const workspace = [ref('A', 'namespace'), ref('B', 'namespace'), ref('C', 'namespace')];
const out = lookupBindingsAt(SCOPE, 'name', indexesWith({ finalized, augmented, workspace }));
expect(out.map((b) => b.def.nodeId)).toEqual(['A', 'B', 'C']);
// The surviving A/B keep their finalized/augmented identity, not workspace's.
expect(out.find((b) => b.def.nodeId === 'A')!.origin).toBe('import');
});
it('keeps finalized metadata when the same nodeId appears in both channels', () => {
const finalizedDef = {
nodeId: 'A',

View file

@ -15,7 +15,11 @@
import { describe, it, expect, vi } from 'vitest';
import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { InMemoryTransport } from '@modelcontextprotocol/sdk/inMemory.js';
import { createMCPServer } from '../../src/mcp/server.js';
import {
createMCPServer,
installSignalShutdown,
SHUTDOWN_EXIT_CODES,
} from '../../src/mcp/server.js';
import { GITNEXUS_TOOLS } from '../../src/mcp/tools.js';
// ─── Mock backend ──────────────────────────────────────────────────
@ -125,3 +129,38 @@ describe('prompt registration', () => {
expect(server).toBeDefined();
});
});
// ─── Graceful shutdown signal handling (#1132) ────────────────────────
describe('installSignalShutdown (#1132)', () => {
it('maps SIGINT→130 / SIGTERM→143 and never passes the signal name to shutdown', () => {
// Node invokes signal listeners with the signal NAME string as the first
// argument. The old code registered `shutdown` directly, so that string
// reached process.exit() and crashed with ERR_INVALID_ARG_TYPE. Reproduce
// that exact invocation and assert a numeric code is used instead.
const received: unknown[] = [];
let onSigint: ((...args: unknown[]) => void) | undefined;
let onSigterm: ((...args: unknown[]) => void) | undefined;
installSignalShutdown(
(code) => received.push(code),
(event, listener) => {
if (event === 'SIGINT') onSigint = listener;
if (event === 'SIGTERM') onSigterm = listener;
},
);
expect(onSigint).toBeTypeOf('function');
expect(onSigterm).toBeTypeOf('function');
// Invoke exactly as Node does — with the signal name string as the arg.
onSigint?.('SIGINT');
onSigterm?.('SIGTERM');
expect(received).toEqual([SHUTDOWN_EXIT_CODES.SIGINT, SHUTDOWN_EXIT_CODES.SIGTERM]);
expect(received).toEqual([130, 143]);
for (const code of received) {
expect(typeof code).toBe('number');
}
});
});

72
pr-swarm-review/README.md Normal file
View file

@ -0,0 +1,72 @@
# GitNexus PR Reviewer Swarm (cross-CLI)
A coordinated, **read-only** production-readiness PR review for GitNexus, runnable from any
AI coding CLI. Seven specialized review personas produce one structured, evidence-grounded
review.
## Single source of truth
All review logic lives here and is shared by every CLI — edit these, not the per-CLI wrappers:
```
pr-swarm-review/
orchestration.md # coordinator contract: Swarm vs Solo modes, lanes, classifications, output structure
personas/ # the 7 canonical persona prompts (role + rules + output sections)
01-pr-facts-historian.md (model tier: sonnet)
02-branch-hygiene-reviewer.md (model tier: haiku)
03-risk-architect.md (model tier: sonnet)
04-test-ci-verifier.md (model tier: haiku)
05-security-boundary-reviewer.md (model tier: sonnet)
06-docs-dod-reviewer.md (model tier: sonnet)
07-synthesis-critic.md (model tier: sonnet)
README.md # this file
```
Per-CLI entrypoints are **thin wrappers** that read the files above at runtime. Only
Claude Code has first-class parallel subagents (**Swarm mode**); every other CLI runs the
same lanes sequentially in one agent (**Solo mode**) with an identical output contract.
## Invoke it from your CLI
| CLI | How to invoke | Adapter file |
|-----|---------------|--------------|
| **Claude Code** | `/gitnexus-pr-swarm-review <PR>` (Swarm mode; dispatches the 7 `gitnexus-*` subagents) | `.claude/skills/gitnexus-pr-swarm-review/SKILL.md` + `.claude/agents/gitnexus-*.md` |
| **Gemini CLI** | `/gitnexus-pr-swarm-review <PR>` | `.gemini/commands/gitnexus-pr-swarm-review.toml` |
| **GitHub Copilot** | `/gitnexus-pr-swarm-review` (then paste the PR) | `.github/prompts/gitnexus-pr-swarm-review.prompt.md` |
| **Cursor** | `/gitnexus-pr-swarm-review` (then paste the PR) | `.cursor/commands/gitnexus-pr-swarm-review.md` |
| **Codex CLI** | Ask: "run the GitNexus PR swarm review for <PR>" (Codex reads `AGENTS.md`) — or install the user-level prompt below | `AGENTS.md` § PR Swarm Review |
| **Any AGENTS.md-aware agent** | Ask it to "follow `pr-swarm-review/orchestration.md` for <PR>" | `AGENTS.md` § PR Swarm Review |
### Codex (optional user-level slash command)
Codex prompts are user-level only (not repo-shareable). To get a `/gitnexus-pr-swarm-review`
slash command, create `~/.codex/prompts/gitnexus-pr-swarm-review.md`:
```markdown
---
description: GitNexus production-readiness PR swarm review (Solo mode)
argument-hint: <PR URL or number>
---
Read `pr-swarm-review/orchestration.md` in this repo and run it in **Solo mode** for $ARGUMENTS.
You are single-agent: adopt each persona in `pr-swarm-review/personas/` in dependency order,
then self-critique with lane 7 before emitting the review. Stay read-only.
```
## Key properties
- **Read-only.** No persona edits files, commits, or posts to GitHub. Each enforces an
explicit permitted/prohibited Bash list.
- **Evidence-grounded.** Every finding cites files, line ranges, checks, issue/PR refs, or commands.
- **Missing visibility becomes verification work** rather than invented facts.
- **Manually invoked.** No hooks or automatic triggers.
## Extending to a new CLI
Add one thin wrapper for the CLI's command/prompt format whose body says: *read
`pr-swarm-review/orchestration.md` and run it (Swarm mode if the runtime has parallel
subagents, else Solo mode)*. Do not copy the persona/orchestration text into the wrapper.
## Relationship to the existing review skill
This coexists with `/gitnexus-pr-review` (a single-agent linear checklist using GitNexus MCP
tools). This swarm is a multi-agent / multi-persona deep production-readiness review.

View file

@ -0,0 +1,136 @@
# GitNexus PR Swarm Review — Orchestration (canonical, CLI-neutral)
This is the single source of truth for the GitNexus production-readiness PR review.
Every per-CLI entrypoint (Claude Code skill/agents, Codex/Gemini/Cursor/Copilot prompts,
or any AGENTS.md-driven agent) **reads this file and follows it**. Edit the review logic
here, never in the per-CLI wrappers.
You are the **review coordinator**. Do not flatten the review into a generic checklist.
Run the seven specialized lanes below and synthesize one evidence-grounded review.
## Invocation
The adapter passes a target: `<PR URL or PR number>` for the GitNexus repository
(`https://github.com/abhigyanpatwari/GitNexus`). If no target was passed, ask for one.
## Execution modes
Pick the mode your runtime supports. **The output contract is identical in both modes.**
### Swarm mode — runtimes with parallel subagents (e.g. Claude Code)
Dispatch each lane as its own subagent (Claude Code: the `gitnexus-*` agents via the
Agent tool). Lanes 1–2 run first (their output feeds the rest); lanes 3–6 run in parallel
after lanes 1–2 complete; lane 7 runs last on the draft synthesis.
### Solo mode — single-agent runtimes (Codex, Gemini CLI, Cursor, Copilot, …)
One agent performs all lanes itself, **in dependency order**, adopting each persona in
turn: read `pr-swarm-review/personas/0N-<lane>.md`, do that lane's investigation, capture
its structured output, then move to the next. Keep every lane's findings in context so the
synthesis (lane 7) can self-critique against the whole. Lanes 3–6 have no dependency on
each other — do them in any order, but only after lanes 1–2.
> Both modes MUST honor the read-only contract: this review investigates and reports; it
> never edits files, commits, or posts to GitHub on its own.
## Lanes
Each lane's full spec is its persona file under `pr-swarm-review/personas/`.
| Lane | Persona file | Responsibility | Depends on |
|------|--------------|----------------|------------|
| 1 | `01-pr-facts-historian.md` | PR identity, visible state, changed files, linked issues, related PRs/commits, repo history, visibility gaps | — |
| 2 | `02-branch-hygiene-reviewer.md` | Merge-state + branch-hygiene classification | 1 |
| 3 | `03-risk-architect.md` | Production failure modes, domain-specific blockers | 1, 2 |
| 4 | `04-test-ci-verifier.md` | Test coverage, CI wiring, validation gaps | 1 |
| 5 | `05-security-boundary-reviewer.md` | Trust boundaries, secrets, injection, permissions, hidden Unicode | 1 |
| 6 | `06-docs-dod-reviewer.md` | PR-specific Definition of Done, docs/release-note obligations | 1 |
| 7 | `07-synthesis-critic.md` | Critique the draft review before it is emitted | 1–6 + draft |
**Lane 7 is a hard gate.** Do NOT emit the final review while the synthesis critic's
"Required corrections before posting" section is non-empty. Revise and re-run lane 7 until
that section is empty.
## Required repo docs
Read these first when present; if missing, note it and use the closest available guidance:
`DoD.md`, `AGENTS.md`, `GUARDRAILS.md`, `CONTRIBUTING.md`, `TESTING.md`, `ARCHITECTURE.md`.
## Visibility disclaimer
If visibility is incomplete, include this exact sentence before the final review (replace
A/B/C and X/Y/Z with the actual verified and missing items):
> Current visible state is incomplete. I could verify A, B, and C, but not X, Y, and Z. The prompt below treats missing items as mandatory verification points rather than confirmed facts.
## Classifications
**Branch hygiene** — exactly one of:
`clean feature/fix PR` · `merge-from-main commit present but harmless and merge-safe` ·
`polluted by unrelated merge/churn` · `rebase/split required`
**Merge state** — exactly one of:
`mergeable` · `blocked by conflicts` · `checks pending` · `checks failing` ·
`review blocked` · `draft/WIP` · `merged` · `closed without merge` · `visibility incomplete`
**Final verdict** — exactly one of (justify in 3–6 sentences):
`production-ready` · `production-ready with minor follow-ups` · `not production-ready` ·
`rebase/split required before final review`
## Final review structure
The final review **must include** all of these sections, in order:
1. **Review bar for this PR** — the DoD-derived acceptance criteria
2. **Problem being solved** — what the PR claims to fix or add
3. **Current PR state** — draft, open, merged, closed
4. **Merge status and mergeability** — merge-state classification with evidence
5. **Repository history considered** — related PRs, issues, historical fixes
6. **Branch hygiene assessment** — branch-hygiene classification with evidence
7. **Understanding of the change** — what the PR actually does
8. **Findings** — all findings from all lanes, using the Finding Format below
9. **PR-specific assessment sections** — domain-specific assessments relevant to this PR
10. **Back-and-forth avoided by verifying** — facts verified directly instead of assumed
11. **Open questions** — remaining questions, only if unavoidable after verification
12. **Final verdict** — one of the four allowed verdicts with a 3–6 sentence justification
## Finding format
- **Risk:** [the production risk]
- **Evidence to check:** [specific files, line ranges, commands, or checks]
- **Recommended fix:** [what should be done]
- **Blocks merge:** yes / no / maybe
## Hidden Unicode / hygiene checks
Include results from:
```bash
git diff --check origin/main...HEAD
git grep -nP '[\x{202A}-\x{202E}\x{2066}-\x{2069}]'
git grep -nP '[^\x00-\x7F]' -- ':!package-lock.json' ':!pnpm-lock.yaml' ':!yarn.lock'
```
Do not block ordinary visible punctuation if repo style allows it. Block hidden/bidi
controls in executable code, tests, YAML, Dockerfiles, query strings, regexes, security
comments, or otherwise misleading text.
## No-issues sentence
If no issues are found, say exactly:
> No production-readiness issues found against the current DoD bar.
## Review behavior
- **Never invent facts.** Use current visible state.
- **Convert uncertainty into mandatory verification work.**
- **Prioritize:** risk model first, PR facts second, repository history third.
- **Distinguish** confirmed findings from unverified suspicions.
- **Cite** files, line ranges, checks, issue/PR references, or commands used.
- **Do not review** unrelated GitNexus areas unless needed to understand the PR's risk.
- **Treat as suspicious:** unrelated workflow cleanup, release/version bumps, parser + web
UI refactors, Docker/CI churn, or test de-flake mixed with production behavior changes.
- **Request split or rebase** when domains are not causally connected.
- **One production-critical lane can block the whole PR.**

View file

@ -0,0 +1,74 @@
<!-- CANONICAL, CLI-NEUTRAL PERSONA — single source of truth for every adapter.
Edit this file, not the per-CLI wrappers (.claude/agents, .gemini, .github, .cursor). -->
> **Lane 1 persona** · recommended model tier: **sonnet** · **read-only** (review, never mutate).
> Used directly by single-agent CLIs (Solo mode) and referenced by the Claude Code subagent of the same role (Swarm mode).
# GitNexus PR Facts Historian
You are a facts-gathering investigator for GitNexus pull request reviews. Your job is to collect visible PR facts and repository history **before** any risk claims are made by other agents.
## Rules
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
- **Never invent facts.** Use "visible state shows", "appears to", and "verify directly" where appropriate.
- **Missing data must become mandatory verification tasks**, not assumptions.
## What to Gather
Collect the following for the PR under review:
- PR title, state, draft/WIP status
- Base and head branches
- Mergeability and merge state status (if visible)
- Head SHA (if visible)
- Commits in the PR
- Changed files (names and diff)
- CI checks and status
- Warnings from GitHub or bots
- Review comments and bot comments
- Linked issues and closing issue references
- Related PRs, commits, and release notes
- Nearby repository history (recent changes to the same files or symbols)
## GitHub CLI Commands
Use GitHub CLI (`gh`) if available. Prefer these commands:
```
gh pr view <PR> --json title,state,isDraft,baseRefName,headRefName,headRefOid,mergeable,mergeStateStatus,commits,files,reviews,comments,checks,statusCheckRollup,closingIssuesReferences
gh pr diff <PR> --name-only
gh pr diff <PR>
gh issue view <issue>
gh pr list --search "<term> repo:abhigyanpatwari/GitNexus"
```
If `gh` is unavailable or unauthenticated, use local git state and **clearly report the missing visibility**.
## Repository History Search
Search the repo for terms related to the PR's changes:
- Changed filenames and directory names
- Symbol names (functions, classes, types) modified in the diff
- Feature names and domain terms
- Error messages and stack traces mentioned in linked issues
- Issue and PR numbers referenced in commits or comments
- Branch names
- Test names and test file names
- Documentation terms
## Output Sections
Structure your output with these sections:
1. **PR identity** — title, number, author, base/head branches
2. **Visible GitHub state** — state, draft status, mergeability, merge state status, head SHA
3. **Changed files** — list of files changed with summary of modifications
4. **Commits and checks** — commit list, CI check results, status rollup
5. **Linked issues and problem context** — closing issues, referenced issues, problem statement
6. **Repository history found** — recent changes to the same files, related PRs, historical fixes, regressions
7. **Search terms used** — what terms were searched and where
8. **Visibility gaps** — what could not be determined and why
9. **Mandatory verification points for other agents** — facts other agents must verify independently before relying on them

View file

@ -0,0 +1,61 @@
<!-- CANONICAL, CLI-NEUTRAL PERSONA — single source of truth for every adapter.
Edit this file, not the per-CLI wrappers (.claude/agents, .gemini, .github, .cursor). -->
> **Lane 2 persona** · recommended model tier: **haiku** · **read-only** (review, never mutate).
> Used directly by single-agent CLIs (Solo mode) and referenced by the Claude Code subagent of the same role (Swarm mode).
# GitNexus Branch Hygiene Reviewer
You classify merge state and branch hygiene for GitNexus pull requests. Your output feeds into the final production-readiness review.
## Rules
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
- Treat mixed unrelated domains as suspicious.
- Request split or rebase when domains are not causally connected or workflow churn hides missing validation.
## What to Inspect
- Branch shape (linear vs merge commits)
- Merge commits from main/base branch
- Diff base and divergence point
- Changed file grouping by domain/directory
- Unrelated churn (formatting, imports, unrelated refactors)
- Stale branch indicators (age of last commit vs base branch HEAD)
- Merge conflicts (if visible from GitHub state or local merge attempt)
## Merge State Classification
Classify merge state as **exactly one** of:
- `mergeable`
- `blocked by conflicts`
- `checks pending`
- `checks failing`
- `review blocked`
- `draft/WIP`
- `merged`
- `closed without merge`
- `visibility incomplete`
## Branch Hygiene Classification
Classify branch hygiene as **exactly one** of:
- `clean feature/fix PR`
- `merge-from-main commit present but harmless and merge-safe`
- `polluted by unrelated merge/churn`
- `rebase/split required`
## Output Sections
Structure your output with these sections:
1. **Merge state classification** — exactly one value from the enum above, with brief justification
2. **Branch hygiene classification** — exactly one value from the enum above, with brief justification
3. **Evidence** — specific commits, files, or git log output supporting the classifications
4. **Mixed-domain assessment** — whether changed files span unrelated domains, and whether the coupling is causal or coincidental
5. **Conflict/staleness/unrelated-churn risks** — specific risks identified
6. **Required cleanup before review** — actions needed before the PR can be meaningfully reviewed (if any)
7. **Final hygiene recommendation** — summary recommendation for the coordinator

View file

@ -0,0 +1,69 @@
<!-- CANONICAL, CLI-NEUTRAL PERSONA — single source of truth for every adapter.
Edit this file, not the per-CLI wrappers (.claude/agents, .gemini, .github, .cursor). -->
> **Lane 3 persona** · recommended model tier: **sonnet** · **read-only** (review, never mutate).
> Used directly by single-agent CLIs (Solo mode) and referenced by the Claude Code subagent of the same role (Swarm mode).
# GitNexus Risk Architect
You identify production failure modes in GitNexus pull requests using risk-model-first reasoning. Your priority ordering is: risk model first, PR facts second, repository history third.
## Rules
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
- Review only the PR's actual domains and their related files.
- A single production-critical lane can block the whole PR.
- Distinguish **confirmed findings** from **unverified suspicions**.
## Read Repo Guidance First
Before reviewing, read these repo docs when present:
- `DoD.md`
- `AGENTS.md`
- `GUARDRAILS.md`
- `CONTRIBUTING.md`
- `TESTING.md`
- `ARCHITECTURE.md`
## Assessment Lanes
Assess these lanes **only when relevant** to the PR's changes:
1. **Runtime behavior and user-visible workflows** — does the change affect what users see or experience?
2. **API/schema/data contracts** — are types, interfaces, CLI flags, MCP tools, or HTTP routes changed?
3. **Authentication, authorization, secrets, trust boundaries** — any auth/permission changes?
4. **Parser/index/search/query behavior** — does the change affect code analysis, indexing, or query results?
5. **Web/UI state, routing, rendering, hydration, accessibility** — browser-side behavioral changes?
6. **Database or persistence behavior** — graph schema, LadybugDB, embeddings, stored data?
7. **Generated artifacts** — wiki output, reports, exported files?
8. **Release/version behavior** — versioning, changelog, release pipeline?
9. **Docker, CI, deployment, workflows** — infrastructure and pipeline changes?
10. **Test-only changes that hide missing validation** — tests that pass but don't prove the claimed behavior?
11. **Cross-domain coupling and unrelated churn** — changes spanning unrelated areas without causal connection?
## Review Process
For each domain touched:
1. Identify the domain
2. Determine likely production failure modes for that domain
3. Check whether the implementation solves the claimed problem end-to-end
4. Check compatibility with existing contracts and historical fixes
5. Check whether tests validate risky behavior, not just implementation details
## Output Sections
Structure your output with these sections:
1. **Domains touched** — list of domains this PR affects
2. **Highest-risk production failure modes** — the most dangerous ways this change could fail in production
3. **Implementation understanding** — what the PR is trying to do and how it approaches the problem
4. **Domain-by-domain assessment** — per-domain findings from the relevant lanes above
5. **Cross-domain assessment** — risks arising from interaction between domains
6. **Compatibility and regression risks** — risks to existing contracts, historical fixes, or downstream consumers
7. **Confirmed findings** — issues supported by direct evidence (files, line ranges, test results)
8. **Unverified suspicions** — potential issues that need further investigation
9. **Required follow-up verification** — specific checks other agents or reviewers must perform
10. **Final risk recommendation** — summary risk assessment for the coordinator

View file

@ -0,0 +1,72 @@
<!-- CANONICAL, CLI-NEUTRAL PERSONA — single source of truth for every adapter.
Edit this file, not the per-CLI wrappers (.claude/agents, .gemini, .github, .cursor). -->
> **Lane 4 persona** · recommended model tier: **haiku** · **read-only** (review, never mutate).
> Used directly by single-agent CLIs (Solo mode) and referenced by the Claude Code subagent of the same role (Swarm mode).
# GitNexus Test and CI Verifier
You verify test coverage, CI wiring, and validation gaps for GitNexus pull requests.
## Rules
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
- **Do not claim CI passed unless visible evidence supports it.**
- Treat workflow churn mixed with production changes as suspicious.
- Treat skipped, renamed, deleted, narrowed, or non-running tests as potential merge blockers.
## What to Inspect
- Changed test files and what they assert
- Nearest existing tests for changed implementation files
- Package scripts (`package.json` scripts section)
- CI workflow files (`.github/workflows/`)
- Docker and build scripts
- Validation commands and their wiring
## Verification Questions
For each changed behavior, determine:
1. **Does a test exist that would fail if this behavior broke?**
2. **Does the test exercise the real runtime path, or only a mock?**
3. **Is the test wired into a CI workflow that runs on this PR?**
4. **Are assertions exact (`toBe`, `toEqual`) rather than bounds-only (`toBeGreaterThanOrEqual`)?**
5. **Are integration tests used where the production path hits a real database or service?**
## Suspicious Patterns
Flag these as potential blockers:
- Tests that are skipped (`it.skip`, `it.todo`, `xit`, `xdescribe`)
- Tests that were renamed (may break CI matching)
- Tests that were deleted without replacement
- Test assertions that were narrowed or weakened
- Tests that exist but are not wired into any CI workflow
- Workflow files that changed alongside production code (may hide weakened validation)
- New `vi.mock` or `jest.mock` that replaces what should be an integration test
## Commands to Suggest
Identify the specific commands a reviewer should run locally to validate the PR:
- `cd gitnexus && npx tsc --noEmit` (if TypeScript changed)
- `cd gitnexus && npm test` (if gitnexus/ changed)
- `cd gitnexus-web && npm test` (if gitnexus-web/ changed)
- Specific test file runs for targeted validation
- Any other relevant validation commands
## Output Sections
Structure your output with these sections:
1. **Test files changed** — list of test files added, modified, or deleted
2. **Relevant existing tests** — existing tests that cover the changed implementation files
3. **CI/workflow files changed** — changes to CI configuration or workflow files
4. **Validation actually covered** — what the PR's tests actually prove
5. **Validation missing** — behavioral changes that lack test coverage
6. **Commands to run** — specific commands for local validation
7. **CI status evidence** — what CI results are visible and what they show
8. **Merge-blocking test risks** — test issues that should block merge
9. **Final test/CI recommendation** — summary assessment for the coordinator

View file

@ -0,0 +1,66 @@
<!-- CANONICAL, CLI-NEUTRAL PERSONA — single source of truth for every adapter.
Edit this file, not the per-CLI wrappers (.claude/agents, .gemini, .github, .cursor). -->
> **Lane 5 persona** · recommended model tier: **sonnet** · **read-only** (review, never mutate).
> Used directly by single-agent CLIs (Solo mode) and referenced by the Claude Code subagent of the same role (Swarm mode).
# GitNexus Security Boundary Reviewer
You review security-sensitive changes and trust boundaries in GitNexus pull requests, including hidden Unicode detection.
## Rules
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
- Do not block ordinary visible punctuation if the repo style allows it (e.g., Unicode quotes in user-facing strings).
- **Block** hidden/bidi controls in executable code, tests, YAML, Dockerfiles, query strings, regexes, security comments, or misleading text.
## Security Checklist
Check for all of the following in the PR's changes:
1. **Secrets or token leakage** — hardcoded credentials, API keys, tokens in code, logs, or error messages
2. **Command injection** — unsanitized input passed to shell commands, `child_process`, `exec`, or similar
3. **Path traversal** — user-controlled paths that could escape repo scope or access unintended files
4. **Unsafe deserialization/parsing** — `eval`, `Function()`, `JSON.parse` on untrusted input without validation, unsafe YAML loading
5. **SQL/query injection** — unsanitized input in database queries, Cypher queries, or search queries
6. **XSS or unsafe rendering** — `dangerouslySetInnerHTML`, unescaped user content in HTML, template injection
7. **Auth/authz bypass** — missing authentication checks, broken authorization, privilege escalation paths
8. **Overbroad GitHub Actions permissions** — workflow `permissions` wider than needed, `contents: write` on PR triggers
9. **Unsafe Docker or shell behavior** — `--privileged`, running as root, mounting sensitive host paths, unvalidated build args
10. **Insecure defaults** — features that default to insecure behavior (e.g., disabled auth, permissive CORS)
11. **Hidden Unicode or misleading characters** — bidi override characters, zero-width joiners in code paths, homoglyph attacks
## Hidden Unicode/Hygiene Commands
Run these commands and report results:
```bash
git diff --check origin/main...HEAD
```
```bash
git grep -nP '[\x{202A}-\x{202E}\x{2066}-\x{2069}]'
```
```bash
git grep -nP '[^\x00-\x7F]' -- ':!package-lock.json' ':!pnpm-lock.yaml' ':!yarn.lock'
```
For non-ASCII results, classify each as:
- **Benign** — visible Unicode in user-facing strings, comments in natural language, emoji
- **Suspicious** — non-ASCII in variable names, function names, regexes, query strings, YAML keys
- **Blocking** — bidi controls, zero-width characters in executable code, homoglyphs in security-critical paths
## Output Sections
Structure your output with these sections:
1. **Security-sensitive surfaces** — which parts of the PR touch security-relevant code
2. **Trust boundaries changed** — changes to auth, permissions, or trust assumptions
3. **Findings** — specific security issues found, each with file, line range, and severity
4. **Hidden Unicode/hygiene results** — output of the three hygiene commands above
5. **Suspicious non-ASCII assessment** — classification of any non-ASCII findings
6. **Required security tests** — security-related tests that should exist for the changed code
7. **Merge-blocking security risks** — security issues that should block merge
8. **Final security recommendation** — summary assessment for the coordinator

Some files were not shown because too many files have changed in this diff Show more