GitNexus/gitnexus/test/helpers/parse-source-safe-mock.ts
Hector Prats d69eadfb7f
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Release Candidate / Check if release candidate should run (push) Waiting to run
Release Candidate / ci (push) Blocked by required conditions
Release Candidate / Publish release candidate to npm (push) Blocked by required conditions
Release Candidate / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV (#1433)
* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV

tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:

  - captures.ts (C# scope extraction)
  - namespace-siblings.ts (extractFileStructure)
  - parse-worker.ts (worker thread parse path)
  - parsing-processor.ts (sequential parse fallback)

Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.

Additional C# scope fixes:
  - scope-tree.ts: Module scopes may share the same range as a top-level
    namespace_declaration (files with no leading `using` directives). The
    rangeStrictlyContains check rejects equal ranges. Added
    rangeNonStrictlyContains for Module parents.
  - scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
    same Module == Namespace range case caused orphaned scopes. Added
    moduleAwareContains helper.
  - scope-extractor-bridge.ts: empty captures from ERROR-root files still
    called extractScope -> "no Module scope found" warning. Added early
    return for empty/non-array captures.
  - namespace-siblings.ts: three sites pushed onto binding arrays frozen by
    finalize-algorithm. Fixed with spread-copy before mutation.

lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.

Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).

* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV

LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.

Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.

This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.

* fix: address codeql findings on PR #1433

The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.

Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).

* fix(windows): replace 32767-char truncation with chunked-input parsing

The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.

Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.

Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.

* fix(windows): cover all parse sites and correct vector-extension state

Address adversarial review on PR #1433:

1. Extend parseSourceSafe to all remaining parser.parse() call sites that
   handle full file content. The first commit only converted the four
   sites with active truncation hacks; cache-miss paths in
   call-processor (x2), heritage-processor (x2), import-processor, and
   the Go/Python/TypeScript captures + Go range-binding still called
   parser.parse() directly. On Windows those would still SIGSEGV for
   files > 32767 chars.

2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
   in lbug-adapter.ts. The flag means "successfully loaded" and is
   checked by an early-return at the top of loadVectorExtension; setting
   it on the skip path made the second call return true and let
   QUERY_VECTOR_INDEX run against a DB without the extension.

3. Drop the placeholder issues/... URL in the same comment.

4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
   (direct/callback boundary), the 32 767 Windows crash boundary,
   single-line > chunk size, CRLF near boundary, and large all-Chinese
   source. Confirms the callback path is correct for non-ASCII content,
   which is also exercised by the existing csharp-captures large-file
   test.

Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.

* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement

Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.

Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
  so group/ and embeddings/ can import without crossing into ingestion
  internals. core/tree-sitter/ already houses parser-loader.ts and is
  the natural shared facade. All 11 existing importers updated; no shim
  left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
  http-route, include, tree-sitter-scanner) and the embeddings
  ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
  module-load grammar smoke test parsing a 36-char literal. Trivially
  safe by inspection, intentionally direct, filtered out by the lint
  rule via the string-literal-arg skip.

Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
  parseSourceSafe. The spy is what catches a regression: parser.parse
  on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
  assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
  gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
  inside each mock factory so vitest's hoister does not race the
  static import binding.

Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
  gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
  ...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
  Skips JSON/URL/marked/Number/Math, string-literal first args
  (smoke tests), test files, and the helper itself. Auto-fix rewrites
  the call site only; the developer adds the import after tsc
  surfaces the missing identifier — same tradeoff as
  unused-imports/no-unused-imports.

Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md

* fix(test): use mkdtempSync in http-route-extractor regression test

Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-10 16:00:36 +01:00

53 lines
2.3 KiB
TypeScript

import { vi } from 'vitest';
import type * as SafeParseModule from '../../src/core/tree-sitter/safe-parse.js';
/**
* Build a vitest mock module for `gitnexus/src/core/tree-sitter/safe-parse.ts`
* that spies on `parseSourceSafe` while still delegating to the real
* implementation.
*
* Background: tests that feed >32 767-char inputs through extractors,
* chunkers, or any parse caller need to assert the call routed through
* `parseSourceSafe` rather than `parser.parse(string, ...)` directly. A
* direct call SIGSEGVs on Windows for inputs that size; on Linux/macOS it
* succeeds, so a "no throw" assertion alone silently passes with the
* bypass reintroduced. The spy assertion is what actually catches the
* regression.
*
* Why the test still has to call `vi.mock` with a literal path: vitest's
* hoister static-analyzes the first argument of `vi.mock`, and the path
* varies by directory depth across test files. Everything else — the
* `vi.importActual` round-trip, the spy installation, and the merged
* module shape — lives here.
*
* Why the test still has to dynamic-`import()` this helper inside the
* `vi.mock` factory: `vi.mock` is hoisted above static imports, so the
* factory closure cannot reference statically-imported helpers (they are
* uninitialized at hoist time). The factory body, however, is async and
* runs only when the mocked module is first consumed — by which point
* the helper resolves cleanly via dynamic `import()`.
*
* Usage:
*
* const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
*
* vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
* const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
* return buildSafeParseMock(parseSourceSafeSpy);
* });
*
* it('routes large input through parseSourceSafe', async () => {
* parseSourceSafeSpy.mockClear();
* // ... call extractor with >40 000-char input ...
* expect(parseSourceSafeSpy).toHaveBeenCalled();
* });
*/
export async function buildSafeParseMock(
spy: ReturnType<typeof vi.fn>,
): Promise<typeof SafeParseModule> {
const actual = await vi.importActual<typeof SafeParseModule>(
'../../src/core/tree-sitter/safe-parse.js',
);
spy.mockImplementation(actual.parseSourceSafe);
return { ...actual, parseSourceSafe: spy };
}