GitNexus/gitnexus/test/unit/ast-utils.test.ts
Hector Prats d69eadfb7f
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Release Candidate / Check if release candidate should run (push) Waiting to run
Release Candidate / ci (push) Blocked by required conditions
Release Candidate / Publish release candidate to npm (push) Blocked by required conditions
Release Candidate / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV (#1433)
* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV

tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:

  - captures.ts (C# scope extraction)
  - namespace-siblings.ts (extractFileStructure)
  - parse-worker.ts (worker thread parse path)
  - parsing-processor.ts (sequential parse fallback)

Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.

Additional C# scope fixes:
  - scope-tree.ts: Module scopes may share the same range as a top-level
    namespace_declaration (files with no leading `using` directives). The
    rangeStrictlyContains check rejects equal ranges. Added
    rangeNonStrictlyContains for Module parents.
  - scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
    same Module == Namespace range case caused orphaned scopes. Added
    moduleAwareContains helper.
  - scope-extractor-bridge.ts: empty captures from ERROR-root files still
    called extractScope -> "no Module scope found" warning. Added early
    return for empty/non-array captures.
  - namespace-siblings.ts: three sites pushed onto binding arrays frozen by
    finalize-algorithm. Fixed with spread-copy before mutation.

lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.

Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).

* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV

LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.

Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.

This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.

* fix: address codeql findings on PR #1433

The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.

Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).

* fix(windows): replace 32767-char truncation with chunked-input parsing

The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.

Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.

Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.

* fix(windows): cover all parse sites and correct vector-extension state

Address adversarial review on PR #1433:

1. Extend parseSourceSafe to all remaining parser.parse() call sites that
   handle full file content. The first commit only converted the four
   sites with active truncation hacks; cache-miss paths in
   call-processor (x2), heritage-processor (x2), import-processor, and
   the Go/Python/TypeScript captures + Go range-binding still called
   parser.parse() directly. On Windows those would still SIGSEGV for
   files > 32767 chars.

2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
   in lbug-adapter.ts. The flag means "successfully loaded" and is
   checked by an early-return at the top of loadVectorExtension; setting
   it on the skip path made the second call return true and let
   QUERY_VECTOR_INDEX run against a DB without the extension.

3. Drop the placeholder issues/... URL in the same comment.

4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
   (direct/callback boundary), the 32 767 Windows crash boundary,
   single-line > chunk size, CRLF near boundary, and large all-Chinese
   source. Confirms the callback path is correct for non-ASCII content,
   which is also exercised by the existing csharp-captures large-file
   test.

Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.

* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement

Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.

Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
  so group/ and embeddings/ can import without crossing into ingestion
  internals. core/tree-sitter/ already houses parser-loader.ts and is
  the natural shared facade. All 11 existing importers updated; no shim
  left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
  http-route, include, tree-sitter-scanner) and the embeddings
  ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
  module-load grammar smoke test parsing a 36-char literal. Trivially
  safe by inspection, intentionally direct, filtered out by the lint
  rule via the string-literal-arg skip.

Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
  parseSourceSafe. The spy is what catches a regression: parser.parse
  on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
  assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
  gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
  inside each mock factory so vitest's hoister does not race the
  static import binding.

Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
  gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
  ...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
  Skips JSON/URL/marked/Number/Math, string-literal first args
  (smoke tests), test files, and the helper itself. Auto-fix rewrites
  the call site only; the developer adds the import after tsc
  surfaces the missing identifier — same tradeoff as
  unused-imports/no-unused-imports.

Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md

* fix(test): use mkdtempSync in http-route-extractor regression test

Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-10 16:00:36 +01:00

102 lines
4.1 KiB
TypeScript

import { beforeEach, describe, expect, it, vi } from 'vitest';
const { createParserForLanguage, getLanguageFromFilename, parseSourceSafeSpy } = vi.hoisted(() => ({
createParserForLanguage: vi.fn(),
getLanguageFromFilename: vi.fn((filePath: string) =>
filePath.endsWith('.py') ? 'python' : 'typescript',
),
parseSourceSafeSpy: vi.fn(),
}));
vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
createParserForLanguage,
isLanguageAvailable: vi.fn().mockReturnValue(true),
resolveLanguageKey: vi.fn((language: string, filePath?: string) =>
language === 'typescript' && filePath?.endsWith('.tsx') ? 'typescript:tsx' : language,
),
}));
vi.mock('../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
vi.mock('gitnexus-shared', () => ({
getLanguageFromFilename,
}));
describe('ensureAndParse', () => {
beforeEach(() => {
vi.resetModules();
createParserForLanguage.mockReset();
getLanguageFromFilename.mockClear();
});
it('reuses the parser for the same grammar key across interleaved languages', async () => {
const tsParse = vi
.fn()
.mockReturnValueOnce({ lang: 'ts', content: 'first' })
.mockReturnValueOnce({ lang: 'ts', content: 'second' });
const pyParse = vi.fn().mockReturnValue({ lang: 'py', content: 'middle' });
createParserForLanguage.mockImplementation(async (language: string, filePath?: string) => {
if (language === 'typescript') return { parse: tsParse, key: filePath };
if (language === 'python') return { parse: pyParse, key: filePath };
throw new Error(`unexpected language ${language}`);
});
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
const tsFirst = await ensureAndParse('const one = 1;', 'first.ts');
const pyMiddle = await ensureAndParse('value = 1', 'middle.py');
const tsSecond = await ensureAndParse('const two = 2;', 'second.ts');
expect(tsFirst).toEqual({ lang: 'ts', content: 'first' });
expect(pyMiddle).toEqual({ lang: 'py', content: 'middle' });
expect(tsSecond).toEqual({ lang: 'ts', content: 'second' });
expect(createParserForLanguage).toHaveBeenCalledTimes(2);
expect(tsParse).toHaveBeenCalledTimes(2);
expect(pyParse).toHaveBeenCalledTimes(1);
});
it('uses separate parser instances for .ts and .tsx', async () => {
const tsParse = vi.fn().mockReturnValue({ lang: 'ts' });
const tsxParse = vi.fn().mockReturnValue({ lang: 'tsx' });
createParserForLanguage.mockImplementation(async (_language: string, filePath?: string) => {
if (filePath?.endsWith('.tsx')) return { parse: tsxParse };
return { parse: tsParse };
});
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
await ensureAndParse('const value = 1;', 'plain.ts');
await ensureAndParse('export const View = <div />;', 'view.tsx');
await ensureAndParse('const other = 2;', 'other.ts');
expect(createParserForLanguage).toHaveBeenCalledTimes(2);
expect(tsParse).toHaveBeenCalledTimes(2);
expect(tsxParse).toHaveBeenCalledTimes(1);
});
// Windows SIGSEGV regression: ensureAndParse must route through parseSourceSafe
// so >32 767-char inputs do not crash the process. Direct parser.parse(content)
// on strings that size SIGSEGVs on Windows; the spy assertion is what catches
// a bypass since parser.parse(40 000 chars) succeeds on Linux/macOS.
it('routes >32 767-char input through parseSourceSafe', async () => {
parseSourceSafeSpy.mockClear();
const fakeParse = vi.fn().mockReturnValue({ rootNode: { type: 'module' } });
createParserForLanguage.mockResolvedValue({ parse: fakeParse });
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
const largeInput = 'const x = 1;\n'.repeat(4000); // ~52 000 chars
expect(largeInput.length).toBeGreaterThan(40_000);
const result = await ensureAndParse(largeInput, 'big.ts');
expect(parseSourceSafeSpy).toHaveBeenCalled();
expect(result).not.toBeNull();
});
});