fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV (#1433)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Release Candidate / Check if release candidate should run (push) Waiting to run
Release Candidate / ci (push) Blocked by required conditions
Release Candidate / Publish release candidate to npm (push) Blocked by required conditions
Release Candidate / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run

* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV

tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:

  - captures.ts (C# scope extraction)
  - namespace-siblings.ts (extractFileStructure)
  - parse-worker.ts (worker thread parse path)
  - parsing-processor.ts (sequential parse fallback)

Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.

Additional C# scope fixes:
  - scope-tree.ts: Module scopes may share the same range as a top-level
    namespace_declaration (files with no leading `using` directives). The
    rangeStrictlyContains check rejects equal ranges. Added
    rangeNonStrictlyContains for Module parents.
  - scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
    same Module == Namespace range case caused orphaned scopes. Added
    moduleAwareContains helper.
  - scope-extractor-bridge.ts: empty captures from ERROR-root files still
    called extractScope -> "no Module scope found" warning. Added early
    return for empty/non-array captures.
  - namespace-siblings.ts: three sites pushed onto binding arrays frozen by
    finalize-algorithm. Fixed with spread-copy before mutation.

lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.

Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).

* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV

LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.

Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.

This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.

* fix: address codeql findings on PR #1433

The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.

Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).

* fix(windows): replace 32767-char truncation with chunked-input parsing

The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.

Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.

Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.

* fix(windows): cover all parse sites and correct vector-extension state

Address adversarial review on PR #1433:

1. Extend parseSourceSafe to all remaining parser.parse() call sites that
   handle full file content. The first commit only converted the four
   sites with active truncation hacks; cache-miss paths in
   call-processor (x2), heritage-processor (x2), import-processor, and
   the Go/Python/TypeScript captures + Go range-binding still called
   parser.parse() directly. On Windows those would still SIGSEGV for
   files > 32767 chars.

2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
   in lbug-adapter.ts. The flag means "successfully loaded" and is
   checked by an early-return at the top of loadVectorExtension; setting
   it on the skip path made the second call return true and let
   QUERY_VECTOR_INDEX run against a DB without the extension.

3. Drop the placeholder issues/... URL in the same comment.

4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
   (direct/callback boundary), the 32 767 Windows crash boundary,
   single-line > chunk size, CRLF near boundary, and large all-Chinese
   source. Confirms the callback path is correct for non-ASCII content,
   which is also exercised by the existing csharp-captures large-file
   test.

Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.

* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement

Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.

Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
  so group/ and embeddings/ can import without crossing into ingestion
  internals. core/tree-sitter/ already houses parser-loader.ts and is
  the natural shared facade. All 11 existing importers updated; no shim
  left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
  http-route, include, tree-sitter-scanner) and the embeddings
  ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
  module-load grammar smoke test parsing a 36-char literal. Trivially
  safe by inspection, intentionally direct, filtered out by the lint
  rule via the string-literal-arg skip.

Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
  parseSourceSafe. The spy is what catches a regression: parser.parse
  on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
  assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
  gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
  inside each mock factory so vitest's hoister does not race the
  static import binding.

Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
  gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
  ...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
  Skips JSON/URL/marked/Number/Math, string-literal first args
  (smoke tests), test files, and the helper itself. Auto-fix rewrites
  the call site only; the developer adds the import after tsc
  surfaces the missing identifier — same tradeoff as
  unused-imports/no-unused-imports.

Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md

* fix(test): use mkdtempSync in http-route-extractor regression test

Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
This commit is contained in:
Hector Prats 2026-05-10 17:00:36 +02:00 committed by GitHub
parent 2620b704e0
commit d69eadfb7f
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
29 changed files with 533 additions and 29 deletions

View file

@ -0,0 +1,95 @@
/**
* Custom ESLint rule: require `parseSourceSafe(parser, content, ...)` instead
* of direct `<parser>.parse(<content>, ...)` calls.
*
* Background: tree-sitter's Node.js native binding crashes with SIGSEGV on
* Windows when handed a JS string longer than 32 767 chars. The crash happens
* inside the binding's V8 string-to-buffer conversion and cannot be intercepted
* by JavaScript `try/catch`. `parseSourceSafe` (in
* `gitnexus/src/core/tree-sitter/safe-parse.ts`) routes large inputs through
* the chunked-callback overload of `parser.parse(input, ...)` which bypasses
* the broken conversion path. PR #1433 fixed every direct call site at the
* time; this rule prevents new direct calls from creeping in.
*
* The rule is auto-fixable for the call-site rewrite. It does NOT auto-add the
* import (computing the correct relative path per file is brittle); after the
* call rewrite runs, the consumer file's `tsc` will complain about an
* undefined identifier and the developer adds the import. This is the same
* tradeoff `unused-imports/no-unused-imports` makes in the opposite direction.
*
* False-positive suppression:
* - Skips calls whose receiver is a known non-tree-sitter library (`JSON`,
* `URL`, `marked`, `Number`).
* - Skips calls whose first argument is a string-literal (grammar-load smoke
* tests like `_testParser.parse('service X { rpc Y (R) returns (R); }')`).
* - Skips test files (`.test.ts`/`.test.tsx`/`.spec.ts`).
* - Skips the `safe-parse.ts` helper itself.
*/
const SKIPPED_RECEIVERS = new Set(['JSON', 'URL', 'marked', 'Number', 'Math']);
export default {
meta: {
type: 'problem',
docs: {
description:
'Require parseSourceSafe instead of direct tree-sitter `<parser>.parse(content, ...)` calls (Windows SIGSEGV protection)',
recommended: true,
},
fixable: 'code',
schema: [],
messages: {
useSafeParse:
'Direct `{{receiver}}.parse(...)` can SIGSEGV on Windows for inputs > 32 767 chars (uncatchable from JS). Use `parseSourceSafe({{receiver}}, ...)` from `core/tree-sitter/safe-parse.js`. Auto-fix rewrites the call; add the missing import yourself.',
},
},
create(context) {
const filename = context.filename ?? context.getFilename();
// Don't lint the helper itself or test files.
if (filename.includes('safe-parse')) return {};
if (/[.](?:test|spec)\.tsx?$/.test(filename)) return {};
const sourceCode = context.sourceCode ?? context.getSourceCode();
return {
CallExpression(node) {
const callee = node.callee;
if (callee.type !== 'MemberExpression') return;
if (callee.computed) return;
if (callee.property.type !== 'Identifier') return;
if (callee.property.name !== 'parse') return;
// Skip known non-tree-sitter receivers.
if (callee.object.type === 'Identifier' && SKIPPED_RECEIVERS.has(callee.object.name)) {
return;
}
// Smoke tests pass a string literal directly; those are trivially safe.
const firstArg = node.arguments[0];
if (!firstArg) return;
if (firstArg.type === 'Literal' && typeof firstArg.value === 'string') return;
if (firstArg.type === 'TemplateLiteral' && firstArg.expressions.length === 0) return;
const receiverText = sourceCode.getText(callee.object);
// Receiver-text-shape skip: anything matching well-known JS APIs that
// happen to have a `.parse(<expr>)` shape but aren't tree-sitter.
if (
/^(JSON|URL|marked|Number|Math|Date|globalThis\.JSON)\b/.test(receiverText) ||
/\bjson\.parse\b/i.test(receiverText)
) {
return;
}
context.report({
node,
messageId: 'useSafeParse',
data: { receiver: receiverText },
fix(fixer) {
const argsText = node.arguments.map((arg) => sourceCode.getText(arg)).join(', ');
return fixer.replaceText(node, `parseSourceSafe(${receiverText}, ${argsText})`);
},
});
},
};
},
};

View file

@ -3,6 +3,15 @@ import tsParser from '@typescript-eslint/parser';
import unusedImports from 'eslint-plugin-unused-imports';
import reactHooks from 'eslint-plugin-react-hooks';
import prettierConfig from 'eslint-config-prettier';
import requireSafeParse from './eslint-rules/require-safe-parse.mjs';
// Local plugin hosting custom rules that enforce GitNexus-specific invariants
// (currently: the Windows-SIGSEGV-safe parser entrypoint).
const gitnexusLocalPlugin = {
rules: {
'require-safe-parse': requireSafeParse,
},
};
// Selectors that protect MCP-reachable code from corrupting the JSON-RPC
// stdio frame stream. The MCP-reachable block below uses these directly;
@ -135,6 +144,23 @@ export default [
},
},
// Windows SIGSEGV protection: every tree-sitter parse in `core/` must route
// through parseSourceSafe. Direct `<parser>.parse(content, ...)` crashes on
// Windows for inputs > 32 767 chars (V8 string-conversion bug, uncatchable
// from JS). The rule auto-fixes the call site; the developer adds the
// missing import after the fix runs. Out of scope: tests (skipped by the
// rule), the helper itself (`safe-parse.ts`), and the `grpc-patterns/proto.ts`
// grammar-load smoke test (filtered by string-literal-arg skip in the rule).
{
files: ['gitnexus/src/core/**/*.ts'],
plugins: {
gitnexus: gitnexusLocalPlugin,
},
rules: {
'gitnexus/require-safe-parse': 'error',
},
},
// React-specific rules for gitnexus-web
{
files: ['gitnexus-web/src/**/*.{ts,tsx}'],

View file

@ -10,6 +10,7 @@ import {
isLanguageAvailable,
resolveLanguageKey,
} from '../tree-sitter/parser-loader.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
const parserCache = new Map<string, any>();
@ -29,7 +30,7 @@ export const ensureAndParse = async (content: string, filePath: string): Promise
parserCache.set(parserKey, parserInstance);
}
return parserInstance.parse(content);
return parseSourceSafe(parserInstance, content);
};
const FUNCTION_LIKE_TYPES = new Set([

View file

@ -5,6 +5,7 @@ import { createIgnoreFilter } from '../../../config/ignore-service.js';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import { logger } from '../../logger.js';
import {
GRPC_SCAN_GLOB,
@ -428,7 +429,7 @@ export class GrpcExtractor implements ContractExtractor {
let detections: GrpcDetection[] = [];
try {
parser.setLanguage(plugin.language);
const tree = parser.parse(content);
const tree = parseSourceSafe(parser, content);
detections = plugin.scan(tree);
} catch {
continue;

View file

@ -5,6 +5,7 @@ import { createIgnoreFilter } from '../../../config/ignore-service.js';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-patterns/index.js';
/**
@ -172,7 +173,7 @@ export class HttpRouteExtractor implements ContractExtractor {
}
try {
parser.setLanguage(plugin.language);
const tree = parser.parse(content);
const tree = parseSourceSafe(parser, content);
const detections = plugin.scan(tree);
cachedDetections.set(rel, detections);
return detections;

View file

@ -10,6 +10,7 @@ import { readSafe } from './fs-utils.js';
import { buildSuffixIndex, type SuffixIndex } from '../../ingestion/import-resolvers/utils.js';
import { createIgnoreFilter } from '../../../config/ignore-service.js';
import { getMaxFileSizeBytes } from '../../ingestion/utils/max-file-size.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import { logger } from '../../logger.js';
/**
@ -505,7 +506,7 @@ export class IncludeExtractor implements ContractExtractor {
let extractionSource: 'tree_sitter' | 'regex_fallback';
try {
parser.setLanguage(lang);
const tree = parser.parse(content);
const tree = parseSourceSafe(parser, content);
let matches: Parser.QueryMatch[];
try {
matches = query.matches(tree.rootNode);

View file

@ -3,6 +3,7 @@ import Parser from 'tree-sitter';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import {
getPluginForFile,
THRIFT_SCAN_GLOB,
@ -311,7 +312,7 @@ export class ThriftExtractor implements ContractExtractor {
let detections: ThriftDetection[] = [];
try {
parser.setLanguage(plugin.language);
const tree = parser.parse(content);
const tree = parseSourceSafe(parser, content);
detections = plugin.scan(tree);
} catch {
continue;

View file

@ -1,4 +1,5 @@
import Parser from 'tree-sitter';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
/**
* Shared, language-agnostic tree-sitter scanning utilities used by group
@ -155,7 +156,7 @@ export function scanFile<TMeta>(
let tree: Parser.Tree;
try {
parser.setLanguage(plugin.language);
tree = parser.parse(content);
tree = parseSourceSafe(parser, content);
} catch {
return [];
}

View file

@ -40,6 +40,7 @@ import { getLanguageFromFilename, SupportedLanguages } from 'gitnexus-shared';
import { isRegistryPrimary } from './registry-primary-flag.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
import {
FUNCTION_NODE_TYPES,
findEnclosingClassId,
@ -771,7 +772,7 @@ export const processCalls = async (
if (!tree) {
const parseContent = provider.preprocessSource?.(file.content, file.path) ?? file.content;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch (parseError) {
@ -3283,7 +3284,7 @@ export const extractFetchCallsFromFiles = async (
if (!tree) {
const parseContent = provider.preprocessSource?.(file.content, file.path) ?? file.content;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch {

View file

@ -22,6 +22,7 @@ import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, type NodeLabel, type SupportedLanguages } from 'gitnexus-shared';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
import { getProvider } from './languages/index.js';
import { getTreeSitterBufferSize } from './constants.js';
import type {
@ -224,7 +225,7 @@ export const processHeritage = async (
// re-parses see the same input as the cached AST.
const parseContent = provider.preprocessSource?.(file.content, file.path) ?? file.content;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch (parseError) {
@ -419,7 +420,7 @@ export async function extractExtractedHeritageFromFiles(
if (!tree) {
const parseContent = provider.preprocessSource?.(file.content, file.path) ?? file.content;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch {

View file

@ -8,6 +8,7 @@ import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename } from 'gitnexus-shared';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
import type { ExtractedImport } from './workers/parse-worker.js';
import { getTreeSitterBufferSize } from './constants.js';
import { loadImportConfigs } from './language-config.js';
@ -307,7 +308,7 @@ export const processImports = async (
if (!tree) {
const parseContent = provider.preprocessSource?.(file.content, file.path) ?? file.content;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch (parseError) {

View file

@ -24,6 +24,7 @@ import { synthesizeCsharpReceiverBinding } from './receiver-binding.js';
import { getCsharpParser, getCsharpScopeQuery } from './query.js';
import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
/** Declaration anchors that carry function-like arity metadata. */
const FUNCTION_DECL_TAGS = [
@ -86,7 +87,7 @@ export function emitCsharpScopeCaptures(
// the LanguageProvider contract layer; cast here at the use site.
let tree = cachedTree as ReturnType<ReturnType<typeof getCsharpParser>['parse']> | undefined;
if (tree === undefined) {
tree = getCsharpParser().parse(sourceText, undefined, {
tree = parseSourceSafe(getCsharpParser(), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
recordCacheMiss();

View file

@ -36,6 +36,7 @@ import type { BindingRef, ParsedFile, Scope, ScopeId, SymbolDefinition } from 'g
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { getCsharpParser } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
interface CsharpFileStructure {
/** Declared namespace names in file source order. Empty array means
@ -56,7 +57,7 @@ function extractFileStructure(content: string, cachedTree: unknown): CsharpFileS
type CsharpTree = ReturnType<ReturnType<typeof getCsharpParser>['parse']>;
const tree =
(cachedTree as CsharpTree | undefined) ??
getCsharpParser().parse(content, undefined, {
parseSourceSafe(getCsharpParser(), content, undefined, {
bufferSize: getTreeSitterBufferSize(content),
});
const namespaces: string[] = [];
@ -359,7 +360,7 @@ export function populateCsharpNamespaceSiblings(
const q = def.qualifiedName ?? '';
const key = q.includes('.') ? q.slice(q.lastIndexOf('.') + 1) : q;
if (key === '') continue;
const arr = defsByName.get(key) ?? [];
const arr = [...(defsByName.get(key) ?? [])];
arr.push(def);
defsByName.set(key, arr);
}

View file

@ -12,6 +12,7 @@ import { splitGoImportStatement } from './import-decomposer.js';
import { synthesizeGoReceiverBinding } from './receiver-binding.js';
import { synthesizeGoTypeBindings } from './type-binding.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
export function emitGoScopeCaptures(
sourceText: string,
@ -20,7 +21,7 @@ export function emitGoScopeCaptures(
): readonly CaptureMatch[] {
let tree = cachedTree as ReturnType<ReturnType<typeof getGoParser>['parse']> | undefined;
if (tree === undefined) {
tree = getGoParser().parse(sourceText, undefined, {
tree = parseSourceSafe(getGoParser(), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
recordGoCacheMiss();

View file

@ -2,6 +2,7 @@ import type { ParsedFile, Scope, TypeRef } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { getGoParser } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
export function populateGoRangeBindings(
parsedFiles: readonly ParsedFile[],
@ -20,7 +21,7 @@ export function populateGoRangeBindings(
const cachedTree = ctx.treeCache?.get(parsed.filePath);
const tree =
(cachedTree as ReturnType<typeof parser.parse> | undefined) ??
parser.parse(sourceText, undefined, {
parseSourceSafe(parser, sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
const moduleScope = parsed.scopes.find((s) => s.kind === 'Module');

View file

@ -24,6 +24,7 @@ import { synthesizeReceiverTypeBinding } from './receiver-binding.js';
import { computePythonArityMetadata } from './arity-metadata.js';
import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
import { pythonFunctionDefinitionLabel } from './simple-hooks.js';
export function emitPythonScopeCaptures(
@ -39,7 +40,7 @@ export function emitPythonScopeCaptures(
let tree = cachedTree as ReturnType<ReturnType<typeof getPythonParser>['parse']> | undefined;
if (tree === undefined) {
try {
tree = getPythonParser().parse(sourceText, undefined, {
tree = parseSourceSafe(getPythonParser(), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
} catch (err) {

View file

@ -38,6 +38,7 @@ import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { synthesizeTsReceiverBinding } from './receiver-binding.js';
import { computeTsArityMetadata } from './arity-metadata.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
/** tree-sitter-typescript node types for function-like scopes that may
* carry a synthesized `this` binding. Kept in sync with the
@ -134,7 +135,7 @@ export function emitTsScopeCaptures(
tree = undefined;
}
if (tree === undefined) {
tree = getTsParser(filePath).parse(sourceText, undefined, {
tree = parseSourceSafe(getTsParser(filePath), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
recordCacheMiss();

View file

@ -5,12 +5,11 @@ import { loadParser, loadLanguage, isLanguageAvailable } from '../tree-sitter/pa
import { getProvider } from './languages/index.js';
import { generateId } from '../../lib/utils.js';
import type { SymbolTableReader, SymbolTableWriter, ExtractedHeritage } from './model/index.js';
// SymbolTableReader is used for the FieldExtractorContext stub; the
// parsing functions themselves need Writer because they call .add().
import { ASTCache } from './ast-cache.js';
import { getLanguageFromFilename, SupportedLanguages } from 'gitnexus-shared';
import { extractVueScript, isVueSetupTopLevel } from './vue-sfc-extractor.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import {
getDefinitionNodeFromCaptures,
@ -384,7 +383,7 @@ const processParsingSequential = async (
let tree: Parser.Tree;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch (parseError) {

View file

@ -20,6 +20,7 @@ import {
getTreeSitterContentByteLength,
TREE_SITTER_MAX_BUFFER,
} from '../constants.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import type { SymbolTableReader } from '../model/symbol-table.js';
import type { ExtractedHeritage } from '../model/heritage-map.js';
@ -1416,7 +1417,7 @@ const processFileGroup = (
let tree;
try {
tree = parser.parse(parseContent, undefined, {
tree = parseSourceSafe(parser, parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent),
});
} catch (err) {

View file

@ -1248,6 +1248,13 @@ export const loadVectorExtension = async (
): Promise<boolean> => {
const useModuleState = targetConn === undefined;
if (useModuleState && vectorExtensionLoaded) return true;
// INSTALL VECTOR crashes with SIGSEGV on Windows: the KuzuDB native extension
// installer has an unhandled error path on Windows that raises a fatal signal
// that JS try/catch cannot intercept. Skip loading — vector/embedding search
// is unavailable but all graph index queries still work. Do NOT set
// vectorExtensionLoaded here: the flag means "successfully loaded", and a
// subsequent call would otherwise short-circuit to `return true` at the top.
if (process.platform === 'win32') return false;
if (!isVectorExtensionSupportedByPlatform()) return false;
const c: lbug.Connection | null = targetConn ?? conn;

View file

@ -420,7 +420,17 @@ async function doInitLbug(repoId: string, dbPath: string): Promise<void> {
// install; analyze owns extension installation. If LOAD fails, search
// features degrade gracefully and the user-facing query path proceeds.
if (!shared.ftsLoaded) {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
// Windows guard: LOAD EXTENSION fts crashes with SIGSEGV on Windows when
// the FTS extension binary is not installed locally (@ladybugdb/core native
// bug — the extension loader hits an unhandled error path that signals SIGSEGV
// rather than throwing a JS exception, so try/catch cannot protect here).
// Skip the load on Windows; bm25-index.js catches the resulting Kuzu catalog
// errors and returns empty BM25 results gracefully. Graph queries are unaffected.
if (process.platform === 'win32') {
shared.ftsLoaded = true;
} else {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
}
}
// Register pool entry only after all connections are pre-warmed and FTS is
@ -484,8 +494,13 @@ export async function initLbugWithDb(
// Load FTS extension if not already loaded on this Database.
// policy: 'load-only' — same contract as initLbug above; the read pool
// must not block on a network install during query execution.
// Windows guard: same SIGSEGV risk as doInitLbug above — skip on Windows.
if (!shared.ftsLoaded) {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
if (process.platform === 'win32') {
shared.ftsLoaded = true;
} else {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
}
}
pool.set(repoId, {

View file

@ -0,0 +1,40 @@
import type Parser from 'tree-sitter';
/**
* tree-sitter 0.21.x's Node native binding crashes (SIGSEGV) on Windows when
* `parser.parse(string, …)` is handed a JS string longer than 32 767 chars.
* The crash happens inside the binding's V8 string-to-buffer conversion and
* cannot be intercepted from JavaScript. The callback (`Parser.Input`) overload
* pulls source in fixed-size chunks via repeated callback invocations and
* bypasses that conversion path entirely.
*
* Chunk size is comfortably below the boundary; any value < 32 767 works.
*/
const SAFE_PARSE_CHUNK_CHARS = 16 * 1024;
/**
* Files at or below this length skip the callback machinery and use the
* direct string overload the bug only manifests above the int16 boundary,
* so small inputs save the cost of N callback invocations per parse.
*/
const DIRECT_PARSE_LIMIT_CHARS = 16 * 1024;
/**
* Parse `sourceText` safely on every platform. See {@link SAFE_PARSE_CHUNK_CHARS}
* for the underlying tree-sitter binding bug this works around.
*/
export function parseSourceSafe(
parser: Parser,
sourceText: string,
oldTree?: Parser.Tree,
options?: Parser.Options,
): Parser.Tree {
if (sourceText.length <= DIRECT_PARSE_LIMIT_CHARS) {
return parser.parse(sourceText, oldTree, options);
}
const input: Parser.Input = (index) => {
if (index >= sourceText.length) return null;
return sourceText.slice(index, index + SAFE_PARSE_CHUNK_CHARS);
};
return parser.parse(input, oldTree, options);
}

View file

@ -0,0 +1,53 @@
import { vi } from 'vitest';
import type * as SafeParseModule from '../../src/core/tree-sitter/safe-parse.js';
/**
* Build a vitest mock module for `gitnexus/src/core/tree-sitter/safe-parse.ts`
* that spies on `parseSourceSafe` while still delegating to the real
* implementation.
*
* Background: tests that feed >32 767-char inputs through extractors,
* chunkers, or any parse caller need to assert the call routed through
* `parseSourceSafe` rather than `parser.parse(string, ...)` directly. A
* direct call SIGSEGVs on Windows for inputs that size; on Linux/macOS it
* succeeds, so a "no throw" assertion alone silently passes with the
* bypass reintroduced. The spy assertion is what actually catches the
* regression.
*
* Why the test still has to call `vi.mock` with a literal path: vitest's
* hoister static-analyzes the first argument of `vi.mock`, and the path
* varies by directory depth across test files. Everything else the
* `vi.importActual` round-trip, the spy installation, and the merged
* module shape lives here.
*
* Why the test still has to dynamic-`import()` this helper inside the
* `vi.mock` factory: `vi.mock` is hoisted above static imports, so the
* factory closure cannot reference statically-imported helpers (they are
* uninitialized at hoist time). The factory body, however, is async and
* runs only when the mocked module is first consumed by which point
* the helper resolves cleanly via dynamic `import()`.
*
* Usage:
*
* const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
*
* vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
* const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
* return buildSafeParseMock(parseSourceSafeSpy);
* });
*
* it('routes large input through parseSourceSafe', async () => {
* parseSourceSafeSpy.mockClear();
* // ... call extractor with >40 000-char input ...
* expect(parseSourceSafeSpy).toHaveBeenCalled();
* });
*/
export async function buildSafeParseMock(
spy: ReturnType<typeof vi.fn>,
): Promise<typeof SafeParseModule> {
const actual = await vi.importActual<typeof SafeParseModule>(
'../../src/core/tree-sitter/safe-parse.js',
);
spy.mockImplementation(actual.parseSourceSafe);
return { ...actual, parseSourceSafe: spy };
}

View file

@ -1,10 +1,11 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
const { createParserForLanguage, getLanguageFromFilename } = vi.hoisted(() => ({
const { createParserForLanguage, getLanguageFromFilename, parseSourceSafeSpy } = vi.hoisted(() => ({
createParserForLanguage: vi.fn(),
getLanguageFromFilename: vi.fn((filePath: string) =>
filePath.endsWith('.py') ? 'python' : 'typescript',
),
parseSourceSafeSpy: vi.fn(),
}));
vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
@ -15,6 +16,11 @@ vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
),
}));
vi.mock('../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
vi.mock('gitnexus-shared', () => ({
getLanguageFromFilename,
}));
@ -72,4 +78,25 @@ describe('ensureAndParse', () => {
expect(tsParse).toHaveBeenCalledTimes(2);
expect(tsxParse).toHaveBeenCalledTimes(1);
});
// Windows SIGSEGV regression: ensureAndParse must route through parseSourceSafe
// so >32 767-char inputs do not crash the process. Direct parser.parse(content)
// on strings that size SIGSEGVs on Windows; the spy assertion is what catches
// a bypass since parser.parse(40 000 chars) succeeds on Linux/macOS.
it('routes >32 767-char input through parseSourceSafe', async () => {
parseSourceSafeSpy.mockClear();
const fakeParse = vi.fn().mockReturnValue({ rootNode: { type: 'module' } });
createParserForLanguage.mockResolvedValue({ parse: fakeParse });
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
const largeInput = 'const x = 1;\n'.repeat(4000); // ~52 000 chars
expect(largeInput.length).toBeGreaterThan(40_000);
const result = await ensureAndParse(largeInput, 'big.ts');
expect(parseSourceSafeSpy).toHaveBeenCalled();
expect(result).not.toBeNull();
});
});

View file

@ -1,8 +1,15 @@
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import * as fs from 'node:fs';
import fsp from 'node:fs/promises';
import * as path from 'node:path';
import * as os from 'node:os';
const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
import {
GrpcExtractor,
buildProtoMap,
@ -677,6 +684,33 @@ stub = leaked_pb2_grpc.LeakedServiceStub(channel)`,
).toBe(false);
});
});
describe('Windows SIGSEGV regression — large input must route through parseSourceSafe', () => {
it('routes >32 767-char source file through parseSourceSafe (not direct parser.parse)', async () => {
parseSourceSafeSpy.mockClear();
// Synthesize a >40 000-char source file in a language whose grpc plugin
// is always available (Go has no optional grammar — the Go plugin is
// unconditionally wired in grpc-patterns/index.ts). Direct
// parser.parse(content) on an input this size SIGSEGVs the process on
// Windows; parseSourceSafe routes through the chunked-callback path and
// works on every platform. The spy assertion is what catches the
// regression — a "no throw" assertion alone is satisfied by the bypass
// on Linux/macOS where parser.parse(40 000 chars) succeeds.
const padding = Array.from(
{ length: 600 },
(_, i) => `func helper${i}() string { return "padding-${i}-aaaaaaaaaaaaaaaaaaaaaa" }\n`,
).join('');
const largeGo = `package big\n\n${padding}\n`;
expect(largeGo.length).toBeGreaterThan(40_000);
writeFile('server/big.go', largeGo);
await extractor.extract(null, tmpDir, makeRepo(tmpDir));
expect(parseSourceSafeSpy).toHaveBeenCalled();
});
});
});
describe('buildProtoMap', () => {

View file

@ -1,7 +1,15 @@
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
import { HttpRouteExtractor } from '../../../src/core/group/extractors/http-route-extractor.js';
import type { RepoHandle } from '../../../src/core/group/types.js';
@ -815,4 +823,34 @@ export default r;
expect(contracts.some((c) => c.symbolRef?.filePath?.startsWith('mentor_env/'))).toBe(false);
});
});
describe('Windows SIGSEGV regression — large input must route through parseSourceSafe', () => {
it('routes >32 767-char source file through parseSourceSafe (not direct parser.parse)', async () => {
parseSourceSafeSpy.mockClear();
// >40 000-char Java controller file. Direct parser.parse(content) on
// an input this size SIGSEGVs the process on Windows. The spy assertion
// is what catches the regression — a "no throw" assertion alone is
// satisfied by the bypass on Linux/macOS where parser.parse(40 000 chars)
// succeeds.
const padding = Array.from(
{ length: 600 },
(_, i) => ` public String helper${i}() { return "padding-${i}-aaaaaaaaaaaaaaaaaaa"; }\n`,
).join('');
const largeJava = `package com.example;\n\n@RestController\npublic class BigController {\n${padding}}\n`;
expect(largeJava.length).toBeGreaterThan(40_000);
// Use mkdtempSync rather than a fixed subdir name: satisfies CodeQL's
// js/insecure-temporary-file rule by generating a unique random suffix
// instead of relying on the parent tmpDir's predictable Date.now() name.
const dir = fs.mkdtempSync(path.join(tmpDir, 'large-input-'));
fs.mkdirSync(path.join(dir, 'src/controller'), { recursive: true });
fs.writeFileSync(path.join(dir, 'src/controller/BigController.java'), largeJava);
const mockDbExecutor = async (_query: string) => [];
await extractor.extract(mockDbExecutor, dir, makeRepo(dir));
expect(parseSourceSafeSpy).toHaveBeenCalled();
});
});
});

View file

@ -1,7 +1,15 @@
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
import { IncludeExtractor } from '../../../src/core/group/extractors/include-extractor.js';
import type { RepoHandle } from '../../../src/core/group/types.js';
import { normalizeContractId } from '../../../src/core/group/matching.js';
@ -560,4 +568,35 @@ int auto_main() { return 0; }`,
}
});
});
describe('Windows SIGSEGV regression — large input must route through parseSourceSafe', () => {
it('routes >32 767-char header file through parseSourceSafe (not direct parser.parse)', async () => {
parseSourceSafeSpy.mockClear();
// Bump the file-size cap so the >40 000-char file isn't filtered before
// it ever reaches the parser. Direct parser.parse(content) on a string
// this size SIGSEGVs the process on Windows. The spy assertion catches
// the regression — a "no throw" assertion alone is satisfied by the
// bypass on Linux/macOS where parser.parse(40 000 chars) succeeds.
const previousLimit = process.env.GITNEXUS_MAX_FILE_SIZE;
process.env.GITNEXUS_MAX_FILE_SIZE = '512';
try {
const includes = Array.from(
{ length: 1500 },
(_, i) => `#include "lib/header_${i}.h"\n`,
).join('');
const largeHeader = `#pragma once\n${includes}\nstruct Big {};\n`;
expect(largeHeader.length).toBeGreaterThan(40_000);
writeFile('big/big.cpp', largeHeader);
await extractor.extract(null, tmpDir, makeRepo(tmpDir));
expect(parseSourceSafeSpy).toHaveBeenCalled();
} finally {
if (previousLimit === undefined) delete process.env.GITNEXUS_MAX_FILE_SIZE;
else process.env.GITNEXUS_MAX_FILE_SIZE = previousLimit;
}
});
});
});

View file

@ -1,8 +1,16 @@
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import * as fs from 'node:fs';
import fsp from 'node:fs/promises';
import * as path from 'node:path';
import * as os from 'node:os';
const { parseSourceSafeSpy } = vi.hoisted(() => ({ parseSourceSafeSpy: vi.fn() }));
vi.mock('../../../src/core/tree-sitter/safe-parse.js', async () => {
const { buildSafeParseMock } = await import('../../helpers/parse-source-safe-mock.js');
return buildSafeParseMock(parseSourceSafeSpy);
});
import {
ThriftExtractor,
buildThriftContext,
@ -580,6 +588,39 @@ class PaymentWorkflow {
expect(contracts).toEqual([]);
});
describe('Windows SIGSEGV regression — large input must route through parseSourceSafe', () => {
it('routes >32 767-char source file through parseSourceSafe (not direct parser.parse)', async () => {
parseSourceSafeSpy.mockClear();
// Need a base .thrift file so buildThriftContext finds at least one
// service to scan; without it the source-scan loop short-circuits.
writeFile(
'idl/order.thrift',
`namespace java billing.v1
service OrderService {
PlaceOrderResponse PlaceOrder(1: PlaceOrderRequest request)
}`,
);
// >40 000-char Java client file. Direct parser.parse(content) on a
// string this size SIGSEGVs the process on Windows. The spy assertion
// catches the regression — a "no throw" assertion alone is satisfied
// by the bypass on Linux/macOS where parser.parse(40 000 chars) succeeds.
const padding = Array.from(
{ length: 600 },
(_, i) => ` public String helper${i}() { return "padding-${i}-aaaaaaaaaaaaaaaaaaa"; }\n`,
).join('');
const largeJava = `package com.example;\n\nimport billing.v1.OrderService;\n\npublic class BigClient {\n private OrderService.Iface client;\n${padding}}\n`;
expect(largeJava.length).toBeGreaterThan(40_000);
writeFile('src/main/java/com/example/BigClient.java', largeJava);
await extractor.extract(null, tmpDir, makeRepo(tmpDir));
expect(parseSourceSafeSpy).toHaveBeenCalled();
});
});
});
describe('buildThriftContext', () => {

View file

@ -0,0 +1,74 @@
import { describe, it, expect } from 'vitest';
import Parser from 'tree-sitter';
import Python from 'tree-sitter-python';
import { parseSourceSafe } from '../../src/core/tree-sitter/safe-parse.js';
const makeParser = (): Parser => {
const p = new Parser();
p.setLanguage(Python);
return p;
};
const buildSource = (chars: number, lineLen = 80): string => {
const line = 'x = 1' + ' '.repeat(Math.max(0, lineLen - 6)) + '\n';
const lines = Math.ceil(chars / line.length);
return line.repeat(lines).slice(0, chars);
};
describe('parseSourceSafe', () => {
it('parses small ASCII sources via the direct path', () => {
const tree = parseSourceSafe(makeParser(), 'x = 1\n');
expect(tree.rootNode.type).toBe('module');
expect(tree.rootNode.hasError).toBe(false);
});
it('parses sources at the direct/callback boundary (16 KiB)', () => {
const src = buildSource(16 * 1024);
const tree = parseSourceSafe(makeParser(), src);
expect(tree.rootNode.hasError).toBe(false);
expect(tree.rootNode.endIndex).toBe(src.length);
});
it('parses sources just above the boundary via the callback path', () => {
const src = buildSource(16 * 1024 + 1);
const tree = parseSourceSafe(makeParser(), src);
expect(tree.rootNode.hasError).toBe(false);
expect(tree.rootNode.endIndex).toBe(src.length);
});
it('parses sources at and around the 32 767-char Windows crash boundary', () => {
for (const len of [32_766, 32_767, 32_768]) {
const src = buildSource(len);
const tree = parseSourceSafe(makeParser(), src);
expect(tree.rootNode.hasError, `len=${len}`).toBe(false);
expect(tree.rootNode.endIndex, `len=${len}`).toBe(src.length);
}
});
it('parses a single line longer than the chunk size (no newlines)', () => {
const src = '"' + 'a'.repeat(20_000) + '"\n';
const tree = parseSourceSafe(makeParser(), src);
expect(tree.rootNode.hasError).toBe(false);
expect(tree.rootNode.endIndex).toBe(src.length);
});
it('parses sources with CRLF line endings near a chunk boundary', () => {
const line = 'x = 1' + ' '.repeat(75) + '\r\n';
const src = line.repeat(Math.ceil(20_000 / line.length));
const tree = parseSourceSafe(makeParser(), src);
expect(tree.rootNode.hasError).toBe(false);
expect(tree.rootNode.endIndex).toBe(src.length);
});
it('parses a large all-non-ASCII source identically to the direct path', () => {
const small = '# ' + '漢'.repeat(50) + '\n';
const direct = makeParser().parse(small);
const safe = parseSourceSafe(makeParser(), small);
expect(safe.rootNode.toString()).toBe(direct.rootNode.toString());
const large = ('# ' + '漢'.repeat(8_000) + '\n').repeat(3);
const tree = parseSourceSafe(makeParser(), large);
expect(tree.rootNode.hasError).toBe(false);
expect(tree.rootNode.endIndex).toBe(large.length);
});
});