mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-08-28 05:25:25 +00:00
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Release Candidate / Check if release candidate should run (push) Waiting to run
Release Candidate / ci (push) Blocked by required conditions
Release Candidate / Publish release candidate to npm (push) Blocked by required conditions
Release Candidate / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV
tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:
- captures.ts (C# scope extraction)
- namespace-siblings.ts (extractFileStructure)
- parse-worker.ts (worker thread parse path)
- parsing-processor.ts (sequential parse fallback)
Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.
Additional C# scope fixes:
- scope-tree.ts: Module scopes may share the same range as a top-level
namespace_declaration (files with no leading `using` directives). The
rangeStrictlyContains check rejects equal ranges. Added
rangeNonStrictlyContains for Module parents.
- scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
same Module == Namespace range case caused orphaned scopes. Added
moduleAwareContains helper.
- scope-extractor-bridge.ts: empty captures from ERROR-root files still
called extractScope -> "no Module scope found" warning. Added early
return for empty/non-array captures.
- namespace-siblings.ts: three sites pushed onto binding arrays frozen by
finalize-algorithm. Fixed with spread-copy before mutation.
lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.
Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).
* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV
LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.
Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.
This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.
* fix: address codeql findings on PR #1433
The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.
Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).
* fix(windows): replace 32767-char truncation with chunked-input parsing
The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.
Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.
Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.
* fix(windows): cover all parse sites and correct vector-extension state
Address adversarial review on PR #1433:
1. Extend parseSourceSafe to all remaining parser.parse() call sites that
handle full file content. The first commit only converted the four
sites with active truncation hacks; cache-miss paths in
call-processor (x2), heritage-processor (x2), import-processor, and
the Go/Python/TypeScript captures + Go range-binding still called
parser.parse() directly. On Windows those would still SIGSEGV for
files > 32767 chars.
2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
in lbug-adapter.ts. The flag means "successfully loaded" and is
checked by an early-return at the top of loadVectorExtension; setting
it on the skip path made the second call return true and let
QUERY_VECTOR_INDEX run against a DB without the extension.
3. Drop the placeholder issues/... URL in the same comment.
4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
(direct/callback boundary), the 32 767 Windows crash boundary,
single-line > chunk size, CRLF near boundary, and large all-Chinese
source. Confirms the callback path is correct for non-ASCII content,
which is also exercised by the existing csharp-captures large-file
test.
Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.
* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement
Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.
Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
so group/ and embeddings/ can import without crossing into ingestion
internals. core/tree-sitter/ already houses parser-loader.ts and is
the natural shared facade. All 11 existing importers updated; no shim
left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
http-route, include, tree-sitter-scanner) and the embeddings
ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
module-load grammar smoke test parsing a 36-char literal. Trivially
safe by inspection, intentionally direct, filtered out by the lint
rule via the string-literal-arg skip.
Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
parseSourceSafe. The spy is what catches a regression: parser.parse
on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
inside each mock factory so vitest's hoister does not race the
static import binding.
Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
Skips JSON/URL/marked/Number/Math, string-literal first args
(smoke tests), test files, and the helper itself. Auto-fix rewrites
the call site only; the developer adds the import after tsc
surfaces the missing identifier — same tradeoff as
unused-imports/no-unused-imports.
Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md
* fix(test): use mkdtempSync in http-route-extractor regression test
Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
102 lines
4.1 KiB
TypeScript
102 lines
4.1 KiB
TypeScript
import { beforeEach, describe, expect, it, vi } from 'vitest';
|
|
|
|
const { createParserForLanguage, getLanguageFromFilename, parseSourceSafeSpy } = vi.hoisted(() => ({
|
|
createParserForLanguage: vi.fn(),
|
|
getLanguageFromFilename: vi.fn((filePath: string) =>
|
|
filePath.endsWith('.py') ? 'python' : 'typescript',
|
|
),
|
|
parseSourceSafeSpy: vi.fn(),
|
|
}));
|
|
|
|
vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
|
|
createParserForLanguage,
|
|
isLanguageAvailable: vi.fn().mockReturnValue(true),
|
|
resolveLanguageKey: vi.fn((language: string, filePath?: string) =>
|
|
language === 'typescript' && filePath?.endsWith('.tsx') ? 'typescript:tsx' : language,
|
|
),
|
|
}));
|
|
|
|
vi.mock('../../src/core/tree-sitter/safe-parse.js', async () => {
|
|
const { buildSafeParseMock } = await import('../helpers/parse-source-safe-mock.js');
|
|
return buildSafeParseMock(parseSourceSafeSpy);
|
|
});
|
|
|
|
vi.mock('gitnexus-shared', () => ({
|
|
getLanguageFromFilename,
|
|
}));
|
|
|
|
describe('ensureAndParse', () => {
|
|
beforeEach(() => {
|
|
vi.resetModules();
|
|
createParserForLanguage.mockReset();
|
|
getLanguageFromFilename.mockClear();
|
|
});
|
|
|
|
it('reuses the parser for the same grammar key across interleaved languages', async () => {
|
|
const tsParse = vi
|
|
.fn()
|
|
.mockReturnValueOnce({ lang: 'ts', content: 'first' })
|
|
.mockReturnValueOnce({ lang: 'ts', content: 'second' });
|
|
const pyParse = vi.fn().mockReturnValue({ lang: 'py', content: 'middle' });
|
|
|
|
createParserForLanguage.mockImplementation(async (language: string, filePath?: string) => {
|
|
if (language === 'typescript') return { parse: tsParse, key: filePath };
|
|
if (language === 'python') return { parse: pyParse, key: filePath };
|
|
throw new Error(`unexpected language ${language}`);
|
|
});
|
|
|
|
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
|
|
|
|
const tsFirst = await ensureAndParse('const one = 1;', 'first.ts');
|
|
const pyMiddle = await ensureAndParse('value = 1', 'middle.py');
|
|
const tsSecond = await ensureAndParse('const two = 2;', 'second.ts');
|
|
|
|
expect(tsFirst).toEqual({ lang: 'ts', content: 'first' });
|
|
expect(pyMiddle).toEqual({ lang: 'py', content: 'middle' });
|
|
expect(tsSecond).toEqual({ lang: 'ts', content: 'second' });
|
|
expect(createParserForLanguage).toHaveBeenCalledTimes(2);
|
|
expect(tsParse).toHaveBeenCalledTimes(2);
|
|
expect(pyParse).toHaveBeenCalledTimes(1);
|
|
});
|
|
|
|
it('uses separate parser instances for .ts and .tsx', async () => {
|
|
const tsParse = vi.fn().mockReturnValue({ lang: 'ts' });
|
|
const tsxParse = vi.fn().mockReturnValue({ lang: 'tsx' });
|
|
|
|
createParserForLanguage.mockImplementation(async (_language: string, filePath?: string) => {
|
|
if (filePath?.endsWith('.tsx')) return { parse: tsxParse };
|
|
return { parse: tsParse };
|
|
});
|
|
|
|
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
|
|
|
|
await ensureAndParse('const value = 1;', 'plain.ts');
|
|
await ensureAndParse('export const View = <div />;', 'view.tsx');
|
|
await ensureAndParse('const other = 2;', 'other.ts');
|
|
|
|
expect(createParserForLanguage).toHaveBeenCalledTimes(2);
|
|
expect(tsParse).toHaveBeenCalledTimes(2);
|
|
expect(tsxParse).toHaveBeenCalledTimes(1);
|
|
});
|
|
|
|
// Windows SIGSEGV regression: ensureAndParse must route through parseSourceSafe
|
|
// so >32 767-char inputs do not crash the process. Direct parser.parse(content)
|
|
// on strings that size SIGSEGVs on Windows; the spy assertion is what catches
|
|
// a bypass since parser.parse(40 000 chars) succeeds on Linux/macOS.
|
|
it('routes >32 767-char input through parseSourceSafe', async () => {
|
|
parseSourceSafeSpy.mockClear();
|
|
|
|
const fakeParse = vi.fn().mockReturnValue({ rootNode: { type: 'module' } });
|
|
createParserForLanguage.mockResolvedValue({ parse: fakeParse });
|
|
|
|
const { ensureAndParse } = await import('../../src/core/embeddings/ast-utils.js');
|
|
|
|
const largeInput = 'const x = 1;\n'.repeat(4000); // ~52 000 chars
|
|
expect(largeInput.length).toBeGreaterThan(40_000);
|
|
|
|
const result = await ensureAndParse(largeInput, 'big.ts');
|
|
|
|
expect(parseSourceSafeSpy).toHaveBeenCalled();
|
|
expect(result).not.toBeNull();
|
|
});
|
|
});
|