GitNexus/gitnexus/test/unit/embedding-chunking.test.ts
Karl Lehenbauer 639eb04b31
fix(swift): preprocess indented conditional directives so class bodies survive parsing (#2771)
* fix(swift): preprocess indented conditional directives so class bodies survive parsing

* fix(swift): make conditional-directive blanking comment-, string- and brace-aware (#2771)

Addresses the review findings on PR #2771. The transform fired
unconditionally, which turned valid Swift into parse errors while missing the
most common shape it was written for.

- The blank/keep decision now consults `blockCommentDepth`, so `  #endif */` —
  the result of commenting out a conditional block — keeps its comment
  terminator. Previously `hasError` went raw=false -> preprocessed=true and the
  rest of the file was swallowed.
- The decision keys on the scanner's brace depth instead of indentation. A
  column-0 `#if` inside a class body is blanked (6 of 7 body shapes previously
  still lost the enclosing declaration) and an indented file-scope directive is
  not — matching what the doc comment already claimed. Bare-CR line endings,
  NBSP/ideographic indentation and a leading BOM are recognized too.
- A group is blanked only when every branch is brace-balanced. An `#if`/`#else`
  that splits a declaration header leaves one unmatched `{` once both branches
  survive, which collapsed five top-level nodes into one and gave unrelated
  types fabricated `NetworkClient.` qualified names. Such a group now degrades
  to the pre-fix behavior.
- Multiline strings honour `\"""` escapes, and a plain `"""` closes even when a
  `#` follows it, so the scanner no longer wedges in string state and silently
  stops blanking for the rest of the file.
- The pound run is counted once per position and skipped. It was quadratic:
  10.6s for one 64k-`#` line, well inside the 512 KB walker limit.
- Extended regex literals (`#/.../#`) no longer open a phantom block comment.
- Directive-free files return early, matching `stripUeMacros`.

Worker parity: `emitSwiftScopeCaptures` and `emitCppScopeCaptures` re-apply
their provider's `preprocessSource` on the parse-cache-miss path — Dart already
did this — and the embedding parse in `ensureAndParse` applies the hook as
well. Before this the worker and the scope-capture/embedding halves analyzed
different programs, turning a consistent degradation into cold-run/warm-run
non-determinism. A new parity test pins the equivalence for every provider that
defines the hook.

SCHEMA_BUMP 37 -> 38: this changes parse semantics, the chunk key hashes raw
on-disk bytes, and `preprocessSource` runs after the key is computed — so a
same-package-version warm cache would replay pre-fix Swift results verbatim,
including across `--force`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(ingestion): apply preprocessSource once in the scope bridge (#2771)

Follow-up cleanup on the review fixes. The previous commit re-applied each
provider's `preprocessSource` inside `emitSwiftScopeCaptures` and
`emitCppScopeCaptures`, mirroring what Dart already did — three copies of the
same rule, and a contract that asked every future emitter to remember it.

`extractParsedFile` is the single funnel every `emitScopeCaptures` caller
passes through (parse worker, scope-resolution run, Vue script extraction), and
it already receives the provider. Applying the hook there on the cache-miss
path covers all three languages and every future one, names no language in
shared code, and drops Dart's unconditional transform on the cache-hit path.
Verified the three emitters use `sourceText` for nothing but the parse, so the
substitution is output-identical — which the parity test asserts directly.

Also from the cleanup pass:

- the parity test derives its language list from the provider registry, so a
  new provider adopting the hook fails until it adds a fixture
- `ensureAndParse` resolves the provider from the language it already computed,
  instead of a second extension table (`getProviderForFile`)
- the preprocessor returns `sourceText` unchanged when no group was blanked,
  which is the common case for files whose only directives are top-level
- `split(/(\r\n|\n|\r)/)` replaces the hand-rolled line splitter, and the
  per-group brace bookkeeping is two scalars instead of an array
- the hint regex is derived from the line regex so the two cannot drift
- unit assertions compare the WHOLE preprocessed file against the expected
  blanking, replacing per-line spot checks; the pipeline tests share one
  `runFixture` helper and `getNodesForFile` in the resolver test helpers
- `LanguageProvider.preprocessSource` documents the real call sites and says
  plainly that the set is not closed — `populateRangeBindings` still hands
  language helpers raw text

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-01 20:10:58 +00:00

308 lines
10 KiB
TypeScript

/**
* Integration test: Embedding chunking pipeline
*
* Tests the chunking + text generation pipeline together.
*/
import { describe, it, expect, vi } from 'vitest';
import { characterChunk } from '../../src/core/embeddings/character-chunk.js';
import { generateEmbeddingText } from '../../src/core/embeddings/text-generator.js';
import type { EmbeddableNode } from '../../src/core/embeddings/types.js';
const { createParserForLanguage } = vi.hoisted(() => ({
createParserForLanguage: vi.fn(),
}));
vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
createParserForLanguage,
isLanguageAvailable: vi.fn().mockReturnValue(true),
resolveLanguageKey: vi.fn((language: string) => language),
}));
const { getLanguageFromFilename } = vi.hoisted(() => ({
getLanguageFromFilename: vi.fn().mockReturnValue('typescript'),
}));
// Partial mock: `ast-utils` now resolves the LanguageProvider registry to apply
// `preprocessSource`, and that graph needs the real shared exports (#2771).
vi.mock('gitnexus-shared', async (importOriginal) => ({
...(await importOriginal<typeof import('gitnexus-shared')>()),
getLanguageFromFilename,
}));
import { chunkNode } from '../../src/core/embeddings/chunker.js';
const CLASS_PREV_TAIL_SAMPLE = 30;
const STRUCT_PREV_TAIL_SAMPLE = 20;
type FakeNode = {
type: string;
startIndex: number;
endIndex: number;
namedChildCount: number;
namedChild: (index: number) => FakeNode | null;
childForFieldName?: (name: string) => FakeNode | null;
};
const makeFakeNode = (
type: string,
startIndex: number,
endIndex: number,
children: FakeNode[] = [],
fields: Record<string, FakeNode> = {},
): FakeNode => ({
type,
startIndex,
endIndex,
namedChildCount: children.length,
namedChild: (index: number) => children[index] ?? null,
childForFieldName: (name: string) => fields[name] ?? null,
});
const makeDeclarationTree = (
nodeType: string,
bodyType: string,
content: string,
members: Array<{ text: string; type: string }>,
) => {
let searchFrom = 0;
const memberNodes = members.map((member) => {
const startIndex = content.indexOf(member.text, searchFrom);
if (startIndex < 0) {
throw new Error(`Unable to locate declaration member text: ${member.text}`);
}
searchFrom = startIndex + member.text.length;
return makeFakeNode(member.type, startIndex, startIndex + member.text.length);
});
const bodyStart = content.indexOf('{');
const bodyEnd = content.lastIndexOf('}') + 1;
const bodyNode = makeFakeNode(bodyType, bodyStart, bodyEnd, memberNodes);
const declNode = makeFakeNode(nodeType, 0, bodyEnd, [bodyNode], { body: bodyNode });
return {
rootNode: makeFakeNode('program', 0, content.length, [declNode]),
};
};
describe('embedding-chunking integration', () => {
const makeNode = (overrides: Partial<EmbeddableNode>): EmbeddableNode => ({
id: 'Function:src/test.ts:test',
name: 'test',
label: 'Function',
filePath: 'src/test.ts',
content: '',
startLine: 1,
endLine: 10,
...overrides,
});
it('short function produces single chunk with metadata', () => {
const node = makeNode({
content: 'function hello() { return "world"; }',
isExported: true,
repoName: 'my-project',
serverName: 'my-service',
});
const chunks = characterChunk(node.content, 1, 3, 1200, 120);
expect(chunks).toHaveLength(1);
const text = generateEmbeddingText(node, chunks[0].text);
expect(text).toContain('Function: test');
expect(text).toContain('function hello()');
// #2333: verbose metadata is no longer part of embedding text.
expect(text).not.toContain('Repo: my-project');
expect(text).not.toContain('Server: my-service');
expect(text).not.toContain('Export: true');
});
it('long function produces multiple chunks', () => {
const longContent = Array.from({ length: 100 }, (_, i) => ` const line${i} = ${i};`).join(
'\n',
);
const node = makeNode({
content: `function longFn() {\n${longContent}\n}`,
startLine: 1,
endLine: 102,
});
const chunks = characterChunk(node.content, 1, 102, 1200, 120);
expect(chunks.length).toBeGreaterThan(1);
expect(chunks[0].chunkIndex).toBe(0);
});
it('short labels (TypeAlias) skip chunking and embed directly', () => {
const node = makeNode({
label: 'TypeAlias',
name: 'Result',
content: 'type Result<T> = Success<T> | Error;',
});
const chunks = characterChunk(node.content, 1, 1, 1200, 120);
expect(chunks).toHaveLength(1);
const text = generateEmbeddingText(node, chunks[0].text);
expect(text).toContain('TypeAlias: Result');
expect(text).toContain('type Result<T> = Success<T> | Error;');
});
it('long enum uses character fallback', () => {
const enumContent = Array.from(
{ length: 200 },
(_, i) => ` Value${i} = "${'x'.repeat(20)}${i}",`,
).join('\n');
const node = makeNode({
label: 'Enum',
name: 'LargeEnum',
content: `enum LargeEnum {\n${enumContent}\n}`,
startLine: 1,
endLine: 202,
});
const chunks = characterChunk(node.content, 1, 202, 1200, 120);
expect(chunks.length).toBeGreaterThan(1);
});
it('long class produces member-aware chunks with structural metadata', async () => {
const node = makeNode({
label: 'Class',
name: 'Parser',
methodNames: ['parseJSON', 'validate'],
fieldNames: ['options', 'cache'],
content: `class Parser {
options: ParserOptions;
cache: Map<string, any>;
parseJSON(text: string) { return JSON.parse(text); }
validate() { return true; }
}`,
startLine: 20,
endLine: 25,
});
createParserForLanguage.mockResolvedValue({
parse: vi.fn().mockReturnValue(
makeDeclarationTree('class_declaration', 'class_body', node.content, [
{ text: 'options: ParserOptions;', type: 'field_definition' },
{ text: 'cache: Map<string, any>;', type: 'field_definition' },
{
text: 'parseJSON(text: string) { return JSON.parse(text); }',
type: 'method_definition',
},
{ text: 'validate() { return true; }', type: 'method_definition' },
]),
),
});
const chunks = await chunkNode(node.label, node.content, node.filePath, 20, 25, 90, 0);
expect(chunks).toHaveLength(2);
const secondText = generateEmbeddingText(
node,
chunks[1].text,
{},
chunks[1].chunkIndex,
chunks[0].text.slice(-CLASS_PREV_TAIL_SAMPLE),
);
expect(secondText).toContain('Class: Parser');
expect(secondText).toContain('Container: class Parser {');
expect(secondText).toContain('[preceding context]: ...');
expect(secondText).not.toContain('Methods: parseJSON, validate');
expect(secondText).not.toContain('Properties: options, cache');
expect(secondText).toContain('parseJSON(text: string)');
});
it('interface chunks retain structural metadata and signatures', async () => {
const node = makeNode({
label: 'Interface',
name: 'Handler',
methodNames: ['handle', 'validate'],
fieldNames: ['name'],
content: `interface Handler {
handle(event: Event): void;
validate(input: string): boolean;
readonly name: string;
}`,
startLine: 30,
endLine: 34,
});
createParserForLanguage.mockResolvedValue({
parse: vi.fn().mockReturnValue(
makeDeclarationTree('interface_declaration', 'object_type', node.content, [
{ text: 'handle(event: Event): void;', type: 'method_definition' },
{ text: 'validate(input: string): boolean;', type: 'method_definition' },
{ text: 'readonly name: string;', type: 'property_signature' },
]),
),
});
const chunks = await chunkNode(node.label, node.content, node.filePath, 30, 34, 500, 0);
expect(chunks).toHaveLength(1);
const text = generateEmbeddingText(node, chunks[0].text);
expect(text).toContain('Interface: Handler');
expect(text).toContain('Methods: handle, validate');
expect(text).toContain('Container: interface Handler {');
expect(text).toContain('readonly name: string;');
});
it('struct chunks retain structural container context', async () => {
getLanguageFromFilename.mockReturnValue('rust');
const node = makeNode({
label: 'Struct',
name: 'User',
fieldNames: ['name', 'email', 'age', 'address'],
content: `struct User {
name: String,
email: String,
age: u32,
address: String,
}`,
startLine: 40,
endLine: 45,
filePath: 'src/user.rs',
});
createParserForLanguage.mockResolvedValue({
parse: vi.fn().mockReturnValue(
makeDeclarationTree('struct_item', 'declaration_list', node.content, [
{ text: 'name: String,', type: 'field_definition' },
{ text: 'email: String,', type: 'field_definition' },
{ text: 'age: u32,', type: 'field_definition' },
{ text: 'address: String,', type: 'field_definition' },
]),
),
});
const chunks = await chunkNode(node.label, node.content, node.filePath, 40, 45, 45, 0);
expect(chunks).toHaveLength(2);
const secondText = generateEmbeddingText(
node,
chunks[1].text,
{},
chunks[1].chunkIndex,
chunks[0].text.slice(-STRUCT_PREV_TAIL_SAMPLE),
);
expect(secondText).toContain('Struct: User');
expect(secondText).toContain('Container: struct User {');
expect(secondText).not.toContain('Properties: name, email, age, address');
expect(secondText).toContain('age: u32,');
});
it('header is present in every chunk', () => {
const longContent = 'x'.repeat(3000);
const node = makeNode({
content: longContent,
repoName: 'test-repo',
});
const chunks = characterChunk(node.content, 1, 100, 1200, 120);
expect(chunks.length).toBeGreaterThan(1);
for (const chunk of chunks) {
const text = generateEmbeddingText(node, chunk.text);
// The compact header (name, + description when present) repeats on every
// chunk so each chunk keeps its identity; #2333 dropped the metadata lines.
expect(text).toContain('Function: test');
expect(text).not.toContain('Repo: test-repo');
expect(text).not.toContain('Path: src/test.ts');
}
});
});