mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-08-28 05:25:25 +00:00
* fix(swift): preprocess indented conditional directives so class bodies survive parsing * fix(swift): make conditional-directive blanking comment-, string- and brace-aware (#2771) Addresses the review findings on PR #2771. The transform fired unconditionally, which turned valid Swift into parse errors while missing the most common shape it was written for. - The blank/keep decision now consults `blockCommentDepth`, so ` #endif */` — the result of commenting out a conditional block — keeps its comment terminator. Previously `hasError` went raw=false -> preprocessed=true and the rest of the file was swallowed. - The decision keys on the scanner's brace depth instead of indentation. A column-0 `#if` inside a class body is blanked (6 of 7 body shapes previously still lost the enclosing declaration) and an indented file-scope directive is not — matching what the doc comment already claimed. Bare-CR line endings, NBSP/ideographic indentation and a leading BOM are recognized too. - A group is blanked only when every branch is brace-balanced. An `#if`/`#else` that splits a declaration header leaves one unmatched `{` once both branches survive, which collapsed five top-level nodes into one and gave unrelated types fabricated `NetworkClient.` qualified names. Such a group now degrades to the pre-fix behavior. - Multiline strings honour `\"""` escapes, and a plain `"""` closes even when a `#` follows it, so the scanner no longer wedges in string state and silently stops blanking for the rest of the file. - The pound run is counted once per position and skipped. It was quadratic: 10.6s for one 64k-`#` line, well inside the 512 KB walker limit. - Extended regex literals (`#/.../#`) no longer open a phantom block comment. - Directive-free files return early, matching `stripUeMacros`. Worker parity: `emitSwiftScopeCaptures` and `emitCppScopeCaptures` re-apply their provider's `preprocessSource` on the parse-cache-miss path — Dart already did this — and the embedding parse in `ensureAndParse` applies the hook as well. Before this the worker and the scope-capture/embedding halves analyzed different programs, turning a consistent degradation into cold-run/warm-run non-determinism. A new parity test pins the equivalence for every provider that defines the hook. SCHEMA_BUMP 37 -> 38: this changes parse semantics, the chunk key hashes raw on-disk bytes, and `preprocessSource` runs after the key is computed — so a same-package-version warm cache would replay pre-fix Swift results verbatim, including across `--force`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(ingestion): apply preprocessSource once in the scope bridge (#2771) Follow-up cleanup on the review fixes. The previous commit re-applied each provider's `preprocessSource` inside `emitSwiftScopeCaptures` and `emitCppScopeCaptures`, mirroring what Dart already did — three copies of the same rule, and a contract that asked every future emitter to remember it. `extractParsedFile` is the single funnel every `emitScopeCaptures` caller passes through (parse worker, scope-resolution run, Vue script extraction), and it already receives the provider. Applying the hook there on the cache-miss path covers all three languages and every future one, names no language in shared code, and drops Dart's unconditional transform on the cache-hit path. Verified the three emitters use `sourceText` for nothing but the parse, so the substitution is output-identical — which the parity test asserts directly. Also from the cleanup pass: - the parity test derives its language list from the provider registry, so a new provider adopting the hook fails until it adds a fixture - `ensureAndParse` resolves the provider from the language it already computed, instead of a second extension table (`getProviderForFile`) - the preprocessor returns `sourceText` unchanged when no group was blanked, which is the common case for files whose only directives are top-level - `split(/(\r\n|\n|\r)/)` replaces the hand-rolled line splitter, and the per-group brace bookkeeping is two scalars instead of an array - the hint regex is derived from the line regex so the two cannot drift - unit assertions compare the WHOLE preprocessed file against the expected blanking, replacing per-line spot checks; the pipeline tests share one `runFixture` helper and `getNodesForFile` in the resolver test helpers - `LanguageProvider.preprocessSource` documents the real call sites and says plainly that the set is not closed — `populateRangeBindings` still hands language helpers raw text Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(autofix): apply prettier + eslint fixes via /autofix command --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
308 lines
10 KiB
TypeScript
308 lines
10 KiB
TypeScript
/**
|
|
* Integration test: Embedding chunking pipeline
|
|
*
|
|
* Tests the chunking + text generation pipeline together.
|
|
*/
|
|
import { describe, it, expect, vi } from 'vitest';
|
|
import { characterChunk } from '../../src/core/embeddings/character-chunk.js';
|
|
import { generateEmbeddingText } from '../../src/core/embeddings/text-generator.js';
|
|
import type { EmbeddableNode } from '../../src/core/embeddings/types.js';
|
|
|
|
const { createParserForLanguage } = vi.hoisted(() => ({
|
|
createParserForLanguage: vi.fn(),
|
|
}));
|
|
|
|
vi.mock('../../src/core/tree-sitter/parser-loader.js', () => ({
|
|
createParserForLanguage,
|
|
isLanguageAvailable: vi.fn().mockReturnValue(true),
|
|
resolveLanguageKey: vi.fn((language: string) => language),
|
|
}));
|
|
|
|
const { getLanguageFromFilename } = vi.hoisted(() => ({
|
|
getLanguageFromFilename: vi.fn().mockReturnValue('typescript'),
|
|
}));
|
|
|
|
// Partial mock: `ast-utils` now resolves the LanguageProvider registry to apply
|
|
// `preprocessSource`, and that graph needs the real shared exports (#2771).
|
|
vi.mock('gitnexus-shared', async (importOriginal) => ({
|
|
...(await importOriginal<typeof import('gitnexus-shared')>()),
|
|
getLanguageFromFilename,
|
|
}));
|
|
|
|
import { chunkNode } from '../../src/core/embeddings/chunker.js';
|
|
|
|
const CLASS_PREV_TAIL_SAMPLE = 30;
|
|
const STRUCT_PREV_TAIL_SAMPLE = 20;
|
|
|
|
type FakeNode = {
|
|
type: string;
|
|
startIndex: number;
|
|
endIndex: number;
|
|
namedChildCount: number;
|
|
namedChild: (index: number) => FakeNode | null;
|
|
childForFieldName?: (name: string) => FakeNode | null;
|
|
};
|
|
|
|
const makeFakeNode = (
|
|
type: string,
|
|
startIndex: number,
|
|
endIndex: number,
|
|
children: FakeNode[] = [],
|
|
fields: Record<string, FakeNode> = {},
|
|
): FakeNode => ({
|
|
type,
|
|
startIndex,
|
|
endIndex,
|
|
namedChildCount: children.length,
|
|
namedChild: (index: number) => children[index] ?? null,
|
|
childForFieldName: (name: string) => fields[name] ?? null,
|
|
});
|
|
|
|
const makeDeclarationTree = (
|
|
nodeType: string,
|
|
bodyType: string,
|
|
content: string,
|
|
members: Array<{ text: string; type: string }>,
|
|
) => {
|
|
let searchFrom = 0;
|
|
const memberNodes = members.map((member) => {
|
|
const startIndex = content.indexOf(member.text, searchFrom);
|
|
if (startIndex < 0) {
|
|
throw new Error(`Unable to locate declaration member text: ${member.text}`);
|
|
}
|
|
searchFrom = startIndex + member.text.length;
|
|
return makeFakeNode(member.type, startIndex, startIndex + member.text.length);
|
|
});
|
|
|
|
const bodyStart = content.indexOf('{');
|
|
const bodyEnd = content.lastIndexOf('}') + 1;
|
|
const bodyNode = makeFakeNode(bodyType, bodyStart, bodyEnd, memberNodes);
|
|
const declNode = makeFakeNode(nodeType, 0, bodyEnd, [bodyNode], { body: bodyNode });
|
|
return {
|
|
rootNode: makeFakeNode('program', 0, content.length, [declNode]),
|
|
};
|
|
};
|
|
|
|
describe('embedding-chunking integration', () => {
|
|
const makeNode = (overrides: Partial<EmbeddableNode>): EmbeddableNode => ({
|
|
id: 'Function:src/test.ts:test',
|
|
name: 'test',
|
|
label: 'Function',
|
|
filePath: 'src/test.ts',
|
|
content: '',
|
|
startLine: 1,
|
|
endLine: 10,
|
|
...overrides,
|
|
});
|
|
|
|
it('short function produces single chunk with metadata', () => {
|
|
const node = makeNode({
|
|
content: 'function hello() { return "world"; }',
|
|
isExported: true,
|
|
repoName: 'my-project',
|
|
serverName: 'my-service',
|
|
});
|
|
|
|
const chunks = characterChunk(node.content, 1, 3, 1200, 120);
|
|
expect(chunks).toHaveLength(1);
|
|
|
|
const text = generateEmbeddingText(node, chunks[0].text);
|
|
expect(text).toContain('Function: test');
|
|
expect(text).toContain('function hello()');
|
|
// #2333: verbose metadata is no longer part of embedding text.
|
|
expect(text).not.toContain('Repo: my-project');
|
|
expect(text).not.toContain('Server: my-service');
|
|
expect(text).not.toContain('Export: true');
|
|
});
|
|
|
|
it('long function produces multiple chunks', () => {
|
|
const longContent = Array.from({ length: 100 }, (_, i) => ` const line${i} = ${i};`).join(
|
|
'\n',
|
|
);
|
|
const node = makeNode({
|
|
content: `function longFn() {\n${longContent}\n}`,
|
|
startLine: 1,
|
|
endLine: 102,
|
|
});
|
|
|
|
const chunks = characterChunk(node.content, 1, 102, 1200, 120);
|
|
expect(chunks.length).toBeGreaterThan(1);
|
|
expect(chunks[0].chunkIndex).toBe(0);
|
|
});
|
|
|
|
it('short labels (TypeAlias) skip chunking and embed directly', () => {
|
|
const node = makeNode({
|
|
label: 'TypeAlias',
|
|
name: 'Result',
|
|
content: 'type Result<T> = Success<T> | Error;',
|
|
});
|
|
|
|
const chunks = characterChunk(node.content, 1, 1, 1200, 120);
|
|
expect(chunks).toHaveLength(1);
|
|
|
|
const text = generateEmbeddingText(node, chunks[0].text);
|
|
expect(text).toContain('TypeAlias: Result');
|
|
expect(text).toContain('type Result<T> = Success<T> | Error;');
|
|
});
|
|
|
|
it('long enum uses character fallback', () => {
|
|
const enumContent = Array.from(
|
|
{ length: 200 },
|
|
(_, i) => ` Value${i} = "${'x'.repeat(20)}${i}",`,
|
|
).join('\n');
|
|
const node = makeNode({
|
|
label: 'Enum',
|
|
name: 'LargeEnum',
|
|
content: `enum LargeEnum {\n${enumContent}\n}`,
|
|
startLine: 1,
|
|
endLine: 202,
|
|
});
|
|
|
|
const chunks = characterChunk(node.content, 1, 202, 1200, 120);
|
|
expect(chunks.length).toBeGreaterThan(1);
|
|
});
|
|
|
|
it('long class produces member-aware chunks with structural metadata', async () => {
|
|
const node = makeNode({
|
|
label: 'Class',
|
|
name: 'Parser',
|
|
methodNames: ['parseJSON', 'validate'],
|
|
fieldNames: ['options', 'cache'],
|
|
content: `class Parser {
|
|
options: ParserOptions;
|
|
cache: Map<string, any>;
|
|
parseJSON(text: string) { return JSON.parse(text); }
|
|
validate() { return true; }
|
|
}`,
|
|
startLine: 20,
|
|
endLine: 25,
|
|
});
|
|
createParserForLanguage.mockResolvedValue({
|
|
parse: vi.fn().mockReturnValue(
|
|
makeDeclarationTree('class_declaration', 'class_body', node.content, [
|
|
{ text: 'options: ParserOptions;', type: 'field_definition' },
|
|
{ text: 'cache: Map<string, any>;', type: 'field_definition' },
|
|
{
|
|
text: 'parseJSON(text: string) { return JSON.parse(text); }',
|
|
type: 'method_definition',
|
|
},
|
|
{ text: 'validate() { return true; }', type: 'method_definition' },
|
|
]),
|
|
),
|
|
});
|
|
|
|
const chunks = await chunkNode(node.label, node.content, node.filePath, 20, 25, 90, 0);
|
|
expect(chunks).toHaveLength(2);
|
|
|
|
const secondText = generateEmbeddingText(
|
|
node,
|
|
chunks[1].text,
|
|
{},
|
|
chunks[1].chunkIndex,
|
|
chunks[0].text.slice(-CLASS_PREV_TAIL_SAMPLE),
|
|
);
|
|
expect(secondText).toContain('Class: Parser');
|
|
expect(secondText).toContain('Container: class Parser {');
|
|
expect(secondText).toContain('[preceding context]: ...');
|
|
expect(secondText).not.toContain('Methods: parseJSON, validate');
|
|
expect(secondText).not.toContain('Properties: options, cache');
|
|
expect(secondText).toContain('parseJSON(text: string)');
|
|
});
|
|
|
|
it('interface chunks retain structural metadata and signatures', async () => {
|
|
const node = makeNode({
|
|
label: 'Interface',
|
|
name: 'Handler',
|
|
methodNames: ['handle', 'validate'],
|
|
fieldNames: ['name'],
|
|
content: `interface Handler {
|
|
handle(event: Event): void;
|
|
validate(input: string): boolean;
|
|
readonly name: string;
|
|
}`,
|
|
startLine: 30,
|
|
endLine: 34,
|
|
});
|
|
createParserForLanguage.mockResolvedValue({
|
|
parse: vi.fn().mockReturnValue(
|
|
makeDeclarationTree('interface_declaration', 'object_type', node.content, [
|
|
{ text: 'handle(event: Event): void;', type: 'method_definition' },
|
|
{ text: 'validate(input: string): boolean;', type: 'method_definition' },
|
|
{ text: 'readonly name: string;', type: 'property_signature' },
|
|
]),
|
|
),
|
|
});
|
|
|
|
const chunks = await chunkNode(node.label, node.content, node.filePath, 30, 34, 500, 0);
|
|
expect(chunks).toHaveLength(1);
|
|
|
|
const text = generateEmbeddingText(node, chunks[0].text);
|
|
expect(text).toContain('Interface: Handler');
|
|
expect(text).toContain('Methods: handle, validate');
|
|
expect(text).toContain('Container: interface Handler {');
|
|
expect(text).toContain('readonly name: string;');
|
|
});
|
|
|
|
it('struct chunks retain structural container context', async () => {
|
|
getLanguageFromFilename.mockReturnValue('rust');
|
|
const node = makeNode({
|
|
label: 'Struct',
|
|
name: 'User',
|
|
fieldNames: ['name', 'email', 'age', 'address'],
|
|
content: `struct User {
|
|
name: String,
|
|
email: String,
|
|
age: u32,
|
|
address: String,
|
|
}`,
|
|
startLine: 40,
|
|
endLine: 45,
|
|
filePath: 'src/user.rs',
|
|
});
|
|
createParserForLanguage.mockResolvedValue({
|
|
parse: vi.fn().mockReturnValue(
|
|
makeDeclarationTree('struct_item', 'declaration_list', node.content, [
|
|
{ text: 'name: String,', type: 'field_definition' },
|
|
{ text: 'email: String,', type: 'field_definition' },
|
|
{ text: 'age: u32,', type: 'field_definition' },
|
|
{ text: 'address: String,', type: 'field_definition' },
|
|
]),
|
|
),
|
|
});
|
|
|
|
const chunks = await chunkNode(node.label, node.content, node.filePath, 40, 45, 45, 0);
|
|
expect(chunks).toHaveLength(2);
|
|
|
|
const secondText = generateEmbeddingText(
|
|
node,
|
|
chunks[1].text,
|
|
{},
|
|
chunks[1].chunkIndex,
|
|
chunks[0].text.slice(-STRUCT_PREV_TAIL_SAMPLE),
|
|
);
|
|
expect(secondText).toContain('Struct: User');
|
|
expect(secondText).toContain('Container: struct User {');
|
|
expect(secondText).not.toContain('Properties: name, email, age, address');
|
|
expect(secondText).toContain('age: u32,');
|
|
});
|
|
|
|
it('header is present in every chunk', () => {
|
|
const longContent = 'x'.repeat(3000);
|
|
const node = makeNode({
|
|
content: longContent,
|
|
repoName: 'test-repo',
|
|
});
|
|
|
|
const chunks = characterChunk(node.content, 1, 100, 1200, 120);
|
|
expect(chunks.length).toBeGreaterThan(1);
|
|
|
|
for (const chunk of chunks) {
|
|
const text = generateEmbeddingText(node, chunk.text);
|
|
// The compact header (name, + description when present) repeats on every
|
|
// chunk so each chunk keeps its identity; #2333 dropped the metadata lines.
|
|
expect(text).toContain('Function: test');
|
|
expect(text).not.toContain('Repo: test-repo');
|
|
expect(text).not.toContain('Path: src/test.ts');
|
|
}
|
|
});
|
|
});
|