GitNexus/gitnexus/test/integration/fts-fullfile-search.test.ts
Parafee41 316aaed928
fix(indexing): keep full text file content searchable (#2323)
* fix(indexing): keep full text file content searchable

* fix(indexing): flush CSV chunks by byte size

* fix(fts): flatten newlines/tabs in indexed content so multiline files are searchable (#2317)

The end-to-end FTS test (review follow-up F2) exposed that removing the 10KB
cap alone does NOT fix #2317: Ladybug's FTS tokenizer splits ONLY on the space
character — \n, \r, and \t are not delimiters. So multiline file/symbol content
indexes as a few giant cross-line tokens that no word query matches; full
content is stored but stays unsearchable. (Verified: identical 8KB content is
fully searchable when space-separated and entirely unsearchable when
newline-separated.) The existing fts-description-search test never caught this
because all its seed content is single-line.

Collapse \r\n\t -> single space in the FTS-indexed text (extractContent's File
and snippet content, plus the description column) via normalizeFtsText. This
rewrites the stored column too, so File content returned via the graph API is
space-flattened — an accepted trade for making file/symbol text searchable.

Add the real end-to-end guard test/integration/fts-fullfile-search.test.ts:
write a >16KB file, load it through the real streamAllCSVsToDisk -> COPY ->
createSearchFTSIndexes path, and assert searchFTSFromLbug returns a needle past
10KB (plus a short-content no-regression and a stored-cell-not-truncated
guard). It drives the COPY path a Cypher-seed test would bypass, reusing
withTestLbugDB's FTS-availability gating via a new before-FTS load hook.

* docs(lbug): note the deliberate File-unbounded / snippet-capped asymmetry

The File branch returns full content (whitespace-normalized for FTS, bounded
upstream by the walker cap) while the symbol snippet path 11 lines down stays
MAX_SNIPPET-capped. Comment the intent so the uncapped File branch doesn't read
as a forgotten guard. No behavior change.

* test(lbug): update #2203 overlap round-trip for FTS whitespace normalization

The newline/tab→space normalization (a170915a, #2317) flattens stored File
content, so the #2203 overlap test's "File content == original multiline
source" assertion no longer holds. The test's actual invariant — overlap path
== serial path, byte-for-byte — is unchanged and still asserted; BasicBlock
text (not FTS-indexed) still round-trips raw. Update only the File-content
expectation to the whitespace-flattened form and document why.

* fix(lbug): collapse CSV flush to a single byte threshold

BufferedCSVWriter flushed on row-count (FLUSH_EVERY=500) OR byte-count
(FLUSH_BYTES=8MB) — two independent triggers for one job. Byte count is
the only one tied to the actual risk (an unbounded buffer.join('\n')
string), so drop FLUSH_EVERY and make shouldFlushCSVBuffer single-arg.

Rather than tune FLUSH_BYTES by guesswork or expose it as an env knob,
derive its safety margin from constants the codebase already hard-enforces:
a single row is capped at TREE_SITTER_MAX_BUFFER (32MB, clamped regardless
of GITNEXUS_MAX_FILE_SIZE) and at most doubled by escapeCSVField's
quote-escaping, so the worst-case joined chunk (FLUSH_BYTES + 2 *
TREE_SITTER_MAX_BUFFER ≈ 72MB) sits >7x under Node's MAX_STRING_LENGTH
(~512MB) — the ceiling that throws RangeError: Invalid string length.
A new test pins that margin numerically so it can't erode unnoticed, which
covers the "configurable" alternative better than a knob would: there's no
evidence any deployment needs a different value, and an unbounded env var
would let an operator silently walk the margin back into the danger zone.

Also updates the two tests tied to the removed row-count path: the
FLUSH_EVERY-boundary integration test now crosses FLUSH_BYTES with real
oversized File content instead of relying on row count, and the
shouldFlushCSVBuffer unit test drops to the new single-arg signature.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-07-01 08:27:09 +01:00

86 lines
4.2 KiB
TypeScript

/**
* End-to-end FTS searchability for full File content (#2317 / PR #2323).
*
* PR #2323 removed the 10KB `MAX_FILE_CONTENT` cap so full text-file content
* reaches `file_fts`. The PR's own test proves the late needle lands in the
* generated `file.csv`; it does NOT prove an FTS *search* returns content past
* 10KB. This closes that gap through the REAL pipeline:
*
* write >10KB file on disk → loadGraphToLbug (streamAllCSVsToDisk → COPY of
* the multiline quoted cell) → createFTSIndex(file_fts) → searchFTSFromLbug.
*
* A Cypher CREATE seed would bypass COPY and pass even if COPY truncated the
* cell — the exact thing #2317 must guarantee — so this uses `loadGraphToLbug`
* via the harness's `beforeFTS` hook (which runs before the gated FTS build),
* reusing `withTestLbugDB`'s offline-skip / GITNEXUS_REQUIRE_FTS gating.
*/
import { describe, it, expect } from 'vitest';
import fs from 'node:fs/promises';
import path from 'node:path';
import { withTestLbugDB } from '../helpers/test-indexed-db.js';
import { buildTestGraph } from '../helpers/test-graph.js';
import { searchFTSFromLbug } from '../../src/core/search/bm25-index.js';
// A token near the top (< 10KB) and a distinctive token past ~20KB. Both are
// unique lowercase-alphabetic non-stopwords so they tokenize cleanly under the
// `porter` stemmer and never collide with the printable-ASCII filler (keeping
// the first 1000 chars text, so isBinaryContent does not swap in its sentinel).
const earlyWord = 'sentinelalpha';
const lateNeedle = 'zarquonbeacon';
const FILLER = 'filler line for full file content indexing\n'; // ~43 chars
const FILE_BODY =
`${earlyWord} appears near the very top of the file\n` +
FILLER.repeat(500) + // ~21.5KB of filler → lateNeedle lands well past 10KB
`${lateNeedle} appears far past the old ten kilobyte cutoff\n`;
withTestLbugDB(
'fts-fullfile-search',
() => {
describe('full File content past 10KB is FTS-searchable (#2317)', () => {
it('returns the file for a needle located past the old 10KB cutoff', async () => {
const { results } = await searchFTSFromLbug(lateNeedle, 20);
expect(results.map((r) => r.filePath)).toContain('large.txt');
});
it('persists the full multiline cell through COPY — past 10KB, not the binary sentinel', async () => {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const rows = await adapter.executeQuery(
"MATCH (f:File {filePath: 'large.txt'}) RETURN f.content AS content",
);
const stored = String(rows[0].content);
expect(stored.length).toBeGreaterThan(10240);
expect(stored).not.toContain('[Binary file');
expect(stored).toContain(lateNeedle);
});
it('still finds a token within the first 10KB (no short-content regression)', async () => {
const { results } = await searchFTSFromLbug(earlyWord, 20);
expect(results.map((r) => r.filePath)).toContain('large.txt');
});
});
},
{
// Triggers the FTS-availability probe + offline-skip / GITNEXUS_REQUIRE_FTS
// gating, and builds file_fts over the COPY'd File rows (after beforeFTS).
ftsIndexes: [{ table: 'File', indexName: 'file_fts', columns: ['name', 'content'] }],
// No Cypher `seed`; no pool adapter → searchFTSFromLbug routes through the
// core-adapter connection loadGraphToLbug + createFTSIndex wrote to.
beforeFTS: async (dbPath) => {
// Colocate scratch dirs under the suite temp root so they're auto-cleaned.
const root = path.dirname(dbPath);
const repoDir = path.join(root, 'repo');
const storageDir = path.join(root, 'storage');
await fs.mkdir(repoDir, { recursive: true });
await fs.mkdir(storageDir, { recursive: true });
await fs.writeFile(path.join(repoDir, 'large.txt'), FILE_BODY);
// extractContent reads File content from disk, so the on-disk file is the
// source of the COPY'd cell. loadGraphToLbug runs the real emit + COPY.
const graph = buildTestGraph([
{ id: 'file:large.txt', label: 'File', name: 'large.txt', filePath: 'large.txt' },
]);
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
await adapter.loadGraphToLbug(graph, repoDir, storageDir);
},
},
);