mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-08 03:08:13 +00:00
* docs(parse): record why dispatchGroups is a required interface member Review finding #10 argued dispatchGroups should be optional to match `getQuarantinedPaths?` / `getStats?`. Those are compatibility accommodation for WorkerPool shapes that predate them, not a convention for new members; optional here would force a `?.` plus an unreachable fallback at the single production call site. Documenting the decision so the next reader does not re-litigate it from the neighbouring optional markers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit addaab647377f3c4553f752fa3ca1388bcb9ca81) * refactor(parse): simplify round accounting and dispatch setup Simplification pass over the dispatch-rounds change. Behavior preserved: identical graph on a full analyze (51,286 nodes / 163,092 edges). - Drop `roundMissBytes`. `roundBufferedBytes` counts the same bytes plus the cache hits, so it is always the greater of the two and the first disjunct of the close condition could never fire on its own. One counter, one reset, one check. - Measure round bytes with `Buffer.byteLength(content, 'utf8')` instead of `String.length`. UTF-16 code units undercount non-ASCII source by up to 3x, so the cap meant to bound main-thread retention was letting a CJK-heavy repo hold well past its nominal budget. Matches `estimateItemBytes` in the pool. - Reset the durable ParsedFile directories for a round's chunks concurrently. Each targets its own chunk-hash directory, and running them serially put N round trips of fs work on the critical path the round exists to shorten. The try/catch stays inside the mapped callback, so one failure still degrades that chunk alone. - Skip the quarantine filter entirely when nothing is quarantined, which is every run without a worker death. It was an identity copy of every group. - `dispatchChunkParseRound` takes `DispatchGroup<...>` rather than re-declaring that shape inline; the type was already imported and used in its body. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 527d5b6e0ca8ae7bbc6a414c5ac7e27fd85e9995) * refactor(parse): count round misses with the same idiom startRound uses `drainRound` hand-rolled a reduce to count 'miss' entries while `startRound`, one function above, filters the same predicate over the same union. Same integer, one idiom. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 7eaa193b0cb5fa515844f36ae1401d6fb2fed7b8) * fix(parse): honor GITNEXUS_WORKER_POOL_SIZE above the auto sizing cap The auto pool size is bounded by source bytes so a tiny repo does not spawn a full idle pool. That bound was also clamping the operator's env override, because the env value is read inside `resolveAutoPoolSize()` and the result went through `Math.min(..., workProportionalCap)`. `DEFAULT_POOL_SIZE_CAP`'s own comment offers `GITNEXUS_WORKER_POOL_SIZE` and `--workers <N>` as equivalent escape hatches for operators on bigger machines. They were not. Measured on a 30MB corpus, where the byte-derived cap is 16: --workers 24 -> pool: 24/24 active GITNEXUS_WORKER_POOL_SIZE=24 -> pool: 16/16 active (silently ignored) Both are deliberate operator input, so both now bypass the work-proportional cap, which goes back to bounding only the auto default. After the fix, on the same corpus, with identical graph output (51,286 nodes / 163,092 edges): GITNEXUS_WORKER_POOL_SIZE=24 -> pool: 24/24 active GITNEXUS_WORKER_POOL_SIZE=4 -> pool: 4/4 active unset -> pool: 16/16 active Verified by hand against the pool's own throughput log; not covered by an automated regression test, since the pool size is only observable through that log line and not through the progress stream a test can read. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 17ed08608c878079b2927da25cfd39c1608a02a2) * fix(parse): bound the durable-reset fan-out and pin the pool-size override Review follow-ups on #3200. The round's durable ParsedFile directory resets went out as one unbounded `Promise.all` — one recursive rm + mkdir per miss chunk, all at once. A round can hold hundreds of small packs, and those resets compete for descriptors with the chunk prefetch this loop already has in flight. `readFileContents` degrades a losing read SILENTLY by documented contract, so a dropped file would vanish from the chunk, from the graph, and from the chunk hash — shipping a narrowed index with exit 0. Now routed through `mapConcurrent` at the same width the file reads use, which keeps the pipelining win and caps in-flight descriptors. An operator's pool size is now also bounded by the number of files there are to parse, so `GITNEXUS_WORKER_POOL_SIZE=100000` on a five-file repo cannot become the literal thread count. This applies to `--workers` and the env var alike, so the parity the previous commit established is intact. It does NOT shrink an incremental re-analyze: `totalParseable` counts every parseable file in the scan, not the changed ones. Adds the regression test a reviewer asked for. The existing coverage (`worker-pool-resilience` calling `resolveAutoPoolSize` directly, `analyze-worker-pool-size` mocking `runFullAnalysis`) never reaches `runChunkedParseAndResolve`'s `effectivePoolSize`, so both stayed green through a revert of the fix. The new test drives the real parse phase with a worker double that writes a per-`threadId` marker, and counts them: verified it fails on the reverted line with `expected [ 'worker-1' ] to have a length of 3 but got 1`, and passes on HEAD. Also corrects the `GITNEXUS_PARSE_ROUND_BYTES` docstring, which still described the cache-miss counter deleted two commits ago. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(parse): skip caching a chunk with a stale durable generation; warn on over-subscription Closes the two findings left open by the review of #3200. When `prepareDurableParsedFileChunk` fails, the previous generation's shards are still on disk, so a later warm hit would union them with the new ones. The chunk is now recorded and its parse-cache write skipped -- the same posture `finalizeWorkerChunk` already takes for a quarantined chunk, and for the same reason: do not cache what we cannot vouch for. The next run re-dispatches into a directory it can actually clear. Bounding the reset fan-out removed the correlated trigger; this closes the individual case. Pool size over-subscription now warns rather than caps. Silently capping is precisely what the override exists to prevent, so an operator's number is still honored -- but an exported GITNEXUS_WORKER_POOL_SIZE applies to every analyze in a long-lived caller (watch auto-sync, the MCP server), including small incremental ones, and that is easy to set once and forget. The warning names the host's usable core count, so it is a hardware fact rather than an invented threshold. `resolveHostParallelism` is extracted from `resolveAutoPoolSize` rather than re-deriving the cgroup-aware fallback at the new call site. Tests: the stale-generation skip is pinned by a new case asserting nothing is written under any key; verified it fails without the guard with `expected 1 to be +0`. 60 unit and 49 integration tests pass across the affected suites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(parse): guard dispatch-round cadence with a bench, not a wall-clock budget Round boundaries are deliberately invisible to graph output — batching that changed output would be a bug — so nothing in the repo could see the #3196 win regress. It would have come back as a silent ~1.5x on every cold analyze. Two earlier attempts to pin it as a unit test failed for that exact reason: one scraped a logger line the progress stream does not carry, the other asserted graph content that is identical either way. Extracts the round-close fold into `createRoundBudget`, so the decision is a shared unit the bench measures rather than a copy that drifts. The parse loop is streaming and cannot know chunk sizes up front, so an accumulator is the honest shape — not a planner. Four deterministic arms, one ratio, no millisecond gate: - layout_fingerprint — pack membership. Every cache key derives from it, so drift needs a SCHEMA_BUMP, never a lone re-baseline. - packs / single_file_packs — the FLOOR. `rounds` only asserts something while the corpus over-splits (774 packs where the byte budget needs 5). This is bench/import-target's lesson, where four heap arms read 0 B and passed every ceiling: a ceiling says "not too big", nothing said "still measuring". - rounds — the regression signal, both directions. - cjk_rounds vs ascii_rounds — pins UTF-8 byte accounting. The two corpora share a UTF-16 length and differ only in encoded size, so String.length collapses them to equal. This is the arm no unit test could be. - pack_scaling_ratio — (t_4n/t_n)/4, min-of-15. A ratio because wall-clock is runner-speed-dependent and this repo has the scar: callable-value-flow's ms gate failed twice at 2.07 and 1.975 against 1.9 with correct code, on a sub-11ms measurement. Every arm verified to fail before being recorded: close-every-chunk reads 774 rounds, disabling the close reads 1, reverting roundFileBytes to String.length takes cjk_rounds 8 -> 3, and shrinking the corpus trips the shape floor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(bench): record the analyze phase breakdown and the rejected optimizations Where analyze time actually goes, measured while landing #3194/#3196/#3200, plus the two optimizations that looked compelling and were measured away. The headline is that the parse work is done: a one-file-edit re-analyze is 36.5s, of which parse is 2.8s (8%). scopeResolution is 40% and the unlogged graph emit + FTS rebuild is 49% — neither is incremental, and the ~18s sits outside the phase runner so every phase log is blind to it. Also records the trap that invalidated an earlier measurement: a non-git corpus never records a schema fingerprint, so every run is a forced rebuild and any "warm" number taken that way is fiction. Rejected, with numbers: more workers (16/20/24 land inside run-to-run spread) and bundling the worker entry (~250ms on a normal filesystem; the 8.6s that motivated it was a 9p-mount artifact). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
576 lines
24 KiB
TypeScript
576 lines
24 KiB
TypeScript
/**
|
|
* Regression coverage for the warm-cache ParsedFile gap (#2038, abhigyanpatwari
|
|
* review on parse-cache.ts).
|
|
*
|
|
* #1983 made the parse worker the sole parser: it serializes ParsedFiles to a
|
|
* disk store that scope-resolution streams back, so the main thread NEVER
|
|
* re-parses (the unbounded tree-sitter 0.21.1 native leak → OOM on huge repos).
|
|
* The gap: on a WARM re-analyze where every chunk is a parse-cache HIT, no
|
|
* worker runs, the run-scoped store is cleared at parse start, and the cached
|
|
* `ParseWorkerResult` carries no ParsedFiles — so scope-resolution would find an
|
|
* empty store and fall back to main-thread `extractParsedFile`, re-opening the
|
|
* OOM.
|
|
*
|
|
* The fix: workers ALSO write a durable, content-addressed ParsedFile store
|
|
* keyed by chunk hash (`parsedfile-cache/`); a warm hit LOADS those shards
|
|
* in place (no copy into the run-scoped store) so scope-resolution streams
|
|
* them exactly as on a cold run — zero re-parse, byte-identical.
|
|
*
|
|
* Two layers of coverage:
|
|
* (1) Store-level — the durable persist → `loadParsedFilesForPaths`
|
|
* round-trip at the EXACT seam scope-resolution consumes (phase.ts:255),
|
|
* plus the index version gate and the prune-coherence rule. Build-free.
|
|
* (2) Integration — a two-run `runChunkedParseAndResolve`: run #1 (all miss)
|
|
* populates the durable store; run #2 (all hits) spawns NO worker and
|
|
* loads full coverage from durable shards; the coherence gate re-dispatches
|
|
* when durable shards are absent; and a mixed-mode run (one file changed)
|
|
* hits the unchanged chunk while re-parsing the changed one.
|
|
*/
|
|
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
|
|
import fs from 'node:fs';
|
|
import os from 'node:os';
|
|
import path from 'node:path';
|
|
import { pathToFileURL } from 'node:url';
|
|
|
|
// Partial mock: lets one test make prepareDurableParsedFileChunk fail without
|
|
// touching the worker-side persist path (which shares the same directory).
|
|
const prepareOverride = vi.hoisted(() => ({
|
|
impl: undefined as undefined | (() => Promise<void>),
|
|
}));
|
|
const persistOverride = vi.hoisted(() => ({
|
|
impl: undefined as undefined | (() => Promise<boolean>),
|
|
}));
|
|
vi.mock('../../src/storage/parsedfile-store.js', async (importOriginal) => {
|
|
const real = await importOriginal<typeof import('../../src/storage/parsedfile-store.js')>();
|
|
return {
|
|
...real,
|
|
prepareDurableParsedFileChunk: (durableDir: string, chunkHash: string) =>
|
|
prepareOverride.impl
|
|
? prepareOverride.impl()
|
|
: real.prepareDurableParsedFileChunk(durableDir, chunkHash),
|
|
persistParsedFileChunk: (
|
|
storagePath: string,
|
|
shardId: string,
|
|
parsedFiles: readonly ParsedFile[],
|
|
) =>
|
|
persistOverride.impl
|
|
? persistOverride.impl()
|
|
: real.persistParsedFileChunk(storagePath, shardId, parsedFiles),
|
|
};
|
|
});
|
|
|
|
import { createKnowledgeGraph } from '../../src/core/graph/graph.js';
|
|
import { runChunkedParseAndResolve } from '../../src/core/ingestion/pipeline-phases/parse-impl.js';
|
|
import {
|
|
computeChunkHash,
|
|
fileContentHash,
|
|
PARSE_CACHE_VERSION,
|
|
} from '../../src/storage/parse-cache.js';
|
|
import {
|
|
getDurableParsedFileDir,
|
|
getParsedFileStoreDir,
|
|
prepareDurableParsedFileChunk,
|
|
persistDurableParsedFileShardSync,
|
|
durableChunkHasShards,
|
|
loadParsedFilesForPaths,
|
|
loadDurableParsedFileIndex,
|
|
pruneAndSaveDurableParsedFileStore,
|
|
clearParsedFileStore,
|
|
} from '../../src/storage/parsedfile-store.js';
|
|
import type { ParseWorkerResult } from '../../src/core/ingestion/workers/parse-worker.js';
|
|
import type { ParsedFile } from 'gitnexus-shared';
|
|
|
|
// A structurally-minimal ParsedFile. `loadParsedFilesForPaths` keys on
|
|
// `filePath`; the rest are empty so a restored shard is byte-stable and
|
|
// scope-resolution has nothing to resolve (no edges) but no malformed input.
|
|
const mkParsedFile = (filePath: string): ParsedFile =>
|
|
({
|
|
filePath,
|
|
moduleScope: '',
|
|
scopes: [],
|
|
parsedImports: [],
|
|
localDefs: [],
|
|
referenceSites: [],
|
|
}) as unknown as ParsedFile;
|
|
|
|
// ─── Layer 1: durable store mechanics (build-free) ──────────────────────────
|
|
|
|
describe('durable ParsedFile store — content-addressed warm-cache coverage', () => {
|
|
let tempDir = '';
|
|
beforeEach(() => {
|
|
tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'durable-parsedfile-store-'));
|
|
});
|
|
afterEach(() => {
|
|
if (tempDir) fs.rmSync(tempDir, { recursive: true, force: true });
|
|
});
|
|
|
|
it('persist → loadParsedFilesForPaths gives full coverage (the warm seam)', async () => {
|
|
const durableDir = getDurableParsedFileDir(tempDir);
|
|
const chunkHash = 'a'.repeat(64);
|
|
const files = ['src/a.ts', 'src/b.ts'];
|
|
|
|
// A worker would write this at flush on a cache MISS.
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 7, 0, files.map(mkParsedFile));
|
|
await clearParsedFileStore(tempDir);
|
|
const wanted = new Set(files);
|
|
expect(await durableChunkHasShards(tempDir, chunkHash, wanted)).toBe(true);
|
|
const loaded = await loadParsedFilesForPaths(tempDir, wanted);
|
|
expect([...loaded.keys()].sort()).toEqual([...files].sort());
|
|
});
|
|
|
|
it('durableChunkHasShards is false when the chunk has no durable shards', async () => {
|
|
expect(await durableChunkHasShards(tempDir, 'b'.repeat(64), new Set(['missing.ts']))).toBe(
|
|
false,
|
|
);
|
|
});
|
|
|
|
it('prepares a fresh durable generation without retaining old worker shards', async () => {
|
|
const durableDir = getDurableParsedFileDir(tempDir);
|
|
const chunkHash = 'f'.repeat(64);
|
|
const chunkDir = path.join(durableDir, chunkHash);
|
|
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 1, 0, [mkParsedFile('old.ts')]);
|
|
await prepareDurableParsedFileChunk(durableDir, chunkHash);
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 1, 0, [mkParsedFile('new-a.ts')]);
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 2, 0, [mkParsedFile('new-b.ts')]);
|
|
|
|
const shards = fs
|
|
.readdirSync(chunkDir)
|
|
.filter((name) => name.endsWith('.v8'))
|
|
.sort();
|
|
expect(shards).toEqual([`${chunkHash}-w1-0.v8`, `${chunkHash}-w2-0.v8`]);
|
|
const wanted = new Set(['old.ts', 'new-a.ts', 'new-b.ts']);
|
|
expect(await durableChunkHasShards(tempDir, chunkHash, new Set(['new-a.ts', 'new-b.ts']))).toBe(
|
|
true,
|
|
);
|
|
const files = await loadParsedFilesForPaths(tempDir, wanted);
|
|
expect([...files.keys()].sort()).toEqual(['new-a.ts', 'new-b.ts']);
|
|
});
|
|
|
|
it('index load is version-gated (PARSE_CACHE_VERSION mismatch ⇒ empty)', async () => {
|
|
const durableDir = getDurableParsedFileDir(tempDir);
|
|
const chunkHash = 'c'.repeat(64);
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 1, 0, [mkParsedFile('x.ts')]);
|
|
await pruneAndSaveDurableParsedFileStore(durableDir, PARSE_CACHE_VERSION, new Set([chunkHash]));
|
|
|
|
expect(await loadDurableParsedFileIndex(durableDir, PARSE_CACHE_VERSION)).toEqual(
|
|
new Map([[chunkHash, new Set(['x.ts'])]]),
|
|
);
|
|
// A schema bump (different version) invalidates the whole durable store.
|
|
expect(await loadDurableParsedFileIndex(durableDir, '999+9.9.9')).toEqual(new Map());
|
|
});
|
|
|
|
it('prune keeps only keepKeys subdirs with ≥1 shard, drops the rest, and re-indexes', async () => {
|
|
const durableDir = getDurableParsedFileDir(tempDir);
|
|
const keep = 'd'.repeat(64);
|
|
const drop = 'e'.repeat(64); // present on disk but NOT in keepKeys (e.g. quarantined / stale)
|
|
persistDurableParsedFileShardSync(durableDir, keep, 1, 0, [mkParsedFile('keep.ts')]);
|
|
persistDurableParsedFileShardSync(durableDir, drop, 1, 0, [mkParsedFile('drop.ts')]);
|
|
|
|
await pruneAndSaveDurableParsedFileStore(durableDir, PARSE_CACHE_VERSION, new Set([keep]));
|
|
|
|
expect(fs.existsSync(path.join(durableDir, keep))).toBe(true);
|
|
expect(fs.existsSync(path.join(durableDir, drop))).toBe(false);
|
|
expect(await loadDurableParsedFileIndex(durableDir, PARSE_CACHE_VERSION)).toEqual(
|
|
new Map([[keep, new Set(['keep.ts'])]]),
|
|
);
|
|
});
|
|
});
|
|
|
|
// ─── Layer 2: parse-impl integration (injected worker, build-free) ───────────
|
|
|
|
// A test worker that mirrors the production flush contract: it writes a
|
|
// run-scoped V8 shard AND a durable, content-addressed V8 shard (when the
|
|
// flush carries a chunk hash) using the SAME directory layout as the real worker.
|
|
const writeStoreWorker = (workerPath: string, markerPath: string): void => {
|
|
fs.writeFileSync(
|
|
workerPath,
|
|
`
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const v8 = require('node:v8');
|
|
const { createHash } = require('node:crypto');
|
|
const { parentPort, threadId, workerData } = require('node:worker_threads');
|
|
const storePath = workerData && workerData.parsedFileStoreStoragePath;
|
|
const durablePath = workerData && workerData.durableParsedFileStoragePath;
|
|
let shardSeq = 0;
|
|
fs.writeFileSync(${JSON.stringify(markerPath)}, 'spawned');
|
|
parentPort.postMessage({ type: 'ready' });
|
|
const writeV8 = (filePath, graph, paths) => {
|
|
const payload = v8.serialize(graph);
|
|
const listing = paths.some((p) => /[\\r\\n\\0]/.test(p))
|
|
? Buffer.alloc(0)
|
|
: Buffer.from(paths.length + '\\n' + paths.join('\\n') + '\\n', 'utf8');
|
|
const MAGIC = Buffer.from('GNXV8CF1');
|
|
const v8ver = Buffer.from(process.versions.v8, 'utf8');
|
|
const nodeMajor = Number.parseInt(process.versions.node.split('.')[0], 10);
|
|
const header = Buffer.allocUnsafe(16 + v8ver.length + 12);
|
|
MAGIC.copy(header, 0);
|
|
header.writeUInt32LE(5, 8);
|
|
header.writeUInt16LE(nodeMajor, 12);
|
|
header.writeUInt16LE(v8ver.length, 14);
|
|
v8ver.copy(header, 16);
|
|
let off = 16 + v8ver.length;
|
|
header.writeUInt32LE(listing.length === 0 ? 0 : paths.length, off);
|
|
header.writeUInt32LE(listing.length, off + 4);
|
|
header.writeUInt32LE(payload.length, off + 8);
|
|
fs.mkdirSync(path.dirname(filePath), { recursive: true });
|
|
const payloadHash = createHash('sha256').update(listing).update(payload).digest();
|
|
fs.writeFileSync(filePath, Buffer.concat([header, listing, payload, payloadHash]));
|
|
};
|
|
const reset = () => ({
|
|
nodes: [], relationships: [], symbols: [], imports: [], calls: [], assignments: [], heritage: [],
|
|
routes: [], fetchCalls: [], fetchWrapperDefs: [], decoratorRoutes: [], routerIncludes: [],
|
|
routerImports: [], toolDefs: [], ormQueries: [], constructorBindings: [], fileScopeBindings: [],
|
|
parsedFiles: [], skippedLanguages: {}, fileCount: 0,
|
|
scopeExtractionFailures: [],
|
|
});
|
|
let accumulated = reset();
|
|
parentPort.on('message', (msg) => {
|
|
if (msg && msg.type === 'sub-batch') {
|
|
for (const file of msg.files) {
|
|
const filePath = file.path;
|
|
const name = filePath.split('/').pop().replace(/\\.ts$/, '');
|
|
accumulated.nodes.push({
|
|
id: 'Function:' + filePath + ':' + name,
|
|
label: 'Function',
|
|
properties: { name, filePath, startLine: 1, endLine: 1, language: 'typescript' },
|
|
});
|
|
accumulated.parsedFiles.push({
|
|
filePath, moduleScope: '', scopes: [], parsedImports: [], localDefs: [], referenceSites: [],
|
|
});
|
|
if (filePath.includes('broken')) accumulated.scopeExtractionFailures.push(filePath);
|
|
accumulated.fileCount++;
|
|
}
|
|
parentPort.postMessage({ type: 'progress', filesProcessed: accumulated.fileCount });
|
|
parentPort.postMessage({ type: 'sub-batch-done' });
|
|
return;
|
|
}
|
|
if (msg && msg.type === 'flush') {
|
|
if ((storePath || durablePath) && accumulated.parsedFiles.length > 0) {
|
|
const seq = shardSeq++;
|
|
const paths = accumulated.parsedFiles.map((pf) => pf.filePath);
|
|
let wroteStore = false;
|
|
if (durablePath && typeof msg.chunkHash === 'string') {
|
|
writeV8(
|
|
path.join(durablePath, msg.chunkHash, msg.chunkHash + '-w' + threadId + '-' + seq + '.v8'),
|
|
accumulated.parsedFiles,
|
|
paths,
|
|
);
|
|
}
|
|
if (storePath) {
|
|
writeV8(
|
|
path.join(storePath, 'parsedfile-store', 'w' + threadId + '-' + seq + '.v8'),
|
|
accumulated.parsedFiles,
|
|
paths,
|
|
);
|
|
wroteStore = true;
|
|
}
|
|
const keepForMain = accumulated.parsedFiles.some((pf) =>
|
|
pf.filePath.includes('persist-fallback')
|
|
);
|
|
if (wroteStore && !keepForMain) accumulated.parsedFiles = [];
|
|
}
|
|
parentPort.postMessage({ type: 'result', data: accumulated });
|
|
accumulated = reset();
|
|
}
|
|
});
|
|
`,
|
|
);
|
|
};
|
|
|
|
describe('parse-impl warm-cache ParsedFile coverage (#2038)', () => {
|
|
let tempDir = '';
|
|
let repoDir = '';
|
|
let storageDir = '';
|
|
let workerPath = '';
|
|
let markerPath = '';
|
|
|
|
beforeEach(() => {
|
|
tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'warm-cache-coverage-'));
|
|
repoDir = path.join(tempDir, 'repo');
|
|
storageDir = path.join(tempDir, 'storage');
|
|
fs.mkdirSync(repoDir, { recursive: true });
|
|
fs.mkdirSync(storageDir, { recursive: true });
|
|
workerPath = path.join(tempDir, 'store-worker.js');
|
|
markerPath = path.join(tempDir, 'worker-spawned.marker');
|
|
writeStoreWorker(workerPath, markerPath);
|
|
});
|
|
afterEach(() => {
|
|
if (tempDir) fs.rmSync(tempDir, { recursive: true, force: true });
|
|
prepareOverride.impl = undefined;
|
|
persistOverride.impl = undefined;
|
|
});
|
|
|
|
const writeFile = (rel: string, content: string): { path: string; size: number } => {
|
|
const full = path.join(repoDir, rel);
|
|
fs.mkdirSync(path.dirname(full), { recursive: true });
|
|
fs.writeFileSync(full, content);
|
|
return { path: rel, size: fs.statSync(full).size };
|
|
};
|
|
|
|
const newCache = () => ({
|
|
version: PARSE_CACHE_VERSION,
|
|
entries: new Map<string, ParseWorkerResult[]>(),
|
|
usedKeys: new Set<string>(),
|
|
storagePath: storageDir,
|
|
onDiskKeys: new Set<string>(),
|
|
});
|
|
|
|
// The post-run orchestrator step (run-analyze) — persist parse cache + prune
|
|
// the durable store to the surviving keys. Mirrored here so run #2 sees an
|
|
// index, exactly like a real second invocation.
|
|
const persistCaches = async (cache: ReturnType<typeof newCache>): Promise<void> => {
|
|
const { saveParseCache, pruneCache } = await import('../../src/storage/parse-cache.js');
|
|
pruneCache(cache, cache.usedKeys);
|
|
const saved = await saveParseCache(storageDir, cache);
|
|
await pruneAndSaveDurableParsedFileStore(
|
|
getDurableParsedFileDir(storageDir),
|
|
PARSE_CACHE_VERSION,
|
|
new Set(saved),
|
|
);
|
|
};
|
|
|
|
const run = async (
|
|
cache: ReturnType<typeof newCache>,
|
|
files: { path: string; size: number }[],
|
|
chunkByteBudget?: number,
|
|
): Promise<Awaited<ReturnType<typeof runChunkedParseAndResolve>>> => {
|
|
const rels = files.map((f) => f.path);
|
|
return runChunkedParseAndResolve(
|
|
createKnowledgeGraph(),
|
|
files,
|
|
rels,
|
|
files.length,
|
|
repoDir,
|
|
Date.now(),
|
|
() => {},
|
|
{
|
|
workerUrlForTest: pathToFileURL(workerPath),
|
|
workerPoolSize: 1,
|
|
parseCache: cache,
|
|
...(chunkByteBudget !== undefined ? { chunkByteBudget } : {}),
|
|
},
|
|
);
|
|
};
|
|
|
|
it('run #1 (miss) populates the durable store, keyed by chunk hash', async () => {
|
|
const f = writeFile('src/cached.ts', 'export function cached() { return 1; }\n');
|
|
const chunkHash = computeChunkHash([
|
|
{
|
|
filePath: f.path,
|
|
contentHash: fileContentHash(fs.readFileSync(path.join(repoDir, f.path), 'utf-8')),
|
|
},
|
|
]);
|
|
const cache = newCache();
|
|
|
|
await run(cache, [f]);
|
|
|
|
expect(fs.existsSync(markerPath)).toBe(true); // worker ran (miss)
|
|
const chunkDir = path.join(getDurableParsedFileDir(storageDir), chunkHash);
|
|
expect(fs.existsSync(chunkDir)).toBe(true);
|
|
expect(fs.readdirSync(chunkDir).filter((n) => n.endsWith('.v8')).length).toBeGreaterThan(0);
|
|
expect(cache.usedKeys.has(chunkHash)).toBe(true);
|
|
});
|
|
|
|
it('a failing durable-generation reset degrades instead of failing the analyze', async () => {
|
|
const f = writeFile('src/degrade.ts', 'export function degrade() { return 1; }\n');
|
|
prepareOverride.impl = () => Promise.reject(new Error('EACCES: simulated cache failure'));
|
|
try {
|
|
await expect(run(newCache(), [f])).resolves.toBeDefined();
|
|
} finally {
|
|
prepareOverride.impl = undefined;
|
|
}
|
|
});
|
|
|
|
it('does not cache a chunk whose durable generation could not be reset', async () => {
|
|
// The reset failed, so the previous generation's shards are still on disk.
|
|
// Caching this chunk would let a future warm hit union those stale shards
|
|
// with the new ones. Same posture as a quarantined chunk: leave it uncached
|
|
// so the next run re-dispatches into a directory it can actually clear.
|
|
const f = writeFile('src/stale-generation.ts', 'export function stale() { return 1; }\n');
|
|
const cache = newCache();
|
|
prepareOverride.impl = () => Promise.reject(new Error('EACCES: simulated cache failure'));
|
|
try {
|
|
await expect(run(cache, [f])).resolves.toBeDefined();
|
|
} finally {
|
|
prepareOverride.impl = undefined;
|
|
}
|
|
|
|
// Nothing was written under any key -- neither on disk nor in memory.
|
|
expect(cache.onDiskKeys.size + cache.entries.size).toBe(0);
|
|
});
|
|
|
|
it('retains worker ParsedFiles when the main-thread run-store write fails', async () => {
|
|
const f = writeFile(
|
|
'src/persist-fallback.ts',
|
|
'export function persistFallback() { return 1; }\n',
|
|
);
|
|
persistOverride.impl = () => Promise.resolve(false);
|
|
|
|
const result = await run(newCache(), [f]);
|
|
|
|
expect(result.parsedFiles.map((parsed) => parsed.filePath)).toContain(f.path);
|
|
});
|
|
|
|
it('does not snapshot durable shards when the parse-cache payload is missing', async () => {
|
|
const f = writeFile('src/orphan.ts', 'export function orphan() { return 1; }\n');
|
|
const chunkHash = computeChunkHash([
|
|
{
|
|
filePath: f.path,
|
|
contentHash: fileContentHash(fs.readFileSync(path.join(repoDir, f.path), 'utf-8')),
|
|
},
|
|
]);
|
|
const durableDir = getDurableParsedFileDir(storageDir);
|
|
persistDurableParsedFileShardSync(durableDir, chunkHash, 1, 0, [mkParsedFile(f.path)]);
|
|
await pruneAndSaveDurableParsedFileStore(durableDir, PARSE_CACHE_VERSION, new Set([chunkHash]));
|
|
|
|
await run(newCache(), [f]);
|
|
|
|
const runShards = fs
|
|
.readdirSync(getParsedFileStoreDir(storageDir))
|
|
.filter((name) => name.endsWith('.v8'));
|
|
expect(runShards.length).toBeGreaterThan(0);
|
|
expect(runShards.every((name) => !name.startsWith(chunkHash))).toBe(true);
|
|
});
|
|
|
|
it('a repeated cache miss replaces the durable chunk generation', async () => {
|
|
const f = writeFile('src/repeated.ts', 'export function repeated() { return 1; }\n');
|
|
const chunkHash = computeChunkHash([
|
|
{
|
|
filePath: f.path,
|
|
contentHash: fileContentHash(fs.readFileSync(path.join(repoDir, f.path), 'utf-8')),
|
|
},
|
|
]);
|
|
|
|
await run(newCache(), [f]);
|
|
await run(newCache(), [f]);
|
|
|
|
const chunkDir = path.join(getDurableParsedFileDir(storageDir), chunkHash);
|
|
const shards = fs.readdirSync(chunkDir).filter((name) => name.endsWith('.v8'));
|
|
expect(shards).toHaveLength(1);
|
|
const shard = shards[0];
|
|
if (!shard) throw new Error('expected one durable V8 shard');
|
|
const { tryLoadV8Cache } = await import('../../src/storage/v8-sidecar.js');
|
|
const hit = await tryLoadV8Cache(path.join(chunkDir, shard));
|
|
expect(hit?.kind).toBe('hit');
|
|
if (hit?.kind !== 'hit') return;
|
|
const parsed = hit.value as Array<{ filePath: string }>;
|
|
expect(parsed.map((item) => item.filePath)).toEqual(['src/repeated.ts']);
|
|
});
|
|
|
|
it('run #2 (all hits) spawns NO worker — the warm path is served from caches', async () => {
|
|
const f = writeFile('src/cached.ts', 'export function cached() { return 1; }\n');
|
|
const cache = newCache();
|
|
|
|
await run(cache, [f]); // miss → populates
|
|
await persistCaches(cache);
|
|
|
|
// Reload caches from disk for the warm run, like a fresh invocation.
|
|
const { loadParseCache } = await import('../../src/storage/parse-cache.js');
|
|
const warm = await loadParseCache(storageDir);
|
|
fs.rmSync(markerPath, { force: true }); // reset the spawn marker
|
|
|
|
await run(warm as ReturnType<typeof newCache>, [f]);
|
|
|
|
expect(fs.existsSync(markerPath)).toBe(false); // NO worker spawned on the warm hit
|
|
});
|
|
|
|
it('replays scope-extraction failures from a warm parse-cache hit', async () => {
|
|
const f = writeFile('src/broken.ts', 'export function broken() { return 1; }\n');
|
|
const cache = newCache();
|
|
|
|
const cold = await run(cache, [f]);
|
|
expect(cold.scopeExtractionFailures).toEqual([f.path]);
|
|
await persistCaches(cache);
|
|
|
|
const { loadParseCache } = await import('../../src/storage/parse-cache.js');
|
|
const warm = await loadParseCache(storageDir);
|
|
fs.rmSync(markerPath, { force: true });
|
|
|
|
const replayed = await run(warm as ReturnType<typeof newCache>, [f]);
|
|
expect(fs.existsSync(markerPath)).toBe(false);
|
|
expect(replayed.scopeExtractionFailures).toEqual([f.path]);
|
|
});
|
|
|
|
it('coherence gate: a parse-cache hit with NO durable shards re-dispatches the worker', async () => {
|
|
const f = writeFile('src/cached.ts', 'export function cached() { return 1; }\n');
|
|
const cache = newCache();
|
|
|
|
await run(cache, [f]); // miss → populates parse cache + durable
|
|
await persistCaches(cache);
|
|
|
|
// Wipe ONLY the durable store, leaving the parse cache intact — simulates a
|
|
// first run after the durable store was introduced, or a pruned shard.
|
|
fs.rmSync(getDurableParsedFileDir(storageDir), { recursive: true, force: true });
|
|
|
|
const { loadParseCache } = await import('../../src/storage/parse-cache.js');
|
|
const warm = await loadParseCache(storageDir);
|
|
fs.rmSync(markerPath, { force: true });
|
|
|
|
await run(warm as ReturnType<typeof newCache>, [f]);
|
|
|
|
// The gate must NOT silently skip — it falls through to a worker re-dispatch
|
|
// (which repopulates the durable store), never a main-thread re-extract.
|
|
expect(fs.existsSync(markerPath)).toBe(true);
|
|
});
|
|
|
|
it('coherence gate: a parse-cache hit with a corrupt durable shard re-dispatches', async () => {
|
|
const f = writeFile('src/corrupt.ts', 'export function corrupt() { return 1; }\n');
|
|
const cache = newCache();
|
|
|
|
await run(cache, [f]);
|
|
await persistCaches(cache);
|
|
const chunkHash = computeChunkHash([
|
|
{
|
|
filePath: f.path,
|
|
contentHash: fileContentHash('export function corrupt() { return 1; }\n'),
|
|
},
|
|
]);
|
|
const chunkDir = path.join(getDurableParsedFileDir(storageDir), chunkHash);
|
|
const shard = fs.readdirSync(chunkDir).find((name) => name.endsWith('.v8'));
|
|
expect(shard).toBeDefined();
|
|
if (!shard) return;
|
|
fs.writeFileSync(path.join(chunkDir, shard), Buffer.from([0, 1, 2]));
|
|
|
|
const { loadParseCache } = await import('../../src/storage/parse-cache.js');
|
|
const warm = await loadParseCache(storageDir);
|
|
fs.rmSync(markerPath, { force: true });
|
|
|
|
await run(warm as ReturnType<typeof newCache>, [f]);
|
|
|
|
expect(fs.existsSync(markerPath)).toBe(true);
|
|
});
|
|
|
|
it('mixed-mode: changing one file re-parses its chunk while the unchanged chunk restores', async () => {
|
|
// Force one file per chunk (chunkByteBudget: 1) so a and b hash to DISTINCT
|
|
// chunks — the true mixed-mode the pr-2038 mixed-mode gap warns about.
|
|
const a = writeFile('src/a.ts', 'export function a() { return 1; }\n');
|
|
const b = writeFile('src/b.ts', 'export function b() { return 2; }\n');
|
|
const cache = newCache();
|
|
|
|
await run(cache, [a, b], 1); // both miss → both durable subdirs populated
|
|
await persistCaches(cache);
|
|
|
|
const aHash = computeChunkHash([
|
|
{ filePath: a.path, contentHash: fileContentHash('export function a() { return 1; }\n') },
|
|
]);
|
|
// run #1 populated a's durable subdir (the chunk that will HIT on run #2).
|
|
expect(fs.existsSync(path.join(getDurableParsedFileDir(storageDir), aHash))).toBe(true);
|
|
|
|
// Change b's content → b's chunk hash changes → b misses, a still hits.
|
|
fs.writeFileSync(path.join(repoDir, b.path), 'export function b() { return 999; }\n');
|
|
const b2 = { path: b.path, size: fs.statSync(path.join(repoDir, b.path)).size };
|
|
|
|
const { loadParseCache } = await import('../../src/storage/parse-cache.js');
|
|
const warm = await loadParseCache(storageDir);
|
|
fs.rmSync(markerPath, { force: true });
|
|
|
|
await run(warm as ReturnType<typeof newCache>, [a, b2], 1);
|
|
|
|
// The worker spawned (for the changed file b); a was loaded from durable.
|
|
expect(fs.existsSync(markerPath)).toBe(true);
|
|
// a's UNCHANGED chunk is still a hit served from the durable store.
|
|
expect((warm as ReturnType<typeof newCache>).usedKeys.has(aHash)).toBe(true);
|
|
});
|
|
});
|