mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-07 02:58:02 +00:00
* fix(schema): declare the full scope-resolution relation cross product (#2792) `RELATION_SCHEMA` was hand-listed, and every prior fix added only the FROM/TO pair named in a crash report — `Const→Method` in #2769, the Swift/Rust member pairs before it. So `analyze` kept aborting at `assertDeclaredPair` on the next codebase whose edges happened to land on a different pair; #2792 reports `Class→Variable` on Java. Audit the surface instead of the symptom. `buildGraphNodeLookup` skips any node whose label is not in `isLinkableLabel`, so the lookup holds only linkable-labelled nodes — and both endpoints of every graph-bridge edge resolve through that lookup. The emittable surface is therefore exactly: FROM LINKABLE_LABELS + File (the module-level caller fallback) TO LINKABLE_LABELS + CALL_TARGET_TYPES `isCallerAnchorLabel` is a strict subset of linkable and contributes nothing on top. `CALL_TARGET_TYPES` contributes `Delegate`, which `tryEmitEdgeWithExplicitTargetId` can emit without going through the lookup at all. Generate that 14x14 block into the DDL rather than listing it: 223 -> 322 declared pairs, and no future pair from these sets can be missing by construction. The containment/inheritance/DI/route/cluster/PDG pairs stay hand-declared — no single predicate describes them. Both label sets live in the ingestion layer, which `core/lbug` must not import, so schema.ts carries twin lists. test/unit/schema-pair-coverage.ts derives the requirement from the originals and fails CI when either set grows without the pairs landing here — the piecemeal loop this fix ends. Measured before widening: at 322 pairs the cost is inside noise (1.09s vs 1.12s per 300 anchored queries on a 32-table DB), but the full 32x32 cross product is ~1.8x on untyped-endpoint anchored queries. The audited subset is the right scope, not "declare everything". INCREMENTAL_SCHEMA_VERSION 34 -> 35: LadybugDB fixes endpoint pairs when the rel table is created, so a pre-v35 database physically cannot store these edges. Closes #2792 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(schema): declare the non-bridge structural pairs COBOL and Vue emit The generated scope-resolution block closed the half of RELATION_SCHEMA a label predicate can describe. The hand-declared half was still stale: with #2791's Function->Variable fix applied, `analyze` continued to abort on this repo's own test/fixtures/lang-resolution with Relationship label pair Module→Property is not declared A full sweep (assertDeclaredPair patched to log-and-skip, run over the whole fixture corpus) found 13 undeclared pairs over 106 edges. This branch already covered 3 of them via the cross product; the remaining 10 come from emitters outside the graph bridge: - cobol-processor.ts mints Module / Namespace / Record / Property / CodeElement and wires them with CONTAINS, CALLS and ACCESSES (9 pairs) - vue-sfc-extractor.ts emits BINDS_EVENT_HANDLER from a handler Function to the child component's File, the only edge whose target is a File (1 pair) CodeElement, Namespace, Record and File are in neither scope-bridge label set, so neither the generated block nor schema-pair-coverage.test.ts can reach them. Adds test/integration/structural-pair-coverage.test.ts, which derives the requirement from a corpus instead of a predicate: it runs the real pipeline over the non-bridge fixtures and requires every FROM/TO pair they produce to be declared. Mutation-checked — dropping `FROM Function TO File` fails it with exactly Function|File. Verified: cobol-app, vue-basic and php-transitive-traits now index instead of aborting; the full lang-resolution corpus completes at 10,876 nodes / 18,517 edges; scrypster/muninndb at 0b7a4272 (the #2789 repro) completes at 20,069 nodes / 71,580 edges, matching #2791 exactly, so this supersedes that PR. * refactor(test): simplify the structural pair coverage guard Cleanup pass over the previous commit. No behaviour change to the schema. - reuse `FIXTURES` and `runPipelineFromRepo` from resolvers/helpers.ts instead of re-deriving the fixture root and importing pipeline.js directly - gate on `distWorkerExists()` like every other integration test that passes `workerUrlForTest`, so a missing dist skips rather than fails - run the three fixtures with `it.concurrent.each`; they share nothing and the cost is almost all worker spawn plus grammar load, which overlaps well (tests phase 21-24s -> 5.6s measured) - replace the sentinel-in-a-Set filter with a plain `.filter()` chain, matching the sibling unit test, and move the declared/table lookups off the per-edge path onto the deduped set - move the pure string pin out of the integration tier into schema-pair-coverage.test.ts, where the identical construct already lives, so it needs no build and survives fixture deletion - trim the schema and test prose that restated the code, and correct the BINDS_EVENT_HANDLER attribution: it is emitted by languages/vue/scope-resolver.ts, not vue-sfc-extractor.ts - amend the v35 comment to mention the 10 structural pairs it now also stamps Still mutation-checked: dropping `FROM Function TO File` now fails both the integration sweep and the unit pin with exactly Function|File. 89 tests green. * fix(schema): generate the attachment pair surface and close four analyze aborts Review of the generated scope-bridge cross product found four `analyze` hard-aborts still live at head, each reproduced end-to-end on the default user path (`analyze --index-only --skip-git`): Method→Annotation Spring `@Bean` + `@ConditionalOnMissingBean` (Java + Kotlin) Method→File Vue Options-API `methods:` handler bound to a child event Namespace→Record COBOL `DECLARATIVES` / `USE AFTER STANDARD ERROR ON <file>` Class→Tool `@mcp.tool()` applied to a class All four are pre-existing on main, and both existing guards were structurally blind to them: the unit guard derives from LINKABLE_LABELS ∪ CALL_TARGET_TYPES (none of Annotation/Tool/Record/File-as-target is a member) and the corpus guard ran three fixtures that exercise none of these emitters. All 16 tests passed while all four crashes were live. The PR's model — "bridge endpoint × structural endpoint" — does not fit: Namespace→Record is structural on both sides. The property that does hold is that the ANCHOR is a lookup result, not a literal at the emit site, so the emitter cannot constrain its label. That gives a second closed-form rule: DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS DEFINITION_ANCHOR_LABELS is derived from NODE_TABLES by subtraction, so a new node table joins automatically. 332 → 450 declared pairs. Sized against a committed harness (gitnexus/bench/schema-pairs), real @ladybugdb/core, identical data: 450 costs 0.93–1.05× of 332 on untyped-endpoint anchored queries — inside noise — versus 1.22–1.43× at 641 and 2.03–2.34× at 1024. The harness reproduces the known #2792 cliff, which is what makes the 450 figure trustworthy. Also in this change: - Delete the 161 hand-declared pairs the rules already generate (233 → 72). The declared set is byte-identical at 450; those lines were load-bearing shadow, because the generator suppresses anything already declared structurally, so narrowing a rule later would silently keep pairs alive. A new guard fails CI if a hand-declared pair is ever re-added inside a rule. - Import LINKABLE_LABELS / CALL_TARGET_TYPES instead of hand-copying them. The twins' stated justification ("the ingestion layer must not be imported here") is false: csv-generator.ts and lbug-adapter.ts, siblings in the same directory, already do, and no rule in AGENTS.md / ARCHITECTURE.md / CONTRIBUTING.md / GUARDRAILS.md states otherwise. - Resolve `resolveStreamGraphEmit` after the guards that rebind `options.force`, not at function entry. It gates on `force`, and every freshness guard runs ~360 lines later, so the v34→v35 bump would have pushed every existing index down the non-streamed emit path — losing the #2680 memory streaming added for the #2649 kernel-scale OOM, for exactly the population most likely to be memory-constrained. - `UndeclaredRelationPairError` now carries the relationship type, both node ids and the source file, with a matching CLI branch. The old message named only the abstract label pair, which a user could not act on. Found through the cause chain, since pipeline-phases/runner.ts rewraps every phase failure. - Share one classifier (`relPairKeyFor`) across the router, both emit sinks and the corpus guard, which previously hand-mirrored the router's skip rule; one cause-chain walker in lib/utils.ts; one exported pair-matching regex. - Corpus guard: four new fixtures reproducing the aborts, per-fixture sentinel pairs so a fixture that stops emitting fails loudly instead of passing vacuously on an empty graph. The per-edge path stays allocation-free: the failure context is passed positionally and the message is built only inside the throw. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX * test(bench): re-baseline the COBOL capture fingerprint for the new fixture `bench/scope-capture` globs `lang-resolution/cobol-*`, so the `cobol-declaratives` fixture added in81daf370e(to reproduce the `Namespace→Record` analyze abort) joined that corpus and shifted the fingerprint — 14 → 15 files. Verified corpus-only, not a capture change: with that one fixture moved aside the fingerprint is byte-identical to the prior baseline (d45bb091…), and81daf370etouches no COBOL capture code. The new value reproduces CI's reported hash exactly. Scaling 0.677 < 1.5 budget. `bench/scope-capture/measure.mjs --check` → PASS (15 languages). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
352 lines
14 KiB
TypeScript
352 lines
14 KiB
TypeScript
/**
|
||
* PdgEmitSink unit tests (issue #2202 U2).
|
||
*
|
||
* Verifies the streaming PDG emit sink:
|
||
* - routes BasicBlock nodes + PDG edges to bounded CSV-on-disk;
|
||
* - delegates structural nodes/edges + the whole-program TAINT_PATH edge to
|
||
* the real graph (never streamed);
|
||
* - is byte-identical (set-wise) to the whole-graph `streamAllCSVsToDisk`
|
||
* emit for the same node/edge set (the issue's byte-identity acceptance);
|
||
* - never accumulates the PDG layer in the in-memory graph (the RSS bound).
|
||
*/
|
||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
|
||
import fs from 'node:fs';
|
||
import fsp from 'node:fs/promises';
|
||
import os from 'node:os';
|
||
import path from 'node:path';
|
||
|
||
import { createKnowledgeGraph } from '../../../src/core/graph/graph.js';
|
||
import { streamAllCSVsToDisk, buildBasicBlockRow } from '../../../src/core/lbug/csv-generator.js';
|
||
import { PdgEmitSink } from '../../../src/core/lbug/pdg-emit-sink.js';
|
||
import type { GraphNode, GraphRelationship } from 'gitnexus-shared';
|
||
|
||
const bbNode = (fp: string, idx: number, line: number): GraphNode => ({
|
||
id: `BasicBlock:${fp}:1:0:${idx}`,
|
||
label: 'BasicBlock',
|
||
properties: { name: '', filePath: fp, startLine: line, endLine: line + 1, text: `blk ${idx}` },
|
||
});
|
||
|
||
const pdgEdge = (
|
||
fp: string,
|
||
from: number,
|
||
to: number,
|
||
type: GraphRelationship['type'],
|
||
reason: string,
|
||
): GraphRelationship => ({
|
||
id: `${type}:${fp}:${from}->${to}`,
|
||
sourceId: `BasicBlock:${fp}:1:0:${from}`,
|
||
targetId: `BasicBlock:${fp}:1:0:${to}`,
|
||
type,
|
||
confidence: 1,
|
||
reason,
|
||
});
|
||
|
||
/** Sorted non-empty lines of a CSV file (order-independent comparison). */
|
||
const sortedLines = async (csvPath: string): Promise<string[]> => {
|
||
const text = await fsp.readFile(csvPath, 'utf8');
|
||
return text
|
||
.split('\n')
|
||
.filter((l) => l.length > 0)
|
||
.sort();
|
||
};
|
||
|
||
let tmpRoot: string;
|
||
|
||
beforeEach(() => {
|
||
tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'pdg-sink-'));
|
||
});
|
||
|
||
afterEach(() => {
|
||
fs.rmSync(tmpRoot, { recursive: true, force: true });
|
||
});
|
||
|
||
describe('PdgEmitSink — routing', () => {
|
||
it('routes BasicBlock nodes and PDG edges to CSV, never to the real graph', () => {
|
||
const real = createKnowledgeGraph();
|
||
const sink = new PdgEmitSink(real, path.join(tmpRoot, 'pdg-csv'));
|
||
|
||
sink.addNode(bbNode('a.ts', 0, 1));
|
||
sink.addNode(bbNode('a.ts', 1, 5));
|
||
sink.addRelationship(pdgEdge('a.ts', 0, 1, 'CFG', 'seq'));
|
||
sink.addRelationship(pdgEdge('a.ts', 0, 1, 'REACHING_DEF', 'x:1:0'));
|
||
|
||
// PDG layer must not land in the in-memory graph (the RSS bound).
|
||
expect(real.nodeCount).toBe(0);
|
||
expect(real.relationshipCount).toBe(0);
|
||
expect(sink.nodeCount).toBe(0);
|
||
|
||
sink.finalize();
|
||
});
|
||
|
||
it('throws rather than silently dropping an undeclared endpoint-label pair (#2769)', () => {
|
||
// PDG edges are always BasicBlock|BasicBlock in production, but the sink
|
||
// applies no declared-pair gate of its own — this pins that the shared
|
||
// `assertDeclaredPair` guard (also wired into GraphEmitSink) still fires
|
||
// here rather than reaching COPY and silently dropping the edge.
|
||
const real = createKnowledgeGraph();
|
||
const sink = new PdgEmitSink(real, path.join(tmpRoot, 'pdg-csv'));
|
||
|
||
const undeclared: GraphRelationship = {
|
||
id: 'CFG:Static:a->Static:b',
|
||
sourceId: 'Static:src/a.ts:A',
|
||
targetId: 'Static:src/a.ts:B',
|
||
type: 'CFG',
|
||
confidence: 1,
|
||
reason: 'seq',
|
||
};
|
||
expect(() => sink.addRelationship(undeclared)).toThrow(
|
||
/Relationship label pair Static → Static is not declared/,
|
||
);
|
||
});
|
||
|
||
it('delegates structural nodes, CALLS, and the whole-program TAINT_PATH edge to the real graph', () => {
|
||
const real = createKnowledgeGraph();
|
||
const sink = new PdgEmitSink(real, path.join(tmpRoot, 'pdg-csv'));
|
||
|
||
sink.addNode({
|
||
id: 'Function:a.ts:fn:1',
|
||
label: 'Function',
|
||
properties: { name: 'fn', filePath: 'a.ts', startLine: 1, endLine: 9 },
|
||
});
|
||
sink.addRelationship({
|
||
id: 'CALLS:1',
|
||
sourceId: 'Function:a.ts:fn:1',
|
||
targetId: 'Function:a.ts:fn2:9',
|
||
type: 'CALLS',
|
||
confidence: 1,
|
||
reason: '',
|
||
});
|
||
// TAINT_PATH is a whole-program (Function→Function) edge — NOT streamed.
|
||
sink.addRelationship({
|
||
id: 'TAINT_PATH:1',
|
||
sourceId: 'Function:a.ts:fn:1',
|
||
targetId: 'Function:a.ts:fn2:9',
|
||
type: 'TAINT_PATH',
|
||
confidence: 0.9,
|
||
reason: 'src->sink',
|
||
});
|
||
|
||
expect(real.nodeCount).toBe(1);
|
||
expect(real.relationshipCount).toBe(2);
|
||
expect(real.getNode('Function:a.ts:fn:1')).toBeDefined();
|
||
|
||
const manifest = sink.finalize();
|
||
// No BasicBlock node CSV was created (no BasicBlock nodes were routed).
|
||
expect(manifest.nodeFiles.size).toBe(0);
|
||
expect(manifest.relsByPair.size).toBe(0);
|
||
});
|
||
});
|
||
|
||
describe('PdgEmitSink — byte-identity vs whole-graph emit', () => {
|
||
it('streamed CSV line set equals streamAllCSVsToDisk for the same nodes/edges', async () => {
|
||
const fp = 'a.ts';
|
||
const nodes = [bbNode(fp, 0, 1), bbNode(fp, 1, 5), bbNode(fp, 2, 9)];
|
||
const edges: GraphRelationship[] = [
|
||
pdgEdge(fp, 0, 1, 'CFG', 'seq'),
|
||
pdgEdge(fp, 1, 2, 'CFG', 'cond-true'),
|
||
pdgEdge(fp, 0, 2, 'REACHING_DEF', 'x:1:0'),
|
||
pdgEdge(fp, 1, 2, 'CDG', 'T'),
|
||
pdgEdge(fp, 0, 1, 'POST_DOMINATE', ''),
|
||
pdgEdge(fp, 0, 2, 'TAINTED', 'taint'),
|
||
pdgEdge(fp, 1, 2, 'SANITIZES', 'clean'),
|
||
];
|
||
|
||
// Whole-graph path: add to a plain graph, run streamAllCSVsToDisk.
|
||
const wholeGraph = createKnowledgeGraph();
|
||
for (const n of nodes) wholeGraph.addNode(n);
|
||
for (const e of edges) wholeGraph.addRelationship(e);
|
||
const wholeDir = path.join(tmpRoot, 'csv');
|
||
await streamAllCSVsToDisk(wholeGraph, path.join(tmpRoot, 'no-such-repo'), wholeDir);
|
||
|
||
// Streamed path: route the same set through the sink.
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
for (const n of nodes) sink.addNode(n);
|
||
for (const e of edges) sink.addRelationship(e);
|
||
const manifest = sink.finalize();
|
||
|
||
// BasicBlock node CSV: identical line set.
|
||
expect(await sortedLines(path.join(pdgDir, 'basicblock.csv'))).toEqual(
|
||
await sortedLines(path.join(wholeDir, 'basicblock.csv')),
|
||
);
|
||
|
||
// PDG edges all route to the BasicBlock|BasicBlock pair file: identical set.
|
||
expect(await sortedLines(path.join(pdgDir, 'rel_BasicBlock_BasicBlock.csv'))).toEqual(
|
||
await sortedLines(path.join(wholeDir, 'rel_BasicBlock_BasicBlock.csv')),
|
||
);
|
||
|
||
// Manifest reports the streamed files + row counts.
|
||
expect(manifest.nodeFiles.get('BasicBlock')?.rows).toBe(nodes.length);
|
||
expect(manifest.relsByPair.get('BasicBlock|BasicBlock')?.rows).toBe(edges.length);
|
||
});
|
||
|
||
it('emits rows via the shared builder (buildBasicBlockRow)', async () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
const n = bbNode('a.ts', 0, 3);
|
||
sink.addNode(n);
|
||
sink.finalize();
|
||
const lines = await sortedLines(path.join(pdgDir, 'basicblock.csv'));
|
||
// header + one data row; the data row is exactly buildBasicBlockRow(n).
|
||
expect(lines).toContain(buildBasicBlockRow(n));
|
||
});
|
||
});
|
||
|
||
describe('PdgEmitSink — bounded retention', () => {
|
||
it('flushes incrementally so the graph never holds the PDG layer', async () => {
|
||
const real = createKnowledgeGraph();
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const CHUNK = 2;
|
||
const sink = new PdgEmitSink(real, pdgDir, CHUNK); // tiny chunk to force flushes
|
||
|
||
// TOTAL is intentionally NOT a multiple of CHUNK so the final partial chunk
|
||
// is genuinely still buffered (unflushed) at the mid-stream read. With a
|
||
// multiple (e.g. 50 % 2 === 0) the last addRow's flush would have written
|
||
// every row and the "mid-stream" assertion would prove nothing (#2202
|
||
// review #7).
|
||
const TOTAL = 51;
|
||
const REMAINDER = TOTAL % CHUNK; // 1 — must be non-zero
|
||
expect(REMAINDER).toBeGreaterThan(0);
|
||
for (let i = 0; i < TOTAL; i++) sink.addNode(bbNode('a.ts', i, i));
|
||
|
||
// Mid-stream (before finalize): exactly the whole flushed chunks are on
|
||
// disk; the partial last chunk (REMAINDER rows) is still buffered in memory,
|
||
// proving the writer streams to the OS and never buffers the whole layer.
|
||
const midText = fs.readFileSync(path.join(pdgDir, 'basicblock.csv'), 'utf8');
|
||
const midDataRows = midText.split('\n').filter((l) => l.length > 0).length - 1; // minus header
|
||
expect(midDataRows).toBe(TOTAL - REMAINDER); // 50 flushed, 1 still buffered
|
||
expect(TOTAL - midDataRows).toBe(REMAINDER); // exactly the unflushed remainder
|
||
expect(TOTAL - midDataRows).toBeLessThanOrEqual(CHUNK); // unflushed is bounded by one chunk
|
||
|
||
// The in-memory graph never received a single BasicBlock.
|
||
expect(real.nodeCount).toBe(0);
|
||
|
||
const manifest = sink.finalize();
|
||
expect(manifest.nodeFiles.get('BasicBlock')?.rows).toBe(TOTAL);
|
||
const finalRows = (await sortedLines(path.join(pdgDir, 'basicblock.csv'))).length - 1;
|
||
expect(finalRows).toBe(TOTAL); // finalize flushed the buffered remainder
|
||
});
|
||
|
||
it('finalize twice throws', () => {
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), path.join(tmpRoot, 'pdg-csv'));
|
||
sink.finalize();
|
||
expect(() => sink.finalize()).toThrow(/twice/);
|
||
});
|
||
});
|
||
|
||
describe('PdgEmitSink — pass-through contract (dedup is the caller’s)', () => {
|
||
// The sink does NOT dedup by id (that would retain every id → O(total ids)
|
||
// memory, undermining the O(chunk) bound). Cross-pass dedup is done upstream,
|
||
// per file, in run.ts (a file imported by two language passes is emitted
|
||
// once). The sink is a faithful pass-through: it writes every id it is given
|
||
// and must not be fed duplicates. See #2202 finding #1 + the run-loop / Vue+TS
|
||
// integration coverage for the cross-pass dedup itself.
|
||
it('writes every BasicBlock it is given (no id dedup in the sink)', async () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
const n = bbNode('a.ts', 0, 1);
|
||
sink.addNode(n);
|
||
sink.addNode(n); // sink does not dedup — both rows are written
|
||
const manifest = sink.finalize();
|
||
expect(manifest.nodeFiles.get('BasicBlock')?.rows).toBe(2);
|
||
expect((await sortedLines(path.join(pdgDir, 'basicblock.csv'))).length - 1).toBe(2);
|
||
});
|
||
|
||
it('writes every PDG edge it is given (no id dedup in the sink)', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
const e = pdgEdge('a.ts', 0, 1, 'CFG', 'seq');
|
||
sink.addRelationship(e);
|
||
sink.addRelationship(e);
|
||
const manifest = sink.finalize();
|
||
expect(manifest.relsByPair.get('BasicBlock|BasicBlock')?.rows).toBe(2);
|
||
});
|
||
|
||
it('skips a PDG edge whose endpoint label is not a node table', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
// sourceId prefix "Bogus" is not in NODE_TABLES → skipped (mirrors RelPairRouter).
|
||
sink.addRelationship({
|
||
id: 'CFG:bogus',
|
||
sourceId: 'Bogus:a.ts:0',
|
||
targetId: 'BasicBlock:a.ts:1:0:1',
|
||
type: 'CFG',
|
||
confidence: 1,
|
||
reason: 'seq',
|
||
});
|
||
const manifest = sink.finalize();
|
||
expect(manifest.relsByPair.size).toBe(0);
|
||
});
|
||
});
|
||
|
||
describe('PdgEmitSink — IO failure poisoning (#2202 review #4/#6)', () => {
|
||
// A streamed-write failure is an IO fault, not the CFG-logic error the emit
|
||
// loop's per-file try/catch is built to swallow. The sink poisons the failing
|
||
// writer (or records an open failure) so finalize fails loudly instead of
|
||
// returning a truncated manifest that the bulk COPY would silently load.
|
||
|
||
afterEach(() => {
|
||
vi.restoreAllMocks();
|
||
});
|
||
|
||
it('poisons the writer on a mid-stream write failure (disk-full) → finalize throws', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir, 1); // flush every row
|
||
|
||
// The writer opens fine (real openSync); the flush's writeSync fails.
|
||
const spy = vi.spyOn(fs, 'writeSync').mockImplementation(() => {
|
||
throw new Error('ENOSPC: no space left on device');
|
||
});
|
||
// The throw propagates to the immediate caller (the emit loop, which would
|
||
// swallow it as a per-file CFG error) — that is exactly why finalize must
|
||
// re-check poison below.
|
||
expect(() => sink.addNode(bbNode('a.ts', 0, 1))).toThrow(/ENOSPC/);
|
||
spy.mockRestore();
|
||
|
||
expect(() => sink.finalize()).toThrow(/IO error|ENOSPC/);
|
||
});
|
||
|
||
it('records an openSync failure (EMFILE) → finalize throws even if the caller swallowed it', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv'); // ctor mkdir/rm run before the spy
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir);
|
||
|
||
const spy = vi.spyOn(fs, 'openSync').mockImplementation(() => {
|
||
throw new Error('EMFILE: too many open files');
|
||
});
|
||
expect(() => sink.addNode(bbNode('a.ts', 0, 1))).toThrow(/EMFILE/);
|
||
spy.mockRestore();
|
||
|
||
expect(() => sink.finalize()).toThrow(/IO error|EMFILE/);
|
||
});
|
||
|
||
it('surfaces a final-flush IO failure from finalize (rows buffered, never mid-flushed)', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir, 1000); // big chunk → no mid flush
|
||
sink.addNode(bbNode('a.ts', 0, 1)); // buffered only (1 < 1000)
|
||
|
||
// The only writeSync happens during the final flush inside finalize → close.
|
||
const spy = vi.spyOn(fs, 'writeSync').mockImplementation(() => {
|
||
throw new Error('ENOSPC: disk full at close');
|
||
});
|
||
// close() never throws (it records poison); finalize reports it.
|
||
expect(() => sink.finalize()).toThrow(/IO error|ENOSPC/);
|
||
spy.mockRestore();
|
||
});
|
||
|
||
it('a poisoned writer stops accepting rows (no unbounded buffering on a dead fd)', () => {
|
||
const pdgDir = path.join(tmpRoot, 'pdg-csv');
|
||
const sink = new PdgEmitSink(createKnowledgeGraph(), pdgDir, 1);
|
||
|
||
const spy = vi.spyOn(fs, 'writeSync').mockImplementation(() => {
|
||
throw new Error('ENOSPC');
|
||
});
|
||
expect(() => sink.addNode(bbNode('a.ts', 0, 1))).toThrow(/ENOSPC/); // poisons the writer
|
||
// Subsequent rows are dropped silently at the writer (it is dead) — they do
|
||
// not re-throw and do not accumulate; the run still fails at finalize.
|
||
expect(() => sink.addNode(bbNode('a.ts', 1, 2))).not.toThrow();
|
||
expect(() => sink.addNode(bbNode('a.ts', 2, 3))).not.toThrow();
|
||
spy.mockRestore();
|
||
|
||
expect(() => sink.finalize()).toThrow(/IO error|ENOSPC/);
|
||
});
|
||
});
|