mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-07 02:58:02 +00:00
* perf(lbug): add PROF_LBUG_LOAD persistence-path timing breakdown (#2203 U1) loadGraphToLbug is un-timed today; the analyze 'emit' number is the scope-resolution emit bucket, not the CSV->COPY persistence path. Add a zero-cost-when-off per-stage breakdown (csv-emit/copy-nodes/rel-split/ copy-rels/fallback/total + node/rel counts) gated by PROF_LBUG_LOAD=1, mirroring the PROF_SCOPE_RESOLUTION pattern. Document the flag in README. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * perf(lbug): route relationships to per-pair CSVs in the emit pass (#2203 U2) Relationships were written once to a monolithic relations.csv, then re-read line-by-line (regex per edge) and re-split into per-FROM->TO-label-pair files before COPY — writing and reading the entire ~1M-edge set twice. Route each edge to its pair file directly during the single emit pass via a shared RelPairRouter, eliminating the monolithic write + re-read + per-edge regex. The router applies the SAME getNodeLabel + validTables filter as the legacy splitRelCsvByLabelPair, which is retained as a differential oracle. A new differential test asserts the direct-emit per-pair files are byte-for-byte identical to the oracle's, with identical skip/total accounting. The prof line (U1) drops its rel-split stage (routing now folds into csv-emit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * perf(lbug): skip per-row microtask tick in BufferedCSVWriter (#2203 U3) addRow awaited an already-resolved promise on every buffered row, scheduling a microtask per node even when nothing flushed (millions at scale). It now returns a promise ONLY when it flushes; the node-emit loop awaits once per iteration after the switch. Flush/drain semantics are unchanged, so backpressure on the rows that actually write is preserved and the emitted CSV bytes are byte-identical (covered by the determinism + differential tests). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * bench(lbug): emit throughput + byte-identity gate for the persistence path (#2203 U4) Build-free bench (bench/emit-persistence/measure.mjs) times streamAllCSVsToDisk on a synthetic graph at two scales and gates: (1) an order-independent sha256 fingerprint over every emitted CSV line — the byte-identity guard for the U2/U3 emit optimisations — and (2) a scaling-ratio budget catching an O(n^2) emit re-regression. Wired into ci-tests.yml alongside the cfg/scope-capture benches. The LadybugDB COPY half needs a real DB, so its timing stays in PROF_LBUG_LOAD + the integration round-trip tests (documented in the bench README, with the deferred COPY-parallelism follow-up). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(review): apply autofix feedback (#2203) - P1: router backpressure drain-await rejected with a generic AbortError, masking the real EMFILE/disk-full error. Expose RelPairRouter.lastError and rethrow it in the emit catch — mirrors the oracle's throw streamError ?? err. - P1: cover RelPairRouter error + backpressure + teardown paths with a new unit test (test/unit/rel-pair-routing.test.ts) using an injected mock stream. - P2: wrap streamAllCSVsToDisk body in try/finally so the setMaxListeners bump is always restored (the U2 rel-routing throw path could leak it). - P2: dedup WriteStreamFactory — re-export the canonical type from rel-pair-routing instead of a second identical declaration. - P2: annotate splitRelCsvByLabelPair @internal as the retained differential oracle so a future dead-code sweep doesn't delete the byte-identity guard. - P3: differential test now covers the proc_ prefix + clears GITNEXUS_SORT_GRAPH_OUTPUT to prevent env-leak desync. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(lbug): scope byte-identity to quote-free ids + lock the quote-in-id divergence (#2215 review) The 'byte-identical' claim was unconditional, but the router derives labels from the raw id while the retained splitRelCsvByLabelPair oracle re-derives them via a regex over the escaped row — so for an id containing a double-quote they diverge (the router is the more-correct path). Soften the wording in rel-pair-routing.ts, the bench README, and the differential-test comment to document the exception, and add a differential test asserting the intended divergence (router routes the quote-in-id edge; oracle drops it) so a future change can't silently revert to the buggy regex semantics. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * bench(lbug): per-file fingerprint so the gate catches pair-file mis-routing (#2215 review) fingerprintEmit flattened every line of every per-pair file into one array, sorted globally, and hashed — losing file boundaries, so a row routed to the WRONG pair file produced an identical fingerprint. Hash a per-file digest (filename + sha256(file bytes)) and combine the sorted entry list, so mis-routing (and within-file row reordering) now changes the fingerprint. Baseline regenerated; the new scheme yields a different hash on byte-identical emit, confirming it is sensitive to file structure the old flatten ignored. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * bench(lbug): add absolute large-scale wall-time backstop to the emit gate (#2215 review) The scaling-ratio gate only compares large/small, so a uniform Nx slowdown at both scales passes with ratio ~1.0. Add an opt-in max_ms_large ceiling (1000ms vs observed ~200ms — generous, host-noise-tolerant) that --check enforces alongside the ratio, catching a gross absolute regression the ratio misses. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(lbug): cover the sorted-output path in the byte-identity differential (#2215 review) The differential test only exercised the default insertion-order emit path. Add a case under GITNEXUS_SORT_GRAPH_OUTPUT=1 that feeds the oracle the same id-sorted order orderedRelationships() uses and asserts per-pair byte-identity, so within-pair row reordering on the sorted path can't slip past the gate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(lbug): cover the invalid-TO-label skip branch (#2215 review) Only an invalid-FROM label was exercised; the validTables skip is an OR over both endpoints, so the invalid-TO branch was untested (an inverted && would have slipped through). Add a valid-FROM/invalid-TO edge to the differential test and the router unit test, asserting it's skipped identically by router and oracle. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(lbug): exercise the BufferedCSVWriter FLUSH_EVERY boundary in vitest (#2215 review) The U3 addRow change (returns a flush promise only on flush; undefined when buffered) and the loop's `if (pending) await pending` were only crossed by the bench, never vitest (all fixtures are <500 nodes). Add a 600-node graph through streamAllCSVsToDisk asserting all rows land exactly once across the 500-row flush boundary — no drops, dups, or corruption. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(lbug): drop redundant step cast in buildRelRow (#2215 review) GraphRelationship.step is already typed number?, so (rel as { step?: number }).step was a no-op structural cast that obscured the shared-type coupling. Use rel.step directly. Byte-identical — bench fingerprint unchanged, differential test green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(lbug): make the unknown-label node drop explicit (#2215 review) With the U3 `let pending` switch idiom, a node whose label matches neither codeWriterMap nor multiLangWriters left `pending` undefined and was silently dropped — a footgun for a future node type. Add an explicit else with a comment documenting that unknown labels are intentionally not persisted and that a new type must be wired into a writer map. No behavior change (byte-identity + tests unchanged). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(lbug): drop the unused WriteStreamFactory re-export (#2215 review) The type was re-exported from lbug-adapter 'to preserve this module's surface,' but no external code imports it by name from here (the only test reference is a comment). Keep the import from rel-pair-routing.ts (its canonical home, still used by splitRelCsvByLabelPair's signature) and drop the dead re-export. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
176 lines
6.1 KiB
TypeScript
176 lines
6.1 KiB
TypeScript
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
|
import { EventEmitter } from 'events';
|
|
import fs from 'fs';
|
|
import path from 'path';
|
|
import os from 'os';
|
|
import { RelPairRouter, getNodeLabel } from '../../src/core/lbug/rel-pair-routing.js';
|
|
|
|
/**
|
|
* Unit tests for RelPairRouter (#2203 U2) — the production per-pair emit path.
|
|
*
|
|
* Mirrors test/unit/rel-csv-split.test.ts: drives the router with an injected
|
|
* mock WriteStream factory so the error, backpressure, and teardown paths are
|
|
* exercised without LadybugDB or real disk streams. These paths are otherwise
|
|
* unreachable in the integration suite (which only hits the no-backpressure
|
|
* happy path), so this is the coverage for the router's failure modes.
|
|
*/
|
|
|
|
// Controllable backpressure + error injection (same shape as the split oracle's mock).
|
|
class MockWriteStream extends EventEmitter {
|
|
public chunks: string[] = [];
|
|
public destroyed = false;
|
|
public ended = false;
|
|
public blocked = false;
|
|
public maxDrainListenersSeen = 0;
|
|
// State flags + events so `stream/promises.finished(ws)` (used by the
|
|
// router's close()) resolves against this mock instead of hanging.
|
|
public writable = true;
|
|
public writableEnded = false;
|
|
public writableFinished = false;
|
|
|
|
write(chunk: string): boolean {
|
|
this.chunks.push(chunk);
|
|
const count = this.listenerCount('drain');
|
|
if (count > this.maxDrainListenersSeen) this.maxDrainListenersSeen = count;
|
|
return !this.blocked;
|
|
}
|
|
|
|
end(cb?: (err?: Error) => void): this {
|
|
this.ended = true;
|
|
this.writableEnded = true;
|
|
this.writableFinished = true;
|
|
this.writable = false;
|
|
if (cb) cb();
|
|
queueMicrotask(() => {
|
|
this.emit('finish');
|
|
this.emit('close');
|
|
});
|
|
return this;
|
|
}
|
|
|
|
destroy(): this {
|
|
this.destroyed = true;
|
|
return this;
|
|
}
|
|
|
|
unblock(): void {
|
|
this.blocked = false;
|
|
this.emit('drain');
|
|
}
|
|
|
|
triggerError(err: Error): void {
|
|
this.emit('error', err);
|
|
}
|
|
}
|
|
|
|
const HEADER = '"from","to","type","confidence","reason","step"';
|
|
const VALID = new Set<string>(['File', 'Function', 'Community', 'Process']);
|
|
|
|
const row = (from: string, to: string, type = 'CALLS'): string =>
|
|
`"${from}","${to}","${type}",1.0,"auto",0`;
|
|
|
|
function mockFactory(streams: MockWriteStream[], opts?: { blocked?: boolean }) {
|
|
return (() => {
|
|
const ws = new MockWriteStream();
|
|
if (opts?.blocked) ws.blocked = true;
|
|
streams.push(ws);
|
|
return ws;
|
|
}) as unknown as (filePath: string) => import('fs').WriteStream;
|
|
}
|
|
|
|
let tmpDir: string;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'rel-pair-routing-test-'));
|
|
});
|
|
|
|
afterEach(() => {
|
|
fs.rmSync(tmpDir, { recursive: true, force: true, maxRetries: 5, retryDelay: 50 });
|
|
});
|
|
|
|
describe('getNodeLabel', () => {
|
|
it('maps comm_/proc_ prefixes and otherwise splits on the first colon', () => {
|
|
expect(getNodeLabel('comm_42')).toBe('Community');
|
|
expect(getNodeLabel('proc_7')).toBe('Process');
|
|
expect(getNodeLabel('Function:src/a.ts:f:1')).toBe('Function');
|
|
expect(getNodeLabel('File:src/a.ts')).toBe('File');
|
|
});
|
|
});
|
|
|
|
describe('RelPairRouter', () => {
|
|
it('routes valid edges to per-pair files (header first) and skips invalid-label edges', async () => {
|
|
const streams: MockWriteStream[] = [];
|
|
const router = new RelPairRouter(tmpDir, HEADER, VALID, mockFactory(streams));
|
|
|
|
const route = async (from: string, to: string) => {
|
|
const p = router.route(from, to, row(from, to));
|
|
if (p) await p;
|
|
};
|
|
await route('File:a', 'Function:a:f:1');
|
|
await route('File:a', 'Function:a:g:2'); // same pair
|
|
await route('Function:a:f:1', 'Function:a:g:2'); // different pair
|
|
await route('Bogus:x', 'File:a'); // invalid FROM label → skipped
|
|
await route('File:a', 'Bogus:y'); // invalid TO label → skipped (other branch)
|
|
await router.close();
|
|
|
|
expect(router.skipped).toBe(2);
|
|
expect(router.total).toBe(3);
|
|
expect([...router.byPair.keys()].sort()).toEqual(['File|Function', 'Function|Function']);
|
|
expect(router.byPair.get('File|Function')!.rows).toBe(2);
|
|
// Header is the first chunk written to each pair stream.
|
|
expect(streams[0].chunks[0]).toBe(HEADER + '\n');
|
|
expect(streams.every((s) => s.ended)).toBe(true);
|
|
});
|
|
|
|
it('returns a drain promise under backpressure and completes once unblocked', async () => {
|
|
const streams: MockWriteStream[] = [];
|
|
const router = new RelPairRouter(
|
|
tmpDir,
|
|
HEADER,
|
|
VALID,
|
|
mockFactory(streams, { blocked: true }),
|
|
);
|
|
|
|
const pending = router.route('File:a', 'Function:a:f:1', row('File:a', 'Function:a:f:1'));
|
|
expect(pending).toBeInstanceOf(Promise); // header write hit backpressure
|
|
streams[0].unblock();
|
|
await pending;
|
|
|
|
expect(streams[0].maxDrainListenersSeen).toBeLessThanOrEqual(1);
|
|
expect(streams[0].chunks[0]).toBe(HEADER + '\n');
|
|
expect(router.total).toBe(1);
|
|
});
|
|
|
|
it('on a stream error: route() throws the real error, lastError exposes it, close() rejects + destroys', async () => {
|
|
const streams: MockWriteStream[] = [];
|
|
const router = new RelPairRouter(tmpDir, HEADER, VALID, mockFactory(streams));
|
|
|
|
const first = router.route('File:a', 'Function:a:f:1', row('File:a', 'Function:a:f:1'));
|
|
if (first) await first;
|
|
|
|
const err = new Error('EMFILE: too many open files');
|
|
streams[0].triggerError(err);
|
|
|
|
// The next route surfaces the REAL error, not a generic AbortError.
|
|
expect(() => router.route('File:a', 'Function:a:g:2', row('File:a', 'Function:a:g:2'))).toThrow(
|
|
'EMFILE',
|
|
);
|
|
expect(router.lastError).toBe(err);
|
|
await expect(router.close()).rejects.toThrow('EMFILE');
|
|
expect(streams[0].destroyed).toBe(true);
|
|
});
|
|
|
|
it('destroy() tears down every open pair stream', async () => {
|
|
const streams: MockWriteStream[] = [];
|
|
const router = new RelPairRouter(tmpDir, HEADER, VALID, mockFactory(streams));
|
|
|
|
const a = router.route('File:a', 'Function:a:f:1', row('File:a', 'Function:a:f:1'));
|
|
if (a) await a;
|
|
const b = router.route('Community:1', 'Community:2', row('Community:1', 'Community:2'));
|
|
if (b) await b;
|
|
|
|
router.destroy();
|
|
expect(streams.length).toBe(2);
|
|
expect(streams.every((s) => s.destroyed)).toBe(true);
|
|
});
|
|
});
|