GitNexus/gitnexus/bench/emit-persistence/measure.mjs
Gergő Magyar df08ecc397
perf(lbug): cut graph-DB emit/persistence wall time (#2203) (#2215)
* perf(lbug): add PROF_LBUG_LOAD persistence-path timing breakdown (#2203 U1)

loadGraphToLbug is un-timed today; the analyze 'emit' number is the
scope-resolution emit bucket, not the CSV->COPY persistence path. Add a
zero-cost-when-off per-stage breakdown (csv-emit/copy-nodes/rel-split/
copy-rels/fallback/total + node/rel counts) gated by PROF_LBUG_LOAD=1,
mirroring the PROF_SCOPE_RESOLUTION pattern. Document the flag in README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(lbug): route relationships to per-pair CSVs in the emit pass (#2203 U2)

Relationships were written once to a monolithic relations.csv, then re-read
line-by-line (regex per edge) and re-split into per-FROM->TO-label-pair files
before COPY — writing and reading the entire ~1M-edge set twice. Route each
edge to its pair file directly during the single emit pass via a shared
RelPairRouter, eliminating the monolithic write + re-read + per-edge regex.

The router applies the SAME getNodeLabel + validTables filter as the legacy
splitRelCsvByLabelPair, which is retained as a differential oracle. A new
differential test asserts the direct-emit per-pair files are byte-for-byte
identical to the oracle's, with identical skip/total accounting. The prof
line (U1) drops its rel-split stage (routing now folds into csv-emit).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(lbug): skip per-row microtask tick in BufferedCSVWriter (#2203 U3)

addRow awaited an already-resolved promise on every buffered row, scheduling
a microtask per node even when nothing flushed (millions at scale). It now
returns a promise ONLY when it flushes; the node-emit loop awaits once per
iteration after the switch. Flush/drain semantics are unchanged, so
backpressure on the rows that actually write is preserved and the emitted
CSV bytes are byte-identical (covered by the determinism + differential tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): emit throughput + byte-identity gate for the persistence path (#2203 U4)

Build-free bench (bench/emit-persistence/measure.mjs) times streamAllCSVsToDisk
on a synthetic graph at two scales and gates: (1) an order-independent sha256
fingerprint over every emitted CSV line — the byte-identity guard for the U2/U3
emit optimisations — and (2) a scaling-ratio budget catching an O(n^2) emit
re-regression. Wired into ci-tests.yml alongside the cfg/scope-capture benches.
The LadybugDB COPY half needs a real DB, so its timing stays in PROF_LBUG_LOAD
+ the integration round-trip tests (documented in the bench README, with the
deferred COPY-parallelism follow-up).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback (#2203)

- P1: router backpressure drain-await rejected with a generic AbortError,
  masking the real EMFILE/disk-full error. Expose RelPairRouter.lastError and
  rethrow it in the emit catch — mirrors the oracle's throw streamError ?? err.
- P1: cover RelPairRouter error + backpressure + teardown paths with a new
  unit test (test/unit/rel-pair-routing.test.ts) using an injected mock stream.
- P2: wrap streamAllCSVsToDisk body in try/finally so the setMaxListeners bump
  is always restored (the U2 rel-routing throw path could leak it).
- P2: dedup WriteStreamFactory — re-export the canonical type from
  rel-pair-routing instead of a second identical declaration.
- P2: annotate splitRelCsvByLabelPair @internal as the retained differential
  oracle so a future dead-code sweep doesn't delete the byte-identity guard.
- P3: differential test now covers the proc_ prefix + clears
  GITNEXUS_SORT_GRAPH_OUTPUT to prevent env-leak desync.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(lbug): scope byte-identity to quote-free ids + lock the quote-in-id divergence (#2215 review)

The 'byte-identical' claim was unconditional, but the router derives labels from
the raw id while the retained splitRelCsvByLabelPair oracle re-derives them via a
regex over the escaped row — so for an id containing a double-quote they diverge
(the router is the more-correct path). Soften the wording in rel-pair-routing.ts,
the bench README, and the differential-test comment to document the exception,
and add a differential test asserting the intended divergence (router routes the
quote-in-id edge; oracle drops it) so a future change can't silently revert to
the buggy regex semantics.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): per-file fingerprint so the gate catches pair-file mis-routing (#2215 review)

fingerprintEmit flattened every line of every per-pair file into one array,
sorted globally, and hashed — losing file boundaries, so a row routed to the
WRONG pair file produced an identical fingerprint. Hash a per-file digest
(filename + sha256(file bytes)) and combine the sorted entry list, so mis-routing
(and within-file row reordering) now changes the fingerprint. Baseline
regenerated; the new scheme yields a different hash on byte-identical emit,
confirming it is sensitive to file structure the old flatten ignored.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): add absolute large-scale wall-time backstop to the emit gate (#2215 review)

The scaling-ratio gate only compares large/small, so a uniform Nx slowdown at
both scales passes with ratio ~1.0. Add an opt-in max_ms_large ceiling (1000ms
vs observed ~200ms — generous, host-noise-tolerant) that --check enforces
alongside the ratio, catching a gross absolute regression the ratio misses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): cover the sorted-output path in the byte-identity differential (#2215 review)

The differential test only exercised the default insertion-order emit path. Add
a case under GITNEXUS_SORT_GRAPH_OUTPUT=1 that feeds the oracle the same
id-sorted order orderedRelationships() uses and asserts per-pair byte-identity,
so within-pair row reordering on the sorted path can't slip past the gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): cover the invalid-TO-label skip branch (#2215 review)

Only an invalid-FROM label was exercised; the validTables skip is an OR over
both endpoints, so the invalid-TO branch was untested (an inverted && would
have slipped through). Add a valid-FROM/invalid-TO edge to the differential
test and the router unit test, asserting it's skipped identically by router and
oracle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): exercise the BufferedCSVWriter FLUSH_EVERY boundary in vitest (#2215 review)

The U3 addRow change (returns a flush promise only on flush; undefined when
buffered) and the loop's `if (pending) await pending` were only crossed by the
bench, never vitest (all fixtures are <500 nodes). Add a 600-node graph through
streamAllCSVsToDisk asserting all rows land exactly once across the 500-row
flush boundary — no drops, dups, or corruption.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): drop redundant step cast in buildRelRow (#2215 review)

GraphRelationship.step is already typed number?, so (rel as { step?: number }).step
was a no-op structural cast that obscured the shared-type coupling. Use rel.step
directly. Byte-identical — bench fingerprint unchanged, differential test green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): make the unknown-label node drop explicit (#2215 review)

With the U3 `let pending` switch idiom, a node whose label matches neither
codeWriterMap nor multiLangWriters left `pending` undefined and was silently
dropped — a footgun for a future node type. Add an explicit else with a comment
documenting that unknown labels are intentionally not persisted and that a new
type must be wired into a writer map. No behavior change (byte-identity + tests
unchanged).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): drop the unused WriteStreamFactory re-export (#2215 review)

The type was re-exported from lbug-adapter 'to preserve this module's surface,'
but no external code imports it by name from here (the only test reference is a
comment). Keep the import from rel-pair-routing.ts (its canonical home, still
used by splitRelCsvByLabelPair's signature) and drop the dead re-export.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 19:40:59 +01:00

216 lines
8 KiB
JavaScript

/**
* Build-free emit-path throughput + byte-identity bench for the graph-DB
* persistence pipeline (issue #2203).
*
* Measures `streamAllCSVsToDisk` — the CSV-generation half of the persistence
* path that U2 (direct per-pair relationship routing) and U3 (per-row
* microtask elimination) optimised. The LadybugDB `COPY` half needs a real DB
* connection, so its timing lives in the runtime `PROF_LBUG_LOAD` breakdown +
* the integration round-trip tests, NOT here (see README.md).
*
* For a synthetic KnowledgeGraph at two scales it reports:
* - elapsed_ms_small / elapsed_ms_large (median over REPS) + a scaling ratio
* `(t_large/t_small)/(LARGE/SMALL)`: ~1.0 linear, ~3.x quadratic;
* - an order-independent sha256 fingerprint over every emitted CSV line
* (node CSVs + per-FROM→TO-label-pair rel CSVs), as the byte-identity gate
* guarding the issue's "byte-identical graph content" requirement.
*
* Build-free: imports the `.ts` hotpaths through tsx
* (`node --import tsx bench/emit-persistence/measure.mjs`). Static `.ts`
* imports work; a top-level `await import()` breaks tsx's lexer.
*
* Without args: prints one JSON object per scenario.
* With `--check`: asserts the fingerprint == the committed baseline AND the
* scaling ratio < the recorded budget; exits non-zero on drift/regression.
*/
import fs from 'node:fs';
import fsp from 'node:fs/promises';
import os from 'node:os';
import path from 'node:path';
import crypto from 'node:crypto';
import { fileURLToPath } from 'node:url';
import { createKnowledgeGraph } from '../../src/core/graph/graph.ts';
import { streamAllCSVsToDisk } from '../../src/core/lbug/csv-generator.ts';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const BASELINE_PATH = path.resolve(__dirname, 'baselines.json');
// ---- synthetic graph generation (deterministic — no randomness) ----
/**
* Build a graph of `entityCount` files, each with 2 functions + 1 class and 4
* relationships. ids carry valid table-label prefixes (File:/Function:/Class:)
* so edges route to real pairs (File→Function, File→Class, Function→Function)
* — exercising the U2 router across multiple label pairs. Node `content` is
* never populated (no backing files), so emit cost reflects the CSV machinery
* (routing, escaping, buffering, disk writes), not content extraction.
*/
function generateGraph(entityCount) {
const graph = createKnowledgeGraph();
for (let i = 0; i < entityCount; i++) {
const fp = `src/e${i}.ts`;
const fileId = `File:${fp}`;
const fnA = `Function:${fp}:fnA:1`;
const fnB = `Function:${fp}:fnB:10`;
const cls = `Class:${fp}:C:20`;
graph.addNode({ id: fileId, label: 'File', properties: { name: `e${i}.ts`, filePath: fp } });
graph.addNode({
id: fnA,
label: 'Function',
properties: { name: 'fnA', filePath: fp, startLine: 1, endLine: 5, isExported: true },
});
graph.addNode({
id: fnB,
label: 'Function',
properties: { name: 'fnB', filePath: fp, startLine: 10, endLine: 15, isExported: false },
});
graph.addNode({
id: cls,
label: 'Class',
properties: { name: 'C', filePath: fp, startLine: 20, endLine: 30, isExported: true },
});
graph.addRelationship({
id: `${fileId}->${fnA}`,
sourceId: fileId,
targetId: fnA,
type: 'CONTAINS',
confidence: 1,
reason: '',
});
graph.addRelationship({
id: `${fileId}->${fnB}`,
sourceId: fileId,
targetId: fnB,
type: 'CONTAINS',
confidence: 1,
reason: '',
});
graph.addRelationship({
id: `${fileId}->${cls}`,
sourceId: fileId,
targetId: cls,
type: 'CONTAINS',
confidence: 1,
reason: '',
});
graph.addRelationship({
id: `${fnA}->${fnB}`,
sourceId: fnA,
targetId: fnB,
type: 'CALLS',
confidence: 1,
reason: '',
});
}
return graph;
}
// ---- byte-identity fingerprint (order-independent) ----
/** sha256 over every non-empty line of every emitted CSV file, sorted so the
* digest is a pure function of the emitted line SET (insertion-order agnostic). */
async function fingerprintEmit(graph, dir) {
await streamAllCSVsToDisk(graph, path.join(dir, 'no-such-repo'), dir);
// Per-file digest bound to the filename: a row routed to the WRONG pair file
// (or a header written to the wrong file) changes the fingerprint — a global
// line-flatten could not catch that. File bytes are hashed as-written (so it
// also catches within-file row reordering); the entry list is sorted so
// readdir order doesn't matter.
const entries = [];
for (const name of fs.readdirSync(dir)) {
if (!name.endsWith('.csv')) continue;
const bytes = await fsp.readFile(path.join(dir, name));
entries.push(`${name}\n${crypto.createHash('sha256').update(bytes).digest('hex')}`);
}
return crypto.createHash('sha256').update(entries.sort().join('\n')).digest('hex');
}
// ---- timing ----
function median(xs) {
const s = [...xs].sort((a, b) => a - b);
const m = Math.floor(s.length / 2);
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
}
async function timeEmit(graph, dir, reps) {
await streamAllCSVsToDisk(graph, path.join(dir, 'no-such-repo'), dir); // warmup (not counted)
const samples = [];
for (let i = 0; i < reps; i++) {
const start = process.hrtime.bigint();
await streamAllCSVsToDisk(graph, path.join(dir, 'no-such-repo'), dir);
samples.push(Number(process.hrtime.bigint() - start) / 1e6);
}
return median(samples);
}
const SMALL = 600;
const LARGE = 2400;
const REPS = 5;
async function measure() {
const tmpRoot = path.join(os.tmpdir(), `gitnexus-emit-bench-${process.pid}`);
await fsp.mkdir(tmpRoot, { recursive: true });
try {
const smallGraph = generateGraph(SMALL);
const largeGraph = generateGraph(LARGE);
const fingerprint = await fingerprintEmit(largeGraph, path.join(tmpRoot, 'fp'));
const small = await timeEmit(smallGraph, path.join(tmpRoot, 'small'), REPS);
const large = await timeEmit(largeGraph, path.join(tmpRoot, 'large'), REPS);
const scalingRatio = small > 0 ? large / small / (LARGE / SMALL) : 0;
return {
scenario: 'streamAllCSVsToDisk',
entities_small: SMALL,
entities_large: LARGE,
nodes_large: LARGE * 4,
rels_large: LARGE * 4,
elapsed_ms_small: Number(small.toFixed(2)),
elapsed_ms_large: Number(large.toFixed(2)),
scaling_ratio: Number(scalingRatio.toFixed(3)),
fingerprint,
};
} finally {
await fsp.rm(tmpRoot, { recursive: true, force: true }).catch(() => {});
}
}
// ---- run ----
const CHECK = process.argv.includes('--check');
const result = await measure();
if (!CHECK) {
process.stdout.write(JSON.stringify(result) + '\n');
} else {
const base = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf8'));
const failures = [];
if (result.fingerprint !== base.fingerprint) {
failures.push(
`byte-identity fingerprint drift (got ${result.fingerprint}, expected ${base.fingerprint})`,
);
}
if (result.scaling_ratio >= base.scaling_budget) {
failures.push(
`scaling ratio ${result.scaling_ratio} >= budget ${base.scaling_budget} ` +
`(${SMALL}->${LARGE} entities, ms ${result.elapsed_ms_small}->${result.elapsed_ms_large})`,
);
}
// Absolute backstop: the scaling ratio alone passes a uniform Nx slowdown (it
// only compares large/small). A generous, host-noise-tolerant ceiling catches
// a gross absolute regression. Opt-in (only enforced when max_ms_large is set).
if (base.max_ms_large !== undefined && result.elapsed_ms_large >= base.max_ms_large) {
failures.push(
`absolute wall-time regression: elapsed_ms_large ${result.elapsed_ms_large}ms >= budget ` +
`${base.max_ms_large}ms (coarse backstop, not a tight SLA)`,
);
}
process.stdout.write(JSON.stringify(result) + '\n');
if (failures.length > 0) {
for (const f of failures) process.stderr.write(`[emit-persistence --check] FAIL: ${f}\n`);
process.exit(1);
}
process.stderr.write('[emit-persistence --check] PASS\n');
}