perf(ingestion): add memory + disk growth gates to the CFG benchmark (#2081)

Extend bench/cfg/measure.mjs beyond wall-time to the two other scalability
dimensions that matter at kernel scale:

- DISK growth: utf8 byte size of the serialized cfgSideChannel — exactly what a
  --pdg run writes onto every ParsedFile shard (durable store + parse cache).
- MEMORY growth: retained JS heap of the cfgSideChannel payload, measured by the
  release-delta method (heap held minus heap after dropping it) — robust to
  pre-existing garbage and dead-stable run-to-run. Needs `node --expose-gc`;
  without it the heap metric is null and its gate is skipped (local runs still
  work). ci-tests.yml now passes --expose-gc so the heap gate runs in CI.

Both gated on linear scaling in baselines.json (disk_bytes_budget / heap_budget
1.2-1.3). Measured: disk ~1.0-1.04, retained heap ~0.87-1.0 — both linear
(~1KB/function each; ~2MB heap / 1.6MB disk at 2000 functions, --pdg only).
Bumped REPS 7->15 to stabilize the noisier time signal and widened the coarse
time tripwire budgets (the disk/heap gates carry the tight regression detection).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Gergo Magyar 2026-06-09 05:28:17 +00:00
parent e6f04fae7b
commit ebfb62c1a9
3 changed files with 96 additions and 36 deletions

View file

@ -265,12 +265,14 @@ jobs:
run: node --import tsx bench/scope-capture/measure.mjs --check
working-directory: gitnexus
- name: CFG construction fingerprint + scaling guards (#2081 M1)
- name: CFG construction time / disk / memory guards (#2081 M1)
# Build-free: asserts collectFunctionCfgs output is unchanged
# (fingerprint) and stays sub-quadratic for the straight-line /
# many-functions / branchy scenarios. Catches an O(n^2) re-regression
# in the per-function CFG builder (e.g. an extendBlock concat chain).
run: node --import tsx bench/cfg/measure.mjs --check
# (fingerprint) and that wall-time, cfgSideChannel disk bytes, AND
# retained heap all stay sub-quadratic for the straight-line /
# many-functions / branchy scenarios. Catches an O(n^2) re-regression in
# the per-function CFG builder (e.g. an extendBlock concat chain) and a
# memory/disk blow-up. --expose-gc enables the retained-heap measurement.
run: node --expose-gc --import tsx bench/cfg/measure.mjs --check
working-directory: gitnexus
- name: Cross-language pipeline benchmarks (GITNEXUS_BENCH, serial)

View file

@ -1,20 +1,23 @@
{
"straight-line": {
"fingerprint": "f5524690b5b7d484573710938c5e9a28e08ef0882fea95111f01575c71f4a66a",
"scaling_budget": 1.8,
"bytes_budget": 1.2,
"_note": "#2081 M1: ONE function, N coalescing statements (extendBlock path). Time ~1.25-1.39 after the O(n)-fragment-join fix; bytes ~1.03 (linear). Budget 1.8 catches a return to ~4.0 quadratic; small absolute time → noisier, hence the wider time budget. Re-baseline the fingerprint only on an intentional CFG-shape change."
"scaling_budget": 2.0,
"disk_bytes_budget": 1.2,
"heap_budget": 1.3,
"_note": "#2081 M1: ONE function, N coalescing statements (extendBlock path). Time ~1.4-1.5 after the O(n)-fragment-join fix; disk ~1.03, retained heap ~0.87 (both linear/sub-linear). Time budget 2.0 is a coarse O(n²) tripwire (a true quadratic is ~4.0) with CI-noise headroom; the disk/heap budgets carry the tight regression detection. Re-baseline the fingerprint only on an intentional CFG-shape change."
},
"many-functions": {
"fingerprint": "c167ccd83086254e2b71eca153ca4a833be14b2d2a3827ab76b49f643aad13d5",
"scaling_budget": 1.5,
"bytes_budget": 1.2,
"_note": "#2081 M1: N small branchy functions (collect walk + per-function build). Time ~1.0 (clean linear); bytes ~1.01."
"disk_bytes_budget": 1.2,
"heap_budget": 1.3,
"_note": "#2081 M1: N small branchy functions (collect walk + per-function build). Time ~1.0, disk ~1.01, retained heap ~1.0 (~1KB/function; ~2MB at 2000 fns)."
},
"branchy": {
"fingerprint": "944ab56ffc70e195f74d8533a8aadf4930d37d13bcfa47cc4feff29e74ddca5c",
"scaling_budget": 1.6,
"bytes_budget": 1.2,
"_note": "#2081 M1: ONE function, N sequential ifs (block/edge growth in one CFG). Time ~1.1; bytes ~1.04."
"scaling_budget": 1.8,
"disk_bytes_budget": 1.2,
"heap_budget": 1.3,
"_note": "#2081 M1: ONE function, N sequential ifs (block/edge growth in one CFG). Time ~1.1-1.25 (REPS=15 median; noisiest scenario), disk ~1.04, retained heap ~1.0. Time budget 1.8 absorbs noise while catching ~4.0 quadratic."
}
}

View file

@ -11,24 +11,30 @@
* - `branchy`: ONE function with N sequential `if`s — stresses block/edge
* growth within a single CFG.
*
* For each scenario it reports `elapsed_ms` at small/large and a scaling ratio
* `(t_large/t_small)/(N_large/N_small)`: ~1.0 is linear, ~4.0 is quadratic (the
* O(n²) shape the M1 perf review flagged for `extendBlock`'s concat chain).
* For each scenario it reports three scaling ratios at small→large
* (`(metric_large/metric_small)/(N_large/N_small)`: ~1.0 is linear, ~4.0 is the
* O(n²) shape the M1 perf review flagged for `extendBlock`'s concat chain):
* - TIME — wall-clock of `collectFunctionCfgs` (median of reps);
* - DISK — utf8 byte size of the serialized `cfgSideChannel` (what a `--pdg`
* run writes onto every ParsedFile shard);
* - MEMORY — retained JS heap of the `cfgSideChannel` payload, by the
* release-delta method (heap held minus heap after dropping it). Requires
* `node --expose-gc`; without it the heap metric is null and its gate skips.
* It also computes an order-independent sha256 fingerprint over the emitted
* blocks/edges of a fixed-size source — the correctness gate that a structural
* speedup must leave behavior-identical.
*
* Build-free: imports the `.ts` hotpaths through tsx
* (`node --import tsx bench/cfg/measure.mjs`). Parsing happens ONCE per size and
* the tree is reused across reps so the measurement isolates CFG build cost, not
* tree-sitter parse time. `maxFunctionLines` is 0 (no cap) here on purpose — the
* bench measures the algorithm; the production default cap is a separate safety
* net (and would otherwise skip the large straight-line function).
* (`node --expose-gc --import tsx bench/cfg/measure.mjs`). Parsing happens ONCE
* per size and the tree is reused across reps so the time measurement isolates
* CFG build cost, not tree-sitter parse time. `maxFunctionLines` is 0 (no cap)
* here on purpose — the bench measures the algorithm; the production default cap
* is a separate safety net (and would otherwise skip the large straight-line fn).
*
* Without args: prints one JSON object per scenario.
* With `--check`: asserts each scenario's fingerprint == its committed baseline
* (baselines.json) AND scaling_ratio < its recorded budget; exits non-zero on
* any drift/regression.
* (baselines.json) AND each of the time / disk / heap ratios is below its
* recorded budget; exits non-zero on any drift/regression.
*/
import fs from 'node:fs';
import path from 'node:path';
@ -90,7 +96,7 @@ const SCENARIOS = [
const SMALL = 500;
const LARGE = 2000; // 4× — O(n) ⇒ ratio ~1, O(n²) ⇒ ratio ~4
const REPS = 7;
const REPS = 15; // median over more reps → stabler time signal at small absolute ms
const FP_SIZE = 15; // fixed size for the behavior fingerprint
const NO_CAP = 0; // measure the algorithm, not the production safety cap
@ -115,14 +121,40 @@ function measureCollect(src, file, reps) {
return {
ms: median(samples),
blockCount: out.cfgs.reduce((a, c) => a + c.blocks.length, 0),
// Serialized size of the cfgSideChannel payload — what rides on every
// ParsedFile through the disk store + parse cache. Should scale linearly
// with source (O(source covered)); a super-linear bytes ratio means the
// CFG is duplicating text and will bloat warm-cache shards at scale.
bytes: JSON.stringify(out.cfgs).length,
// DISK growth: utf8 byte size of the serialized cfgSideChannel — exactly
// what a --pdg run writes onto every ParsedFile shard in the durable store
// + parse cache (the field is plain JSON, so this is the on-disk delta).
// Should scale linearly with source covered; a super-linear ratio means the
// CFG duplicates text and bloats warm-cache shards at scale.
diskBytes: Buffer.byteLength(JSON.stringify(out.cfgs), 'utf8'),
};
}
// ---- memory growth: retained heap of the cfgSideChannel payload ----
// Needs `node --expose-gc` to force collection for a clean delta; without it the
// heap metric is reported as null and its --check gate is skipped (so a local
// run without the flag still works).
const GC = typeof global.gc === 'function' ? () => (global.gc(), global.gc()) : null;
function retainedHeapBytes(src, file) {
if (!GC) return null;
// Retained-size-by-RELEASE: measure the heap with the CFGs held, drop them,
// GC, measure again. The drop isolates exactly the JS heap the cfgSideChannel
// payload retains (the extra RAM a --pdg run carries per file until the shard
// is flushed) — robust to pre-existing garbage, which is constant across both
// measurements. The parse tree is a temporary (its native memory isn't on the
// JS heap); block text strings are fresh copies, so they count here.
let cfgs = collectFunctionCfgs(parse(src).rootNode, visitor, file, NO_CAP).cfgs;
GC();
const withCfgs = process.memoryUsage().heapUsed;
if (cfgs.length < 0) throw new Error('unreachable'); // keep cfgs live past withCfgs
cfgs = null;
GC();
const withoutCfgs = process.memoryUsage().heapUsed;
return Math.max(0, withCfgs - withoutCfgs);
}
// ---- correctness fingerprint (order-independent over blocks + edges) ----
function canonicalizeCfg(cfg) {
@ -149,15 +181,27 @@ function measureScenario(scenario) {
const large = measureCollect(scenario.gen(LARGE), `${scenario.name}.ts`, REPS);
const sizeRatio = LARGE / SMALL;
const scalingRatio = small.ms > 0 ? large.ms / small.ms / sizeRatio : 0;
const bytesRatio = small.bytes > 0 ? large.bytes / small.bytes / sizeRatio : 0;
const diskRatio = small.diskBytes > 0 ? large.diskBytes / small.diskBytes / sizeRatio : 0;
// Memory growth (only when --expose-gc gave us a forced GC).
const heapSmall = retainedHeapBytes(scenario.gen(SMALL), `${scenario.name}.ts`);
const heapLarge = retainedHeapBytes(scenario.gen(LARGE), `${scenario.name}.ts`);
const heapRatio =
heapSmall !== null && heapLarge !== null && heapSmall > 0
? heapLarge / heapSmall / sizeRatio
: null;
return {
scenario: scenario.name,
elapsed_ms_small: Number(small.ms.toFixed(3)),
elapsed_ms_large: Number(large.ms.toFixed(3)),
scaling_ratio: Number(scalingRatio.toFixed(3)),
bytes_small: small.bytes,
bytes_large: large.bytes,
bytes_ratio: Number(bytesRatio.toFixed(3)),
disk_bytes_small: small.diskBytes,
disk_bytes_large: large.diskBytes,
disk_bytes_ratio: Number(diskRatio.toFixed(3)),
heap_bytes_small: heapSmall,
heap_bytes_large: heapLarge,
heap_ratio: heapRatio === null ? null : Number(heapRatio.toFixed(3)),
blocks_small: small.blockCount,
blocks_large: large.blockCount,
...fingerprint(scenario),
@ -191,10 +235,21 @@ if (!CHECK) {
`(${SMALL}->${LARGE} stmts/fns, ms ${r.elapsed_ms_small}->${r.elapsed_ms_large})`,
);
}
if (base.bytes_budget !== undefined && r.bytes_ratio >= base.bytes_budget) {
if (base.disk_bytes_budget !== undefined && r.disk_bytes_ratio >= base.disk_bytes_budget) {
failures.push(
`${r.scenario}: cfgSideChannel bytes ratio ${r.bytes_ratio} >= budget ${base.bytes_budget} ` +
`(bytes ${r.bytes_small}->${r.bytes_large})`,
`${r.scenario}: cfgSideChannel disk-bytes ratio ${r.disk_bytes_ratio} >= budget ` +
`${base.disk_bytes_budget} (bytes ${r.disk_bytes_small}->${r.disk_bytes_large})`,
);
}
// Heap gate only when measured (--expose-gc present) AND a budget exists.
if (
base.heap_budget !== undefined &&
r.heap_ratio !== null &&
r.heap_ratio >= base.heap_budget
) {
failures.push(
`${r.scenario}: retained-heap ratio ${r.heap_ratio} >= budget ${base.heap_budget} ` +
`(heap ${r.heap_bytes_small}->${r.heap_bytes_large})`,
);
}
process.stdout.write(JSON.stringify(r) + '\n');