test(bench): add wide-merge scenario + tighten deep-nest facts floor (#2201 review R7)

wide-merge: N bindings, each assigned in a 3-way branch (a wide multi-operand φ
per binding) inside a loop, then all used after the merge. Unlike dense-bindings
(one chained redef per `if`), every binding fans into its own wide φ, so the
scenario exercises φ-placement + renaming + the reachByScc condensation across
many independent wide merges. N bindings × constant arms ⇒ O(N) facts, so the
gate is rd_scaling LINEARITY (measured ~1.07; budget 2.0 catches a regression to
the per-binding-rescan O(N²) class the reachByScc alias path guards against). It
runs the production SSA path (10007 blocks + a loop) and computes all facts under
the blocks×64 budget (facts_large_min 24000 of a measured 26008 + the
rd_all_computed gate).

deep-nest: tighten facts_large_min 100 → 150 (measured 164) so a partial-
truncation regression that still cleared 100 — but lost facts — now fails, with
~9% headroom for noise.

bench --check PASS (9 scenarios) under --expose-gc; all existing CFG fingerprints
unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Gergo Magyar 2026-06-15 17:33:21 +00:00
parent 375dea99f5
commit d1d83ff60e
2 changed files with 38 additions and 2 deletions

View file

@ -37,8 +37,17 @@
"scaling_budget": 1.8,
"disk_bytes_budget": 1.2,
"rd_scaling_budget": 2.0,
"facts_large_min": 100,
"_note": "#2201: N nested loops carrying ONE variable end-to-end (depth 40->160) -- the pathology the dense worklist is superlinear on and whose block-visit total drives it past the blocks×64 ceiling (it would TRUNCATE to empty). rd is measured under the PRODUCTION blocks×64 budget (rdProductionBudget) to prove the ceiling stops firing: the depth-INDEPENDENT SSA solver (phi-nodes capture loop merges statically; no fixpoint iteration) computes the full facts (164 at large; floor 100 catches a regression to truncation/zero) with rd_scaling ~0.90 (linear in depth; measured ~0.42ms at depth 160). rd_scaling_budget 2.0 catches a regression back to superlinear. No heap_budget -- the deep-nest CFG payload is tiny and the retained-heap delta is GC-noise-dominated. Re-baseline the fingerprint only on an intentional CFG/visitor change."
"facts_large_min": 150,
"_note": "#2201: N nested loops carrying ONE variable end-to-end (depth 40->160) -- the pathology the dense worklist is superlinear on and whose block-visit total drives it past the blocks×64 ceiling (it would TRUNCATE to empty). rd is measured under the PRODUCTION blocks×64 budget (rdProductionBudget) to prove the ceiling stops firing: the depth-INDEPENDENT SSA solver (phi-nodes capture loop merges statically; no fixpoint iteration) computes the full facts (measured 164 at large) with rd_scaling ~0.68 (linear in depth; measured ~0.57ms at depth 160). facts_large_min tightened 100->150 (#2201 review R7): a partial-truncation regression that still cleared the old floor of 100 (but lost facts of the measured 164) now fails, with ~9% headroom under 164 for noise; the companion rd_all_computed gate also catches any non-'computed' status. rd_scaling_budget 2.0 catches a regression back to superlinear. No heap_budget -- the deep-nest CFG payload is tiny and the retained-heap delta is GC-noise-dominated. Re-baseline the fingerprint only on an intentional CFG/visitor change."
},
"wide-merge": {
"fingerprint": "7a66a844ee3994bd930c1e34bad3d7b410a762220e927c94fb3787f38d745280",
"scaling_budget": 1.8,
"disk_bytes_budget": 1.2,
"heap_budget": 1.3,
"rd_scaling_budget": 2.0,
"facts_large_min": 24000,
"_note": "#2201 review R7: N bindings, EACH assigned in a 3-way branch (a wide multi-operand phi per binding) inside a loop, then all used after the merge. Distinct from dense-bindings (one CHAINED redef per `if`): every binding fans into its OWN wide phi, so this exercises phi-placement + renaming + the reachByScc condensation across MANY independent wide merges. N bindings x constant arms => O(N) facts (measured 26008 at the large size), so the gate is rd_scaling LINEARITY: measured ~1.07 (time 9.3->39.8ms over the 4x size step); budget 2.0 catches a regression to the per-binding-rescan O(N^2) class -- the recurring solver antipattern the reachByScc alias fast path (review R2) guards against. rd is measured under the PRODUCTION blocks×64 budget (rdProductionBudget): all functions report 'computed' (the SSA path does not truncate here), and facts_large_min 24000 (measured 26008, ~7% headroom) + the rd_all_computed gate assert the wide merges compute fully. fp_blocks 82 / fp_edges 112 at FP_SIZE=15. Re-baseline the fingerprint only on an intentional CFG/harvest-shape change."
},
"fact-fanout": {
"fingerprint": "83a8243a8aff117f69aeecb39d02a483e6cca70439d75f63e433f4e4ac85578f",

View file

@ -198,6 +198,33 @@ const SCENARIOS = [
return s + ' return x;\n}\n';
},
},
{
name: 'wide-merge',
// #2201 review R7: N bindings, each assigned in a 3-way branch (a WIDE φ
// merge per binding) inside a loop, then all used after the merge. Unlike
// dense-bindings (one chained redef per `if`), every binding here fans into
// its own multi-operand φ — so the scenario stresses φ-placement + renaming +
// the reachByScc condensation across MANY independent wide merges. N bindings
// × constant arms ⇒ O(N) facts, so the gate is rd_scaling LINEARITY: a
// regression to the per-binding-rescan class (O(N²), the recurring solver
// antipattern reachByScc's alias fast path guards against) blows the ratio.
// >=16 blocks + a reachable loop ⇒ the production SSA path.
rdMaxFacts: 0, // measure the algorithm, not the cap
rdProductionBudget: true, // prove the SSA path computes under blocks×64
gen: (n) => {
let s = 'function f(c: number) {\n';
for (let i = 0; i < n; i++) s += ` let v${i} = ${i};\n`;
s += ' while (c > 0) {\n';
for (let i = 0; i < n; i++) {
s +=
` if (c > ${i}) { v${i} = ${i} + c; }` +
` else if (c < ${i}) { v${i} = ${i} - c; }` +
` else { v${i} = c; }\n`;
}
for (let i = 0; i < n; i++) s += ` use(v${i});\n`;
return s + ' c = c - 1;\n }\n return v0;\n}\n';
},
},
{
name: 'fact-fanout',
// #2082 M2: N parallel case-arm defs of one variable + N later uses —