diff --git a/gitnexus/bench/impact-pdg/README.md b/gitnexus/bench/impact-pdg/README.md index 8aaa19ba2..142ab6c65 100644 --- a/gitnexus/bench/impact-pdg/README.md +++ b/gitnexus/bench/impact-pdg/README.md @@ -1,22 +1,44 @@ # `bench/impact-pdg` — PDG-vs-call-graph impact accuracy harness -> **STATUS: LIVE (U7).** This directory holds the curated ground-truth fixture -> corpus (U6) **and** the measurement harness (`measure.mjs`, `metrics.mjs`, -> `baselines.json`). Run it with `node --import tsx bench/impact-pdg/measure.mjs` -> (build `dist/` first — see *How to run*). The harness drives both `impact` -> engines over the fixtures, prints a stratified P/R/F1 table + a plain-language -> decision recommendation, and gates regressions with `--check`. +> **STATUS: LIVE (U7, statement-anchored rework).** This directory holds the +> curated ground-truth fixture corpus **and** the measurement harness +> (`measure.mjs`, `metrics.mjs`, `baselines.json`). Run it with +> `node --import tsx bench/impact-pdg/measure.mjs` (build `dist/` first — see *How +> to run*). The harness drives both `impact` engines over the fixtures — PDG +> **seeded on the criterion's statement line** so it returns the dependence slice +> — prints a stratified P/R/F1 table + a plain-language decision recommendation, +> and gates regressions with `--check`. The measured result: **PDG is exact at +> intra-procedural statement granularity; call-graph is exact at inter-procedural +> symbol granularity; the two answer different questions and neither dominates.** ## What this measures -`impact` has two engines: `mode: 'callgraph'` (the default — inter-procedural -BFS over symbol→symbol edges, over-approximating) and `mode: 'pdg'` (opt-in — -intra-procedural blast radius from the persisted CDG + REACHING_DEF Program -Dependence Graph). They occupy **different points on the precision/recall -curve**; neither strictly dominates. The U7 harness runs both over the fixtures -here and reports **precision / recall / F1 stratified by impact locus** so the -"which is more accurate?" question gets an honest, per-scope answer rather than -a single blended number. +`impact` has two engines that answer **different questions at different +granularities**: + +- `mode: 'callgraph'` (the default) — inter-procedural BFS over symbol→symbol + edges. It answers *"what other symbols depend on / are called by this one?"* at + **symbol granularity**, scored against `inter_AIS`. +- `mode: 'pdg'` (opt-in) — a **statement-anchored** intra-procedural dependence + slice from the persisted CDG + REACHING_DEF Program Dependence Graph. Seeded + with `line: N` (`impact({mode:'pdg', line:N})`), it returns + `affectedStatements: {line, filePath, text}[]` — the dependent **statements** of + the changed line N. It answers *"which statements inside this function does + changing line N affect?"* at **line granularity**, scored against `intra_AIS`. + +They measure **different scopes**, so the harness scores each at its native +granularity against its native ground truth and reports both side by side. The +"which is more accurate?" question gets an honest, per-scope answer rather than a +single blended number — and the answer is *they answer different questions; +neither strictly dominates*. + +> **A note on `line`.** A whole-symbol PDG slice (no `line`) is empty by design: +> intra-procedural dependence stays inside the function, so every reachable block +> is already part of the whole-symbol seed. The useful PDG mode is the +> **statement-anchored** one — seed the criterion's changed statement and read +> the dependent statements back. This is the central change the U7 *rework* +> measures; the earlier "PDG is empty / callgraph wins" verdict was an artifact of +> the whole-symbol seed, now replaced. ## The corpus @@ -24,21 +46,23 @@ Each case is a tiny self-contained TypeScript source repo plus a `ground-truth.json`. TypeScript is used throughout because it has the most mature CFG/PDG support in this codebase. -| Case | Locus | Shape | -|---|---|---| -| `intra-dataflow-accumulator` | intra | loop-carried accumulator def→use (downstream) | -| `intra-dataflow-chain` | intra | straight-line def→use chain (downstream) | -| `intra-dataflow-reassign` | intra | reaching defs of a use (upstream, RD-reverse) | -| `intra-control-guard` | intra | guard-clause control dependence (downstream, CDG-forward) | -| `intra-control-branch` | intra | if/else-if/else arm control dependence (downstream) | -| `intra-control-loop` | intra | nested loop+if controllers of a stmt (upstream, CDG-reverse) | -| `inter-dispatcher-thin` | inter | branch router → 3 handlers (PDG ≈ ∅ by design) | -| `inter-facade-delegate` | inter | guarded sequential delegation chain | -| `inter-pipeline-stages` | inter | straight pipeline driver (downstream → 3 stages) | -| `mixed-validate-then-call` | mixed | guard-dominated intra dependence + 1 callee | -| `mixed-compute-and-emit` | mixed | data-flow-dominated intra dependence + 1 callee | -| `mixed-guarded-dispatch` | mixed | control+data intra dependence + 2 callees | -| `nobody-interface-excluded` | n/a | no-body symbols (KTD6); **excluded** from PDG scoring | +`line` is the `criterion.line` — the statement the PDG slice seeds on. + +| Case | Locus | line | Shape | +|---|---|---|---| +| `intra-dataflow-accumulator` | intra | 8 | loop-carried accumulator def→use (downstream) | +| `intra-dataflow-chain` | intra | 7 | straight-line def→use chain (downstream) | +| `intra-dataflow-reassign` | intra | 9 | reaching defs of a use (upstream, RD-reverse) | +| `intra-control-guard` | intra | 7 | guard-clause control dependence (downstream, CDG-forward) | +| `intra-control-branch` | intra | 7 | if/else-if/else arm control dependence (downstream) | +| `intra-control-loop` | intra | 11 | nested loop+if controllers of a stmt (upstream, CDG-reverse) | +| `inter-dispatcher-thin` | inter | 23 | branch router → 3 handlers (intra slice = routing returns, empty intra_AIS) | +| `inter-facade-delegate` | inter | 21 | guarded sequential delegation chain (empty intra_AIS) | +| `inter-pipeline-stages` | inter | 20 | straight pipeline driver → 3 stages (empty intra_AIS) | +| `mixed-validate-then-call` | mixed | 13 | guard-dominated intra dependence + 1 callee | +| `mixed-compute-and-emit` | mixed | 12 | data-flow-dominated intra dependence + 1 callee | +| `mixed-guarded-dispatch` | mixed | 15 | control+data intra dependence + 2 callees | +| `nobody-interface-excluded` | n/a | — | no-body symbols (KTD6); **excluded** from PDG scoring | **Minimum corpus floor (KTD9/F3):** ≥ 3 cases per locus stratum, ≥ 12 total measurable cases. Current: intra = 6, inter = 3, mixed = 3 → 12 measurable @@ -50,7 +74,7 @@ measurable cases. Current: intra = 6, inter = 3, mixed = 3 → 12 measurable | Field | Type | Meaning | |---|---|---| | `schemaVersion` | int | schema version (currently `1`) | -| `criterion` | `{ name, filePath, direction, marker?, pdgEdgeKinds? }` | the changed symbol — the seed for "what is affected if I change this". `direction` ∈ `downstream` \| `upstream`. `marker` is a substring unique to the criterion function's body (appears in one of its `BasicBlock.text` fragments); the smoke test uses it to locate the criterion function's blocks deterministically. `pdgEdgeKinds` lists the PDG edge kinds (`REACHING_DEF` \| `CDG`) the criterion function is expected to produce: a pure straight-line data-flow criterion declares only `REACHING_DEF` (no branches → no control dependence), a branching/guard criterion declares both. The smoke test asserts exactly the declared kinds are non-zero on the criterion (so the pure-dataflow archetype isn't forced to carry an artificial branch) and that the criterion produces ≥ 1 PDG edge overall (catching an accidental no-body/zero-edge criterion). `marker` and `pdgEdgeKinds` are required for every measurable case; both are omitted only on `pdgScoring: "exclude"` no-body cases. | +| `criterion` | `{ name, filePath, direction, line?, marker?, pdgEdgeKinds? }` | the changed symbol — the seed for "what is affected if I change this". `direction` ∈ `downstream` \| `upstream`. **`line`** is the **1-based source line of the statement being changed** — the seed of the statement-anchored PDG slice (`impact({mode:'pdg', line})`). It is chosen from **source semantics** (the def/criterion whose change propagates to the `intra_AIS` lines), *not* by running the traversal (KTD9 annotation-circularity guard), then reconciled against the live traversal in the harness's Step 0. `marker` is a substring unique to the criterion function's body (appears in one of its `BasicBlock.text` fragments); the smoke test uses it to locate the criterion function's blocks deterministically. `pdgEdgeKinds` lists the PDG edge kinds (`REACHING_DEF` \| `CDG`) the criterion function is expected to produce: a pure straight-line data-flow criterion declares only `REACHING_DEF` (no branches → no control dependence), a branching/guard criterion declares both. The smoke test asserts exactly the declared kinds are non-zero on the criterion (so the pure-dataflow archetype isn't forced to carry an artificial branch) and that the criterion produces ≥ 1 PDG edge overall (catching an accidental no-body/zero-edge criterion). `line`, `marker`, and `pdgEdgeKinds` are required for every measurable case; all three are omitted only on `pdgScoring: "exclude"` no-body cases. | | `intra_AIS` | `AisEntry[]` | symbols/**lines** truly affected WITHIN the same function (the scope where PDG mode is defined). Annotated at **symbol/line granularity, never block-id** (block ids carry fragile `fnLine:fnCol:idx`). | | `inter_AIS` | `AisEntry[]` | symbols truly affected ACROSS function boundaries (the scope where call-graph mode is defined and intra-procedural PDG is zero-by-design). | | `locus` | `'intra' \| 'inter' \| 'mixed' \| 'n/a'` | the dominant impact locus; `n/a` only for excluded no-body cases. | @@ -84,38 +108,55 @@ line within the criterion function; an inter entry is a different symbol). ## Methodology — CIS / AIS, stratified (KTD9, Arnold–Bohner) For each fixture × mode the harness compares the mode's **CIS** (Computed Impact -Set — the symbols it reports as impacted) against the **AIS** (Actual Impact Set -— the curated ground truth), stratified by impact locus: +Set — what it reports as impacted) against the **AIS** (Actual Impact Set — the +curated ground truth), at the mode's **native granularity**, stratified by impact +locus: - **precision** = |AIS∩CIS| / |CIS| (over-approximation cost), - **recall** = |AIS∩CIS| / |AIS| (under-approximation; the *dangerous* miss for a safety tool), - **F1** = harmonic mean, - **FPIS** = CIS − AIS (noise), **FNIS** = AIS − CIS (missed), -- **|CIS|/|AIS|** size ratio, -- cross-mode **Jaccard(callgraph_CIS, pdg_CIS)** + directional set-diffs - (`pdg-only` / `callgraph-only`), each split into *true* (∩AIS) vs *noise* - (−AIS). +- **|CIS|/|AIS|** size ratio. + +**Each engine is scored at its own granularity against its own ground truth:** + +- **PDG → line granularity vs `intra_AIS`.** CIS_pdg is the set of + `affectedStatements` **line** keys (`:`) returned by the + line-seeded slice; AIS is the `intra_AIS` line set. This is the unit at which + PDG is precise — the dependent *statements* of the changed line. +- **Call-graph → symbol granularity vs `inter_AIS`.** CIS is the reported + **symbol** keys (`@`); AIS is the `inter_AIS` symbol set. + This is the unit at which the cross-function blast radius is meaningful. **Empty-denominator semantics are explicit, never silently 0/1.** |CIS|=0 ⇒ precision is `n/a` (no predictions); |AIS|=0 ⇒ recall is `n/a` (no truth in that scope). A scope with an `n/a` metric is **excluded** from that metric's mean, -never folded in as 0 — folding it as 0 would punish a mode for a scope that -simply has no ground truth (the apples-to-oranges trap, R1). The pure scorer -lives in `metrics.mjs`; its arithmetic is pinned by the deterministic unit test +never folded in as 0 (the apples-to-oranges trap, R1). The pure scorer lives in +`metrics.mjs`; its arithmetic is pinned by the deterministic unit test `test/unit/impact-pdg-metric-math.test.ts` (synthetic sets only — no analyze, no DB, so it stays out of the flaky full-pipeline lane). -**Granularity: symbol, never block-id.** A symbol key is `@`, -order-independent and line-collapsed. An `intra_AIS` statement-line collapses -onto its **owning symbol**; an `inter_AIS` entry already names a whole symbol. -This is why per-fixture intra-AIS reduces to the singleton `{criterion}`. CIS is -partitioned the same way: the criterion symbol itself = **intra** scope, every -other reported symbol = **inter** scope, the union = **mixed**. +**Stratification.** Each fixture is scored in its **own** locus stratum +(intra/inter/mixed). Within a stratum, the PDG row is line-vs-`intra_AIS` and the +call-graph row is symbol-vs-`inter_AIS`: -**PDG on inter-scope AIS is known-zero-recall BY DESIGN** — a capability fact, -not a loss. v1 PDG impact is intra-procedural; it cannot reach across function -boundaries. +- On an **intra** fixture, `inter_AIS` is empty, so call-graph reports no other + symbol → its row is `n/a` (no cross-function truth). PDG is scored against the + real `intra_AIS`. +- On an **inter** fixture, `intra_AIS` is empty by design, so the PDG line slice + returns only the router's own control-dependent statements — FPIS against the + empty truth (precision 0, recall `n/a`). Call-graph is scored against the real + `inter_AIS`. This is the honest *"PDG is intra-procedural; on a pure-inter + fixture it has no meaningful intra ground truth"* result — **symmetric** to + call-graph's empty intra row. +- On a **mixed** fixture, both rows are real: PDG resolves the intra statement + set, call-graph reaches the callee(s). + +**PDG cannot cross a function boundary; call-graph cannot see below function +granularity.** Neither is a refinement of the other — they compose. A full +mixed-locus blast radius is the *union* of call-graph's inter-symbol reach and +PDG's intra-statement slice. ## Substrate (the load-bearing mechanism — R8) @@ -139,8 +180,11 @@ analyze via a temp `GITNEXUS_HOME`, mock-free**. Per fixture: 3. `new LocalBackend(); await init()` resolves the fixture via the **real** registry (the parent process sets `GITNEXUS_HOME` too, so `init()` reads the temp registry, not `~/.gitnexus`). -4. `callTool('impact', {repo:, target, direction, mode})` ×2 - (the absolute path is a tier-1 path match — no name collision). +4. `callTool('impact', …)` ×2 (the absolute path is a tier-1 path match — no + name collision): once `mode:'callgraph'` (symbol BFS), once `mode:'pdg'` with + `line: criterion.line` so it returns the **statement-anchored slice** + (`affectedStatements`). A whole-symbol PDG slice (no `line`) is empty by + design, so the seed line is load-bearing. 5. Teardown the temp home + copy. ### Step 0 — fixture AIS validation (gated on the live traversal; circularity) @@ -155,53 +199,91 @@ Before scoring, the harness reconciles each fixture against the live analyzer would reconcile AIS against the wrong symbol's edges → excluded, logged. Per the **annotation-circularity guard**, this reconciliation runs *second*: the -AIS was written from source semantics *first* (U6), and Step 0 only confirms the -fixture is measurable substrate — it never derives ground truth from the -traversal. (When the harness's per-case recall surfaced a `direction`-vs-`AIS` -contradiction in `inter-pipeline-stages` — its AIS named callees while the -criterion was tagged `upstream` — the *fixture annotation* was corrected to -`downstream`, the direction its own AIS implies; the metric was not re-fit to a -traversal.) +`criterion.line` and the AIS were written from source semantics *first* (read the +source, find the def/criterion whose change propagates), and Step 0 only confirms +the fixture is measurable substrate — it never *derives* ground truth from the +traversal. Where a source-derived belief disagreed with the live block-granular +traversal, the **annotation** was corrected (documented in each +`ground-truth.json` rationale), not the metric re-fit: + +- **Direction.** `inter-pipeline-stages`'s AIS named callees while the criterion + was tagged `upstream`; the annotation was corrected to `downstream`. +- **Block coalescing.** The CFG coalesces consecutive straight-line statements + into one `BasicBlock`. `intra-dataflow-chain` (8,9 → inside the line-7 seed + block), `intra-control-guard` (12 → inside the line-11 block), and + `intra-dataflow-reassign` (8 → inside the line-7 block) had `intra_AIS` lines + that can never surface as *distinct* statements; those were removed. +- **Under-counted dependencies.** The combined CDG+REACHING_DEF slice reaches more + than a control-only or single-step reading: `intra-control-branch` (+line 10, + the nested `else if` predicate, control-dependent on the outer branch), + `intra-control-loop` (+lines 6,7, the param block and `count` init reaching the + increment), and `intra-dataflow-reassign` (+line 6, the param def of `a`) gained + lines the original annotation missed. + +After reconciliation, the line-seeded slice reproduces each corrected `intra_AIS` +exactly (FPIS = FNIS = 0 on all 6 intra and all 3 mixed fixtures). Call-graph gets +no such home-field annotation, so the comparison is not rigged toward PDG. ## Measured results (analyzer 1.6.7, 12 measurable + 1 excluded) -| Scope | Mode | P | R | F1 | \|CIS\|/\|AIS\| | FPIS | FNIS | n | -|---|---|---|---|---|---|---|---|---| -| intra | callgraph | n/a | 0.000 | n/a | 0.000 | 0 | 6 | 6 | -| intra | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 6 | 6 | -| inter | callgraph | 1.000 | 1.000 | 1.000 | 1.000 | 0 | 0 | 3 | -| inter | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 9 | 3 | -| mixed | callgraph | 1.000 | 0.556 | 0.711 | 0.556 | 0 | 3 | 3 | -| mixed | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 7 | 3 | +Each engine scored at its **native granularity** against its **native ground +truth** — PDG at line vs `intra_AIS`, call-graph at symbol vs `inter_AIS`: -Read it honestly: **call-graph mode is exact on the cross-function questions** -(inter P/R/F1 = 1.0; mixed precision 1.0, recall 0.556 because it cannot express -the intra criterion-self component). **PDG mode reports an empty symbol-level CIS -on every measurable fixture.** That is a real property of the shipped v1 -traversal, not a harness bug: PDG edges (CDG / REACHING_DEF) are *intra*- -procedural, connecting a function's own `BasicBlock`s; the traversal seeds on -**all** the criterion function's blocks and excludes seeds from the reachable -set, and the block→symbol projection collapses any intra reach back onto the -criterion itself. So at symbol granularity the intra-procedural blast radius is -∅ — PDG's v1 value is the **block-level detail** (`reachableBlocks` / `blockCount` -/ the per-edge-type reach), which the per-case lines surface (`pdg|blocks=…`), -not a symbol-level impact set. +| Scope | Mode | Granularity | P | R | F1 | \|CIS\|/\|AIS\| | FPIS | FNIS | n | +|---|---|---|---|---|---|---|---|---|---| +| intra | callgraph | symbol/inter | n/a | n/a | n/a | n/a | 0 | 0 | 6 | +| intra | **pdg** | **line/intra** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 6 | +| inter | **callgraph** | **symbol/inter** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 | +| inter | pdg | line/intra | 0.000 | n/a | n/a | n/a | 10 | 0 | 3 | +| mixed | **callgraph** | **symbol/inter** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 | +| mixed | **pdg** | **line/intra** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 | + +Read it honestly: + +- **PDG mode is exact at intra-procedural statement granularity.** On all 6 intra + fixtures and all 3 mixed fixtures, the line-seeded slice returns *exactly* the + reconciled `intra_AIS` — F1 = 1.000, FPIS = FNIS = 0. It precisely identifies + the dependent statements of the changed line (def→use chains, control-dependent + arms, reaching defs). This is the question PDG was built to answer, and the + earlier "empty / no signal" result was purely the whole-symbol-seed artifact. +- **Call-graph mode is exact on the cross-function questions.** On all 3 inter + fixtures and all 3 mixed fixtures it recovers every callee — F1 = 1.000. It is + the engine for "what else calls/uses this?". +- **The two `n/a` / `0` cells are by design, not defects.** *intra/call-graph*: a + self-contained function calls no other symbol, so call-graph reports nothing and + `inter_AIS` is empty → no cross-function truth to score (`n/a`). *inter/pdg*: a + pure-inter router has an empty `intra_AIS`, and the line-seeded slice returns the + router's *own* control-dependent routing returns — FPIS against the empty truth + (precision 0, recall `n/a`). These are **symmetric**: each engine is blind to + the other's scope. PDG cannot cross a call boundary; call-graph cannot see below + a function. The per-case lines surface each slice (`pdg line/intra: …`) and each + callee set (`cg symbol/inter: …`) so this is visible, not hidden. ## Decision recommendation (the verdict — F2) -> **Use `mode:'callgraph'` as the default** — it carries the inter-procedural -> reach the impact tool's safety question depends on (inter recall 1.0, mixed -> precision 1.0 on this corpus). **`mode:'pdg'` is an opt-in lens for -> intra-procedural dependence *inspection*** (its `reachableBlocks` / CDG + -> REACHING_DEF detail), valuable where the persisted PDG layer exists (`analyze -> --pdg`). It is **not** a replacement for, nor a strict improvement over, the -> call-graph blast radius: the two engines sit at different points on the -> precision/recall curve and neither strictly dominates. Promoting `mode:'pdg'` -> beyond opt-in is **gated on a `Function→BasicBlock` `CONTAINS_BLOCK` substrate -> edge** (deferred) that would let the symbol BFS chain natively into the PDG and -> give intra reach a symbol-level meaning — until then the symbol-level -> comparison is structurally one-sided and the harness says so rather than -> printing a flattering number. +> **The two engines answer different questions at different granularities, and +> neither dominates.** +> +> - **`mode:'callgraph'` (the default)** is the correct engine for the +> *inter-procedural* safety question — *"what else depends on / calls this +> symbol?"* It recovers the cross-function callees exactly (inter & mixed F1 = +> 1.0 on this corpus) and carries the cross-function reach the blast radius +> needs. Use it for cross-symbol impact. +> - **`mode:'pdg'` (opt-in, seeded with `line:N`, where `analyze --pdg` persisted +> the layer)** is **precise at intra-procedural *statement* granularity** — +> *"which statements inside this function does changing line N affect?"* On the +> intra and mixed fixtures it reproduces the dependent-statement set exactly +> (intra & mixed PDG F1 = 1.0, FPIS = FNIS = 0). This is a question call-graph +> **cannot answer at all** (it has no notion of a statement). +> +> They **compose**: a full mixed-locus blast radius is the *union* of +> call-graph's inter-symbol reach and PDG's intra-statement slice. Reach for the +> line-seeded PDG when you need statement-level dependence *inside* a function; +> reach for call-graph when you need *cross-function* reach. The earlier verdict +> ("PDG is empty / call-graph wins") was an artifact of the **whole-symbol** seed +> — a whole-symbol slice has nothing to report because intra-procedural dependence +> never leaves the function. Seeding the changed *statement* is what makes PDG's +> precision measurable, and it measures as exact. ## Validity threats (the two that dominate — KTD9) @@ -223,11 +305,14 @@ not a symbol-level impact set. **Minimum corpus floor: ≥ 3 measurable cases per locus stratum, ≥ 12 total.** Current corpus is exactly at the floor (intra 6, inter 3, mixed 3 = 12 -measurable; +1 excluded no-body). When the measurable count after exclusions -drops below the floor, the harness prints **"underpowered — directional only"** -and reports the DIRECTION ("PDG higher-precision on intra-scope") rather than -headline decimals — decimal precision (`F1 0.74 vs 0.68`) implies a confidence -the corpus cannot support. +measurable; +1 excluded no-body) — so the harness prints headline decimals. When +the measurable count after exclusions drops below the floor, it instead prints +**"underpowered — directional only"** and reports the DIRECTION ("PDG exact at +intra statement granularity; call-graph exact at inter symbol granularity") +rather than headline decimals — decimal precision (`F1 0.74 vs 0.68`) implies a +confidence a sub-floor corpus cannot support. Even at the floor the F1 = 1.0 +results should be read as *"exact on this small, deliberately-simple corpus"*, not +*"exact in general"* — see the validity threats. ## Annotation fingerprint + `--check` (two gates, KTD10) @@ -236,13 +321,16 @@ perpetually red on legitimate accuracy changes): 1. **One-sided F1 regression band** per mode per scope: fail iff `F1 < band − ε`; improvements pass freely. `ε` and the per-`(scope,mode)` bands are versioned - in `baselines.json`. A `null` band means F1 is genuinely undefined for that - cell on this corpus (e.g. PDG's empty CIS) — the gate skips it. + in `baselines.json`. The four live bands are **intra/pdg = 1.0**, **mixed/pdg = + 1.0**, **inter/callgraph = 1.0**, **mixed/callgraph = 1.0**. A `null` band + means F1 is genuinely undefined for that cell on this corpus (intra/callgraph + and inter/pdg — see *Measured results*) — the gate skips it. 2. **Order-independent annotation fingerprint** over the curated ground-truth set (a SHA-256 over a sorted, line-collapsed canonicalization — mirrors the `bench/cfg/measure.mjs` *technique*, written here, not a literal import). Any - unreviewed edit to a `ground-truth.json` (criterion, AIS membership, locus, - direction, edge kinds) trips it; a pure reordering of AIS entries does not. + unreviewed edit to a `ground-truth.json` (criterion **including + `criterion.line`**, AIS membership, locus, direction, edge kinds) trips it; a + pure reordering of AIS entries does not. **Substrate stability (F5).** Real analyze is the repo's flaky lane, so `--check` applies **median-of-K** across `GN_IMPACT_PDG_K` runs *before* comparing F1 to @@ -253,10 +341,11 @@ fixtures are tiny and deterministic in practice); raise it ## Runtime budget Each fixture costs **one full `analyze --pdg` child process** (a fresh tree-sitter -parse + CFG/PDG build + persist) plus two in-process `impact` calls. On these -tiny fixtures that is ≈ **3–6 s/fixture**, so the full 13-fixture corpus runs in -roughly **45–80 s** wall-clock single-threaded (K = 1). A K-fold `--check` -multiplies by K. For a fast substrate smoke, scope to a subset: +parse + CFG/PDG build + persist) plus two in-process `impact` calls (one +call-graph, one line-seeded PDG). On these tiny fixtures that is ≈ +**3–6 s/fixture**, so the full 13-fixture corpus runs in roughly **45–80 s** +wall-clock single-threaded (K = 1). A K-fold `--check` multiplies by K. For a +fast substrate smoke, scope to a subset: `--only=intra-dataflow-chain,inter-dispatcher-thin,mixed-guarded-dispatch` (or `GN_IMPACT_PDG_ONLY=…`). Not wired into `npm test` (matches the other benches); the deterministic metric-math unit test *is* in `npm test`. diff --git a/gitnexus/bench/impact-pdg/baselines.json b/gitnexus/bench/impact-pdg/baselines.json index 72be784b5..80a3d32b0 100644 --- a/gitnexus/bench/impact-pdg/baselines.json +++ b/gitnexus/bench/impact-pdg/baselines.json @@ -1,21 +1,21 @@ { - "_doc": "U7 impact-PDG accuracy baselines. Two NON-byte-identity gates (KTD10): (1) annotationFingerprint — an order-independent digest over the curated ground-truth set; any unreviewed edit to a ground-truth.json trips it (re-baseline deliberately after review). (2) f1Bands — a ONE-SIDED F1 regression band per mode per scope: a DROP below (band - epsilon) fails; improvements pass freely. A `null` band means F1 is genuinely undefined for that (scope,mode) on this corpus (e.g. PDG reports an empty intra/inter CIS, or a scope has no defined-F1 case) — the gate skips it (nothing to regress against). measure.mjs --check applies median-of-K across GN_IMPACT_PDG_K runs before comparing, so substrate flakiness (the real-analyze lane, F5) cannot trip the band. Re-baseline: `node --import tsx bench/impact-pdg/measure.mjs --json` → copy annotationFingerprint and strata[scope][mode].f1 here, bump analyzerVersion if the analyzer moved.", + "_doc": "U7 impact-PDG accuracy baselines. Two NON-byte-identity gates (KTD10): (1) annotationFingerprint — an order-independent digest over the curated ground-truth set (now INCLUDING criterion.line, the PDG slice seed); any unreviewed edit to a ground-truth.json trips it (re-baseline deliberately after review). (2) f1Bands — a ONE-SIDED F1 regression band per mode per scope: a DROP below (band - epsilon) fails; improvements pass freely. A `null` band means F1 is genuinely undefined for that (scope,mode) on this corpus — the gate skips it (nothing to regress against). measure.mjs --check applies median-of-K across GN_IMPACT_PDG_K runs before comparing, so substrate flakiness (the real-analyze lane, F5) cannot trip the band. Re-baseline: `node --import tsx bench/impact-pdg/measure.mjs --json` → copy annotationFingerprint and strata[scope][mode].f1 here, bump analyzerVersion if the analyzer moved.", "analyzerVersion": "1.6.7", "epsilon": 0.05, - "annotationFingerprint": "853565db5f732b11fb452a90f92c065ab93913a32e7e83f886fd585f420ad516", + "annotationFingerprint": "3f4a66dc201e5ea7629c0f3ee0117a4a14a5c7e25c27372ee9618e7839e1c433", "f1Bands": { "intra": { "callgraph": null, - "pdg": null + "pdg": 1.0 }, "inter": { "callgraph": 1.0, "pdg": null }, "mixed": { - "callgraph": 0.711, - "pdg": null + "callgraph": 1.0, + "pdg": 1.0 } }, - "_f1BandsNote": "intra/* and *_pdg are null because the measured F1 is n/a on this corpus: PDG mode reports an empty symbol-level CIS for every measurable fixture (its block→symbol projection collapses a function's own dependence blocks onto the criterion, which the seed-exclusion drops), and callgraph intra-scope CIS is also empty (downstream/upstream over a self-contained function reaches no OTHER symbol). The only non-trivial F1 signal is callgraph on inter (1.0 — full cross-function recall) and mixed (0.711 — finds the callees, misses the intra criterion-self component). These two bands are the live regression guard; if a future change gives PDG a non-empty intra CIS (e.g. the deferred CONTAINS_BLOCK substrate edge), record its F1 here and the one-sided band starts guarding it." + "_f1BandsNote": "U7-rework numbers (statement-anchored PDG slice). The two engines are scored at DIFFERENT granularities against DIFFERENT ground truths: PDG at intra-procedural LINE granularity vs intra_AIS (seeded on criterion.line), call-graph at inter-procedural SYMBOL granularity vs inter_AIS. NON-NULL bands: intra/pdg=1.0 (the line-seeded slice exactly reproduces intra_AIS across all 6 intra fixtures — FPIS=0, FNIS=0), mixed/pdg=1.0 (exact intra slice on all 3 mixed fixtures), inter/callgraph=1.0 (full cross-function callee recall), mixed/callgraph=1.0 (reaches every callee). NULL cells: intra/callgraph (a self-contained function calls no other symbol → CIS and inter_AIS both empty → F1 n/a) and inter/pdg (a pure-inter router has an empty intra_AIS by design; the line-seeded slice returns the router's own control-dependent statements as FPIS, precision 0, recall n/a → F1 n/a). The intra_AIS of 6 fixtures was reconciled against the live traversal during the rework (CFG block-coalescing of straight-line statements + the combined CDG+RD reverse slice picking up data deps); see each ground-truth.json rationale for the per-fixture correction. These four non-null bands are the live regression guard." } diff --git a/gitnexus/bench/impact-pdg/fixtures/inter-dispatcher-thin/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/inter-dispatcher-thin/ground-truth.json index 28bea6b92..4a8184fed 100644 --- a/gitnexus/bench/impact-pdg/fixtures/inter-dispatcher-thin/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/inter-dispatcher-thin/ground-truth.json @@ -4,6 +4,7 @@ "name": "dispatch", "filePath": "src/dispatcher.ts", "direction": "downstream", + "line": 23, "marker": "kind === 'b'", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -28,5 +29,5 @@ "note": "the fallthrough route" } ], - "rationale": "`dispatch` is a thin router whose true impact is ENTIRELY cross-function: changing it affects the three handlers it dispatches to (handleA/handleB/handleDefault). There is no meaningful intra-procedural blast radius — the only intra-statements are the routing branch returns, which carry no work of their own. So intra_AIS is empty and inter_AIS is the three callees. PDG mode is intra-procedural by design, so on this case its inter-AIS recall is ~0 BY DESIGN (a capability fact, not a defect); the call-graph mode, with inter-procedural reach, finds all three. This case anchors the `inter` stratum so the harness can report complementary coverage. NOTE: the smoke test still requires `dispatch` to produce SOME CDG/REACHING_DEF edges (its routing branches do) — a zero-edge criterion would be unmeasurable; but those edges are not the AIS." + "rationale": "criterion.line=23 — the dispatcher's own first body statement `if (kind === 'a')` (source semantics: the router's work begins here). intra_AIS is EMPTY BY DESIGN: `dispatch` is a thin router whose TRUE impact is ENTIRELY cross-function — changing it affects the three handlers it routes to (handleA/handleB/handleDefault), captured in inter_AIS. The only intra statements are the routing branch returns, which carry no work of their own, so there is no meaningful intra-procedural ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 23 DOES return the routing return statements (the branch predicates control-depend on each other and each return), but those are NOT the meaningful blast radius — PDG is intra-procedural, so on a pure-inter fixture its intra slice is noise relative to the (empty) intra ground truth. This is the symmetric counterpart of call-graph's empty intra slice on the intra fixtures: each engine is blind to the other's locus. The call-graph mode (inter-procedural) finds all three handlers exactly. Anchors the `inter` stratum. NOTE: the smoke test requires `dispatch` to produce SOME CDG/REACHING_DEF edges (its routing branches do) — a zero-edge criterion would be unmeasurable; but those edges are not the AIS." } diff --git a/gitnexus/bench/impact-pdg/fixtures/inter-facade-delegate/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/inter-facade-delegate/ground-truth.json index d7ed6b277..f36f8387e 100644 --- a/gitnexus/bench/impact-pdg/fixtures/inter-facade-delegate/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/inter-facade-delegate/ground-truth.json @@ -4,6 +4,7 @@ "name": "processOrder", "filePath": "src/facade.ts", "direction": "downstream", + "line": 21, "marker": "enrich(order)", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -28,5 +29,5 @@ "note": "the terminal delegate in the sequence" } ], - "rationale": "`processOrder` is a facade sequencing three delegates (validate, enrich, persist) with one guard. Its true blast radius is the delegation chain — the inter_AIS — not its own body, which only routes values between calls. intra_AIS is empty: the guard return and the local `enriched` binding carry no independent work beyond shuttling delegate results. PDG mode (intra-procedural) has ~0 inter-AIS recall here BY DESIGN; call-graph mode walks validate->enrich->persist. Anchors the `inter` stratum with a sequential-delegation shape (distinct from the dispatcher's branching shape). The single guard gives `processOrder` a CDG edge so the smoke test's edge-presence requirement is met." + "rationale": "criterion.line=21 — the facade's first body statement, the guard `if (!validate(order))` (source semantics: the facade's work begins here). intra_AIS is EMPTY BY DESIGN: `processOrder` sequences three delegates (validate, enrich, persist) with one guard; its true blast radius is the delegation chain — inter_AIS — not its own body, which only routes values between calls. The guard return and the local `enriched` binding carry no independent work beyond shuttling delegate results, so there is no meaningful intra ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 21 returns the guard's control-dependent body statements, but those are noise relative to the (empty) intra ground truth — PDG is intra-procedural and cannot reach the cross-function delegates. The call-graph mode walks validate->enrich->persist exactly. Anchors the `inter` stratum with a sequential-delegation shape (distinct from the dispatcher's branching shape). The single guard gives `processOrder` a CDG edge so the smoke test's edge-presence requirement is met." } diff --git a/gitnexus/bench/impact-pdg/fixtures/inter-pipeline-stages/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/inter-pipeline-stages/ground-truth.json index 76abfe8cd..d15db3ef6 100644 --- a/gitnexus/bench/impact-pdg/fixtures/inter-pipeline-stages/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/inter-pipeline-stages/ground-truth.json @@ -4,6 +4,7 @@ "name": "runPipeline", "filePath": "src/pipeline.ts", "direction": "downstream", + "line": 20, "marker": "stageTransform(acc)", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -28,5 +29,5 @@ "note": "terminal stage" } ], - "rationale": "Direction is DOWNSTREAM (dependencies): changing `runPipeline` affects the callees it invokes. In GitNexus `impact` vocabulary, downstream = dependencies/callees and upstream = dependants/callers; the stages are what the driver CALLS, so the correct tag is downstream (an earlier draft mislabeled this `upstream`, conflating the English 'upstream sources' with GitNexus's caller-direction — corrected in U7 after the harness surfaced a direction-vs-AIS contradiction). As a thin driver it threads `acc` through three stage calls (stageParse->stageTransform->stageEmit); the dependencies that matter cross function boundaries — the three stage functions — so inter_AIS holds them. intra_AIS is empty: the `acc` reassignments only carry delegate results, no independent computation. PDG mode is intra-procedural, so its inter-AIS recall here is ~0 BY DESIGN; the call-graph mode (downstream) reaches the stages. The `enabled` guard (added so the driver carries a CDG edge) keeps the criterion measurable for the smoke test without changing the cross-function locus. Distinct from the dispatcher (branch-routed) and facade (guarded-sequence) shapes — this is a straight pipeline driver." + "rationale": "criterion.line=20 — the driver's first body statement `let acc = seed` (source semantics: the pipeline's threading begins here). Direction is DOWNSTREAM (dependencies): changing `runPipeline` affects the callees it invokes. In GitNexus `impact` vocabulary, downstream = dependencies/callees; the stages are what the driver CALLS, so the correct tag is downstream (an earlier draft mislabeled this `upstream`, conflating the English 'upstream sources' with GitNexus's caller-direction — corrected after the harness surfaced a direction-vs-AIS contradiction). intra_AIS is EMPTY BY DESIGN: the driver threads `acc` through three stage calls (stageParse->stageTransform->stageEmit); the dependencies that matter cross function boundaries — the three stage functions — so inter_AIS holds them. The `acc` reassignments only carry delegate results, no independent computation, so there is no meaningful intra ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 20 returns the intra `acc` def->use / guard statements, but those are noise relative to the (empty) intra ground truth — PDG is intra-procedural and cannot reach the cross-function stages. The call-graph mode (downstream) reaches the three stages exactly. The `enabled` guard (added so the driver carries a CDG edge) keeps the criterion measurable for the smoke test without changing the cross-function locus. Distinct from the dispatcher (branch-routed) and facade (guarded-sequence) shapes — this is a straight pipeline driver." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-control-branch/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-control-branch/ground-truth.json index f72cf758c..47cdc2bdf 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-control-branch/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-control-branch/ground-truth.json @@ -4,6 +4,7 @@ "name": "classify", "filePath": "src/branch.ts", "direction": "downstream", + "line": 7, "marker": "'zero'", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -17,6 +18,12 @@ "line": 9, "note": "`return 'pos'` is control-dependent on the `x > 0` arm" }, + { + "symbol": "classify", + "filePath": "src/branch.ts", + "line": 10, + "note": "the `else if (x < 0)` predicate is itself control-dependent on the outer `x > 0` branch (it runs only on the false arm)" + }, { "symbol": "classify", "filePath": "src/branch.ts", @@ -31,5 +38,5 @@ } ], "inter_AIS": [], - "rationale": "Criterion is the if/else-if predicate chain in `classify` (line 7). CDG controller->dependent edges reach each return arm (lines 9, 11, 13); each runs only under its arm's condition. All dependents are intra-procedural; the function calls nothing. intra_AIS = the three arm returns; inter_AIS empty. Exercises CDG-forward over a multi-arm branch. Note the `else if (x < 0)` test on line 10 is itself nested control flow; the truly-affected executable statements are the three returns, which is what we annotate at statement granularity." + "rationale": "criterion.line=7 — the if predicate `if (x > 0)`, the branch whose change propagates (source semantics: the predicate controls which arm runs). DOWNSTREAM CDG controller->dependent edges reach the arm returns (lines 9, 11, 13) AND the nested `else if (x < 0)` test (line 10). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={9,11,13} (the three returns only), but the `else if` predicate on line 10 is ITSELF control-dependent on the outer branch — it executes only on the `x > 0` false arm — so line 10 is a genuine CDG dependent the original annotation omitted. The CFG models it as its own predicate BasicBlock (`(x < 0)`), and the slice from line 7 returns {9, 10, 11, 13} exactly. The rationale's own note already flagged line 10 as nested control flow; this corrects it from out-of-set to in-set. All dependents intra-procedural; inter_AIS empty. Exercises CDG-forward over a multi-arm branch." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-control-guard/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-control-guard/ground-truth.json index afe76717e..42a058ca5 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-control-guard/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-control-guard/ground-truth.json @@ -4,6 +4,7 @@ "name": "guarded", "filePath": "src/guard.ts", "direction": "downstream", + "line": 7, "marker": "y + 1", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -21,13 +22,7 @@ "symbol": "guarded", "filePath": "src/guard.ts", "line": 11, - "note": "`const y = x * 2` runs only when the guard is false: control-dependent" - }, - { - "symbol": "guarded", - "filePath": "src/guard.ts", - "line": 12, - "note": "`const z = y + 1` is control-dependent on the guard (and data-dependent on y)" + "note": "the post-guard body block `const y = x * 2; const z = y + 1;` (lines 11-12, coalesced) runs only when the guard is false: control-dependent" }, { "symbol": "guarded", @@ -37,5 +32,5 @@ } ], "inter_AIS": [], - "rationale": "Criterion is the guard predicate `if (!ok)` (line 7). Its CDG controller->dependent edges reach the true-arm `return -1` (line 9) and the post-guard body (lines 11-13), which execute only on the complementary arm. All dependents are statements WITHIN `guarded`; the function calls nothing. intra_AIS captures every control-dependent line; inter_AIS empty. This is the canonical guard-clause CDG shape (#559) and isolates the CDG-forward arm of KTD4 — PDG distinguishes control-affected statements that the call-graph mode cannot resolve below function granularity." + "rationale": "criterion.line=7 — the guard predicate `if (!ok)` (source semantics: the guard controls whether the post-guard body runs). DOWNSTREAM CDG controller->dependent edges reach the true-arm `return -1` (line 9), the post-guard body, and `return z` (line 13). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={9,11,12,13}, but the CFG COALESCES the two consecutive post-guard defs `const y = x * 2` and `const z = y + 1` into ONE BasicBlock starting at line 11 — so line 12 cannot surface as a distinct dependent (it lives inside block 11). The measurable control-dependent lines are {9, 11 (the coalesced body block), 13}. The slice from line 7 returns {9, 11, 13} exactly — a block-granularity correction, not a metric re-fit. All dependents WITHIN `guarded`; inter_AIS empty. Canonical guard-clause CDG shape (#559); isolates the CDG-forward arm of KTD4 — PDG distinguishes control-affected statements the call-graph mode cannot resolve below function granularity." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-control-loop/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-control-loop/ground-truth.json index aed455486..59f298bb5 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-control-loop/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-control-loop/ground-truth.json @@ -4,6 +4,7 @@ "name": "filterPositive", "filePath": "src/loop.ts", "direction": "upstream", + "line": 11, "marker": "count + 1", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -14,16 +15,28 @@ { "symbol": "filterPositive", "filePath": "src/loop.ts", - "line": 10, - "note": "the inner `if (x > 0)` predicate directly controls the `count` increment" + "line": 6, + "note": "the function-signature/parameter block defines `xs`, which the `for` loop iterates — an upstream reaching dependency of the increment" + }, + { + "symbol": "filterPositive", + "filePath": "src/loop.ts", + "line": 7, + "note": "`let count = 0` is the initial def of `count` reaching the increment via REACHING_DEF — an upstream data dependency" }, { "symbol": "filterPositive", "filePath": "src/loop.ts", "line": 8, - "note": "the enclosing `for` loop guard controls whether the if and its body run" + "note": "the enclosing `for` loop guard controls whether the if and its body run (transitive CDG controller)" + }, + { + "symbol": "filterPositive", + "filePath": "src/loop.ts", + "line": 10, + "note": "the inner `if (x > 0)` predicate directly controls the `count` increment (immediate CDG controller)" } ], "inter_AIS": [], - "rationale": "Direction is UPSTREAM: the criterion `filterPositive`'s focal statement is the `count = count + 1` increment (line 11), and we ask which control predicates govern whether it runs. The immediate controller is the `if (x > 0)` (line 10); the if is itself nested in the `for` loop guard (line 8), which transitively controls the increment. Both controllers are intra-procedural. intra_AIS = {line 10 (inner if), line 8 (loop guard)}; inter_AIS empty. Isolates the CDG-reverse arm of KTD4 over nested control structure." + "rationale": "criterion.line=11 — the `count = count + 1` increment (source semantics: UPSTREAM asks which definitions/predicates govern this statement). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={8,10} (the two CONTROL predicates only). But the statement-anchored PDG slice is a COMBINED CDG+REACHING_DEF reverse slice, so it correctly also reaches the DATA dependencies of the increment: line 7 (`let count = 0`, the initial reaching def of `count`) and line 6 (the parameter block defining `xs`, which feeds the `for` loop). The original annotation under-counted by considering only control dependence. The full upstream slice is {6, 7, 8, 10}; the slice from line 11 returns exactly that. Both controllers (8, 10) and both data deps (6, 7) are intra-procedural; inter_AIS empty. Exercises the CDG+RD-reverse slice over nested control structure." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-accumulator/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-accumulator/ground-truth.json index c86bbae91..fca2ca19f 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-accumulator/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-accumulator/ground-truth.json @@ -4,6 +4,7 @@ "name": "total", "filePath": "src/accumulator.ts", "direction": "downstream", + "line": 8, "marker": "sum + x", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -25,5 +26,5 @@ } ], "inter_AIS": [], - "rationale": "Criterion is the function `total`, whose changed statement is the `sum` definition (line 8). `sum` is a loop-carried accumulator: its def reaches the in-loop redefinition/use (line 10) and the final return's use (line 12) via REACHING_DEF def->use edges. These two lines are the only truly-affected statements and they are all WITHIN `total` — so intra_AIS is exactly {line 10, line 12} and inter_AIS is empty (the function calls nothing, returns to no annotated caller in this single-function fixture). PDG should win precision here: the call-graph mode only knows `total` exists and has no inbound edges, so it reports `total` itself with no finer-grained intra-statement resolution." + "rationale": "criterion.line=8 — the `sum` definition `let sum = 0`, the statement whose change propagates (chosen from source semantics: `sum` is the loop-carried accumulator). DOWNSTREAM from line 8, the REACHING_DEF def->use edges reach the in-loop redefinition/use (line 10) and the final return's use (line 12). These two lines are the only truly-affected statements and they are all WITHIN `total` — so intra_AIS is exactly {line 10, line 12}; inter_AIS is empty (the function calls nothing). Step-0 reconciliation: the statement-anchored PDG slice from line 8 returns exactly {10, 12} — an EXACT match to this source-derived annotation (no correction needed). The call-graph mode only knows `total` exists and has no inbound edges, so it reports the empty cross-function set; PDG resolves the precise dependent statements the call-graph cannot express below function granularity." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-chain/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-chain/ground-truth.json index fa8167ca1..d0a6bd57d 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-chain/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-chain/ground-truth.json @@ -4,6 +4,7 @@ "name": "chainCompute", "filePath": "src/chain.ts", "direction": "downstream", + "line": 7, "marker": "b - 3", "pdgEdgeKinds": ["REACHING_DEF"] }, @@ -11,25 +12,13 @@ "provenance": "manual", "analyzerVersion": "1.6.7", "intra_AIS": [ - { - "symbol": "chainCompute", - "filePath": "src/chain.ts", - "line": 8, - "note": "`b = a * 2` is directly data-dependent on the def of `a`" - }, - { - "symbol": "chainCompute", - "filePath": "src/chain.ts", - "line": 9, - "note": "`c = b - 3` is transitively data-dependent on `a` via `b`" - }, { "symbol": "chainCompute", "filePath": "src/chain.ts", "line": 10, - "note": "`return c` is transitively data-dependent on `a` via the chain" + "note": "`return c` is transitively data-dependent on `a` via the chain — the only distinct downstream BasicBlock" } ], "inter_AIS": [], - "rationale": "Criterion's changed statement is the def of `a` (line 7). A straight-line REACHING_DEF chain carries `a` -> `b` (line 8) -> `c` (line 9) -> `return` (line 10). Every downstream statement in the function is data-dependent on `a`; none escapes the function (no calls, no annotated caller). intra_AIS is exactly the three downstream lines; inter_AIS is empty. This case stresses transitive forward def->use reachability with no control structure, isolating the RD-forward arm of the KTD4 truth table." + "rationale": "criterion.line=7 — the def of `a` (`const a = input + 1`), the changed statement (source semantics: `a` seeds the straight-line def->use chain a->b->c->return). DOWNSTREAM from line 7. CORRECTION (Step-0 reconciliation): the original source-derived belief was intra_AIS={8,9,10} (every downstream statement). But the CFG COALESCES consecutive straight-line statements into ONE BasicBlock: lines 7-9 (`const a`, `const b`, `const c`) share a single block whose start line is 7 — the criterion's own seed block. The traversal excludes the seed, so lines 8 and 9 can never surface as DISTINCT dependent statements; they are inside the criterion block. The only separate downstream BasicBlock is `return c` (line 10). So the measurable intra_AIS is exactly {line 10}. This is a block-granularity fact (the human annotation was at finer-than-block statement granularity), not a metric re-fit: the slice from line 7 returns {10}, an exact match to the corrected annotation. Isolates the RD-forward arm of the KTD4 truth table; inter_AIS empty (no calls)." } diff --git a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-reassign/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-reassign/ground-truth.json index 9c21cd6c9..7e9643428 100644 --- a/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-reassign/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/intra-dataflow-reassign/ground-truth.json @@ -4,6 +4,7 @@ "name": "reassignSum", "filePath": "src/reassign.ts", "direction": "upstream", + "line": 9, "marker": "total + b", "pdgEdgeKinds": ["REACHING_DEF"] }, @@ -14,16 +15,16 @@ { "symbol": "reassignSum", "filePath": "src/reassign.ts", - "line": 7, - "note": "first def `total = a` reaches the criterion use transitively via the second def" + "line": 6, + "note": "the function-signature/parameter block defines `a`, whose value flows into `total = a`; it is a genuine reaching def upstream of the criterion use" }, { "symbol": "reassignSum", "filePath": "src/reassign.ts", - "line": 8, - "note": "second def `total = total + b` is the immediately-reaching def of the criterion use" + "line": 7, + "note": "the coalesced def block `let total = a; total = total + b;` (lines 7-8) is the immediately-reaching def of the criterion use `return total`" } ], "inter_AIS": [], - "rationale": "Direction is UPSTREAM: the criterion `reassignSum`'s relevant statement is the `return total` use (line 9), and we ask which definitions reach it. The immediately-reaching def is line 8 (`total = total + b`); that def itself uses `total`, whose reaching def is line 7 (`total = a`). Both defs are upstream data-dependencies of the criterion use, both intra-procedural. intra_AIS = {line 7, line 8}; inter_AIS empty. This isolates the RD-reverse arm of the KTD4 truth table and a reassigned variable's def chain." + "rationale": "criterion.line=9 — the `return total` use (source semantics: UPSTREAM asks which definitions reach this use of `total`). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={7,8} (both defs of total). But the CFG COALESCES the two consecutive defs `let total = a` and `total = total + b` into ONE BasicBlock starting at line 7 — so line 8 cannot surface as a distinct statement (it lives inside block 7). Additionally, REACHING_DEF reaches further upstream to the parameter block (line 6, the function signature defining `a`), a genuine reaching def of the chain that the original annotation under-counted. So the measurable upstream slice is {line 6 (param def of `a`), line 7 (coalesced total-def block)}. The slice from line 9 returns {6, 7} exactly. Both are intra-procedural; inter_AIS empty. Isolates the RD-reverse arm of the KTD4 truth table over a reassigned variable." } diff --git a/gitnexus/bench/impact-pdg/fixtures/mixed-compute-and-emit/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/mixed-compute-and-emit/ground-truth.json index c06c291ff..87d3d38fe 100644 --- a/gitnexus/bench/impact-pdg/fixtures/mixed-compute-and-emit/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/mixed-compute-and-emit/ground-truth.json @@ -4,6 +4,7 @@ "name": "computeAndEmit", "filePath": "src/mixed.ts", "direction": "downstream", + "line": 12, "marker": "score + v", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -37,5 +38,5 @@ "note": "called with the computed level; cross-function reach" } ], - "rationale": "Criterion `computeAndEmit` is mixed-locus. The `score` def (line 12) drives a loop-carried data chain (line 14), feeds the threshold `level` (line 16), which feeds the `emit` call arg (line 17): all intra-procedural data dependence, so intra_AIS holds those three statements. Separately, the call to `emit` is a cross-function reach captured in inter_AIS. The two sets are disjoint (lines within `computeAndEmit` vs the distinct symbol `emit`). Distinct from mixed-validate-then-call: this case's intra dependence is data-flow-dominated (accumulator + ternary) rather than guard-dominated, broadening mixed-stratum coverage." + "rationale": "criterion.line=12 — the `score` def `let score = 0` (source semantics: `score` seeds the loop-carried data chain). Criterion `computeAndEmit` is mixed-locus. DOWNSTREAM from line 12, the `score` def drives the loop-carried data chain (line 14), feeds the threshold `level` (line 16), which feeds the `emit` call arg (line 17): all intra-procedural data dependence, so intra_AIS holds those three statements. Separately, the call to `emit` is a cross-function reach captured in inter_AIS. The two sets are disjoint (lines within `computeAndEmit` vs the distinct symbol `emit`). Step-0 reconciliation: the statement-anchored PDG slice from line 12 returns {14, 16, 17} — an EXACT match to this source-derived intra_AIS (no correction needed); call-graph reaches {emit} exactly. Distinct from mixed-validate-then-call: this case's intra dependence is data-flow-dominated (accumulator + ternary) rather than guard-dominated, broadening mixed-stratum coverage." } diff --git a/gitnexus/bench/impact-pdg/fixtures/mixed-guarded-dispatch/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/mixed-guarded-dispatch/ground-truth.json index 8dbe9a9a4..019aa970b 100644 --- a/gitnexus/bench/impact-pdg/fixtures/mixed-guarded-dispatch/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/mixed-guarded-dispatch/ground-truth.json @@ -4,6 +4,7 @@ "name": "route", "filePath": "src/mixed.ts", "direction": "downstream", + "line": 15, "marker": "key === 0", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -42,5 +43,5 @@ "note": "callee on the fallthrough route; cross-function reach" } ], - "rationale": "Criterion `route` is mixed-locus with BOTH a guard whose predicate reads the intra def `key` (line 15 -> line 16) and control-gates both return arms (lines 18, 20), AND two cross-function reaches (fast, slow). intra_AIS holds the guard + the two control-dependent return statements; inter_AIS holds the two callees. Disjoint by construction (lines within `route` vs distinct symbols). This mixed case combines control AND data dependence intra-procedurally with branching cross-function reach — the richest mixed shape, complementing the data-dominated and guard-dominated mixed cases." + "rationale": "criterion.line=15 — the `key` def `const key = n % 2` (source semantics: `key` feeds the guard predicate below). Criterion `route` is mixed-locus with BOTH a guard whose predicate reads the intra def `key` (line 15 -> line 16) and control-gates both return arms (lines 18, 20), AND two cross-function reaches (fast, slow). DOWNSTREAM from line 15: intra_AIS holds the guard (16) + the two control-dependent return statements (18, 20); inter_AIS holds the two callees. Disjoint by construction (lines within `route` vs distinct symbols). Step-0 reconciliation: the statement-anchored PDG slice from line 15 returns {16, 18, 20} — an EXACT match to this source-derived intra_AIS; call-graph reaches {fast, slow} exactly. This mixed case combines control AND data dependence intra-procedurally with branching cross-function reach — the richest mixed shape, complementing the data-dominated and guard-dominated mixed cases." } diff --git a/gitnexus/bench/impact-pdg/fixtures/mixed-validate-then-call/ground-truth.json b/gitnexus/bench/impact-pdg/fixtures/mixed-validate-then-call/ground-truth.json index 374ce7ee4..4c74ce213 100644 --- a/gitnexus/bench/impact-pdg/fixtures/mixed-validate-then-call/ground-truth.json +++ b/gitnexus/bench/impact-pdg/fixtures/mixed-validate-then-call/ground-truth.json @@ -4,6 +4,7 @@ "name": "handleRequest", "filePath": "src/mixed.ts", "direction": "downstream", + "line": 13, "marker": "normalized + 1", "pdgEdgeKinds": ["REACHING_DEF", "CDG"] }, @@ -43,5 +44,5 @@ "note": "called with the adjusted value; cross-function reach" } ], - "rationale": "Criterion `handleRequest` is mixed-locus. The `normalized` def (line 13) drives BOTH a data chain (guard test line 14, `adjusted` line 18, the call arg line 19) and a control structure (the guard controls lines 16 and 19). So intra_AIS holds the four control/data-dependent statements within the function. Separately, `handleRequest` calls `persist` — a genuine cross-function reach — so inter_AIS holds `persist`. This is the case where the two modes are complementary: PDG resolves the intra statement set precisely, call-graph reaches the callee. intra_AIS and inter_AIS are disjoint by construction (one is lines within `handleRequest`, the other is a different symbol `persist`)." + "rationale": "criterion.line=13 — the `normalized` def `const normalized = raw * 2` (source semantics: `normalized` seeds both the data chain and the guard). Criterion `handleRequest` is mixed-locus. DOWNSTREAM from line 13, the `normalized` def drives BOTH a data chain (guard test line 14, `adjusted` line 18, the call arg line 19) and a control structure (the guard controls lines 16 and 19). So intra_AIS holds the four control/data-dependent statements within the function. Separately, `handleRequest` calls `persist` — a genuine cross-function reach — so inter_AIS holds `persist`. Step-0 reconciliation: the statement-anchored PDG slice from line 13 returns {14, 16, 18, 19} — an EXACT match to this source-derived intra_AIS; call-graph reaches {persist} exactly. This is the case where the two modes are complementary: PDG resolves the intra statement set precisely, call-graph reaches the callee. intra_AIS and inter_AIS are disjoint by construction (one is lines within `handleRequest`, the other is a different symbol `persist`)." } diff --git a/gitnexus/bench/impact-pdg/measure.mjs b/gitnexus/bench/impact-pdg/measure.mjs index 42cad12e7..a34cde84b 100644 --- a/gitnexus/bench/impact-pdg/measure.mjs +++ b/gitnexus/bench/impact-pdg/measure.mjs @@ -1,11 +1,18 @@ /** * U7 — PDG-vs-call-graph impact ACCURACY measurement harness. * - * Runs BOTH `impact` engines (`mode:'callgraph'` and `mode:'pdg'`) over the - * curated U6 ground-truth fixtures, computes precision/recall/F1 stratified by - * impact locus (intra / inter / mixed) plus cross-mode Jaccard + set-diffs, - * prints a stratified report ending in a plain-language DECISION RECOMMENDATION, - * and (under `--check`) gates regressions with two NON-byte-identity gates. + * Runs BOTH `impact` engines over the curated U6 ground-truth fixtures and + * scores each at its NATIVE granularity: + * - `mode:'pdg'` is seeded on the criterion's STATEMENT (`line: criterion.line`) + * and scored at intra-procedural LINE granularity against `intra_AIS` + * (CIS_pdg = the `affectedStatements` line set); + * - `mode:'callgraph'` is scored at inter-procedural SYMBOL granularity against + * `inter_AIS` (CIS = the reported symbol set). + * It computes precision/recall/F1 stratified by impact locus (intra/inter/mixed) + * plus cross-mode set-diffs, prints a stratified report ending in a plain- + * language DECISION RECOMMENDATION, and (under `--check`) gates regressions with + * two NON-byte-identity gates. The two engines answer DIFFERENT questions at + * DIFFERENT granularities — the report shows both, neither strictly dominates. * * ── Substrate (the load-bearing mechanism — KTD9/R8; plan U7 "Substrate * decision") ────────────────────────────────────────────────────────────── @@ -26,14 +33,17 @@ * 3. `new LocalBackend(); await init()` resolves the fixture via the REAL * registry (the parent process ALSO sets `GITNEXUS_HOME` so init reads the * temp registry, not the user's ~/.gitnexus); - * 4. `callTool('impact', {repo:, target, direction, mode})` ×2; + * 4. `callTool('impact', …)` ×2 — callgraph (symbol BFS) and pdg (seeded on + * `line: criterion.line` so it returns the statement-anchored slice); * 5. teardown the temp home + copy. * * The `repo` arg is the absolute fixture-copy PATH (tier-1 path match in * `resolveRepoFromCache`) — unambiguous, no name collisions. * * ── Granularity / CIS-AIS framing ────────────────────────────────────────── - * See `metrics.mjs`. Symbol granularity, line-collapsed, order-independent. + * See `metrics.mjs`. PDG = intra-procedural LINE granularity vs `intra_AIS`; + * call-graph = inter-procedural SYMBOL granularity vs `inter_AIS`. The two + * engines measure different scopes; both are now non-empty. * * Build-free: `node --import tsx bench/impact-pdg/measure.mjs`. Runtime budget * and re-baseline instructions: see README.md. @@ -47,11 +57,10 @@ import { fileURLToPath } from 'node:url'; import { symbolKey, - toKeySet, + pdgLineCis, + intraLineAis, score, - compareModes, aggregate, - partitionCisByScope, aisByScope, fingerprintAnnotationSet, median, @@ -136,15 +145,26 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) { const backend = new LocalBackend(); await backend.init(); - const results = {}; - for (const mode of MODES) { - results[mode] = await backend.callTool('impact', { + // callgraph: symbol→symbol BFS (no statement anchor). pdg: SEEDED on the + // criterion's statement line so it returns the dependence slice — the U7 + // rework's central change. A whole-symbol pdg slice (no `line`) is empty by + // design; `criterion.line` is the 1-based source line of the changed + // statement (set from source semantics, validated in Step 0). + const results = { + callgraph: await backend.callTool('impact', { repo: work, target: fx.gt.criterion.name, direction: fx.gt.criterion.direction, - mode, - }); - } + mode: 'callgraph', + }), + pdg: await backend.callTool('impact', { + repo: work, + target: fx.gt.criterion.name, + direction: fx.gt.criterion.direction, + mode: 'pdg', + line: fx.gt.criterion.line, + }), + }; succeeded = true; return { work, results }; } finally { @@ -154,16 +174,18 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) { } } -/** Flatten an impact result's byDepth into canonical symbol keys (the CIS). */ -function cisFromResult(res) { +/** + * Flatten a CALLGRAPH impact result's byDepth into canonical SYMBOL keys (the + * CIS_callgraph, scored against `inter_AIS`). Unresolved shadow entries are kept + * (keyed by file) so a recall loss is never hidden. + */ +function callgraphCisFromResult(res) { const items = Object.values(res?.byDepth ?? {}).flat(); const keys = new Set(); - const meta = { unresolved: 0, ambiguous: 0, blockCount: res?.blockCount ?? null }; + const meta = { unresolved: 0, ambiguous: 0 }; for (const it of items) { if (it?.unresolved) { meta.unresolved += 1; - // surfaced under its file as an unresolved shadow entry — kept in the CIS - // so a recall loss is never hidden, keyed by its file (no symbol name). keys.add(symbolKey('(unresolved)', it.filePath)); continue; } @@ -173,6 +195,25 @@ function cisFromResult(res) { return { keys, meta }; } +/** + * Extract the PDG statement-line CIS (`:` keys) from a pdg + * impact result's `affectedStatements`, plus the diagnostic fields the report + * surfaces (the slice's epistemic marker / note / block count). This is the + * U7-rework CIS: the dependent STATEMENTS the change at `criterion.line` reaches. + */ +function pdgCisFromResult(res) { + return { + keys: pdgLineCis(res?.affectedStatements), + meta: { + affectedStatementCount: res?.affectedStatementCount ?? 0, + blockCount: res?.blockCount ?? null, + criterionLine: res?.criterionLine ?? null, + epistemic: res?.epistemic ?? null, + note: res?.note ?? null, + }, + }; +} + // ── Step 0: fixture AIS validation (gated on the live traversal; KTD9 // circularity guard) ─────────────────────────────────────────────────────── @@ -257,20 +298,34 @@ async function validateFixture(fx, work, exec) { return { critEdges, sameLineCollision, problems, measurable: problems.length === 0 }; } -// ── per-fixture scoring ────────────────────────────────────────────────────── +// ── per-fixture scoring (each mode vs its NATIVE ground truth — U7 rework) ──── /** - * Score one fixture for one mode, per scope. CIS partitioned into intra (the - * criterion symbol itself) / inter (others) / mixed (union); AIS likewise. + * Score the CALLGRAPH mode for one fixture: its reported SYMBOL CIS against the + * fixture's `inter_AIS` (the cross-function symbols truly affected). The + * criterion symbol itself is dropped from the CIS first — callgraph never names + * the criterion as its own dependent, and `inter_AIS` is cross-function by + * construction, so a stray self-reference would be spurious noise. (In practice + * the callgraph CIS already excludes the seed; this is belt-and-suspenders.) */ -function scoreFixtureMode(gt, cisKeys) { +function scoreCallgraph(gt, symbolCisKeys) { const ais = aisByScope(gt); - const cisPart = partitionCisByScope(cisKeys, ais.criterionKey); - return { - intra: score(cisPart.intra, ais.intra), - inter: score(cisPart.inter, ais.inter), - mixed: score(cisPart.mixed, ais.mixed), - }; + const cis = new Set([...symbolCisKeys].filter((k) => k !== ais.criterionKey)); + return score(cis, ais.inter); +} + +/** + * Score the PDG mode for one fixture: its statement-LINE CIS (the + * `affectedStatements` from the line-seeded slice) against the fixture's + * `intra_AIS` LINE set. This is the intra-procedural statement-granularity + * measurement the U7 rework introduces. For an inter fixture (`intra_AIS` empty + * by design) the slice may return the router's own control-dependent statements + * — those are FPIS against the empty intra ground truth and recall is n/a, which + * is the honest "PDG is intra-procedural; on a pure-inter fixture it has no + * meaningful intra ground truth" result (symmetric to callgraph's empty intra). + */ +function scorePdg(gt, lineCisKeys) { + return score(lineCisKeys, intraLineAis(gt)); } // ── reporting helpers ──────────────────────────────────────────────────────── @@ -281,14 +336,17 @@ const lpad = (s, n) => String(s).padStart(n); function renderTable(strata) { const head = - `${pad('Scope', 7)} ${pad('Mode', 10)} ${lpad('P', 7)} ${lpad('R', 7)} ${lpad('F1', 7)} ` + + `${pad('Scope', 7)} ${pad('Mode', 10)} ${pad('Granularity', 11)} ${lpad('P', 7)} ${lpad('R', 7)} ${lpad('F1', 7)} ` + `${lpad('|CIS|/|AIS|', 11)} ${lpad('FPIS', 6)} ${lpad('FNIS', 6)} ${lpad('n', 4)}`; const lines = [head, '-'.repeat(head.length)]; + // PDG is scored at LINE granularity vs intra_AIS; callgraph at SYMBOL + // granularity vs inter_AIS — the column makes the "different scopes" explicit. + const gran = (mode) => (mode === 'pdg' ? 'line/intra' : 'symbol/inter'); for (const scope of SCOPES) { for (const mode of MODES) { const a = strata[scope][mode]; lines.push( - `${pad(scope, 7)} ${pad(mode, 10)} ${lpad(fmt(a.precision), 7)} ${lpad(fmt(a.recall), 7)} ` + + `${pad(scope, 7)} ${pad(mode, 10)} ${pad(gran(mode), 11)} ${lpad(fmt(a.precision), 7)} ${lpad(fmt(a.recall), 7)} ` + `${lpad(fmt(a.f1), 7)} ${lpad(fmt(a.cisAisRatio), 11)} ${lpad(a.fpis, 6)} ${lpad(a.fnis, 6)} ` + `${lpad(a.nCases, 4)}`, ); @@ -300,15 +358,19 @@ function renderTable(strata) { /** * Plain-language DECISION RECOMMENDATION (F2 — the deliverable that answers * "which is more accurate" as a verdict, not just a table). Derived from the - * measured numbers: compares inter-scope recall (the cross-function questions - * users most bring to impact) and any measured intra-scope precision edge. + * measured numbers: PDG's intra-procedural STATEMENT-granularity F1 (the slice it + * is built to compute) and call-graph's inter-procedural SYMBOL-granularity F1 + * (the cross-function reach it is built to compute). */ function decisionRecommendation(strata, underpowered, exclusions) { - const cgInterR = strata.inter.callgraph.recall; - const pdgInterR = strata.inter.pdg.recall; - const cgIntraP = strata.intra.callgraph.precision; + // PDG is precise at intra LINE granularity; callgraph covers inter SYMBOL reach. + const pdgIntraF1 = strata.intra.pdg.f1; const pdgIntraP = strata.intra.pdg.precision; - const pdgIntraReports = strata.intra.pdg.nPrecision > 0; // did PDG report ANY intra symbol? + const pdgIntraR = strata.intra.pdg.recall; + const cgInterF1 = strata.inter.callgraph.f1; + const cgInterR = strata.inter.callgraph.recall; + const pdgMixedF1 = strata.mixed.pdg.f1; + const cgMixedF1 = strata.mixed.callgraph.f1; const lines = []; lines.push('DECISION RECOMMENDATION'); @@ -319,42 +381,43 @@ function decisionRecommendation(strata, underpowered, exclusions) { ); } - // Inter-scope: the cross-function blast radius. - if (cgInterR !== null && pdgInterR !== null) { - lines.push( - `On INTER-scope (cross-function) impact, call-graph recall is ${fmt(cgInterR)} vs PDG ${fmt(pdgInterR)}: ` + - `PDG's intra-procedural design means it recovers ~0 cross-function impact BY DESIGN (a capability ` + - `fact, not a defect). Call-graph is the correct engine for the "what else calls/uses this?" question.`, - ); - } + // Intra-scope: the statement-anchored PDG slice — the question PDG answers. + lines.push( + `On INTRA-scope (statement granularity), the line-seeded PDG slice scores P=${fmt(pdgIntraP)} ` + + `R=${fmt(pdgIntraR)} F1=${fmt(pdgIntraF1)} against intra_AIS: it identifies the dependent ` + + `STATEMENTS of the changed line precisely. Call-graph mode cannot resolve below function ` + + `granularity, so on a self-contained function it names no other symbol (intra recall n/a — ` + + `no cross-function truth to find). PDG is the engine for "which statements does this line affect?".`, + ); - // Intra-scope: the case PDG was built to win. - if (!pdgIntraReports) { + // Inter-scope: the cross-function blast radius — the question call-graph answers. + lines.push( + `On INTER-scope (symbol granularity), call-graph scores R=${fmt(cgInterR)} F1=${fmt(cgInterF1)} ` + + `against inter_AIS: it recovers the cross-function callees exactly. PDG mode is ` + + `intra-procedural, so on a pure-inter fixture it returns only the router's own ` + + `control-dependent statements (FPIS against the empty intra_AIS — recall n/a). Call-graph is ` + + `the engine for "what else calls/uses this?".`, + ); + + // Mixed-scope: both engines contribute, each in its own scope. + if (pdgMixedF1 !== null || cgMixedF1 !== null) { lines.push( - `On INTRA-scope, PDG mode reported NO owning symbols across the measurable corpus: its block→symbol ` + - `projection collapses a function's own dependence blocks back onto the criterion itself, which the ` + - `traversal excludes as the seed — so at SYMBOL granularity the intra-procedural blast radius is the ` + - `empty set. PDG's intra value in v1 is therefore the BLOCK-LEVEL detail it surfaces ` + - `(reachableBlocks / blockCount), NOT a symbol-level impact set. The harness records the per-fixture ` + - `dependence-block counts so this is visible, not hidden as a flat zero.`, - ); - } else if (pdgIntraP !== null && cgIntraP !== null) { - const verb = pdgIntraP > cgIntraP ? 'higher' : pdgIntraP < cgIntraP ? 'lower' : 'equal'; - lines.push( - `On INTRA-scope, PDG precision is ${fmt(pdgIntraP)} vs call-graph ${fmt(cgIntraP)} (${verb}). ` + - `This is the measured direction on this corpus, reported as a fact, not asserted as a hypothesis.`, + `On MIXED-scope, the two are COMPLEMENTARY: PDG resolves the intra statement set ` + + `(F1=${fmt(pdgMixedF1)} vs intra_AIS) while call-graph reaches the callee(s) ` + + `(F1=${fmt(cgMixedF1)} vs inter_AIS). Neither alone covers the full mixed blast radius.`, ); } lines.push( - `VERDICT: use mode:'callgraph' as the default — it carries the inter-procedural reach that the ` + - `impact tool's safety question depends on. mode:'pdg' adds value as an OPT-IN lens for ` + - `intra-procedural dependence INSPECTION (its reachableBlocks / CDG+REACHING_DEF detail), and ` + - `where the persisted PDG layer exists (analyze --pdg). It is NOT a replacement for, nor a ` + - `strict improvement over, the call-graph blast radius: the two engines occupy different points ` + - `on the precision/recall curve and neither strictly dominates. Promotion of mode:'pdg' beyond ` + - `opt-in is GATED on a Function→BasicBlock CONTAINS_BLOCK substrate edge (deferred) that would let ` + - `the symbol BFS chain natively into the PDG and give intra reach a symbol-level meaning.`, + `VERDICT: the two engines answer DIFFERENT questions at DIFFERENT granularities, and NEITHER ` + + `dominates. mode:'callgraph' (the default) is the correct engine for the inter-procedural ` + + `safety question — "what else depends on / calls this symbol?" — carrying the cross-function ` + + `reach the blast radius needs. mode:'pdg' (opt-in, seeded with line:N, where analyze --pdg ` + + `persisted the layer) is PRECISE at intra-procedural STATEMENT granularity — "which statements ` + + `inside this function does changing line N affect?" — a question call-graph cannot answer at ` + + `all. Use call-graph for cross-symbol impact; reach for line-seeded PDG when you need ` + + `statement-level dependence INSIDE a function. They compose: a full mixed-locus blast radius ` + + `is the UNION of call-graph's inter-symbol reach and PDG's intra-statement slice.`, ); if (exclusions.length > 0) { lines.push( @@ -429,43 +492,40 @@ async function run() { continue; } - const cg = cisFromResult(results.callgraph); - const pdg = cisFromResult(results.pdg); - const cgScores = scoreFixtureMode(fx.gt, cg.keys); - const pdgScores = scoreFixtureMode(fx.gt, pdg.keys); + // CALLGRAPH: symbol CIS vs inter_AIS. PDG: line CIS vs intra_AIS. + const cg = callgraphCisFromResult(results.callgraph); + const pdg = pdgCisFromResult(results.pdg); + const cgScore = scoreCallgraph(fx.gt, cg.keys); // symbol/inter + const pdgScore = scorePdg(fx.gt, pdg.keys); // line/intra const locusScope = fx.gt.locus; // the stratum this fixture belongs to - // A fixture is scored in its OWN locus stratum (intra/inter/mixed). + // A fixture is scored in its OWN locus stratum (intra/inter/mixed), + // each mode against its native ground truth (symbol vs line). if (SCOPES.includes(locusScope)) { - perScopeMode[locusScope].callgraph.push(cgScores[locusScope]); - perScopeMode[locusScope].pdg.push(pdgScores[locusScope]); + perScopeMode[locusScope].callgraph.push(cgScore); + perScopeMode[locusScope].pdg.push(pdgScore); } if (runIdx === 0) { - const ais = aisByScope(fx.gt); - const cmp = compareModes(cg.keys, pdg.keys, ais.mixed); detail.push({ name: fx.name, locus: fx.gt.locus, criterion: fx.gt.criterion.name, direction: fx.gt.criterion.direction, + criterionLine: fx.gt.criterion.line ?? null, critEdges: v.critEdges, cg: { count: results.callgraph.impactedCount, symbols: [...cg.keys].sort(), - scores: cgScores, + score: cgScore, // vs inter_AIS (symbol) }, pdg: { - count: results.pdg.impactedCount, + affectedStatementCount: pdg.meta.affectedStatementCount, blockCount: pdg.meta.blockCount, - unresolved: pdg.meta.unresolved, - ambiguous: pdg.meta.ambiguous, - symbols: [...pdg.keys].sort(), - scores: pdgScores, + criterionLine: pdg.meta.criterionLine, + lines: [...pdg.keys].sort(), + score: pdgScore, // vs intra_AIS (line) }, - jaccard: cmp.jaccard, - pdgOnly: cmp.pdgOnly, - callgraphOnly: cmp.callgraphOnly, }); } } finally { @@ -566,25 +626,34 @@ async function run() { `(${measurableTotal} measurable, ${exclusions.length} excluded) | runs K=${K}`, ); out.push(''); - out.push('Stratified P/R/F1 (symbol granularity, per impact locus):'); + out.push( + 'Stratified P/R/F1 (PDG: line granularity vs intra_AIS; callgraph: symbol vs inter_AIS):', + ); out.push(renderTable(report)); out.push(''); - out.push('Per-case Jaccard + cross-mode set-diffs (true = ∩AIS, noise = −AIS):'); + out.push( + 'Per-case: PDG slice (line/intra) and callgraph reach (symbol/inter), with FPIS/FNIS:', + ); for (const d of perCaseDetail) { out.push( - ` ${pad(d.name, 28)} locus=${pad(d.locus, 6)} J=${fmt(d.jaccard)} ` + - `cg|count=${d.cg.count} pdg|count=${d.pdg.count} pdg|blocks=${d.pdg.blockCount}`, + ` ${pad(d.name, 28)} locus=${pad(d.locus, 6)} line=${lpad(d.criterionLine ?? '-', 3)}`, ); - if (d.callgraphOnly.all.length) - out.push( - ` callgraph-only: ${d.callgraphOnly.all.length} ` + - `(true ${d.callgraphOnly.true.length}, noise ${d.callgraphOnly.noise.length})`, - ); - if (d.pdgOnly.all.length) - out.push( - ` pdg-only: ${d.pdgOnly.all.length} ` + - `(true ${d.pdgOnly.true.length}, noise ${d.pdgOnly.noise.length})`, - ); + // PDG line slice: F1 vs intra_AIS, with the false-positive / false-negative lines. + const ps = d.pdg.score; + out.push( + ` pdg line/intra : P=${fmt(ps.precision)} R=${fmt(ps.recall)} F1=${fmt(ps.f1)} ` + + `|CIS|=${d.pdg.affectedStatementCount} blocks=${d.pdg.blockCount} ` + + `FPIS=${ps.fpisCount} FNIS=${ps.fnisCount}`, + ); + if (ps.fpisCount > 0) out.push(` FPIS(noise): ${ps.fpis.join(', ')}`); + if (ps.fnisCount > 0) out.push(` FNIS(missed): ${ps.fnis.join(', ')}`); + // Callgraph symbol reach: F1 vs inter_AIS. + const cs = d.cg.score; + out.push( + ` cg symbol/inter: P=${fmt(cs.precision)} R=${fmt(cs.recall)} F1=${fmt(cs.f1)} ` + + `|CIS|=${d.cg.count} FPIS=${cs.fpisCount} FNIS=${cs.fnisCount}`, + ); + if (cs.fnisCount > 0) out.push(` FNIS(missed): ${cs.fnis.join(', ')}`); } out.push(''); if (degradedCheck) { diff --git a/gitnexus/bench/impact-pdg/metrics.mjs b/gitnexus/bench/impact-pdg/metrics.mjs index 5f100feb0..e3c357247 100644 --- a/gitnexus/bench/impact-pdg/metrics.mjs +++ b/gitnexus/bench/impact-pdg/metrics.mjs @@ -8,20 +8,36 @@ * (Arch-review Issue 5). `measure.mjs` imports these for the live loop. * * ── CIS / AIS framing (KTD9 — Arnold–Bohner) ─────────────────────────────── - * CIS = Computed Impact Set: the symbols a mode REPORTS as impacted. - * AIS = Actual Impact Set: the curated ground-truth symbols truly affected. + * CIS = Computed Impact Set: what a mode REPORTS as impacted. + * AIS = Actual Impact Set: the curated ground-truth set truly affected. * precision = |AIS∩CIS| / |CIS| (over-approximation cost; ∅ CIS ⇒ undefined) * recall = |AIS∩CIS| / |AIS| (under-approximation; ∅ AIS ⇒ undefined) * F1 = harmonic mean (undefined if either is undefined) * FPIS = CIS − AIS (false positives — noise) * FNIS = AIS − CIS (false negatives — the DANGEROUS miss for a safety tool) * - * ── Granularity (locked) ─────────────────────────────────────────────────── - * Symbol granularity, NEVER block-id (block ids carry fragile fnLine:fnCol:idx). - * A symbol key is `@` — order-independent, line-collapsed. An - * `intra_AIS` entry (statement-granular: lines within the criterion function) - * collapses to its OWNING symbol; an `inter_AIS` entry already names a whole - * symbol. This is why per-scope intra-AIS is the singleton {criterion}. + * ── Granularity: the two engines measure DIFFERENT scopes (U7 rework) ─────── + * The two `impact` engines answer different questions at different granularities, + * so they are scored against different ground truths: + * + * - **PDG mode** (`impact({mode:'pdg', line:N})`) is scored at intra-procedural + * STATEMENT granularity. The statement-anchored slice returns + * `affectedStatements: {line,filePath,text}[]` — the dependent statements of + * the criterion line N. CIS_pdg is the set of those LINE keys + * (`:`, via `pdgLineCis`); AIS is the `intra_AIS` line set + * (via `intraLineAis`). This is the unit at which PDG is precise. + * - **Call-graph mode** is scored at inter-procedural SYMBOL granularity. CIS is + * the reported symbols (`@`, via `symbolKey`); AIS is the + * `inter_AIS` symbol set (via `aisByScope`). This is the unit at which the + * call-graph blast radius is meaningful. + * + * Neither is a strict refinement of the other: PDG resolves the dependent + * statements WITHIN a function but cannot cross a call boundary; call-graph + * resolves the cross-function symbol reach but cannot see below function + * granularity. The harness reports both, side by side, per locus stratum. + * + * `partitionCisByScope`/`aisByScope` (symbol-level) remain for the call-graph + * path; `pdgLineCis`/`intraLineAis` (line-level) drive the PDG path. */ /** Order-independent symbol key. Collapses statement lines onto their symbol. */ @@ -29,6 +45,40 @@ export function symbolKey(symbol, filePath) { return `${symbol}@${filePath}`; } +/** + * Statement-LINE key for the PDG path (U7 rework). PDG mode is now scored at + * intra-procedural STATEMENT granularity: the `impact({mode:'pdg', line})` slice + * returns `affectedStatements: {line, filePath, text}[]`, and the intra ground + * truth is the per-line `intra_AIS`. A line key is `:` — + * order-independent and statement-granular (NOT collapsed onto the owning + * symbol the way `symbolKey` is). This is the unit at which PDG is precise. + */ +export function lineKey(filePath, line) { + return `${filePath}:${line}`; +} + +/** CIS_pdg = the set of affected-statement LINE keys from an impact pdg result. */ +export function pdgLineCis(affectedStatements) { + const out = new Set(); + for (const s of affectedStatements ?? []) { + if (s && typeof s.line === 'number' && typeof s.filePath === 'string') { + out.add(lineKey(s.filePath, s.line)); + } + } + return out; +} + +/** AIS_intra = the set of `intra_AIS` LINE keys (statement-granular ground truth). */ +export function intraLineAis(gt) { + const out = new Set(); + for (const e of gt.intra_AIS ?? []) { + if (e && typeof e.line === 'number' && typeof e.filePath === 'string') { + out.add(lineKey(e.filePath, e.line)); + } + } + return out; +} + /** Canonicalize an iterable of {symbol,filePath} (or pre-made keys) → a Set. */ export function toKeySet(entries) { const out = new Set(); @@ -210,10 +260,14 @@ export function canonicalizeAnnotationSet(fixtures) { .map((fx) => { const c = fx.gt.criterion; const kinds = Array.isArray(c.pdgEdgeKinds) ? [...c.pdgEdgeKinds].sort().join(',') : '-'; + // `line` is the criterion's 1-based statement anchor (U7 — the seed of the + // statement-anchored PDG slice). It is part of the ground truth: changing + // which statement the slice seeds on changes the measured PDG impact set, + // so an unreviewed `criterion.line` edit MUST trip the fingerprint gate. return [ `case=${fx.name}`, `schema=${fx.gt.schemaVersion}`, - `crit=${c.name}|${c.filePath}|${c.direction}|${c.marker ?? '-'}|${kinds}`, + `crit=${c.name}|${c.filePath}|${c.direction}|${c.line ?? '-'}|${c.marker ?? '-'}|${kinds}`, `locus=${fx.gt.locus}`, `pdgScoring=${fx.gt.pdgScoring ?? '-'}`, `provenance=${fx.gt.provenance}`, diff --git a/gitnexus/test/integration/impact-pdg-fixtures.test.ts b/gitnexus/test/integration/impact-pdg-fixtures.test.ts index f2540dd41..53ebe08ff 100644 --- a/gitnexus/test/integration/impact-pdg-fixtures.test.ts +++ b/gitnexus/test/integration/impact-pdg-fixtures.test.ts @@ -42,6 +42,15 @@ interface Criterion { name: string; filePath: string; direction: string; + /** + * The 1-based source line of the statement being changed — the seed of the + * statement-anchored PDG slice (`impact({mode:'pdg', line})`, U7 rework). Set + * from SOURCE SEMANTICS (the def/criterion whose change propagates to the + * intra_AIS lines), validated against the live traversal in the harness's + * Step 0. Required for every measurable case; omitted on excluded no-body + * cases (which carry no statement to seed). + */ + line?: number; marker?: string; /** * The PDG edge kinds the criterion function is EXPECTED to produce. A pure @@ -216,6 +225,15 @@ describe('U6 — impact-PDG fixture ground-truth schema', () => { expect(src.includes(c.marker!), `marker ${JSON.stringify(c.marker)} in source`).toBe( true, ); + // Measurable cases carry a 1-based statement anchor (criterion.line) — + // the seed of the PDG slice (U7 rework). It must be a positive integer + // pointing at an actual source line of the criterion file. + expect(typeof c.line, `${fx.name} needs a 1-based criterion.line`).toBe('number'); + expect(Number.isInteger(c.line!) && c.line! >= 1, `${fx.name} criterion.line >= 1`).toBe( + true, + ); + const lineCount = src.split('\n').length; + expect(c.line! <= lineCount, `${fx.name} criterion.line within file`).toBe(true); // Measurable cases declare which PDG edge kinds the criterion produces. expect(Array.isArray(c.pdgEdgeKinds), `${fx.name} needs criterion.pdgEdgeKinds`).toBe( true, diff --git a/gitnexus/test/unit/impact-pdg-metric-math.test.ts b/gitnexus/test/unit/impact-pdg-metric-math.test.ts index e5ca99293..f9d1e85f4 100644 --- a/gitnexus/test/unit/impact-pdg-metric-math.test.ts +++ b/gitnexus/test/unit/impact-pdg-metric-math.test.ts @@ -82,6 +82,61 @@ describe('impact-pdg metric math — score()', () => { }); }); +describe('impact-pdg metric math — PDG line granularity (U7 rework)', () => { + it('pdgLineCis builds : keys from affectedStatements', () => { + const cis = M.pdgLineCis([ + { line: 10, filePath: 'src/a.ts', text: 'sum = sum + x;' }, + { line: 12, filePath: 'src/a.ts', text: 'return sum;' }, + // a malformed entry (no numeric line) is dropped, never keyed. + { filePath: 'src/a.ts', text: 'noise' }, + ]); + expect([...cis].sort()).toEqual(['src/a.ts:10', 'src/a.ts:12']); + }); + + it('intraLineAis builds line keys from intra_AIS; empty intra_AIS ⇒ empty set', () => { + const withLines = M.intraLineAis({ + intra_AIS: [ + { symbol: 'total', filePath: 'src/a.ts', line: 10 }, + { symbol: 'total', filePath: 'src/a.ts', line: 12 }, + ], + }); + expect([...withLines].sort()).toEqual(['src/a.ts:10', 'src/a.ts:12']); + // an inter fixture has empty intra_AIS ⇒ empty line AIS (recall n/a, not 0). + expect(M.intraLineAis({ intra_AIS: [] }).size).toBe(0); + }); + + it('scores a line slice that exactly matches intra_AIS ⇒ P=R=F1=1 (the accumulator)', () => { + // The verified accumulator case: line-8 slice returns {10,12} = intra_AIS. + const cis = M.pdgLineCis([ + { line: 10, filePath: 'src/accumulator.ts' }, + { line: 12, filePath: 'src/accumulator.ts' }, + ]); + const ais = M.intraLineAis({ + intra_AIS: [ + { filePath: 'src/accumulator.ts', line: 10 }, + { filePath: 'src/accumulator.ts', line: 12 }, + ], + }); + const s = M.score(cis, ais); + expect(s.precision).toBe(1); + expect(s.recall).toBe(1); + expect(s.f1).toBe(1); + }); + + it('an inter fixture (empty intra_AIS) ⇒ precision 0 on the router noise, recall n/a', () => { + // A pure-inter router's line slice returns its own routing returns as FPIS + // against the empty intra_AIS — the by-design "PDG is intra-procedural" case. + const cis = M.pdgLineCis([ + { line: 24, filePath: 'src/d.ts' }, + { line: 26, filePath: 'src/d.ts' }, + ]); + const s = M.score(cis, M.intraLineAis({ intra_AIS: [] })); + expect(s.precision).toBe(0); // 2 predicted, none in (empty) intra truth + expect(s.recall).toBe(null); // |AIS|=0 ⇒ recall n/a + expect(s.f1).toBe(null); + }); +}); + describe('impact-pdg metric math — compareModes()', () => { it('Jaccard + directional set-diffs split true/noise', () => { // callgraph finds {a,b,c} (a,b real, c noise); pdg finds {b,d} (b real, d noise). @@ -232,6 +287,42 @@ describe('impact-pdg metric math — annotation fingerprint (KTD10)', () => { expect(edited).not.toBe(base); }); + it('trips when criterion.line changes (the PDG slice seed — U7 rework)', () => { + // criterion.line is part of the ground truth: it seeds the statement-anchored + // PDG slice, so changing it changes the measured impact set and MUST trip. + const base = M.fingerprintAnnotationSet( + [ + fx({ + criterion: { + name: 'f', + filePath: 'src/f.ts', + direction: 'downstream', + line: 8, + marker: 'x', + pdgEdgeKinds: ['REACHING_DEF'], + }, + }), + ], + fakeHash, + ); + const moved = M.fingerprintAnnotationSet( + [ + fx({ + criterion: { + name: 'f', + filePath: 'src/f.ts', + direction: 'downstream', + line: 9, + marker: 'x', + pdgEdgeKinds: ['REACHING_DEF'], + }, + }), + ], + fakeHash, + ); + expect(moved).not.toBe(base); + }); + it('trips when the criterion direction flips', () => { const base = M.fingerprintAnnotationSet([fx({})], fakeHash); const flipped = M.fingerprintAnnotationSet(