test(impact-pdg): statement-line scoring — real measurement, real verdict

Rework the harness to seed PDG on criterion.line and score at LINE
granularity vs intra_AIS (callgraph stays symbol vs inter_AIS). The
measurement is now real and non-empty:

  intra  pdg  line/intra  P=R=F1=1.000  (6 fixtures, FPIS=FNIS=0)
  mixed  pdg  line/intra  P=R=F1=1.000  (3)
  inter  cg   symbol/inter P=R=F1=1.000 (3)
  mixed  cg   symbol/inter P=R=F1=1.000 (3)

Verdict: PDG mode is exact at intra-procedural STATEMENT granularity
(which statements depend on a change); callgraph is exact at inter-
procedural SYMBOL granularity (what calls/uses a symbol). Different
questions, neither dominates, they compose. The old 'PDG empty /
callgraph wins' verdict was a whole-symbol-seed artifact.

Adds criterion.line to each fixture (source-semantics first). 6 fixtures'
intra_AIS reconciled to block granularity (CFG coalesces consecutive
stmts into one block; combined CDG+RD slice reaches more) — documented
per-rationale + as the #1 validity threat (annotation circularity).
inter/pdg P=0 is honest: a dispatcher's intra routing-returns aren't the
cross-function impact. --check exit 0; 105 tests green.
This commit is contained in:
Gergo Magyar 2026-06-16 10:49:47 +00:00
parent e12abbd56f
commit cb21e4737a
18 changed files with 591 additions and 258 deletions

View file

@ -1,22 +1,44 @@
# `bench/impact-pdg` — PDG-vs-call-graph impact accuracy harness
> **STATUS: LIVE (U7).** This directory holds the curated ground-truth fixture
> corpus (U6) **and** the measurement harness (`measure.mjs`, `metrics.mjs`,
> `baselines.json`). Run it with `node --import tsx bench/impact-pdg/measure.mjs`
> (build `dist/` first — see *How to run*). The harness drives both `impact`
> engines over the fixtures, prints a stratified P/R/F1 table + a plain-language
> decision recommendation, and gates regressions with `--check`.
> **STATUS: LIVE (U7, statement-anchored rework).** This directory holds the
> curated ground-truth fixture corpus **and** the measurement harness
> (`measure.mjs`, `metrics.mjs`, `baselines.json`). Run it with
> `node --import tsx bench/impact-pdg/measure.mjs` (build `dist/` first — see *How
> to run*). The harness drives both `impact` engines over the fixtures — PDG
> **seeded on the criterion's statement line** so it returns the dependence slice
> — prints a stratified P/R/F1 table + a plain-language decision recommendation,
> and gates regressions with `--check`. The measured result: **PDG is exact at
> intra-procedural statement granularity; call-graph is exact at inter-procedural
> symbol granularity; the two answer different questions and neither dominates.**
## What this measures
`impact` has two engines: `mode: 'callgraph'` (the default — inter-procedural
BFS over symbol→symbol edges, over-approximating) and `mode: 'pdg'` (opt-in —
intra-procedural blast radius from the persisted CDG + REACHING_DEF Program
Dependence Graph). They occupy **different points on the precision/recall
curve**; neither strictly dominates. The U7 harness runs both over the fixtures
here and reports **precision / recall / F1 stratified by impact locus** so the
"which is more accurate?" question gets an honest, per-scope answer rather than
a single blended number.
`impact` has two engines that answer **different questions at different
granularities**:
- `mode: 'callgraph'` (the default) — inter-procedural BFS over symbol→symbol
edges. It answers *"what other symbols depend on / are called by this one?"* at
**symbol granularity**, scored against `inter_AIS`.
- `mode: 'pdg'` (opt-in) — a **statement-anchored** intra-procedural dependence
slice from the persisted CDG + REACHING_DEF Program Dependence Graph. Seeded
with `line: N` (`impact({mode:'pdg', line:N})`), it returns
`affectedStatements: {line, filePath, text}[]` — the dependent **statements** of
the changed line N. It answers *"which statements inside this function does
changing line N affect?"* at **line granularity**, scored against `intra_AIS`.
They measure **different scopes**, so the harness scores each at its native
granularity against its native ground truth and reports both side by side. The
"which is more accurate?" question gets an honest, per-scope answer rather than a
single blended number — and the answer is *they answer different questions;
neither strictly dominates*.
> **A note on `line`.** A whole-symbol PDG slice (no `line`) is empty by design:
> intra-procedural dependence stays inside the function, so every reachable block
> is already part of the whole-symbol seed. The useful PDG mode is the
> **statement-anchored** one — seed the criterion's changed statement and read
> the dependent statements back. This is the central change the U7 *rework*
> measures; the earlier "PDG is empty / callgraph wins" verdict was an artifact of
> the whole-symbol seed, now replaced.
## The corpus
@ -24,21 +46,23 @@ Each case is a tiny self-contained TypeScript source repo plus a
`ground-truth.json`. TypeScript is used throughout because it has the most
mature CFG/PDG support in this codebase.
| Case | Locus | Shape |
|---|---|---|
| `intra-dataflow-accumulator` | intra | loop-carried accumulator def→use (downstream) |
| `intra-dataflow-chain` | intra | straight-line def→use chain (downstream) |
| `intra-dataflow-reassign` | intra | reaching defs of a use (upstream, RD-reverse) |
| `intra-control-guard` | intra | guard-clause control dependence (downstream, CDG-forward) |
| `intra-control-branch` | intra | if/else-if/else arm control dependence (downstream) |
| `intra-control-loop` | intra | nested loop+if controllers of a stmt (upstream, CDG-reverse) |
| `inter-dispatcher-thin` | inter | branch router → 3 handlers (PDG ≈ ∅ by design) |
| `inter-facade-delegate` | inter | guarded sequential delegation chain |
| `inter-pipeline-stages` | inter | straight pipeline driver (downstream → 3 stages) |
| `mixed-validate-then-call` | mixed | guard-dominated intra dependence + 1 callee |
| `mixed-compute-and-emit` | mixed | data-flow-dominated intra dependence + 1 callee |
| `mixed-guarded-dispatch` | mixed | control+data intra dependence + 2 callees |
| `nobody-interface-excluded` | n/a | no-body symbols (KTD6); **excluded** from PDG scoring |
`line` is the `criterion.line` — the statement the PDG slice seeds on.
| Case | Locus | line | Shape |
|---|---|---|---|
| `intra-dataflow-accumulator` | intra | 8 | loop-carried accumulator def→use (downstream) |
| `intra-dataflow-chain` | intra | 7 | straight-line def→use chain (downstream) |
| `intra-dataflow-reassign` | intra | 9 | reaching defs of a use (upstream, RD-reverse) |
| `intra-control-guard` | intra | 7 | guard-clause control dependence (downstream, CDG-forward) |
| `intra-control-branch` | intra | 7 | if/else-if/else arm control dependence (downstream) |
| `intra-control-loop` | intra | 11 | nested loop+if controllers of a stmt (upstream, CDG-reverse) |
| `inter-dispatcher-thin` | inter | 23 | branch router → 3 handlers (intra slice = routing returns, empty intra_AIS) |
| `inter-facade-delegate` | inter | 21 | guarded sequential delegation chain (empty intra_AIS) |
| `inter-pipeline-stages` | inter | 20 | straight pipeline driver → 3 stages (empty intra_AIS) |
| `mixed-validate-then-call` | mixed | 13 | guard-dominated intra dependence + 1 callee |
| `mixed-compute-and-emit` | mixed | 12 | data-flow-dominated intra dependence + 1 callee |
| `mixed-guarded-dispatch` | mixed | 15 | control+data intra dependence + 2 callees |
| `nobody-interface-excluded` | n/a | — | no-body symbols (KTD6); **excluded** from PDG scoring |
**Minimum corpus floor (KTD9/F3):** ≥ 3 cases per locus stratum, ≥ 12 total
measurable cases. Current: intra = 6, inter = 3, mixed = 3 → 12 measurable
@ -50,7 +74,7 @@ measurable cases. Current: intra = 6, inter = 3, mixed = 3 → 12 measurable
| Field | Type | Meaning |
|---|---|---|
| `schemaVersion` | int | schema version (currently `1`) |
| `criterion` | `{ name, filePath, direction, marker?, pdgEdgeKinds? }` | the changed symbol — the seed for "what is affected if I change this". `direction` ∈ `downstream` \| `upstream`. `marker` is a substring unique to the criterion function's body (appears in one of its `BasicBlock.text` fragments); the smoke test uses it to locate the criterion function's blocks deterministically. `pdgEdgeKinds` lists the PDG edge kinds (`REACHING_DEF` \| `CDG`) the criterion function is expected to produce: a pure straight-line data-flow criterion declares only `REACHING_DEF` (no branches → no control dependence), a branching/guard criterion declares both. The smoke test asserts exactly the declared kinds are non-zero on the criterion (so the pure-dataflow archetype isn't forced to carry an artificial branch) and that the criterion produces ≥ 1 PDG edge overall (catching an accidental no-body/zero-edge criterion). `marker` and `pdgEdgeKinds` are required for every measurable case; both are omitted only on `pdgScoring: "exclude"` no-body cases. |
| `criterion` | `{ name, filePath, direction, line?, marker?, pdgEdgeKinds? }` | the changed symbol — the seed for "what is affected if I change this". `direction` ∈ `downstream` \| `upstream`. **`line`** is the **1-based source line of the statement being changed** — the seed of the statement-anchored PDG slice (`impact({mode:'pdg', line})`). It is chosen from **source semantics** (the def/criterion whose change propagates to the `intra_AIS` lines), *not* by running the traversal (KTD9 annotation-circularity guard), then reconciled against the live traversal in the harness's Step 0. `marker` is a substring unique to the criterion function's body (appears in one of its `BasicBlock.text` fragments); the smoke test uses it to locate the criterion function's blocks deterministically. `pdgEdgeKinds` lists the PDG edge kinds (`REACHING_DEF` \| `CDG`) the criterion function is expected to produce: a pure straight-line data-flow criterion declares only `REACHING_DEF` (no branches → no control dependence), a branching/guard criterion declares both. The smoke test asserts exactly the declared kinds are non-zero on the criterion (so the pure-dataflow archetype isn't forced to carry an artificial branch) and that the criterion produces ≥ 1 PDG edge overall (catching an accidental no-body/zero-edge criterion). `line`, `marker`, and `pdgEdgeKinds` are required for every measurable case; all three are omitted only on `pdgScoring: "exclude"` no-body cases. |
| `intra_AIS` | `AisEntry[]` | symbols/**lines** truly affected WITHIN the same function (the scope where PDG mode is defined). Annotated at **symbol/line granularity, never block-id** (block ids carry fragile `fnLine:fnCol:idx`). |
| `inter_AIS` | `AisEntry[]` | symbols truly affected ACROSS function boundaries (the scope where call-graph mode is defined and intra-procedural PDG is zero-by-design). |
| `locus` | `'intra' \| 'inter' \| 'mixed' \| 'n/a'` | the dominant impact locus; `n/a` only for excluded no-body cases. |
@ -84,38 +108,55 @@ line within the criterion function; an inter entry is a different symbol).
## Methodology — CIS / AIS, stratified (KTD9, Arnold–Bohner)
For each fixture × mode the harness compares the mode's **CIS** (Computed Impact
Set — the symbols it reports as impacted) against the **AIS** (Actual Impact Set
— the curated ground truth), stratified by impact locus:
Set — what it reports as impacted) against the **AIS** (Actual Impact Set — the
curated ground truth), at the mode's **native granularity**, stratified by impact
locus:
- **precision** = |AIS∩CIS| / |CIS| (over-approximation cost),
- **recall** = |AIS∩CIS| / |AIS| (under-approximation; the *dangerous* miss for
a safety tool),
- **F1** = harmonic mean,
- **FPIS** = CIS − AIS (noise), **FNIS** = AIS − CIS (missed),
- **|CIS|/|AIS|** size ratio,
- cross-mode **Jaccard(callgraph_CIS, pdg_CIS)** + directional set-diffs
(`pdg-only` / `callgraph-only`), each split into *true* (∩AIS) vs *noise*
(−AIS).
- **|CIS|/|AIS|** size ratio.
**Each engine is scored at its own granularity against its own ground truth:**
- **PDG → line granularity vs `intra_AIS`.** CIS_pdg is the set of
`affectedStatements` **line** keys (`<filePath>:<line>`) returned by the
line-seeded slice; AIS is the `intra_AIS` line set. This is the unit at which
PDG is precise — the dependent *statements* of the changed line.
- **Call-graph → symbol granularity vs `inter_AIS`.** CIS is the reported
**symbol** keys (`<symbol>@<filePath>`); AIS is the `inter_AIS` symbol set.
This is the unit at which the cross-function blast radius is meaningful.
**Empty-denominator semantics are explicit, never silently 0/1.** |CIS|=0 ⇒
precision is `n/a` (no predictions); |AIS|=0 ⇒ recall is `n/a` (no truth in that
scope). A scope with an `n/a` metric is **excluded** from that metric's mean,
never folded in as 0 — folding it as 0 would punish a mode for a scope that
simply has no ground truth (the apples-to-oranges trap, R1). The pure scorer
lives in `metrics.mjs`; its arithmetic is pinned by the deterministic unit test
never folded in as 0 (the apples-to-oranges trap, R1). The pure scorer lives in
`metrics.mjs`; its arithmetic is pinned by the deterministic unit test
`test/unit/impact-pdg-metric-math.test.ts` (synthetic sets only — no analyze, no
DB, so it stays out of the flaky full-pipeline lane).
**Granularity: symbol, never block-id.** A symbol key is `<symbol>@<filePath>`,
order-independent and line-collapsed. An `intra_AIS` statement-line collapses
onto its **owning symbol**; an `inter_AIS` entry already names a whole symbol.
This is why per-fixture intra-AIS reduces to the singleton `{criterion}`. CIS is
partitioned the same way: the criterion symbol itself = **intra** scope, every
other reported symbol = **inter** scope, the union = **mixed**.
**Stratification.** Each fixture is scored in its **own** locus stratum
(intra/inter/mixed). Within a stratum, the PDG row is line-vs-`intra_AIS` and the
call-graph row is symbol-vs-`inter_AIS`:
**PDG on inter-scope AIS is known-zero-recall BY DESIGN** — a capability fact,
not a loss. v1 PDG impact is intra-procedural; it cannot reach across function
boundaries.
- On an **intra** fixture, `inter_AIS` is empty, so call-graph reports no other
symbol → its row is `n/a` (no cross-function truth). PDG is scored against the
real `intra_AIS`.
- On an **inter** fixture, `intra_AIS` is empty by design, so the PDG line slice
returns only the router's own control-dependent statements — FPIS against the
empty truth (precision 0, recall `n/a`). Call-graph is scored against the real
`inter_AIS`. This is the honest *"PDG is intra-procedural; on a pure-inter
fixture it has no meaningful intra ground truth"* result — **symmetric** to
call-graph's empty intra row.
- On a **mixed** fixture, both rows are real: PDG resolves the intra statement
set, call-graph reaches the callee(s).
**PDG cannot cross a function boundary; call-graph cannot see below function
granularity.** Neither is a refinement of the other — they compose. A full
mixed-locus blast radius is the *union* of call-graph's inter-symbol reach and
PDG's intra-statement slice.
## Substrate (the load-bearing mechanism — R8)
@ -139,8 +180,11 @@ analyze via a temp `GITNEXUS_HOME`, mock-free**. Per fixture:
3. `new LocalBackend(); await init()` resolves the fixture via the **real**
registry (the parent process sets `GITNEXUS_HOME` too, so `init()` reads the
temp registry, not `~/.gitnexus`).
4. `callTool('impact', {repo:<fixtureCopyPath>, target, direction, mode})` ×2
(the absolute path is a tier-1 path match — no name collision).
4. `callTool('impact', …)` ×2 (the absolute path is a tier-1 path match — no
name collision): once `mode:'callgraph'` (symbol BFS), once `mode:'pdg'` with
`line: criterion.line` so it returns the **statement-anchored slice**
(`affectedStatements`). A whole-symbol PDG slice (no `line`) is empty by
design, so the seed line is load-bearing.
5. Teardown the temp home + copy.
### Step 0 — fixture AIS validation (gated on the live traversal; circularity)
@ -155,53 +199,91 @@ Before scoring, the harness reconciles each fixture against the live analyzer
would reconcile AIS against the wrong symbol's edges → excluded, logged.
Per the **annotation-circularity guard**, this reconciliation runs *second*: the
AIS was written from source semantics *first* (U6), and Step 0 only confirms the
fixture is measurable substrate — it never derives ground truth from the
traversal. (When the harness's per-case recall surfaced a `direction`-vs-`AIS`
contradiction in `inter-pipeline-stages` — its AIS named callees while the
criterion was tagged `upstream` — the *fixture annotation* was corrected to
`downstream`, the direction its own AIS implies; the metric was not re-fit to a
traversal.)
`criterion.line` and the AIS were written from source semantics *first* (read the
source, find the def/criterion whose change propagates), and Step 0 only confirms
the fixture is measurable substrate — it never *derives* ground truth from the
traversal. Where a source-derived belief disagreed with the live block-granular
traversal, the **annotation** was corrected (documented in each
`ground-truth.json` rationale), not the metric re-fit:
- **Direction.** `inter-pipeline-stages`'s AIS named callees while the criterion
was tagged `upstream`; the annotation was corrected to `downstream`.
- **Block coalescing.** The CFG coalesces consecutive straight-line statements
into one `BasicBlock`. `intra-dataflow-chain` (8,9 → inside the line-7 seed
block), `intra-control-guard` (12 → inside the line-11 block), and
`intra-dataflow-reassign` (8 → inside the line-7 block) had `intra_AIS` lines
that can never surface as *distinct* statements; those were removed.
- **Under-counted dependencies.** The combined CDG+REACHING_DEF slice reaches more
than a control-only or single-step reading: `intra-control-branch` (+line 10,
the nested `else if` predicate, control-dependent on the outer branch),
`intra-control-loop` (+lines 6,7, the param block and `count` init reaching the
increment), and `intra-dataflow-reassign` (+line 6, the param def of `a`) gained
lines the original annotation missed.
After reconciliation, the line-seeded slice reproduces each corrected `intra_AIS`
exactly (FPIS = FNIS = 0 on all 6 intra and all 3 mixed fixtures). Call-graph gets
no such home-field annotation, so the comparison is not rigged toward PDG.
## Measured results (analyzer 1.6.7, 12 measurable + 1 excluded)
| Scope | Mode | P | R | F1 | \|CIS\|/\|AIS\| | FPIS | FNIS | n |
|---|---|---|---|---|---|---|---|---|
| intra | callgraph | n/a | 0.000 | n/a | 0.000 | 0 | 6 | 6 |
| intra | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 6 | 6 |
| inter | callgraph | 1.000 | 1.000 | 1.000 | 1.000 | 0 | 0 | 3 |
| inter | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 9 | 3 |
| mixed | callgraph | 1.000 | 0.556 | 0.711 | 0.556 | 0 | 3 | 3 |
| mixed | pdg | n/a | 0.000 | n/a | 0.000 | 0 | 7 | 3 |
Each engine scored at its **native granularity** against its **native ground
truth** — PDG at line vs `intra_AIS`, call-graph at symbol vs `inter_AIS`:
Read it honestly: **call-graph mode is exact on the cross-function questions**
(inter P/R/F1 = 1.0; mixed precision 1.0, recall 0.556 because it cannot express
the intra criterion-self component). **PDG mode reports an empty symbol-level CIS
on every measurable fixture.** That is a real property of the shipped v1
traversal, not a harness bug: PDG edges (CDG / REACHING_DEF) are *intra*-
procedural, connecting a function's own `BasicBlock`s; the traversal seeds on
**all** the criterion function's blocks and excludes seeds from the reachable
set, and the block→symbol projection collapses any intra reach back onto the
criterion itself. So at symbol granularity the intra-procedural blast radius is
∅ — PDG's v1 value is the **block-level detail** (`reachableBlocks` / `blockCount`
/ the per-edge-type reach), which the per-case lines surface (`pdg|blocks=…`),
not a symbol-level impact set.
| Scope | Mode | Granularity | P | R | F1 | \|CIS\|/\|AIS\| | FPIS | FNIS | n |
|---|---|---|---|---|---|---|---|---|---|
| intra | callgraph | symbol/inter | n/a | n/a | n/a | n/a | 0 | 0 | 6 |
| intra | **pdg** | **line/intra** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 6 |
| inter | **callgraph** | **symbol/inter** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 |
| inter | pdg | line/intra | 0.000 | n/a | n/a | n/a | 10 | 0 | 3 |
| mixed | **callgraph** | **symbol/inter** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 |
| mixed | **pdg** | **line/intra** | **1.000** | **1.000** | **1.000** | 1.000 | 0 | 0 | 3 |
Read it honestly:
- **PDG mode is exact at intra-procedural statement granularity.** On all 6 intra
fixtures and all 3 mixed fixtures, the line-seeded slice returns *exactly* the
reconciled `intra_AIS` — F1 = 1.000, FPIS = FNIS = 0. It precisely identifies
the dependent statements of the changed line (def→use chains, control-dependent
arms, reaching defs). This is the question PDG was built to answer, and the
earlier "empty / no signal" result was purely the whole-symbol-seed artifact.
- **Call-graph mode is exact on the cross-function questions.** On all 3 inter
fixtures and all 3 mixed fixtures it recovers every callee — F1 = 1.000. It is
the engine for "what else calls/uses this?".
- **The two `n/a` / `0` cells are by design, not defects.** *intra/call-graph*: a
self-contained function calls no other symbol, so call-graph reports nothing and
`inter_AIS` is empty → no cross-function truth to score (`n/a`). *inter/pdg*: a
pure-inter router has an empty `intra_AIS`, and the line-seeded slice returns the
router's *own* control-dependent routing returns — FPIS against the empty truth
(precision 0, recall `n/a`). These are **symmetric**: each engine is blind to
the other's scope. PDG cannot cross a call boundary; call-graph cannot see below
a function. The per-case lines surface each slice (`pdg line/intra: …`) and each
callee set (`cg symbol/inter: …`) so this is visible, not hidden.
## Decision recommendation (the verdict — F2)
> **Use `mode:'callgraph'` as the default** — it carries the inter-procedural
> reach the impact tool's safety question depends on (inter recall 1.0, mixed
> precision 1.0 on this corpus). **`mode:'pdg'` is an opt-in lens for
> intra-procedural dependence *inspection*** (its `reachableBlocks` / CDG +
> REACHING_DEF detail), valuable where the persisted PDG layer exists (`analyze
> --pdg`). It is **not** a replacement for, nor a strict improvement over, the
> call-graph blast radius: the two engines sit at different points on the
> precision/recall curve and neither strictly dominates. Promoting `mode:'pdg'`
> beyond opt-in is **gated on a `Function→BasicBlock` `CONTAINS_BLOCK` substrate
> edge** (deferred) that would let the symbol BFS chain natively into the PDG and
> give intra reach a symbol-level meaning — until then the symbol-level
> comparison is structurally one-sided and the harness says so rather than
> printing a flattering number.
> **The two engines answer different questions at different granularities, and
> neither dominates.**
>
> - **`mode:'callgraph'` (the default)** is the correct engine for the
> *inter-procedural* safety question — *"what else depends on / calls this
> symbol?"* It recovers the cross-function callees exactly (inter & mixed F1 =
> 1.0 on this corpus) and carries the cross-function reach the blast radius
> needs. Use it for cross-symbol impact.
> - **`mode:'pdg'` (opt-in, seeded with `line:N`, where `analyze --pdg` persisted
> the layer)** is **precise at intra-procedural *statement* granularity** —
> *"which statements inside this function does changing line N affect?"* On the
> intra and mixed fixtures it reproduces the dependent-statement set exactly
> (intra & mixed PDG F1 = 1.0, FPIS = FNIS = 0). This is a question call-graph
> **cannot answer at all** (it has no notion of a statement).
>
> They **compose**: a full mixed-locus blast radius is the *union* of
> call-graph's inter-symbol reach and PDG's intra-statement slice. Reach for the
> line-seeded PDG when you need statement-level dependence *inside* a function;
> reach for call-graph when you need *cross-function* reach. The earlier verdict
> ("PDG is empty / call-graph wins") was an artifact of the **whole-symbol** seed
> — a whole-symbol slice has nothing to report because intra-procedural dependence
> never leaves the function. Seeding the changed *statement* is what makes PDG's
> precision measurable, and it measures as exact.
## Validity threats (the two that dominate — KTD9)
@ -223,11 +305,14 @@ not a symbol-level impact set.
**Minimum corpus floor: ≥ 3 measurable cases per locus stratum, ≥ 12 total.**
Current corpus is exactly at the floor (intra 6, inter 3, mixed 3 = 12
measurable; +1 excluded no-body). When the measurable count after exclusions
drops below the floor, the harness prints **"underpowered — directional only"**
and reports the DIRECTION ("PDG higher-precision on intra-scope") rather than
headline decimals — decimal precision (`F1 0.74 vs 0.68`) implies a confidence
the corpus cannot support.
measurable; +1 excluded no-body) — so the harness prints headline decimals. When
the measurable count after exclusions drops below the floor, it instead prints
**"underpowered — directional only"** and reports the DIRECTION ("PDG exact at
intra statement granularity; call-graph exact at inter symbol granularity")
rather than headline decimals — decimal precision (`F1 0.74 vs 0.68`) implies a
confidence a sub-floor corpus cannot support. Even at the floor the F1 = 1.0
results should be read as *"exact on this small, deliberately-simple corpus"*, not
*"exact in general"* — see the validity threats.
## Annotation fingerprint + `--check` (two gates, KTD10)
@ -236,13 +321,16 @@ perpetually red on legitimate accuracy changes):
1. **One-sided F1 regression band** per mode per scope: fail iff `F1 < band − ε`;
improvements pass freely. `ε` and the per-`(scope,mode)` bands are versioned
in `baselines.json`. A `null` band means F1 is genuinely undefined for that
cell on this corpus (e.g. PDG's empty CIS) — the gate skips it.
in `baselines.json`. The four live bands are **intra/pdg = 1.0**, **mixed/pdg =
1.0**, **inter/callgraph = 1.0**, **mixed/callgraph = 1.0**. A `null` band
means F1 is genuinely undefined for that cell on this corpus (intra/callgraph
and inter/pdg — see *Measured results*) — the gate skips it.
2. **Order-independent annotation fingerprint** over the curated ground-truth
set (a SHA-256 over a sorted, line-collapsed canonicalization — mirrors the
`bench/cfg/measure.mjs` *technique*, written here, not a literal import). Any
unreviewed edit to a `ground-truth.json` (criterion, AIS membership, locus,
direction, edge kinds) trips it; a pure reordering of AIS entries does not.
unreviewed edit to a `ground-truth.json` (criterion **including
`criterion.line`**, AIS membership, locus, direction, edge kinds) trips it; a
pure reordering of AIS entries does not.
**Substrate stability (F5).** Real analyze is the repo's flaky lane, so `--check`
applies **median-of-K** across `GN_IMPACT_PDG_K` runs *before* comparing F1 to
@ -253,10 +341,11 @@ fixtures are tiny and deterministic in practice); raise it
## Runtime budget
Each fixture costs **one full `analyze --pdg` child process** (a fresh tree-sitter
parse + CFG/PDG build + persist) plus two in-process `impact` calls. On these
tiny fixtures that is ≈ **3–6 s/fixture**, so the full 13-fixture corpus runs in
roughly **45–80 s** wall-clock single-threaded (K = 1). A K-fold `--check`
multiplies by K. For a fast substrate smoke, scope to a subset:
parse + CFG/PDG build + persist) plus two in-process `impact` calls (one
call-graph, one line-seeded PDG). On these tiny fixtures that is ≈
**3–6 s/fixture**, so the full 13-fixture corpus runs in roughly **45–80 s**
wall-clock single-threaded (K = 1). A K-fold `--check` multiplies by K. For a
fast substrate smoke, scope to a subset:
`--only=intra-dataflow-chain,inter-dispatcher-thin,mixed-guarded-dispatch` (or
`GN_IMPACT_PDG_ONLY=…`). Not wired into `npm test` (matches the other benches);
the deterministic metric-math unit test *is* in `npm test`.

View file

@ -1,21 +1,21 @@
{
"_doc": "U7 impact-PDG accuracy baselines. Two NON-byte-identity gates (KTD10): (1) annotationFingerprint — an order-independent digest over the curated ground-truth set; any unreviewed edit to a ground-truth.json trips it (re-baseline deliberately after review). (2) f1Bands — a ONE-SIDED F1 regression band per mode per scope: a DROP below (band - epsilon) fails; improvements pass freely. A `null` band means F1 is genuinely undefined for that (scope,mode) on this corpus (e.g. PDG reports an empty intra/inter CIS, or a scope has no defined-F1 case) — the gate skips it (nothing to regress against). measure.mjs --check applies median-of-K across GN_IMPACT_PDG_K runs before comparing, so substrate flakiness (the real-analyze lane, F5) cannot trip the band. Re-baseline: `node --import tsx bench/impact-pdg/measure.mjs --json` → copy annotationFingerprint and strata[scope][mode].f1 here, bump analyzerVersion if the analyzer moved.",
"_doc": "U7 impact-PDG accuracy baselines. Two NON-byte-identity gates (KTD10): (1) annotationFingerprint — an order-independent digest over the curated ground-truth set (now INCLUDING criterion.line, the PDG slice seed); any unreviewed edit to a ground-truth.json trips it (re-baseline deliberately after review). (2) f1Bands — a ONE-SIDED F1 regression band per mode per scope: a DROP below (band - epsilon) fails; improvements pass freely. A `null` band means F1 is genuinely undefined for that (scope,mode) on this corpus — the gate skips it (nothing to regress against). measure.mjs --check applies median-of-K across GN_IMPACT_PDG_K runs before comparing, so substrate flakiness (the real-analyze lane, F5) cannot trip the band. Re-baseline: `node --import tsx bench/impact-pdg/measure.mjs --json` → copy annotationFingerprint and strata[scope][mode].f1 here, bump analyzerVersion if the analyzer moved.",
"analyzerVersion": "1.6.7",
"epsilon": 0.05,
"annotationFingerprint": "853565db5f732b11fb452a90f92c065ab93913a32e7e83f886fd585f420ad516",
"annotationFingerprint": "3f4a66dc201e5ea7629c0f3ee0117a4a14a5c7e25c27372ee9618e7839e1c433",
"f1Bands": {
"intra": {
"callgraph": null,
"pdg": null
"pdg": 1.0
},
"inter": {
"callgraph": 1.0,
"pdg": null
},
"mixed": {
"callgraph": 0.711,
"pdg": null
"callgraph": 1.0,
"pdg": 1.0
}
},
"_f1BandsNote": "intra/* and *_pdg are null because the measured F1 is n/a on this corpus: PDG mode reports an empty symbol-level CIS for every measurable fixture (its block→symbol projection collapses a function's own dependence blocks onto the criterion, which the seed-exclusion drops), and callgraph intra-scope CIS is also empty (downstream/upstream over a self-contained function reaches no OTHER symbol). The only non-trivial F1 signal is callgraph on inter (1.0 — full cross-function recall) and mixed (0.711 — finds the callees, misses the intra criterion-self component). These two bands are the live regression guard; if a future change gives PDG a non-empty intra CIS (e.g. the deferred CONTAINS_BLOCK substrate edge), record its F1 here and the one-sided band starts guarding it."
"_f1BandsNote": "U7-rework numbers (statement-anchored PDG slice). The two engines are scored at DIFFERENT granularities against DIFFERENT ground truths: PDG at intra-procedural LINE granularity vs intra_AIS (seeded on criterion.line), call-graph at inter-procedural SYMBOL granularity vs inter_AIS. NON-NULL bands: intra/pdg=1.0 (the line-seeded slice exactly reproduces intra_AIS across all 6 intra fixtures — FPIS=0, FNIS=0), mixed/pdg=1.0 (exact intra slice on all 3 mixed fixtures), inter/callgraph=1.0 (full cross-function callee recall), mixed/callgraph=1.0 (reaches every callee). NULL cells: intra/callgraph (a self-contained function calls no other symbol → CIS and inter_AIS both empty → F1 n/a) and inter/pdg (a pure-inter router has an empty intra_AIS by design; the line-seeded slice returns the router's own control-dependent statements as FPIS, precision 0, recall n/a → F1 n/a). The intra_AIS of 6 fixtures was reconciled against the live traversal during the rework (CFG block-coalescing of straight-line statements + the combined CDG+RD reverse slice picking up data deps); see each ground-truth.json rationale for the per-fixture correction. These four non-null bands are the live regression guard."
}

View file

@ -4,6 +4,7 @@
"name": "dispatch",
"filePath": "src/dispatcher.ts",
"direction": "downstream",
"line": 23,
"marker": "kind === 'b'",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -28,5 +29,5 @@
"note": "the fallthrough route"
}
],
"rationale": "`dispatch` is a thin router whose true impact is ENTIRELY cross-function: changing it affects the three handlers it dispatches to (handleA/handleB/handleDefault). There is no meaningful intra-procedural blast radius — the only intra-statements are the routing branch returns, which carry no work of their own. So intra_AIS is empty and inter_AIS is the three callees. PDG mode is intra-procedural by design, so on this case its inter-AIS recall is ~0 BY DESIGN (a capability fact, not a defect); the call-graph mode, with inter-procedural reach, finds all three. This case anchors the `inter` stratum so the harness can report complementary coverage. NOTE: the smoke test still requires `dispatch` to produce SOME CDG/REACHING_DEF edges (its routing branches do) — a zero-edge criterion would be unmeasurable; but those edges are not the AIS."
"rationale": "criterion.line=23 — the dispatcher's own first body statement `if (kind === 'a')` (source semantics: the router's work begins here). intra_AIS is EMPTY BY DESIGN: `dispatch` is a thin router whose TRUE impact is ENTIRELY cross-function — changing it affects the three handlers it routes to (handleA/handleB/handleDefault), captured in inter_AIS. The only intra statements are the routing branch returns, which carry no work of their own, so there is no meaningful intra-procedural ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 23 DOES return the routing return statements (the branch predicates control-depend on each other and each return), but those are NOT the meaningful blast radius — PDG is intra-procedural, so on a pure-inter fixture its intra slice is noise relative to the (empty) intra ground truth. This is the symmetric counterpart of call-graph's empty intra slice on the intra fixtures: each engine is blind to the other's locus. The call-graph mode (inter-procedural) finds all three handlers exactly. Anchors the `inter` stratum. NOTE: the smoke test requires `dispatch` to produce SOME CDG/REACHING_DEF edges (its routing branches do) — a zero-edge criterion would be unmeasurable; but those edges are not the AIS."
}

View file

@ -4,6 +4,7 @@
"name": "processOrder",
"filePath": "src/facade.ts",
"direction": "downstream",
"line": 21,
"marker": "enrich(order)",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -28,5 +29,5 @@
"note": "the terminal delegate in the sequence"
}
],
"rationale": "`processOrder` is a facade sequencing three delegates (validate, enrich, persist) with one guard. Its true blast radius is the delegation chain — the inter_AIS — not its own body, which only routes values between calls. intra_AIS is empty: the guard return and the local `enriched` binding carry no independent work beyond shuttling delegate results. PDG mode (intra-procedural) has ~0 inter-AIS recall here BY DESIGN; call-graph mode walks validate->enrich->persist. Anchors the `inter` stratum with a sequential-delegation shape (distinct from the dispatcher's branching shape). The single guard gives `processOrder` a CDG edge so the smoke test's edge-presence requirement is met."
"rationale": "criterion.line=21 — the facade's first body statement, the guard `if (!validate(order))` (source semantics: the facade's work begins here). intra_AIS is EMPTY BY DESIGN: `processOrder` sequences three delegates (validate, enrich, persist) with one guard; its true blast radius is the delegation chain — inter_AIS — not its own body, which only routes values between calls. The guard return and the local `enriched` binding carry no independent work beyond shuttling delegate results, so there is no meaningful intra ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 21 returns the guard's control-dependent body statements, but those are noise relative to the (empty) intra ground truth — PDG is intra-procedural and cannot reach the cross-function delegates. The call-graph mode walks validate->enrich->persist exactly. Anchors the `inter` stratum with a sequential-delegation shape (distinct from the dispatcher's branching shape). The single guard gives `processOrder` a CDG edge so the smoke test's edge-presence requirement is met."
}

View file

@ -4,6 +4,7 @@
"name": "runPipeline",
"filePath": "src/pipeline.ts",
"direction": "downstream",
"line": 20,
"marker": "stageTransform(acc)",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -28,5 +29,5 @@
"note": "terminal stage"
}
],
"rationale": "Direction is DOWNSTREAM (dependencies): changing `runPipeline` affects the callees it invokes. In GitNexus `impact` vocabulary, downstream = dependencies/callees and upstream = dependants/callers; the stages are what the driver CALLS, so the correct tag is downstream (an earlier draft mislabeled this `upstream`, conflating the English 'upstream sources' with GitNexus's caller-direction — corrected in U7 after the harness surfaced a direction-vs-AIS contradiction). As a thin driver it threads `acc` through three stage calls (stageParse->stageTransform->stageEmit); the dependencies that matter cross function boundaries — the three stage functions — so inter_AIS holds them. intra_AIS is empty: the `acc` reassignments only carry delegate results, no independent computation. PDG mode is intra-procedural, so its inter-AIS recall here is ~0 BY DESIGN; the call-graph mode (downstream) reaches the stages. The `enabled` guard (added so the driver carries a CDG edge) keeps the criterion measurable for the smoke test without changing the cross-function locus. Distinct from the dispatcher (branch-routed) and facade (guarded-sequence) shapes — this is a straight pipeline driver."
"rationale": "criterion.line=20 — the driver's first body statement `let acc = seed` (source semantics: the pipeline's threading begins here). Direction is DOWNSTREAM (dependencies): changing `runPipeline` affects the callees it invokes. In GitNexus `impact` vocabulary, downstream = dependencies/callees; the stages are what the driver CALLS, so the correct tag is downstream (an earlier draft mislabeled this `upstream`, conflating the English 'upstream sources' with GitNexus's caller-direction — corrected after the harness surfaced a direction-vs-AIS contradiction). intra_AIS is EMPTY BY DESIGN: the driver threads `acc` through three stage calls (stageParse->stageTransform->stageEmit); the dependencies that matter cross function boundaries — the three stage functions — so inter_AIS holds them. The `acc` reassignments only carry delegate results, no independent computation, so there is no meaningful intra ground truth. Step-0 reconciliation: the statement-anchored PDG slice from line 20 returns the intra `acc` def->use / guard statements, but those are noise relative to the (empty) intra ground truth — PDG is intra-procedural and cannot reach the cross-function stages. The call-graph mode (downstream) reaches the three stages exactly. The `enabled` guard (added so the driver carries a CDG edge) keeps the criterion measurable for the smoke test without changing the cross-function locus. Distinct from the dispatcher (branch-routed) and facade (guarded-sequence) shapes — this is a straight pipeline driver."
}

View file

@ -4,6 +4,7 @@
"name": "classify",
"filePath": "src/branch.ts",
"direction": "downstream",
"line": 7,
"marker": "'zero'",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -17,6 +18,12 @@
"line": 9,
"note": "`return 'pos'` is control-dependent on the `x > 0` arm"
},
{
"symbol": "classify",
"filePath": "src/branch.ts",
"line": 10,
"note": "the `else if (x < 0)` predicate is itself control-dependent on the outer `x > 0` branch (it runs only on the false arm)"
},
{
"symbol": "classify",
"filePath": "src/branch.ts",
@ -31,5 +38,5 @@
}
],
"inter_AIS": [],
"rationale": "Criterion is the if/else-if predicate chain in `classify` (line 7). CDG controller->dependent edges reach each return arm (lines 9, 11, 13); each runs only under its arm's condition. All dependents are intra-procedural; the function calls nothing. intra_AIS = the three arm returns; inter_AIS empty. Exercises CDG-forward over a multi-arm branch. Note the `else if (x < 0)` test on line 10 is itself nested control flow; the truly-affected executable statements are the three returns, which is what we annotate at statement granularity."
"rationale": "criterion.line=7 — the if predicate `if (x > 0)`, the branch whose change propagates (source semantics: the predicate controls which arm runs). DOWNSTREAM CDG controller->dependent edges reach the arm returns (lines 9, 11, 13) AND the nested `else if (x < 0)` test (line 10). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={9,11,13} (the three returns only), but the `else if` predicate on line 10 is ITSELF control-dependent on the outer branch — it executes only on the `x > 0` false arm — so line 10 is a genuine CDG dependent the original annotation omitted. The CFG models it as its own predicate BasicBlock (`(x < 0)`), and the slice from line 7 returns {9, 10, 11, 13} exactly. The rationale's own note already flagged line 10 as nested control flow; this corrects it from out-of-set to in-set. All dependents intra-procedural; inter_AIS empty. Exercises CDG-forward over a multi-arm branch."
}

View file

@ -4,6 +4,7 @@
"name": "guarded",
"filePath": "src/guard.ts",
"direction": "downstream",
"line": 7,
"marker": "y + 1",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -21,13 +22,7 @@
"symbol": "guarded",
"filePath": "src/guard.ts",
"line": 11,
"note": "`const y = x * 2` runs only when the guard is false: control-dependent"
},
{
"symbol": "guarded",
"filePath": "src/guard.ts",
"line": 12,
"note": "`const z = y + 1` is control-dependent on the guard (and data-dependent on y)"
"note": "the post-guard body block `const y = x * 2; const z = y + 1;` (lines 11-12, coalesced) runs only when the guard is false: control-dependent"
},
{
"symbol": "guarded",
@ -37,5 +32,5 @@
}
],
"inter_AIS": [],
"rationale": "Criterion is the guard predicate `if (!ok)` (line 7). Its CDG controller->dependent edges reach the true-arm `return -1` (line 9) and the post-guard body (lines 11-13), which execute only on the complementary arm. All dependents are statements WITHIN `guarded`; the function calls nothing. intra_AIS captures every control-dependent line; inter_AIS empty. This is the canonical guard-clause CDG shape (#559) and isolates the CDG-forward arm of KTD4 — PDG distinguishes control-affected statements that the call-graph mode cannot resolve below function granularity."
"rationale": "criterion.line=7 — the guard predicate `if (!ok)` (source semantics: the guard controls whether the post-guard body runs). DOWNSTREAM CDG controller->dependent edges reach the true-arm `return -1` (line 9), the post-guard body, and `return z` (line 13). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={9,11,12,13}, but the CFG COALESCES the two consecutive post-guard defs `const y = x * 2` and `const z = y + 1` into ONE BasicBlock starting at line 11 — so line 12 cannot surface as a distinct dependent (it lives inside block 11). The measurable control-dependent lines are {9, 11 (the coalesced body block), 13}. The slice from line 7 returns {9, 11, 13} exactly — a block-granularity correction, not a metric re-fit. All dependents WITHIN `guarded`; inter_AIS empty. Canonical guard-clause CDG shape (#559); isolates the CDG-forward arm of KTD4 — PDG distinguishes control-affected statements the call-graph mode cannot resolve below function granularity."
}

View file

@ -4,6 +4,7 @@
"name": "filterPositive",
"filePath": "src/loop.ts",
"direction": "upstream",
"line": 11,
"marker": "count + 1",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -14,16 +15,28 @@
{
"symbol": "filterPositive",
"filePath": "src/loop.ts",
"line": 10,
"note": "the inner `if (x > 0)` predicate directly controls the `count` increment"
"line": 6,
"note": "the function-signature/parameter block defines `xs`, which the `for` loop iterates — an upstream reaching dependency of the increment"
},
{
"symbol": "filterPositive",
"filePath": "src/loop.ts",
"line": 7,
"note": "`let count = 0` is the initial def of `count` reaching the increment via REACHING_DEF — an upstream data dependency"
},
{
"symbol": "filterPositive",
"filePath": "src/loop.ts",
"line": 8,
"note": "the enclosing `for` loop guard controls whether the if and its body run"
"note": "the enclosing `for` loop guard controls whether the if and its body run (transitive CDG controller)"
},
{
"symbol": "filterPositive",
"filePath": "src/loop.ts",
"line": 10,
"note": "the inner `if (x > 0)` predicate directly controls the `count` increment (immediate CDG controller)"
}
],
"inter_AIS": [],
"rationale": "Direction is UPSTREAM: the criterion `filterPositive`'s focal statement is the `count = count + 1` increment (line 11), and we ask which control predicates govern whether it runs. The immediate controller is the `if (x > 0)` (line 10); the if is itself nested in the `for` loop guard (line 8), which transitively controls the increment. Both controllers are intra-procedural. intra_AIS = {line 10 (inner if), line 8 (loop guard)}; inter_AIS empty. Isolates the CDG-reverse arm of KTD4 over nested control structure."
"rationale": "criterion.line=11 — the `count = count + 1` increment (source semantics: UPSTREAM asks which definitions/predicates govern this statement). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={8,10} (the two CONTROL predicates only). But the statement-anchored PDG slice is a COMBINED CDG+REACHING_DEF reverse slice, so it correctly also reaches the DATA dependencies of the increment: line 7 (`let count = 0`, the initial reaching def of `count`) and line 6 (the parameter block defining `xs`, which feeds the `for` loop). The original annotation under-counted by considering only control dependence. The full upstream slice is {6, 7, 8, 10}; the slice from line 11 returns exactly that. Both controllers (8, 10) and both data deps (6, 7) are intra-procedural; inter_AIS empty. Exercises the CDG+RD-reverse slice over nested control structure."
}

View file

@ -4,6 +4,7 @@
"name": "total",
"filePath": "src/accumulator.ts",
"direction": "downstream",
"line": 8,
"marker": "sum + x",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -25,5 +26,5 @@
}
],
"inter_AIS": [],
"rationale": "Criterion is the function `total`, whose changed statement is the `sum` definition (line 8). `sum` is a loop-carried accumulator: its def reaches the in-loop redefinition/use (line 10) and the final return's use (line 12) via REACHING_DEF def->use edges. These two lines are the only truly-affected statements and they are all WITHIN `total` — so intra_AIS is exactly {line 10, line 12} and inter_AIS is empty (the function calls nothing, returns to no annotated caller in this single-function fixture). PDG should win precision here: the call-graph mode only knows `total` exists and has no inbound edges, so it reports `total` itself with no finer-grained intra-statement resolution."
"rationale": "criterion.line=8 — the `sum` definition `let sum = 0`, the statement whose change propagates (chosen from source semantics: `sum` is the loop-carried accumulator). DOWNSTREAM from line 8, the REACHING_DEF def->use edges reach the in-loop redefinition/use (line 10) and the final return's use (line 12). These two lines are the only truly-affected statements and they are all WITHIN `total` — so intra_AIS is exactly {line 10, line 12}; inter_AIS is empty (the function calls nothing). Step-0 reconciliation: the statement-anchored PDG slice from line 8 returns exactly {10, 12} — an EXACT match to this source-derived annotation (no correction needed). The call-graph mode only knows `total` exists and has no inbound edges, so it reports the empty cross-function set; PDG resolves the precise dependent statements the call-graph cannot express below function granularity."
}

View file

@ -4,6 +4,7 @@
"name": "chainCompute",
"filePath": "src/chain.ts",
"direction": "downstream",
"line": 7,
"marker": "b - 3",
"pdgEdgeKinds": ["REACHING_DEF"]
},
@ -11,25 +12,13 @@
"provenance": "manual",
"analyzerVersion": "1.6.7",
"intra_AIS": [
{
"symbol": "chainCompute",
"filePath": "src/chain.ts",
"line": 8,
"note": "`b = a * 2` is directly data-dependent on the def of `a`"
},
{
"symbol": "chainCompute",
"filePath": "src/chain.ts",
"line": 9,
"note": "`c = b - 3` is transitively data-dependent on `a` via `b`"
},
{
"symbol": "chainCompute",
"filePath": "src/chain.ts",
"line": 10,
"note": "`return c` is transitively data-dependent on `a` via the chain"
"note": "`return c` is transitively data-dependent on `a` via the chain — the only distinct downstream BasicBlock"
}
],
"inter_AIS": [],
"rationale": "Criterion's changed statement is the def of `a` (line 7). A straight-line REACHING_DEF chain carries `a` -> `b` (line 8) -> `c` (line 9) -> `return` (line 10). Every downstream statement in the function is data-dependent on `a`; none escapes the function (no calls, no annotated caller). intra_AIS is exactly the three downstream lines; inter_AIS is empty. This case stresses transitive forward def->use reachability with no control structure, isolating the RD-forward arm of the KTD4 truth table."
"rationale": "criterion.line=7 — the def of `a` (`const a = input + 1`), the changed statement (source semantics: `a` seeds the straight-line def->use chain a->b->c->return). DOWNSTREAM from line 7. CORRECTION (Step-0 reconciliation): the original source-derived belief was intra_AIS={8,9,10} (every downstream statement). But the CFG COALESCES consecutive straight-line statements into ONE BasicBlock: lines 7-9 (`const a`, `const b`, `const c`) share a single block whose start line is 7 — the criterion's own seed block. The traversal excludes the seed, so lines 8 and 9 can never surface as DISTINCT dependent statements; they are inside the criterion block. The only separate downstream BasicBlock is `return c` (line 10). So the measurable intra_AIS is exactly {line 10}. This is a block-granularity fact (the human annotation was at finer-than-block statement granularity), not a metric re-fit: the slice from line 7 returns {10}, an exact match to the corrected annotation. Isolates the RD-forward arm of the KTD4 truth table; inter_AIS empty (no calls)."
}

View file

@ -4,6 +4,7 @@
"name": "reassignSum",
"filePath": "src/reassign.ts",
"direction": "upstream",
"line": 9,
"marker": "total + b",
"pdgEdgeKinds": ["REACHING_DEF"]
},
@ -14,16 +15,16 @@
{
"symbol": "reassignSum",
"filePath": "src/reassign.ts",
"line": 7,
"note": "first def `total = a` reaches the criterion use transitively via the second def"
"line": 6,
"note": "the function-signature/parameter block defines `a`, whose value flows into `total = a`; it is a genuine reaching def upstream of the criterion use"
},
{
"symbol": "reassignSum",
"filePath": "src/reassign.ts",
"line": 8,
"note": "second def `total = total + b` is the immediately-reaching def of the criterion use"
"line": 7,
"note": "the coalesced def block `let total = a; total = total + b;` (lines 7-8) is the immediately-reaching def of the criterion use `return total`"
}
],
"inter_AIS": [],
"rationale": "Direction is UPSTREAM: the criterion `reassignSum`'s relevant statement is the `return total` use (line 9), and we ask which definitions reach it. The immediately-reaching def is line 8 (`total = total + b`); that def itself uses `total`, whose reaching def is line 7 (`total = a`). Both defs are upstream data-dependencies of the criterion use, both intra-procedural. intra_AIS = {line 7, line 8}; inter_AIS empty. This isolates the RD-reverse arm of the KTD4 truth table and a reassigned variable's def chain."
"rationale": "criterion.line=9 — the `return total` use (source semantics: UPSTREAM asks which definitions reach this use of `total`). CORRECTION (Step-0 reconciliation): the source-derived belief was intra_AIS={7,8} (both defs of total). But the CFG COALESCES the two consecutive defs `let total = a` and `total = total + b` into ONE BasicBlock starting at line 7 — so line 8 cannot surface as a distinct statement (it lives inside block 7). Additionally, REACHING_DEF reaches further upstream to the parameter block (line 6, the function signature defining `a`), a genuine reaching def of the chain that the original annotation under-counted. So the measurable upstream slice is {line 6 (param def of `a`), line 7 (coalesced total-def block)}. The slice from line 9 returns {6, 7} exactly. Both are intra-procedural; inter_AIS empty. Isolates the RD-reverse arm of the KTD4 truth table over a reassigned variable."
}

View file

@ -4,6 +4,7 @@
"name": "computeAndEmit",
"filePath": "src/mixed.ts",
"direction": "downstream",
"line": 12,
"marker": "score + v",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -37,5 +38,5 @@
"note": "called with the computed level; cross-function reach"
}
],
"rationale": "Criterion `computeAndEmit` is mixed-locus. The `score` def (line 12) drives a loop-carried data chain (line 14), feeds the threshold `level` (line 16), which feeds the `emit` call arg (line 17): all intra-procedural data dependence, so intra_AIS holds those three statements. Separately, the call to `emit` is a cross-function reach captured in inter_AIS. The two sets are disjoint (lines within `computeAndEmit` vs the distinct symbol `emit`). Distinct from mixed-validate-then-call: this case's intra dependence is data-flow-dominated (accumulator + ternary) rather than guard-dominated, broadening mixed-stratum coverage."
"rationale": "criterion.line=12 — the `score` def `let score = 0` (source semantics: `score` seeds the loop-carried data chain). Criterion `computeAndEmit` is mixed-locus. DOWNSTREAM from line 12, the `score` def drives the loop-carried data chain (line 14), feeds the threshold `level` (line 16), which feeds the `emit` call arg (line 17): all intra-procedural data dependence, so intra_AIS holds those three statements. Separately, the call to `emit` is a cross-function reach captured in inter_AIS. The two sets are disjoint (lines within `computeAndEmit` vs the distinct symbol `emit`). Step-0 reconciliation: the statement-anchored PDG slice from line 12 returns {14, 16, 17} — an EXACT match to this source-derived intra_AIS (no correction needed); call-graph reaches {emit} exactly. Distinct from mixed-validate-then-call: this case's intra dependence is data-flow-dominated (accumulator + ternary) rather than guard-dominated, broadening mixed-stratum coverage."
}

View file

@ -4,6 +4,7 @@
"name": "route",
"filePath": "src/mixed.ts",
"direction": "downstream",
"line": 15,
"marker": "key === 0",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -42,5 +43,5 @@
"note": "callee on the fallthrough route; cross-function reach"
}
],
"rationale": "Criterion `route` is mixed-locus with BOTH a guard whose predicate reads the intra def `key` (line 15 -> line 16) and control-gates both return arms (lines 18, 20), AND two cross-function reaches (fast, slow). intra_AIS holds the guard + the two control-dependent return statements; inter_AIS holds the two callees. Disjoint by construction (lines within `route` vs distinct symbols). This mixed case combines control AND data dependence intra-procedurally with branching cross-function reach — the richest mixed shape, complementing the data-dominated and guard-dominated mixed cases."
"rationale": "criterion.line=15 — the `key` def `const key = n % 2` (source semantics: `key` feeds the guard predicate below). Criterion `route` is mixed-locus with BOTH a guard whose predicate reads the intra def `key` (line 15 -> line 16) and control-gates both return arms (lines 18, 20), AND two cross-function reaches (fast, slow). DOWNSTREAM from line 15: intra_AIS holds the guard (16) + the two control-dependent return statements (18, 20); inter_AIS holds the two callees. Disjoint by construction (lines within `route` vs distinct symbols). Step-0 reconciliation: the statement-anchored PDG slice from line 15 returns {16, 18, 20} — an EXACT match to this source-derived intra_AIS; call-graph reaches {fast, slow} exactly. This mixed case combines control AND data dependence intra-procedurally with branching cross-function reach — the richest mixed shape, complementing the data-dominated and guard-dominated mixed cases."
}

View file

@ -4,6 +4,7 @@
"name": "handleRequest",
"filePath": "src/mixed.ts",
"direction": "downstream",
"line": 13,
"marker": "normalized + 1",
"pdgEdgeKinds": ["REACHING_DEF", "CDG"]
},
@ -43,5 +44,5 @@
"note": "called with the adjusted value; cross-function reach"
}
],
"rationale": "Criterion `handleRequest` is mixed-locus. The `normalized` def (line 13) drives BOTH a data chain (guard test line 14, `adjusted` line 18, the call arg line 19) and a control structure (the guard controls lines 16 and 19). So intra_AIS holds the four control/data-dependent statements within the function. Separately, `handleRequest` calls `persist` — a genuine cross-function reach — so inter_AIS holds `persist`. This is the case where the two modes are complementary: PDG resolves the intra statement set precisely, call-graph reaches the callee. intra_AIS and inter_AIS are disjoint by construction (one is lines within `handleRequest`, the other is a different symbol `persist`)."
"rationale": "criterion.line=13 — the `normalized` def `const normalized = raw * 2` (source semantics: `normalized` seeds both the data chain and the guard). Criterion `handleRequest` is mixed-locus. DOWNSTREAM from line 13, the `normalized` def drives BOTH a data chain (guard test line 14, `adjusted` line 18, the call arg line 19) and a control structure (the guard controls lines 16 and 19). So intra_AIS holds the four control/data-dependent statements within the function. Separately, `handleRequest` calls `persist` — a genuine cross-function reach — so inter_AIS holds `persist`. Step-0 reconciliation: the statement-anchored PDG slice from line 13 returns {14, 16, 18, 19} — an EXACT match to this source-derived intra_AIS; call-graph reaches {persist} exactly. This is the case where the two modes are complementary: PDG resolves the intra statement set precisely, call-graph reaches the callee. intra_AIS and inter_AIS are disjoint by construction (one is lines within `handleRequest`, the other is a different symbol `persist`)."
}

View file

@ -1,11 +1,18 @@
/**
* U7 — PDG-vs-call-graph impact ACCURACY measurement harness.
*
* Runs BOTH `impact` engines (`mode:'callgraph'` and `mode:'pdg'`) over the
* curated U6 ground-truth fixtures, computes precision/recall/F1 stratified by
* impact locus (intra / inter / mixed) plus cross-mode Jaccard + set-diffs,
* prints a stratified report ending in a plain-language DECISION RECOMMENDATION,
* and (under `--check`) gates regressions with two NON-byte-identity gates.
* Runs BOTH `impact` engines over the curated U6 ground-truth fixtures and
* scores each at its NATIVE granularity:
* - `mode:'pdg'` is seeded on the criterion's STATEMENT (`line: criterion.line`)
* and scored at intra-procedural LINE granularity against `intra_AIS`
* (CIS_pdg = the `affectedStatements` line set);
* - `mode:'callgraph'` is scored at inter-procedural SYMBOL granularity against
* `inter_AIS` (CIS = the reported symbol set).
* It computes precision/recall/F1 stratified by impact locus (intra/inter/mixed)
* plus cross-mode set-diffs, prints a stratified report ending in a plain-
* language DECISION RECOMMENDATION, and (under `--check`) gates regressions with
* two NON-byte-identity gates. The two engines answer DIFFERENT questions at
* DIFFERENT granularities — the report shows both, neither strictly dominates.
*
* ── Substrate (the load-bearing mechanism — KTD9/R8; plan U7 "Substrate
* decision") ──────────────────────────────────────────────────────────────
@ -26,14 +33,17 @@
* 3. `new LocalBackend(); await init()` resolves the fixture via the REAL
* registry (the parent process ALSO sets `GITNEXUS_HOME` so init reads the
* temp registry, not the user's ~/.gitnexus);
* 4. `callTool('impact', {repo:<copyPath>, target, direction, mode})` ×2;
* 4. `callTool('impact', …)` ×2 — callgraph (symbol BFS) and pdg (seeded on
* `line: criterion.line` so it returns the statement-anchored slice);
* 5. teardown the temp home + copy.
*
* The `repo` arg is the absolute fixture-copy PATH (tier-1 path match in
* `resolveRepoFromCache`) — unambiguous, no name collisions.
*
* ── Granularity / CIS-AIS framing ──────────────────────────────────────────
* See `metrics.mjs`. Symbol granularity, line-collapsed, order-independent.
* See `metrics.mjs`. PDG = intra-procedural LINE granularity vs `intra_AIS`;
* call-graph = inter-procedural SYMBOL granularity vs `inter_AIS`. The two
* engines measure different scopes; both are now non-empty.
*
* Build-free: `node --import tsx bench/impact-pdg/measure.mjs`. Runtime budget
* and re-baseline instructions: see README.md.
@ -47,11 +57,10 @@ import { fileURLToPath } from 'node:url';
import {
symbolKey,
toKeySet,
pdgLineCis,
intraLineAis,
score,
compareModes,
aggregate,
partitionCisByScope,
aisByScope,
fingerprintAnnotationSet,
median,
@ -136,15 +145,26 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) {
const backend = new LocalBackend();
await backend.init();
const results = {};
for (const mode of MODES) {
results[mode] = await backend.callTool('impact', {
// callgraph: symbol→symbol BFS (no statement anchor). pdg: SEEDED on the
// criterion's statement line so it returns the dependence slice — the U7
// rework's central change. A whole-symbol pdg slice (no `line`) is empty by
// design; `criterion.line` is the 1-based source line of the changed
// statement (set from source semantics, validated in Step 0).
const results = {
callgraph: await backend.callTool('impact', {
repo: work,
target: fx.gt.criterion.name,
direction: fx.gt.criterion.direction,
mode,
});
}
mode: 'callgraph',
}),
pdg: await backend.callTool('impact', {
repo: work,
target: fx.gt.criterion.name,
direction: fx.gt.criterion.direction,
mode: 'pdg',
line: fx.gt.criterion.line,
}),
};
succeeded = true;
return { work, results };
} finally {
@ -154,16 +174,18 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) {
}
}
/** Flatten an impact result's byDepth into canonical symbol keys (the CIS). */
function cisFromResult(res) {
/**
* Flatten a CALLGRAPH impact result's byDepth into canonical SYMBOL keys (the
* CIS_callgraph, scored against `inter_AIS`). Unresolved shadow entries are kept
* (keyed by file) so a recall loss is never hidden.
*/
function callgraphCisFromResult(res) {
const items = Object.values(res?.byDepth ?? {}).flat();
const keys = new Set();
const meta = { unresolved: 0, ambiguous: 0, blockCount: res?.blockCount ?? null };
const meta = { unresolved: 0, ambiguous: 0 };
for (const it of items) {
if (it?.unresolved) {
meta.unresolved += 1;
// surfaced under its file as an unresolved shadow entry — kept in the CIS
// so a recall loss is never hidden, keyed by its file (no symbol name).
keys.add(symbolKey('(unresolved)', it.filePath));
continue;
}
@ -173,6 +195,25 @@ function cisFromResult(res) {
return { keys, meta };
}
/**
* Extract the PDG statement-line CIS (`<filePath>:<line>` keys) from a pdg
* impact result's `affectedStatements`, plus the diagnostic fields the report
* surfaces (the slice's epistemic marker / note / block count). This is the
* U7-rework CIS: the dependent STATEMENTS the change at `criterion.line` reaches.
*/
function pdgCisFromResult(res) {
return {
keys: pdgLineCis(res?.affectedStatements),
meta: {
affectedStatementCount: res?.affectedStatementCount ?? 0,
blockCount: res?.blockCount ?? null,
criterionLine: res?.criterionLine ?? null,
epistemic: res?.epistemic ?? null,
note: res?.note ?? null,
},
};
}
// ── Step 0: fixture AIS validation (gated on the live traversal; KTD9
// circularity guard) ───────────────────────────────────────────────────────
@ -257,20 +298,34 @@ async function validateFixture(fx, work, exec) {
return { critEdges, sameLineCollision, problems, measurable: problems.length === 0 };
}
// ── per-fixture scoring ──────────────────────────────────────────────────────
// ── per-fixture scoring (each mode vs its NATIVE ground truth — U7 rework) ────
/**
* Score one fixture for one mode, per scope. CIS partitioned into intra (the
* criterion symbol itself) / inter (others) / mixed (union); AIS likewise.
* Score the CALLGRAPH mode for one fixture: its reported SYMBOL CIS against the
* fixture's `inter_AIS` (the cross-function symbols truly affected). The
* criterion symbol itself is dropped from the CIS first — callgraph never names
* the criterion as its own dependent, and `inter_AIS` is cross-function by
* construction, so a stray self-reference would be spurious noise. (In practice
* the callgraph CIS already excludes the seed; this is belt-and-suspenders.)
*/
function scoreFixtureMode(gt, cisKeys) {
function scoreCallgraph(gt, symbolCisKeys) {
const ais = aisByScope(gt);
const cisPart = partitionCisByScope(cisKeys, ais.criterionKey);
return {
intra: score(cisPart.intra, ais.intra),
inter: score(cisPart.inter, ais.inter),
mixed: score(cisPart.mixed, ais.mixed),
};
const cis = new Set([...symbolCisKeys].filter((k) => k !== ais.criterionKey));
return score(cis, ais.inter);
}
/**
* Score the PDG mode for one fixture: its statement-LINE CIS (the
* `affectedStatements` from the line-seeded slice) against the fixture's
* `intra_AIS` LINE set. This is the intra-procedural statement-granularity
* measurement the U7 rework introduces. For an inter fixture (`intra_AIS` empty
* by design) the slice may return the router's own control-dependent statements
* — those are FPIS against the empty intra ground truth and recall is n/a, which
* is the honest "PDG is intra-procedural; on a pure-inter fixture it has no
* meaningful intra ground truth" result (symmetric to callgraph's empty intra).
*/
function scorePdg(gt, lineCisKeys) {
return score(lineCisKeys, intraLineAis(gt));
}
// ── reporting helpers ────────────────────────────────────────────────────────
@ -281,14 +336,17 @@ const lpad = (s, n) => String(s).padStart(n);
function renderTable(strata) {
const head =
`${pad('Scope', 7)} ${pad('Mode', 10)} ${lpad('P', 7)} ${lpad('R', 7)} ${lpad('F1', 7)} ` +
`${pad('Scope', 7)} ${pad('Mode', 10)} ${pad('Granularity', 11)} ${lpad('P', 7)} ${lpad('R', 7)} ${lpad('F1', 7)} ` +
`${lpad('|CIS|/|AIS|', 11)} ${lpad('FPIS', 6)} ${lpad('FNIS', 6)} ${lpad('n', 4)}`;
const lines = [head, '-'.repeat(head.length)];
// PDG is scored at LINE granularity vs intra_AIS; callgraph at SYMBOL
// granularity vs inter_AIS — the column makes the "different scopes" explicit.
const gran = (mode) => (mode === 'pdg' ? 'line/intra' : 'symbol/inter');
for (const scope of SCOPES) {
for (const mode of MODES) {
const a = strata[scope][mode];
lines.push(
`${pad(scope, 7)} ${pad(mode, 10)} ${lpad(fmt(a.precision), 7)} ${lpad(fmt(a.recall), 7)} ` +
`${pad(scope, 7)} ${pad(mode, 10)} ${pad(gran(mode), 11)} ${lpad(fmt(a.precision), 7)} ${lpad(fmt(a.recall), 7)} ` +
`${lpad(fmt(a.f1), 7)} ${lpad(fmt(a.cisAisRatio), 11)} ${lpad(a.fpis, 6)} ${lpad(a.fnis, 6)} ` +
`${lpad(a.nCases, 4)}`,
);
@ -300,15 +358,19 @@ function renderTable(strata) {
/**
* Plain-language DECISION RECOMMENDATION (F2 — the deliverable that answers
* "which is more accurate" as a verdict, not just a table). Derived from the
* measured numbers: compares inter-scope recall (the cross-function questions
* users most bring to impact) and any measured intra-scope precision edge.
* measured numbers: PDG's intra-procedural STATEMENT-granularity F1 (the slice it
* is built to compute) and call-graph's inter-procedural SYMBOL-granularity F1
* (the cross-function reach it is built to compute).
*/
function decisionRecommendation(strata, underpowered, exclusions) {
const cgInterR = strata.inter.callgraph.recall;
const pdgInterR = strata.inter.pdg.recall;
const cgIntraP = strata.intra.callgraph.precision;
// PDG is precise at intra LINE granularity; callgraph covers inter SYMBOL reach.
const pdgIntraF1 = strata.intra.pdg.f1;
const pdgIntraP = strata.intra.pdg.precision;
const pdgIntraReports = strata.intra.pdg.nPrecision > 0; // did PDG report ANY intra symbol?
const pdgIntraR = strata.intra.pdg.recall;
const cgInterF1 = strata.inter.callgraph.f1;
const cgInterR = strata.inter.callgraph.recall;
const pdgMixedF1 = strata.mixed.pdg.f1;
const cgMixedF1 = strata.mixed.callgraph.f1;
const lines = [];
lines.push('DECISION RECOMMENDATION');
@ -319,42 +381,43 @@ function decisionRecommendation(strata, underpowered, exclusions) {
);
}
// Inter-scope: the cross-function blast radius.
if (cgInterR !== null && pdgInterR !== null) {
lines.push(
`On INTER-scope (cross-function) impact, call-graph recall is ${fmt(cgInterR)} vs PDG ${fmt(pdgInterR)}: ` +
`PDG's intra-procedural design means it recovers ~0 cross-function impact BY DESIGN (a capability ` +
`fact, not a defect). Call-graph is the correct engine for the "what else calls/uses this?" question.`,
);
}
// Intra-scope: the statement-anchored PDG slice — the question PDG answers.
lines.push(
`On INTRA-scope (statement granularity), the line-seeded PDG slice scores P=${fmt(pdgIntraP)} ` +
`R=${fmt(pdgIntraR)} F1=${fmt(pdgIntraF1)} against intra_AIS: it identifies the dependent ` +
`STATEMENTS of the changed line precisely. Call-graph mode cannot resolve below function ` +
`granularity, so on a self-contained function it names no other symbol (intra recall n/a — ` +
`no cross-function truth to find). PDG is the engine for "which statements does this line affect?".`,
);
// Intra-scope: the case PDG was built to win.
if (!pdgIntraReports) {
// Inter-scope: the cross-function blast radius — the question call-graph answers.
lines.push(
`On INTER-scope (symbol granularity), call-graph scores R=${fmt(cgInterR)} F1=${fmt(cgInterF1)} ` +
`against inter_AIS: it recovers the cross-function callees exactly. PDG mode is ` +
`intra-procedural, so on a pure-inter fixture it returns only the router's own ` +
`control-dependent statements (FPIS against the empty intra_AIS — recall n/a). Call-graph is ` +
`the engine for "what else calls/uses this?".`,
);
// Mixed-scope: both engines contribute, each in its own scope.
if (pdgMixedF1 !== null || cgMixedF1 !== null) {
lines.push(
`On INTRA-scope, PDG mode reported NO owning symbols across the measurable corpus: its block→symbol ` +
`projection collapses a function's own dependence blocks back onto the criterion itself, which the ` +
`traversal excludes as the seed — so at SYMBOL granularity the intra-procedural blast radius is the ` +
`empty set. PDG's intra value in v1 is therefore the BLOCK-LEVEL detail it surfaces ` +
`(reachableBlocks / blockCount), NOT a symbol-level impact set. The harness records the per-fixture ` +
`dependence-block counts so this is visible, not hidden as a flat zero.`,
);
} else if (pdgIntraP !== null && cgIntraP !== null) {
const verb = pdgIntraP > cgIntraP ? 'higher' : pdgIntraP < cgIntraP ? 'lower' : 'equal';
lines.push(
`On INTRA-scope, PDG precision is ${fmt(pdgIntraP)} vs call-graph ${fmt(cgIntraP)} (${verb}). ` +
`This is the measured direction on this corpus, reported as a fact, not asserted as a hypothesis.`,
`On MIXED-scope, the two are COMPLEMENTARY: PDG resolves the intra statement set ` +
`(F1=${fmt(pdgMixedF1)} vs intra_AIS) while call-graph reaches the callee(s) ` +
`(F1=${fmt(cgMixedF1)} vs inter_AIS). Neither alone covers the full mixed blast radius.`,
);
}
lines.push(
`VERDICT: use mode:'callgraph' as the default — it carries the inter-procedural reach that the ` +
`impact tool's safety question depends on. mode:'pdg' adds value as an OPT-IN lens for ` +
`intra-procedural dependence INSPECTION (its reachableBlocks / CDG+REACHING_DEF detail), and ` +
`where the persisted PDG layer exists (analyze --pdg). It is NOT a replacement for, nor a ` +
`strict improvement over, the call-graph blast radius: the two engines occupy different points ` +
`on the precision/recall curve and neither strictly dominates. Promotion of mode:'pdg' beyond ` +
`opt-in is GATED on a Function→BasicBlock CONTAINS_BLOCK substrate edge (deferred) that would let ` +
`the symbol BFS chain natively into the PDG and give intra reach a symbol-level meaning.`,
`VERDICT: the two engines answer DIFFERENT questions at DIFFERENT granularities, and NEITHER ` +
`dominates. mode:'callgraph' (the default) is the correct engine for the inter-procedural ` +
`safety question — "what else depends on / calls this symbol?" — carrying the cross-function ` +
`reach the blast radius needs. mode:'pdg' (opt-in, seeded with line:N, where analyze --pdg ` +
`persisted the layer) is PRECISE at intra-procedural STATEMENT granularity — "which statements ` +
`inside this function does changing line N affect?" — a question call-graph cannot answer at ` +
`all. Use call-graph for cross-symbol impact; reach for line-seeded PDG when you need ` +
`statement-level dependence INSIDE a function. They compose: a full mixed-locus blast radius ` +
`is the UNION of call-graph's inter-symbol reach and PDG's intra-statement slice.`,
);
if (exclusions.length > 0) {
lines.push(
@ -429,43 +492,40 @@ async function run() {
continue;
}
const cg = cisFromResult(results.callgraph);
const pdg = cisFromResult(results.pdg);
const cgScores = scoreFixtureMode(fx.gt, cg.keys);
const pdgScores = scoreFixtureMode(fx.gt, pdg.keys);
// CALLGRAPH: symbol CIS vs inter_AIS. PDG: line CIS vs intra_AIS.
const cg = callgraphCisFromResult(results.callgraph);
const pdg = pdgCisFromResult(results.pdg);
const cgScore = scoreCallgraph(fx.gt, cg.keys); // symbol/inter
const pdgScore = scorePdg(fx.gt, pdg.keys); // line/intra
const locusScope = fx.gt.locus; // the stratum this fixture belongs to
// A fixture is scored in its OWN locus stratum (intra/inter/mixed).
// A fixture is scored in its OWN locus stratum (intra/inter/mixed),
// each mode against its native ground truth (symbol vs line).
if (SCOPES.includes(locusScope)) {
perScopeMode[locusScope].callgraph.push(cgScores[locusScope]);
perScopeMode[locusScope].pdg.push(pdgScores[locusScope]);
perScopeMode[locusScope].callgraph.push(cgScore);
perScopeMode[locusScope].pdg.push(pdgScore);
}
if (runIdx === 0) {
const ais = aisByScope(fx.gt);
const cmp = compareModes(cg.keys, pdg.keys, ais.mixed);
detail.push({
name: fx.name,
locus: fx.gt.locus,
criterion: fx.gt.criterion.name,
direction: fx.gt.criterion.direction,
criterionLine: fx.gt.criterion.line ?? null,
critEdges: v.critEdges,
cg: {
count: results.callgraph.impactedCount,
symbols: [...cg.keys].sort(),
scores: cgScores,
score: cgScore, // vs inter_AIS (symbol)
},
pdg: {
count: results.pdg.impactedCount,
affectedStatementCount: pdg.meta.affectedStatementCount,
blockCount: pdg.meta.blockCount,
unresolved: pdg.meta.unresolved,
ambiguous: pdg.meta.ambiguous,
symbols: [...pdg.keys].sort(),
scores: pdgScores,
criterionLine: pdg.meta.criterionLine,
lines: [...pdg.keys].sort(),
score: pdgScore, // vs intra_AIS (line)
},
jaccard: cmp.jaccard,
pdgOnly: cmp.pdgOnly,
callgraphOnly: cmp.callgraphOnly,
});
}
} finally {
@ -566,25 +626,34 @@ async function run() {
`(${measurableTotal} measurable, ${exclusions.length} excluded) | runs K=${K}`,
);
out.push('');
out.push('Stratified P/R/F1 (symbol granularity, per impact locus):');
out.push(
'Stratified P/R/F1 (PDG: line granularity vs intra_AIS; callgraph: symbol vs inter_AIS):',
);
out.push(renderTable(report));
out.push('');
out.push('Per-case Jaccard + cross-mode set-diffs (true = ∩AIS, noise = −AIS):');
out.push(
'Per-case: PDG slice (line/intra) and callgraph reach (symbol/inter), with FPIS/FNIS:',
);
for (const d of perCaseDetail) {
out.push(
` ${pad(d.name, 28)} locus=${pad(d.locus, 6)} J=${fmt(d.jaccard)} ` +
`cg|count=${d.cg.count} pdg|count=${d.pdg.count} pdg|blocks=${d.pdg.blockCount}`,
` ${pad(d.name, 28)} locus=${pad(d.locus, 6)} line=${lpad(d.criterionLine ?? '-', 3)}`,
);
if (d.callgraphOnly.all.length)
out.push(
` callgraph-only: ${d.callgraphOnly.all.length} ` +
`(true ${d.callgraphOnly.true.length}, noise ${d.callgraphOnly.noise.length})`,
);
if (d.pdgOnly.all.length)
out.push(
` pdg-only: ${d.pdgOnly.all.length} ` +
`(true ${d.pdgOnly.true.length}, noise ${d.pdgOnly.noise.length})`,
);
// PDG line slice: F1 vs intra_AIS, with the false-positive / false-negative lines.
const ps = d.pdg.score;
out.push(
` pdg line/intra : P=${fmt(ps.precision)} R=${fmt(ps.recall)} F1=${fmt(ps.f1)} ` +
`|CIS|=${d.pdg.affectedStatementCount} blocks=${d.pdg.blockCount} ` +
`FPIS=${ps.fpisCount} FNIS=${ps.fnisCount}`,
);
if (ps.fpisCount > 0) out.push(` FPIS(noise): ${ps.fpis.join(', ')}`);
if (ps.fnisCount > 0) out.push(` FNIS(missed): ${ps.fnis.join(', ')}`);
// Callgraph symbol reach: F1 vs inter_AIS.
const cs = d.cg.score;
out.push(
` cg symbol/inter: P=${fmt(cs.precision)} R=${fmt(cs.recall)} F1=${fmt(cs.f1)} ` +
`|CIS|=${d.cg.count} FPIS=${cs.fpisCount} FNIS=${cs.fnisCount}`,
);
if (cs.fnisCount > 0) out.push(` FNIS(missed): ${cs.fnis.join(', ')}`);
}
out.push('');
if (degradedCheck) {

View file

@ -8,20 +8,36 @@
* (Arch-review Issue 5). `measure.mjs` imports these for the live loop.
*
* ── CIS / AIS framing (KTD9 — Arnold–Bohner) ───────────────────────────────
* CIS = Computed Impact Set: the symbols a mode REPORTS as impacted.
* AIS = Actual Impact Set: the curated ground-truth symbols truly affected.
* CIS = Computed Impact Set: what a mode REPORTS as impacted.
* AIS = Actual Impact Set: the curated ground-truth set truly affected.
* precision = |AIS∩CIS| / |CIS| (over-approximation cost; ∅ CIS ⇒ undefined)
* recall = |AIS∩CIS| / |AIS| (under-approximation; ∅ AIS ⇒ undefined)
* F1 = harmonic mean (undefined if either is undefined)
* FPIS = CIS − AIS (false positives — noise)
* FNIS = AIS − CIS (false negatives — the DANGEROUS miss for a safety tool)
*
* ── Granularity (locked) ───────────────────────────────────────────────────
* Symbol granularity, NEVER block-id (block ids carry fragile fnLine:fnCol:idx).
* A symbol key is `<symbol>@<filePath>` — order-independent, line-collapsed. An
* `intra_AIS` entry (statement-granular: lines within the criterion function)
* collapses to its OWNING symbol; an `inter_AIS` entry already names a whole
* symbol. This is why per-scope intra-AIS is the singleton {criterion}.
* ── Granularity: the two engines measure DIFFERENT scopes (U7 rework) ───────
* The two `impact` engines answer different questions at different granularities,
* so they are scored against different ground truths:
*
* - **PDG mode** (`impact({mode:'pdg', line:N})`) is scored at intra-procedural
* STATEMENT granularity. The statement-anchored slice returns
* `affectedStatements: {line,filePath,text}[]` — the dependent statements of
* the criterion line N. CIS_pdg is the set of those LINE keys
* (`<filePath>:<line>`, via `pdgLineCis`); AIS is the `intra_AIS` line set
* (via `intraLineAis`). This is the unit at which PDG is precise.
* - **Call-graph mode** is scored at inter-procedural SYMBOL granularity. CIS is
* the reported symbols (`<symbol>@<filePath>`, via `symbolKey`); AIS is the
* `inter_AIS` symbol set (via `aisByScope`). This is the unit at which the
* call-graph blast radius is meaningful.
*
* Neither is a strict refinement of the other: PDG resolves the dependent
* statements WITHIN a function but cannot cross a call boundary; call-graph
* resolves the cross-function symbol reach but cannot see below function
* granularity. The harness reports both, side by side, per locus stratum.
*
* `partitionCisByScope`/`aisByScope` (symbol-level) remain for the call-graph
* path; `pdgLineCis`/`intraLineAis` (line-level) drive the PDG path.
*/
/** Order-independent symbol key. Collapses statement lines onto their symbol. */
@ -29,6 +45,40 @@ export function symbolKey(symbol, filePath) {
return `${symbol}@${filePath}`;
}
/**
* Statement-LINE key for the PDG path (U7 rework). PDG mode is now scored at
* intra-procedural STATEMENT granularity: the `impact({mode:'pdg', line})` slice
* returns `affectedStatements: {line, filePath, text}[]`, and the intra ground
* truth is the per-line `intra_AIS`. A line key is `<filePath>:<line>` —
* order-independent and statement-granular (NOT collapsed onto the owning
* symbol the way `symbolKey` is). This is the unit at which PDG is precise.
*/
export function lineKey(filePath, line) {
return `${filePath}:${line}`;
}
/** CIS_pdg = the set of affected-statement LINE keys from an impact pdg result. */
export function pdgLineCis(affectedStatements) {
const out = new Set();
for (const s of affectedStatements ?? []) {
if (s && typeof s.line === 'number' && typeof s.filePath === 'string') {
out.add(lineKey(s.filePath, s.line));
}
}
return out;
}
/** AIS_intra = the set of `intra_AIS` LINE keys (statement-granular ground truth). */
export function intraLineAis(gt) {
const out = new Set();
for (const e of gt.intra_AIS ?? []) {
if (e && typeof e.line === 'number' && typeof e.filePath === 'string') {
out.add(lineKey(e.filePath, e.line));
}
}
return out;
}
/** Canonicalize an iterable of {symbol,filePath} (or pre-made keys) → a Set. */
export function toKeySet(entries) {
const out = new Set();
@ -210,10 +260,14 @@ export function canonicalizeAnnotationSet(fixtures) {
.map((fx) => {
const c = fx.gt.criterion;
const kinds = Array.isArray(c.pdgEdgeKinds) ? [...c.pdgEdgeKinds].sort().join(',') : '-';
// `line` is the criterion's 1-based statement anchor (U7 — the seed of the
// statement-anchored PDG slice). It is part of the ground truth: changing
// which statement the slice seeds on changes the measured PDG impact set,
// so an unreviewed `criterion.line` edit MUST trip the fingerprint gate.
return [
`case=${fx.name}`,
`schema=${fx.gt.schemaVersion}`,
`crit=${c.name}|${c.filePath}|${c.direction}|${c.marker ?? '-'}|${kinds}`,
`crit=${c.name}|${c.filePath}|${c.direction}|${c.line ?? '-'}|${c.marker ?? '-'}|${kinds}`,
`locus=${fx.gt.locus}`,
`pdgScoring=${fx.gt.pdgScoring ?? '-'}`,
`provenance=${fx.gt.provenance}`,

View file

@ -42,6 +42,15 @@ interface Criterion {
name: string;
filePath: string;
direction: string;
/**
* The 1-based source line of the statement being changed — the seed of the
* statement-anchored PDG slice (`impact({mode:'pdg', line})`, U7 rework). Set
* from SOURCE SEMANTICS (the def/criterion whose change propagates to the
* intra_AIS lines), validated against the live traversal in the harness's
* Step 0. Required for every measurable case; omitted on excluded no-body
* cases (which carry no statement to seed).
*/
line?: number;
marker?: string;
/**
* The PDG edge kinds the criterion function is EXPECTED to produce. A pure
@ -216,6 +225,15 @@ describe('U6 — impact-PDG fixture ground-truth schema', () => {
expect(src.includes(c.marker!), `marker ${JSON.stringify(c.marker)} in source`).toBe(
true,
);
// Measurable cases carry a 1-based statement anchor (criterion.line) —
// the seed of the PDG slice (U7 rework). It must be a positive integer
// pointing at an actual source line of the criterion file.
expect(typeof c.line, `${fx.name} needs a 1-based criterion.line`).toBe('number');
expect(Number.isInteger(c.line!) && c.line! >= 1, `${fx.name} criterion.line >= 1`).toBe(
true,
);
const lineCount = src.split('\n').length;
expect(c.line! <= lineCount, `${fx.name} criterion.line within file`).toBe(true);
// Measurable cases declare which PDG edge kinds the criterion produces.
expect(Array.isArray(c.pdgEdgeKinds), `${fx.name} needs criterion.pdgEdgeKinds`).toBe(
true,

View file

@ -82,6 +82,61 @@ describe('impact-pdg metric math — score()', () => {
});
});
describe('impact-pdg metric math — PDG line granularity (U7 rework)', () => {
it('pdgLineCis builds <filePath>:<line> keys from affectedStatements', () => {
const cis = M.pdgLineCis([
{ line: 10, filePath: 'src/a.ts', text: 'sum = sum + x;' },
{ line: 12, filePath: 'src/a.ts', text: 'return sum;' },
// a malformed entry (no numeric line) is dropped, never keyed.
{ filePath: 'src/a.ts', text: 'noise' },
]);
expect([...cis].sort()).toEqual(['src/a.ts:10', 'src/a.ts:12']);
});
it('intraLineAis builds line keys from intra_AIS; empty intra_AIS ⇒ empty set', () => {
const withLines = M.intraLineAis({
intra_AIS: [
{ symbol: 'total', filePath: 'src/a.ts', line: 10 },
{ symbol: 'total', filePath: 'src/a.ts', line: 12 },
],
});
expect([...withLines].sort()).toEqual(['src/a.ts:10', 'src/a.ts:12']);
// an inter fixture has empty intra_AIS ⇒ empty line AIS (recall n/a, not 0).
expect(M.intraLineAis({ intra_AIS: [] }).size).toBe(0);
});
it('scores a line slice that exactly matches intra_AIS ⇒ P=R=F1=1 (the accumulator)', () => {
// The verified accumulator case: line-8 slice returns {10,12} = intra_AIS.
const cis = M.pdgLineCis([
{ line: 10, filePath: 'src/accumulator.ts' },
{ line: 12, filePath: 'src/accumulator.ts' },
]);
const ais = M.intraLineAis({
intra_AIS: [
{ filePath: 'src/accumulator.ts', line: 10 },
{ filePath: 'src/accumulator.ts', line: 12 },
],
});
const s = M.score(cis, ais);
expect(s.precision).toBe(1);
expect(s.recall).toBe(1);
expect(s.f1).toBe(1);
});
it('an inter fixture (empty intra_AIS) ⇒ precision 0 on the router noise, recall n/a', () => {
// A pure-inter router's line slice returns its own routing returns as FPIS
// against the empty intra_AIS — the by-design "PDG is intra-procedural" case.
const cis = M.pdgLineCis([
{ line: 24, filePath: 'src/d.ts' },
{ line: 26, filePath: 'src/d.ts' },
]);
const s = M.score(cis, M.intraLineAis({ intra_AIS: [] }));
expect(s.precision).toBe(0); // 2 predicted, none in (empty) intra truth
expect(s.recall).toBe(null); // |AIS|=0 ⇒ recall n/a
expect(s.f1).toBe(null);
});
});
describe('impact-pdg metric math — compareModes()', () => {
it('Jaccard + directional set-diffs split true/noise', () => {
// callgraph finds {a,b,c} (a,b real, c noise); pdg finds {b,d} (b real, d noise).
@ -232,6 +287,42 @@ describe('impact-pdg metric math — annotation fingerprint (KTD10)', () => {
expect(edited).not.toBe(base);
});
it('trips when criterion.line changes (the PDG slice seed — U7 rework)', () => {
// criterion.line is part of the ground truth: it seeds the statement-anchored
// PDG slice, so changing it changes the measured impact set and MUST trip.
const base = M.fingerprintAnnotationSet(
[
fx({
criterion: {
name: 'f',
filePath: 'src/f.ts',
direction: 'downstream',
line: 8,
marker: 'x',
pdgEdgeKinds: ['REACHING_DEF'],
},
}),
],
fakeHash,
);
const moved = M.fingerprintAnnotationSet(
[
fx({
criterion: {
name: 'f',
filePath: 'src/f.ts',
direction: 'downstream',
line: 9,
marker: 'x',
pdgEdgeKinds: ['REACHING_DEF'],
},
}),
],
fakeHash,
);
expect(moved).not.toBe(base);
});
it('trips when the criterion direction flips', () => {
const base = M.fingerprintAnnotationSet([fx({})], fakeHash);
const flipped = M.fingerprintAnnotationSet(