From 80dbca25d09e3006df9e577b4f0edc22bd8e0512 Mon Sep 17 00:00:00 2001 From: Gergo Magyar Date: Mon, 15 Jun 2026 16:56:13 +0000 Subject: [PATCH] docs(cfg): correct the sweepFacts truncation byte-identity mechanism (#2201 review R6) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The outer sweepFacts JSDoc attributed a truncated result's cross-solver byte-identity to the two solvers producing "identical inSets — insertion order included". That is wrong: the dense (RPO fixpoint) and SSA (renaming/SCC) solvers deliberately build a loop-carried use's reaching set in DIFFERENT insertion orders — same set, different order. The actual mechanism is the KTD6 per-use sort that canonicalizes each use's keys by defKey BEFORE the maxFacts cutoff (already documented correctly on the inner comment). Rewrite the outer doc to say so. Documentation only. Co-Authored-By: Claude Opus 4.8 (1M context) --- gitnexus/src/core/ingestion/cfg/reaching-defs.ts | 15 +++++++++++---- 1 file changed, 11 insertions(+), 4 deletions(-) diff --git a/gitnexus/src/core/ingestion/cfg/reaching-defs.ts b/gitnexus/src/core/ingestion/cfg/reaching-defs.ts index 6bcfe666b..44bde21a9 100644 --- a/gitnexus/src/core/ingestion/cfg/reaching-defs.ts +++ b/gitnexus/src/core/ingestion/cfg/reaching-defs.ts @@ -966,10 +966,17 @@ function computeInSetsAuto( /** * Statement sweep — recover statement-granular def→use facts from the per-block * entry reaching lattices, sort them, and apply the maxFacts truncation. SHARED - * by both solvers: the truncated SUBSET depends on the pre-sort emission order - * here (block index, then statement index, then use order, then the reaching - * set's INSERTION order), so producing identical inSets — insertion order - * included — is what makes a truncated result byte-identical across solvers. + * by both solvers, and the maxFacts cutoff is where their (intentionally + * different) reaching-set INSERTION orders would otherwise leak into the output: + * the dense worklist seeds keys in RPO fixpoint order, the SSA solver in + * renaming/SCC order, so a loop-carried use's reaching set is the same SET in a + * different order. The byte-identity of a TRUNCATED result therefore does NOT + * come from matching insertion orders — it comes from the KTD6 per-use + * `useKeys.sort()` BELOW, which canonicalizes each use's keys by defKey before + * the cutoff. (The full, untruncated fact array is re-sorted at the end, so the + * pre-sort is a no-op there; its whole purpose is the truncated prefix.) Outer + * emission order — block index, then statement index, then use order — is shared + * structurally and needs no canonicalization. */ function sweepFacts( blocks: FunctionCfg['blocks'],