mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-03 02:21:44 +00:00
2 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
054641cafa
|
fix(scope-resolution): resolve a package whose directory name repeats higher in the path (#2881) (#2929)
* fix(kotlin): resolve a root-level package whose name repeats higher in the path `getKotlinFileIndex` built its `dirChildren` buckets under two guards inherited from the pre-index per-import scan rather than from anything Kotlin requires: a `startsWith` test that skipped the bucket when the path began with the package name, and an `indexOf` equality that demanded the parent be the FIRST occurrence of `/<name>/` in the path. `s` is taken as `dir.slice(i + 1)` at each `/`, so `dir` ends with `/s` by construction and the file always IS a direct child of a directory named `s`. The guards therefore dropped legitimate buckets: data/src/main/kotlin/com/example/data/Repo.kt (leading, startsWith) top/data/mid/data/Repo.kt (mid-path, indexOf) `import data.helper` resolved to null against both. Only the fan-out tier was affected — `data.Repo` answers from `suffixByStem`, which carries no such guard — which is why the shape looked narrow enough for #2872 to preserve rather than change inside a performance PR. Both guards are removed. The rule stays "the parent directory is named `s`" — a name that appears in the path without being the parent (`top/data/mid/Repo.kt` for `data.something`) is still not a child, and a new case pins that. Widening is filtered downstream for the fan-out tier, which hands the finalize pass a candidate list (#1759), but NOT for the tier-1 fallback, which commits to `children[0]` unfiltered — and that is where most of the change lands: 149 of the 235 moved corpus records are a different first child against 32 wider arrays. Both are deliberate. A narrower bucket for the first-child tier alone would keep its answers identical and would also leave `import data.*` — a wildcard, which strips to `data` and lands on exactly that tier — resolving to null on the very shape this fixes. Both Kotlin benches are re-baselined deliberately, with the drift measured rather than accepted: - bench/kotlin-import-target: 235 of 19968 distinct records moved. 54 null -> resolved (the fix, and exactly the +54 in non_null), 181 answers that changed within a now-larger bucket. Zero buckets lost a member, zero results were dropped, and every reselected answer's parent directory is the queried package segment. The corpus is untouched, so `cases` is unchanged and the fingerprint covers the same surface as the value it replaces. - bench/import-target: the collide arm needed a corpus edit beside the new numbers. Its `d % 7` slice imported `com.example.vendor{d}`, a package that exists nowhere, purely to mirror the unique arm's nested-slice MISS; with that slice now resolving, leaving it would have left collide at 1100 against small's 1153 and broken the same-workload invariant the arm is built on. That assertion is what caught it. The gate controls were re-run against the new baseline, including one the fix makes newly plausible: a HALF fix that drops only `startsWith` and keeps the `indexOf` check still fails the fingerprint, so a partial fix cannot land quietly. Two gates moved with the code rather than being left behind: - kotlin `heap_reading_bytes` and `heap_ceiling_bytes` are re-recorded together as `_heap_reading_note` requires (48073096 -> 48200224, +0.264%, ceiling still 1.5x). The note says why that is small: the heap corpus is built with HEAP_PAD 8, so no path can begin with a suffix of its own directory and the leading-segment half of the old rule is invisible to that arm. - `depth_budget` 2.4 -> 2.2. Deleting two string comparisons per directory component is per-depth work, so the depth band fell from 1.44-1.51 to 1.27-1.40; left at 2.4 the gate's headroom would have drifted from ~1.6x to ~1.8x without anyone deciding to loosen it. `package-dir-index.ts` documents the same first-occurrence rule as universal, and it is not any more: Go, Java and C# still carry it and still have the shape. Fixing them means re-baselining three languages and editing the verbatim pre-change scans that import-target-index-parity.test.ts keeps as the specification, so it is a separate change — the comment now says so instead of describing a rule one of its readers no longer follows. Fixes #2881. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e * fix(scope-resolution): drop the first-occurrence directory rule for Java, Go and C# too #2881 was reported against Kotlin, but the rule it removed was never Kotlin's. It is what the pre-index per-import scan happened to compute — `indexOf` for the package directory, then "nothing after the match holds a slash" — and every resolver built to reproduce that scan inherited it. Three still had it, and all three reproduced the reported defect: java data/src/main/java/com/example/data/Repo.java `import data.*` -> null java top/data/mid/data/Repo.java `import data.*` -> null csharp Models/src/App/Models/User.cs `using Models;` -> null csharp a/Models/b/Models/User.cs `using Models;` -> null go a/internal/auth/b/internal/auth/svc.go import "internal/auth" -> null Controls (`top/data/Repo.java`, `a/b/internal/auth/svc.go`) resolve, so these are the rule firing rather than an unrelated miss. Four sites, all reduced to "the file's parent directory ends with the queried path": - `package-dir-index.ts` `matchingDirs` (Go, Java, C# without csproj): the `indexOf` equality becomes `endsWith`, which also subsumes the length guard it needed — a shorter haystack is false instead of comparing -1 to -1. - `csharp.ts` `matchingDirPositions` (csproj step 3): same, and still deliberately UNANCHORED, so `src/SubModels` keeps answering `Models`. - `csharp.ts` csproj step 2: `indexOf` -> `lastIndexOf`, EXCEPT for an empty `dirPrefix`, which must keep `indexOf`. Its needle is a bare '/', and step 3 answers that query from `singleSegmentDirs` ("exactly one directory deep"), which only the first occurrence expresses; with `lastIndexOf` there, step 2 accepts every `.cs` in any directory and diverges from step 3. The csproj parity test catches it. - `go.ts` `resolveGoPackage`: `indexOf` -> `lastIndexOf`. No production caller, but the parity harness copies it verbatim as its spec. The two C# csproj sites must move together. Fixing only step 3 makes `Lib.Models` return step 3's superset instead of step 2's segment-aligned answer. Risk is not symmetric across the three. Go's consumer is a fan-out list and the finalize pass materializes one IMPORTS edge per element, so widening only ADDS edges. Java and C#-without-csproj commit to a single file through `firstFileDirectlyInPkgDir` with no downstream filter, so a widened bucket can also change which file an already-resolving import binds to — java's collide fingerprints moved while its resolved count did not, which is exactly that. C#'s leg is additionally gated by `csharpSuffixFallbackAllowed` (#1881) before resolution runs. Gates: - Twenty fingerprints re-baselined across go, csharp and java (five arms plus the top-level alias each). resolved 979 -> 1153 small, 4064 -> 4681 large for go and csharp; 1100 -> 1153 / 4456 -> 4681 for java. No `distinct_outcomes` moved. - csharp and java hit the same collide-arm trap Kotlin did: both sent their `d % 7` slice to a namespace that exists nowhere purely to mirror the unique arm's nested-slice MISS, so once that became a hit the arms resolved fewer imports than `small` and the same-workload assertion failed. Both now use their arm's ordinary spelling. - GO WAS NOT GATED AT ALL and the corpus had to change to make it so. Its nested slice repeated only the last segment (`src/pkg{d}/internal/ pkg{d}`) while a Go query addresses the whole package path, so the directory never ended with the query and the rule was never reached — every go arm sat unchanged through the resolver fix. `uniqueDir` and `collideDir` now repeat the shape at the granularity Go queries. `languages.go.heap.path_segments` 13 -> 14 follows from that. - `csharp_csproj`'s heap reading moved -0.79% (stable across runs) and is re-recorded with its ceiling: the step-2 filter decides which lazy `getFilesInDir` maps the probe forces. Everything else stayed within +/-0.03%, which is this box's jitter — `_heap_reading_note`'s claim that the readings reproduce to the byte across processes did not hold here, and the note now says so. The three parity harnesses keep VERBATIM copies of the pre-change scans as their specification, so each copy was updated with the resolver and the cases that pinned the rule now pin its removal. Two of them left the `mustBeNull` set in the shared harness — they resolve now, which holds them to the stronger "pin a winner" bar the rest of that arm uses. Refs #2881. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e * perf(kotlin): intern dirChildren keys per directory and compact the buckets Two optimizations to `getKotlinFileIndex`, both output-identical, kept because they were measured and a third was dropped because it was not. 1. PER-DIRECTORY KEY MEMO. The component walk over `dir` cut one `slice` per component per FILE, and every slice after the first file of a directory is a freshly allocated string that hashes to a key the map already holds and is then dropped. The key list is a pure function of `dir`, so it is interned once per DIRECTORY. Measured -18.4% to -21.7% of the build at 32 000 files; zero retained cost, the memo dies with the frame. 2. BUCKET COMPACTION at the freeze loop. `addChild` mints `[raw]` and pushes, and V8 grows a backing store by `old + old/2 + 16`, so the SECOND child takes a 1-slot store to 17 and every bucket then retains its overshoot. 61 144 buckets at 32 000 files, 52.9% of their slots empty, 88 B each. `slice()` on freeze: -5 397 768 B, -11.20%, and the predicted 5 382 507 B lands within 0.03% of it. Same fix and the same accounting as the python `byBasename` note this repo already carries. `length === 1` is skipped deliberately. A bucket that never grew is already exact, so slicing it allocates a second array to save nothing — on a corpus of single-file packages the unguarded form costs 31% of the build for zero bytes. DROPPED: merging the `dirChildren` walk into the `suffixByStem` walk. It measures -0.10% at 32k, +0.23% at 100k and +0.40% at one file per directory, all inside a base-vs-identical-copy noise floor of -2.3% to +3.1%, and it does not compose usefully with the memo — the second scan it deletes is exactly the scan the memo makes rare. Only its provably free half is kept: `stem.lastIndexOf('/')` in place of `norm.lastIndexOf`, one backwards scan instead of two, exact because an extension carries no '/'. Neither optimization is visible to the correctness fingerprint, which is the point and also the risk: it observes the index only through the four resolver tiers, so a key-order move no corpus query reaches would survive it. Correctness therefore rests on a structural comparison of all three maps — key insertion order, values, bucket contents in order, frozen-ness — over 1234 corpora in both iteration orders, 14 808 comparisons, zero failures. The fingerprint, `cases` and `non_null` are unchanged and MUST NOT be re-baselined by this commit. Gates that did move, both because a reading and its budget move with the code rather than when CI goes red: - `heap_reading_bytes.kotlin` 48 200 224 -> 42 802 456 with its ceiling at 1.5x. A memory WIN passes every arm, so nothing forced this. - `depth_budget` 2.2 -> 2.0. The memo turns a per-file component walk into a per-directory one, which is precisely the per-depth work this arm exists to see: the band went 1.27-1.40 -> 1.20-1.26, and 2.2 held over it would have drifted from ~1.6x headroom to ~1.9x. The gate controls were re-run against the optimized builder, including one this change makes newly plausible: keying the memo on the directory's LAST SEGMENT instead of its full path drifts the fingerprint (36a4e9dad313, non_null 13310 -> 13305). That is the memo's whole safety argument stated as a test — its key decides which key set a directory contributes — and it is the one way this optimization could move an answer. The bucket-cap control was re-run too, since compaction now rewrites the same buckets. Also recorded, from measuring a reuse this repo had been invited to make: replacing `dirChildren` with the shared `package-dir-index` is output-identical (0 divergences over 107 948 answers) and passes every arm of the kotlin bench at 1.37x-1.50x — while costing 409x per fan-out and 8114x on `import data.*` at 200 matching directories on a corpus this bench does not carry. `_blind_spot` in the kotlin baselines now says so, with the memory the trade would have bought (26.2%, 12.18 MiB) and the corpus arm that would have to exist first. Refs #2881. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e * fix(scope-resolution): close the gaps a four-lens review found in the #2881 change Correctness review found no defect in the shipped resolvers — the `endsWith` rewrites, the C# empty-prefix guard, the memo's purity, the `stem` vs `norm` derivation, the Map re-`set` during iteration and Go's `substring` arithmetic were each attacked with running code and each held. Everything below is a gap in what the change ASSERTS, measures or claims. UNGATED BEHAVIOUR, now covered: - `import-resolvers/go.ts` had no test at all. Its rule changed, the shared bench drives the indexed leg rather than this one, and a revert was caught by nothing. `go-package-resolve.test.ts` pins the membership rule and, more usefully, pins that Go's two independent legs agree on it — they disagreed before #2881 and a divergence here means the LanguageProvider hook and the ScopeResolver hook hold different views of a package. - The memo and the compaction are output-identical, so no fingerprint sees them and reverting either leaves both benches green. `kotlin-index-internals .test.ts` asserts them directly: the memo hit path against the miss path, two directories sharing a component-suffix keeping separate buckets (the one way a coarser memo key could move an answer), and that the bucket handed out is the cached, frozen, compacted array on both the sliced and the skipped path. The comment claiming this was "asserted structurally" previously pointed at nothing in the repo. GATES: - Six ratio budgets in `bench/import-target` were slack: the measurements they bound got faster and the numbers were left alone. kotlin depth 3.4 -> 2.8, go 1.6 -> 1.4, csharp 2.2 -> 2.0, java 2.2 -> 2.1, kotlin collide_scaling 1.8 -> 1.65, go 5.5 -> 5.1, each holding the headroom the old value expressed. The absolute ms ceilings are deliberately untouched: they carry runner-contention headroom, and a ratio is runner-speed-invariant where a millisecond is not. This is the failure the branch already fixed one directory over and missed here. - The `csharp_csproj` heap re-baseline is REVERTED. Base and branch both measure ~73.10e6 three runs each; the recorded 73703384 was simply not reproducible, and re-recording it would have dropped that language's derived floor 0.8% for no reason belonging to this change. - kotlin's collide arm was blind to the rule it was re-baselined for — a full revert of the Kotlin guards left both its fingerprints unmoved, because `com/example/models` is not a suffix of `…/models/inner/models`. Deepened to repeat the whole queried path; those two fingerprints are the only ones that moved for it. The same deepening on the java and kotlin UNIQUE arms was measured and REVERTED: ten more fingerprints, java's heap reading up 43%, and no coverage gained, because progressive stripping lands those queries on the same file either way. SIMPLIFICATION: - `go.ts` now states the predicate as ends-with like its three siblings, instead of keeping the `indexOf` shape with `lastIndexOf` swapped in. - C# csproj step 2's direct-child filter is dead for a non-empty prefix — `getFilesInDir`'s keys ARE segment-aligned directory suffixes, so it cannot reject, and measurement agrees over 12 008 pairs. Only the empty-prefix case does work, and only that case remains. - `addChild` had one call site left; inlined. The memo's double read of its own lookup is gone. The V8 byte accounting duplicated verbatim between the resolver comment and the baselines note now lives only in the note. - Four copies of the same ternary in the csproj parity harness collapse onto one hoisted `dirTrail`; two locals in the java harness were named for the branch that was deleted. CLAIMS THAT WERE WRONG: - `package-dir-index.ts` said "the four resolvers agree again". It is six, and the sixth is the evidence: `import-resolvers/jvm.ts` has answered the same question with `lastIndexOf` since #488, so before #2881 Java's and Kotlin's LanguageProvider hook and their ScopeResolver hook disagreed about which files a package holds. - The `uniqueDir` docblock claimed the last segment IS the query granularity for csharp/java/kotlin. They query the whole dotted path first and reach the tail only through stripping — which is why the partial-revert control fires on the go arm alone, now stated instead of implied. - Three parity harnesses described themselves as verbatim copies of the pre-change implementations; they were edited by this branch, so they are re-derivations of the current spec, a weaker claim their headers now make. - The shared harness header still listed the removed rule as current, the `DIRS` docblock still justified shapes by a divergence that no longer exists, and `measure.mjs`'s tier-two docblock plus `_heap_bound_note` still counted nine bounded languages when `HEAP_BOUNDED` derives to three — this branch had dutifully updated a kotlin bound in a list no gate reads. - `_blind_spot` told the next reader to build a repeated-leaf arm that already exists in the sibling bench, with a budget that already fails the swap. Both baselines are also re-serialized to preserve each note's original escaping, undoing ~20 KB of no-op churn an earlier revision introduced by round-tripping the JSON. Refs #2881. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W8QYxYvd5ntikRUvpLZD5e * perf(scope-resolution): drop the string each membership test built per candidate The three `endsWith` membership tests each minted a decorated copy of the directory once per candidate, per import. The decoration cancels: ('/' + D + '/').endsWith('/' + P + '/') <=> D === P || D.endsWith('/' + P) (D + '/').endsWith(P + '/') <=> D.endsWith(P) Verified exhaustively rather than argued — every pair of strings up to length 5 over `{a, b, /}` including the empty string, 132496 pairs, 0 divergences, with the match count reported beside it because two predicates that agree on `false` everywhere also show 0 divergences. `matchingDirs` 32.58 -> 8.22 ns/candidate (3.96x), `matchingDirPositions` 64.9 -> 18.4 ns. C#'s deliberate unanchoredness survives verbatim: `src/SubModels` still answers `Models`. `resolveGoPackage` was the opposite of a win — the rewrite in this branch left the `'/' + path` cons the old `includes` guard used to short-circuit, and the first `endsWith` forces V8 to flatten it once per file. Working on the raw path with an explicit start index is 4.8x faster than that and 1.78x faster than the code before this branch. It also now reuses `resolveGoPackageDir` instead of re-deriving six of its lines. Three claims these files make are corrected while they are open: - `package-dir-index.ts` argued the rule was accidental because a sixth implementation never had it, "wired as `importResolver` by `languages/{java,kotlin}.ts`" and therefore live. It is wired and not read: `provider.importResolver` is consumed only at `import-target-adapter.ts:74-75`, and that module's exports have no importer outside their own unit test, while its docblock claims it is threaded through `finalizeScopeModel`. The argument survives on the pre-index-scan derivation; `jvm.ts` is evidence about how the predicate was written, not about live behaviour. Whether those resolvers should be deleted or wired is left as an open question. - `csharp.ts` derived the empty `dirPrefix` case from "any path whose first slash is its last", which is wrong in both directions: `src/X.cs` satisfies it and emits no empty key, `a//X.cs` violates it and does. The conclusion stands and the filter stays — it is what rejects `a//X.cs`. - Step 2 returns on its first push, so widening it also suppresses step 3's unanchored leg. The narrower answer is the more precise one, but it was an unstated output change. `SuffixIndex.getFilesInDir` now states the segment-alignment its callers rely on, bounded as a guarantee about what may be RETURNED — php's root-anchored index answers only the equality arm. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus * fix(scope-resolution): say what the widened bucket actually does downstream The comment justifying the widening claimed a bucket that is too wide is "filtered downstream" by the finalize pass. It is not, for the edge that matters. `finalize-algorithm.ts` mints one draft per candidate, each keeping its own `targetFile`, and the File->File emitter in `graph-bridge/imports-to-edges.ts` tests only `targetFile === null` and `targetFile === sourceFile` before adding an `IMPORTS` relationship at confidence 1.0 — it never reads `linkStatus`. The `localDefs` filter from #1759 constrains `targetDefId` and the `BindingRef`; every extra bucket member is an unconditional file-level edge regardless. Measured on an Android-shaped layout, one `import data.load` goes from 5 to 6 edges, all six unresolved. No filtering is added here. Whether an unresolved candidate should produce that edge at all is a design question about the graph bridge, not about this bucket. The published drift census — 149 first-child reselections, 32 wider arrays, 54 null -> resolved — has no bucket for a fourth class this change introduces. Tier 3 precedes tier 4, so a bucket the guards used to leave empty returned null and let the progressive strip run; a populated bucket stops tier 4 entirely, turning a bound answer into a candidate list that need not carry the symbol. Re-running the census with a shape classifier finds that class ZERO times over the corpus, and the zero is the finding: the shape reproduces by hand, and this bench's own generator at 4000 repositories hits it 4-12 times per seed. The fingerprint cannot gate what the corpus cannot express — the same blindness the go arm carried until #2881 widened it. Two further claims are brought back in line with what shipped. The memo's docblock said `kotlin-index-internals.test.ts` asserts the key set, key insertion order and bucket order "over the built maps"; that file says it works through the resolver's observable surface and omits key order deliberately. The mutation matrix bounds it honestly: a mis-keyed memo is caught, a deleted one is not, and the compaction's only instrument is the bench heap ceiling. `findKotlinDirectoryChild` no longer claims to return "the same file the scan used to return" — that is precisely what moved. Structural, no behaviour: `let keys` sits with its consumer instead of 33 lines above it, the archaeology moves to the docblock, `tight` -> `compacted`, `dirEnd` -> `lastSlash` (the name three sibling builders use), and the one-use `MutableDirChildren` alias goes with the `addChild` it existed for. `finalize-algorithm.ts` annotates `targetFiles` as `readonly string[]` so `Array.isArray`'s `any[]` predicate can no longer widen a frozen cached bucket into something `.sort()` compiles against. The runtime freeze stays; it is the backstop for every other call site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus * test(scope-resolution): gate the edges #2881 moved but nothing watched Every widened-shape test in the branch used a one-file corpus, so not one of the 149 first-child reselections was pinned — the tier that commits to `children[0]` unfiltered had no test that could see which file it commits to. Kotlin and Java now pin that choice absolutely, in both insertion orders, for the member path (tier 3, both members) and the wildcard path (tier 1, one file) separately, saying plainly that both candidates are valid members and the only tie-break is file-set iteration order. The tier-3-preempts-tier-4 class gets its first gate, with a control that makes it a transition rather than a fact. The bench corpus holds zero instances, so this case is the only thing standing between that behaviour and a silent revert. C# gains three absolute arms, because its differential harness cannot see any of them — the legacy copy was edited in lockstep with production, which the file's own header admits. One pins the empty-`dirPrefix` filter the branch calls load-bearing and which nothing defended: deleting the guard leaves the whole suite green but changes the answer, so the arm was verified to fail with the guard removed and pass with it restored. Java gains the negative control Kotlin already had. `kotlin-index-internals.test.ts` stops implying coverage it does not have. The mutation matrix is recorded in its header: deleting the memo passes every arm (it is output-identical by construction), deleting the compaction's `slice()` passes every arm (a JS array's capacity has no reflective surface), while mis-keying the memo fails three and compacting-but-never-storing fails two. Four arms were added that do fail under those mutations. V8's growth steps were re-measured — 1, 19, 46, 86 with growth at lengths 2, 20, 47, 87 — so the old 1/17/41 model, which under-counted the slack at 40 files by 6x, is gone. `go-package-resolve.test.ts` drops four `as never` casts that were hiding nothing (`GoModuleConfig` is structurally satisfied), and pins vendor/, testdata/ and nested-go.mod directories, which merge into the importing package — a pre-existing unmodelled gap, verified present before #2881 and documented as such rather than blamed on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus * test(bench): gate the bucket compaction, and publish the whole drift taxonomy The compaction shipped with no gate anywhere. Deleting `bucket.slice()` while keeping the freeze moves no fingerprint, no count and no test — only retained heap, 42805256 -> 48184784 B (+12.57%), byte-identical across three runs. Note the direction: compaction reclaims, so losing it makes the reading GROW, which no floor can see. `heap_ceiling_bytes.kotlin` tightens 64203684 -> 46000000 (1.5x -> 1.0747x of the reading), leaving the regression 4.8% clear above the ceiling and the reading 7.5% below it. The band is derived from first principles in `_heap_compaction_gate` (~61000 buckets x 11 spare slots at Node 22's 1->19 step) so it can be re-checked rather than trusted, and the note carries the triage rule: heapUsed accounting drift moves every arm, so kotlin alone over its ceiling is a lost compaction. `_gate_controls` claimed the two optimizations rest on a structural comparison over 1234 corpora in both iteration orders. No such probe exists in the tree. It now names the test that does exist and lists what it actually pins, and says key insertion order is unasserted by design. `_provenance` gains the full shape classification behind the 235 moved records: 149 string -> string, 38 null -> string, 16 null -> array, 32 array grew, and zero of every other transition — including `string -> array`, the resolved-becomes-unresolved class the old taxonomy had no bucket for. The harness was validated byte-exactly first: driven over this corpus the base resolver reproduces ebf1790bf1 / 13256 and head reproduces d91110bee3 / 13310. `measure.mjs` loses a paragraph asserting the C# unique slice repeats the whole queried path, directly above the paragraph explaining it is leaf-only deliberately and the code that makes it so. Acting on the deleted half resolves the csproj arm to zero. While measuring: the csharp collide arm is NOT blind — its fingerprint already moves across #2881 — but both csharp_csproj arms are, because `getFilesInDir` keys on segment-aligned suffixes and neither nested slice is one. Closing that needs a corpus redesign and four re-baselines; recorded, not attempted. One number changes in either baselines file, and it tightens. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Kim7UK3kPhF1ZRZDuuvnus --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
18bc51dfd2
|
perf(import-resolvers): index every scanning resolver, consolidate the memo, gate every registered language (#2911)
* perf(import-resolvers): build buildSuffixIndex's dirMap lazily (#2903) `buildSuffixIndex` eagerly built three maps. `dirMap` is the array-valued one — one entry per directory suffix per file, so O(files x depth) in entries and array churn — and only four call sites ever read it, all via `getFilesInDir`: `import-resolvers/{php,csharp,jvm}.ts` and `import-resolvers/configs/python.ts`. Ruby (through workspace-file-index), the TypeScript scope resolver, Vue's import-target and the include-extractor never ask a directory question, and built it anyway. Since #2880 these indexes are retained for a whole resolution pass rather than rebuilt per import, so that waste is now resident memory. Deferring it to the first `getFilesInDir` call is behaviour-identical — same key, same descending-suffix order, same per-bucket push order, same `substring(lastIndexOf('.'))` extension clamp. The builder assigns the MAP on completion, so a repeated miss cannot rebuild it. Measured on `buildSuffixIndex` alone, 32k paths, index built and `getFilesInDir` never called: C# layout, 13 segments 79,018,680 -> 66,580,488 B -15.74% Ruby layout, 11 segments 60,752,792 -> 48,656,856 B -19.91% and on the whole retained WorkspaceFileIndex the bench measures: csharp 32k 73.62 -> 61.76 MiB ruby 32k 55.26 -> 43.69 MiB When `getFilesInDir` IS called the footprint is unchanged, so the deferral is never a loss. No new retention: all five construction sites already hold both input arrays alive beside the index. The laziness is pinned structurally rather than by timing. The test's corpus is a `string[]` whose elements are accessor properties, so an indexed read is observable and the read count IS the pass count: 14 after construction, still 14 after any number of get/getInsensitive, 28 after the first `getFilesInDir`, 28 after five more. Memoizing the decision instead of the map would read 42. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * perf(php): resolve imports from a per-run index, not a scan per import (#2901) PHP was the last language whose import resolution scanned the workspace per import. Both `resolvePhpImportTarget` and `resolvePhpImportTargetInternal` materialized two full arrays from the Set on every call, then passed `undefined` as the `index` argument — so `resolvePhpImportInternal` fell through to `suffixResolve`'s linear `findIndex`, once per extension per path part. Measured at 20,000 files: 96.40 ms per import. **Handing it the shared SuffixIndex would have moved IMPORTS edges.** All three index-fed sites answer a different question than the scan they short-circuit, each found by differential with a concrete witness: 1. `getInsensitive` — the scan leg is `allFiles.has(path)`, exact whole-path with no case-insensitive counterpart; the shared index answers a ci SUFFIX probe. 2. `getFilesInDir` — the scan is root-anchored `startsWith(nsDir + '/')`; `dirMap` is keyed on every directory SUFFIX, so a vendor copy can win. 3. `suffixResolve` — the scan's `endsWith('/' + S)` matches only a PROPER suffix; `buildSuffixIndex` indexes j=0, so a root-level `Foo.php` starts resolving `use Foo` where it returned null. 3b. the scan's `endsWith(p) || lower.endsWith(lower(p))` has a second disjunct that subsumes the first, so it is purely first-in-Set-order and case-insensitive; `get(S) || getInsensitive(S)` lets a case-exact hit anywhere beat an earlier ci hit. So this is not Ruby's #2880 shape. Both sites take `getWorkspaceFileIndex` for the memoized arrays and hand the internal resolver a PARITY `SuffixIndex` memoized on the same Set identity: `getInsensitive` disabled, `get` implementing the scan's real rule via the shared ci lookup plus one O(files) whole-path correction map, `getFilesInDir` root-anchored in Set order. no composer.json 96.40 -> 0.036 ms/import steady state with composer.json 100.19 -> 0.068 ms/import steady state Also closes PHP's last per-import traversal, in `import-resolvers/php.ts`: its namespace-directory scan ran whenever `getFilesInDir` came back EMPTY, not merely when no index was supplied — despite the comment above it claiming "only when SuffixIndex unavailable". An empty bucket is already the answer, so the scan could only confirm it, at one full pass per import whose namespace matches a PSR-4 prefix but whose directory has no direct `.php` child (measured 11 traversals for 10 imports; now 1). Moving it into the `else` is safe because the bucket is a SUPERSET of what the scan finds — a root-anchored direct child `nsDir/<x>.php` has its directory exactly equal to `nsDir`, and a directory is always one of its own suffixes, so both index shapes contain it. Nine mutations of the new code are caught, including M1 "pass the raw shared index" (the naive fix) at 23 arms. The adapter guard reads 600 instead of 1 under a defensive `new Set(allFilePaths)` — the #1918 P1 hazard the unit differential is structurally blind to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * perf(java): index import resolution instead of scanning per import (#2908) Java scanned the whole workspace twice per import: once for the three-tier direct match, and again INSIDE the progressive prefix-stripping loop — so a single unresolvable import cost one full pass per stripped segment. No WeakMap, no index, and it is registered in `SCOPE_RESOLVERS`, so it ran in production. This is byte-for-byte the C# shape #2878 fixed, so Java now reads the same machinery: `getWorkspaceFileIndex` for `normToRaw` + the segment-suffix index, and a Java-owned `PackageDirIndex` WeakMap over `buildPackageDirIndex(_, n => n.endsWith('.java'))` read through `firstFileDirectlyInPkgDir`. Structure mirrors C#'s `narrowContext` / `resolveDirectMatch` / `resolveByProgressiveStripping`. 20k files, 256 imports, 7-in-8 unresolvable: 8.05 -> 0.62 ms/import steady state once the index is built: 0.0036 ms/import Tie-breaks preserved, and Java's are NOT identical to C#'s: - tier 1 `break`s on the exact match, so an exact whole-path hit wins even when a suffix or directory-child hit came earlier in iteration order — hence `normToRaw.get` before `index.get`, which conflates them; - the stripping loop instead returns at the FIRST hit of `f === tailFile || f.endsWith('/' + tailFile)` and only yields its directory child after the scan completes, so the conflated `index.get` is the correct lookup THERE. Applying tier 1's exact-wins rule inside the loop is a real behaviour change (mutation M6); - `.*` wildcard stripping stays ahead of everything; - `firstFileDirectlyInPkgDir` reproduces Java's at-root/at-nested predicate exactly, including the first-`indexOf` rule — proved algebraically rather than assumed: the `atRoot` branch matches iff `dir === pathLike`, which is `D.indexOf(P) === 0 === D.length - P.length`, and the `atNested` branch's first occurrence in `f` is the first occurrence in `D` shifted by one. Six mutations are caught; a seventh (swapping the two index builds) is a true equivalence and is recorded as such. Hand-derivation also corrected four cases where the legacy code resolves and I had predicted null — including `java.util.List` reaching a local `util/List.java`, because Java has no in-repo-namespace gate like C#'s #1881. That is preserved here and filed separately as #2910; the parity test pins it so the fix is visible. The adapter guard reads 800 instead of 2 under a defensive `new Set(allFilePaths)`. Two traversals is correct: the workspace index and the package-dir index are separate WeakMaps and each iterates the Set once, the same accounting as C#. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * perf(cobol): index COPY resolution instead of two scans per statement (#2908) `cobolScopeResolver.resolveImportTarget` ran two full workspace scans per `COPY`, each calling `path.extname` + `path.basename` + `.toUpperCase()` on every entry: tier 1 over `.cpy`/`.copybook`, tier 2 over `.cbl`/`.cob`/ `.cobol`. No WeakMap, no index, and registered in `SCOPE_RESOLVERS`. Two uppercased-basename maps, one per tier, filled in a SINGLE pass over the Set and memoized on Set identity. Lookup is `copybooks.get(upper) ?? sources.get(upper) ?? null`. 20k files, 500 COPY operands: 3879-4082 -> 10.5-11.7 us/import (~350-369x) steady state once built: 0.253 us/import Tie-breaks preserved: - TIER ORDER. A `.cpy` match beats a `.cbl` match even when the source file appears EARLIER in Set-iteration order. This is the one a naive single-map rewrite silently breaks, so it gets its own fixture. - Within a tier, first in Set-iteration order wins (`if (!tier.has(...))`, mirroring the scans' first-match return). - The key is built with the identical call sequence, `basename(fp, extname(fp).toLowerCase()).toUpperCase()`, so `Foo.CPY` still keys under `FOO.CPY` rather than `FOO`. - `path` stays in the loop rather than hand-rolled `/`-slicing, so backslash handling is unchanged on every platform — pinned by a `dir\sub\BOOK.cpy` case. All six mutations are caught: collapsing the tiers, within-tier last-wins, dropping the target uppercase, dropping the extension lowercase, hand-rolled slicing, and the adapter's defensive copy. The first five are caught by the differential and are invisible to the adapter guard; the sixth is the reverse, which is the layering working as intended — the guard reads 600 instead of 1. `COBOL_SOURCE_EXTENSIONS` was being re-allocated on every call; hoisted to module scope beside `COPYBOOK_EXTENSIONS`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * perf(csharp): index the csproj leg's namespace-directory scan (#2902) #2878 moved C#'s no-csproj leg onto memoized indexes; the csproj leg kept a per-import full scan in `resolveCSharpImportInternal` step 3, measured at ~1.10 ms per import at 50,000 `.cs` files. **The fix the issue proposed would have moved edges.** It suggested skipping the fallback when an exhaustive index is available, on the assumption that step 2's `getFilesInDir` answers the same question. It does not: step 2's `dirMap` is keyed on segment-aligned directory suffixes, while step 3's `normalized.indexOf(dirPrefix + '/')` is an UNANCHORED substring match, so step 3 finds a strict superset — and it runs only when step 2 came back empty, so those extra hits are observable, not shadowed: dirPrefix 'ubModels' step 2 [] step 3 ['src/SubModels/Widget.cs'] dirPrefix 'rc/Models' step 2 [] step 3 src/Models/* AND vendor/mysrc/Models/* So the predicate is kept byte-for-byte and made fast instead. It depends only on the file's directory (the needle ends with `/`, so every occurrence lies wholly inside `D + '/'`), which reduces to the `package-dir-index` formula minus the anchoring leading slash. `PackageDirIndex` itself cannot be reused for the same reason — its matcher is anchored. The index is memoized on the `normalizedFileList` array identity and built lazily at the point step 3 is first reached, so BCL usings — which `continue` out at the root-namespace gate — never pay for it. Candidates come from an exact last-segment bucket when `dirPrefix` contains a slash, a last-segment key sweep when it does not, and `singleSegmentDirs` when it is empty. Positions rather than paths, merged and sorted when several directories match, so file-list order survives. App.Missing @ {App, src} 1103.0 -> 7.6 us (145x, and flat in file count: 7.3 @10k, 7.6 @50k, 8.4 @200k) App.Missing @ {App, ''} 626.7 -> 108.5 us App @ {App, ''} 1077.9 -> 2.0 us (539x) App.Ns8 @ {App, src} 0.6 -> 0.6 us (step-2 hit, untouched) `relative === ''` is preserved exactly, including the no-`projectDir` case where the needle is a bare `/` and the answer is "every `.cs` whose directory has no slash of its own" — `getFilesInDir('', '.cs')` cannot answer that over repo-relative paths, so it has its own arm. 13 of 14 mutations are caught, including M1, the naive skip-when-indexed cleanup, at 9 arms. The survivor drops the empty-prefix fast path and is a true equivalence. M9 initially survived and exposed a real corpus gap — no non-`.cs` file lived inside a directory — now covered. The remaining non-constant term is the slash-free sweep, O(distinct last segments): 456 us at 200k files on a unique-name layout, but 7.9 us on a `SrcN/Models` layout, which is how C# repos are actually laid out. Closing the unique-name case needs a character-suffix map over segments — the O(files x depth) memory shape `package-dir-index.ts` cites #2649 to avoid — so it is documented in the code as a design change rather than tuned here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * test(scope-resolution): assert index reuse for every registered language (#2909) Index reuse was asserted by nine hand-written per-language files, so the guarantee existed exactly for the languages someone remembered — and #2908 is the proof that is not good enough: Java and COBOL were registered, quadratic and unguarded until this branch. `resolveImportTarget` is a required member of `ScopeResolver` with one signature and 16 registrations, so "calling it N times against a stable `allFilePaths` must not traverse the set N times" is a property of the CONTRACT. `import-target-index-reuse.contract.test.ts` drives every entry of `SCOPE_RESOLVERS`, modelled on `construction-syntax-wiring.test.ts` — the established shape here for a property plus a justified inventory. Measured counts, all memoized: c 1 cobol 1 cpp 1 csharp 2 dart 1 go 1 java 2 javascript 2 kotlin 1 php 1 python 1 ruby 1 rust 0 swift 1 typescript 2 vue 2 **`KNOWN_UNINDEXED` is empty.** The audit that produced it also cleared C, C++, Rust, Swift, TypeScript, Vue and JavaScript by hand — Rust's memo lives in `qualified-call.ts::moduleIndexFor`, C's and Swift's loops are inside their WeakMap builders. The empty map stays as a mechanism: a 17th language cannot opt out silently, and the inventory arm fails when a registered resolver has no fixture. Two things the assertion had to get right: - it is `scans(200) === scans(2)`, not `scans === 1`. Per-language counts legitimately differ (C# and Java build two indexes), and comparing two counts needs no per-language expected value. - Rust legitimately scans ZERO times — it answers every leg with `allFilePaths.has(candidate)` probes — so the floor is a per-language `minimumScans`, 1 for fifteen languages and 0 for Rust with the reason on the interface. Paired with a `hitTarget` that must resolve non-null, so the property cannot pass vacuously on a resolver that stopped answering. Miss targets are distinct per import, which defeats the TS/JS/Vue per-target `resolveCache`. Also unifies the instrument. Kotlin and Python counted index BUILDS from production; the other seven count traversals of a `CountingSet`. The build counter is strictly weaker — a scan added BESIDE a reused index moves no build count, which is exactly the mutation `baselines.json` `_blind_spot` records as invisible to every timing arm — and it costs two production modules that ship in the bundle purely for tests, holding module-global state every test must `reset()`. Both guards migrate to `CountingSet`, and `languages/{kotlin,python}/index-stats.ts` plus both call sites are gone, for -59 lines of shipped source. (Mechanical note: the two `index-stats.ts` file deletions appear in the #2901 commit rather than this one. They were staged with `git rm` while a concurrent commit swept the index. The final tree is correct; only that attribution is off, and rewriting a sibling commit to move them was not worth the risk.) Coverage went up in the swap: Kotlin's old "rebuilds when the file set is a different object" arm (3 sets, 3 builds) would have PASSED under a defensive adapter copy. Its replacement fails, as do all six arms across the two files. Verified by mutation: `new Set(allFilePaths)` inserted into the kotlin, python and go adapters fails exactly those three and no others — `python: 200 imports cost 201 traversals, 2 cost 3`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * test(import-target): gate the four newly-indexed resolvers, retighten heap The bench covered go/csharp/dart/ruby/kotlin. The four resolvers indexed on this branch shipped unmeasured, and #2903's memory win was not locked in. **php, java and cobol join the shared corpus**, each with the two load-bearing properties the header requires: imports scale with file count, and most imports MISS so the full cascade runs (resolve rates php 36.0%, java 34.4%, cobol 36.0%). Java's miss families were measured rather than assumed, since it has no in-repo-namespace gate (#2910): `java.*` 1041 imports and `com.google.*` 1006, both resolving 0. COBOL's collide layout repeats a bookname across BOTH extension tiers, so it reaches the copybook-over-source tie-break rather than only the basename map. **`csharp_csproj` is a sixth LANGS entry**, not a new arm dimension — an entry needs five small additions and inherits all five arms and all seven gates, where a context axis would have to be threaded through `buildRepo`, `resolveAll`, `identityPass`, the report shape and every gate. `buildFiles` aliases it to `csharp`, so the two share one corpus by construction and cannot drift. Two configs (`{App, 'src'}`, `{Lib, ''}`) produce all three `dirPrefix` shapes — slashed, slash-free and empty — in five arms instead of ten: App.Ns{d} 30.6% src/Ns{d} step 2 hit App.Missing{n} 25.5% src/Missing{n} step 3, last-segment bucket Lib 14.0% (empty) step 3, singleSegmentDirs Lib.Missing{n} 12.0% Missing{n} step 3, KEY SWEEP — the one non-constant path BCL / Ghost 12.4% — root-namespace-gate control **2221 of 3200 imports reach the indexed leg**, only 12.4% `continue` out. What that arm pins is stated plainly rather than overclaimed: step 3 answers null for all 2221 here (the hits land at step 2), so it gates that leg's COST and its null answers; its positive tie-breaks stay pinned by the unit parity test. **Heap ceilings retightened.** #2903 dropped the measured figures, leaving the 1.5x ceilings at ~1.9x — a straight revert to the old size would have passed: csharp 116,000,000 -> 98,000,000 B (measured 61.76 MiB) ruby 87,000,000 -> 69,000,000 B (measured 43.69 MiB) php new 106,000,000 B (measured 67.29 MiB) java new 154,000,000 B (measured 97.32 MiB, the largest in the file — Maven layout is 18 segments) php and java are gated because both retained NOTHING across imports at BASE and now retain the O(files x depth) suffix index — the same argument that gates C#. cobol is not: two `Map<basename, path>`, O(files) with no depth term, and its retained delta does not clear measurement noise, so a ceiling would gate nothing. `csharp_csproj` is not: same corpus, same index, a duplicate number — its one distinguishing footprint, the lazily-built `dirMap` its `getFilesInDir` forces back, is measured at +20.8% and recorded as a residual instead, because gating it would licence eager-dirMap everywhere. csharp's `depth_ratio` also fell 3.318 -> 2.31 (the no-csproj leg never asks a directory question, so the deep arm stopped paying an eager dirMap build). Budget 5 -> 3.5, restoring the file's 1.5x convention — and `_arms_note` says plainly that 3.5 does NOT lock that win in, because locking it needs ~2.9, which is 1.25x over a 1.05x spread and the kind of tightening `_triage` warns buys flake rather than signal. All five pre-existing languages are byte-identical: 25 cells x 5 fields = 125 values, 0 mismatches. The new arms were proven live by a doctored baseline (cobol ceiling 0.01, php heap 1000 B, java resolved 999) producing three correctly-worded failures and exit 1. Wall-clock 10.9 -> 26.1 s, php and csharp_csproj ~11 s of it — both cascades end in `suffixResolve`'s ~50-extension probe, and both gate the two largest wins on this branch, so neither is a candidate to drop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * perf(javascript): build the suffix index JS resolution never had JavaScript's `PassCache` was TypeScript's minus one field: `index`. So JS called the shared `resolveTsTarget` with `ctx.index === undefined`, and `import-resolvers/standard.ts` fell through to `suffixResolve`'s linear `findIndex` — scanning the materialized path list once per extension (~39) per path part, per import. 2000 files 6448.9 -> 28.5 us/import (TypeScript: 25.0) 8000 files 25972.6 -> 27.4 us/import (TypeScript: 27.0) Per-import scaling over 4x the files: 4.12x -> 1.09x. **Every instrument on this branch was blind to it.** `CountingSet` counts traversals of the Set; this walked the array the adapter had already materialized — the blind spot `counting-file-set.ts` documents in its own header and `baselines.json` records under `_blind_spot`. Under mutation M1, which drops `index` and reproduces the shipped defect exactly, the sixteen- language contract test stays GREEN for javascript, because the pass cache is still reused and `files.scans` reads 2 either way. Two new arms do catch it: a `suffixResolve` linear-branch counter that runs the legacy adapter first as its control (135 entries legacy, 0 now), and a mock-free behavioural assertion that a repo-root module resolves by bare specifier. Adding an index moves output, exactly as it did for PHP in #2901, so it was characterized rather than assumed — 211,200 pairs (400 corpora x 3 importers x 176 targets) plus 184 hand cases. **Two classes move and there is no third:** A null -> repo-root file (108) `require('config')` with root `config.js`. The scan tests `endsWith('/' + suffix)`, so a path with no slash has no proper suffix and was unreachable through that leg — while `./config` from the root already resolved via the exact `Set.has` branch. JS was internally inconsistent. B file -> different file (5679) `import 'app/main'` was resolving to `node_modules/dep0/lib/main.js`; the scan skipped the whole-path candidate at the 2-segment suffix and fell through to the 1-segment `/main.js`, taking the first such file in Set order. C hit -> null ZERO, and impossible: proper-suffix keys are a subset of the index's keys. Both moved classes are JS being wrong. **JS-new agrees with TypeScript on all 211,200 pairs and every corpus case, 0 disagreements** — which is the intended design, since JS delegates to the TS resolver and differed only by this field. Also swaps the single-slot `let cached: PassCache | null` in JS, TS and Vue for a module-level `WeakMap`, matching every other language. Two alternating file sets rebuilt everything on every call: 12.0 -> 1438.2 ms at 4000 files x 400 imports (120x); after, 11.0 -> 15.7 ms. This is LATENT, not live — `pipeline/run.ts:673` builds one Set per provider pass and the three are separate providers — but it is why these were the only languages that could not carry the standard distinct-set guard. They can now: the arm fails on HEAD for all three (`expected 42 to be 2`) and passes after. Six mutations caught, including a global `resolveCache` (M5), which needed a new arm — `expectDistinctFileSetsGetOwnIndex` builds two IDENTICAL corpora, so a stale answer carried between them is also the right answer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * refactor(ingestion): one per-file-set memo primitive, twenty-one call sites Every language that indexes its import resolution hand-rolled the same memo: declare a module-level `WeakMap` keyed on the file-set object, `get`, `if undefined` build and `set`, return. One concept, written twenty-one times, and this branch had just added five more. `import-resolvers/per-file-set.ts` exports it once: perFileSet<K extends object, T extends object>(build: (key: K) => T): (key: K) => T Two decisions, both recorded in the file. `T extends object` rather than `has`-then-`get`: `WeakMap.get` returning `undefined` cannot distinguish "not built" from "built as undefined", and the `has` form needs a cast or a non-null assertion, both banned here — the constraint makes the ambiguous case unrepresentable instead, and a future caller wanting `string | null` gets a compile error pointing at the decision. A throwing build stores nothing and runs again next call, so failures are not memoized and a half-filled index is never published — inert for these pure builders, and the safer direction. `K extends object` rather than `ReadonlySet<string>` is what lets C#'s `readonly string[]`-keyed cache share the helper. Twenty-one sites migrated across `import-resolvers/` and fifteen languages. Every existing doc comment was re-homed onto the new call rather than deleted — several record real invariants (the Set-identity contract, the #1918 pass-through rule, why Rust's memo lives on a different hook). TypeScript, JavaScript and Vue additionally had byte-identical `PassCache` interfaces and builders. `import-resolvers/pass-cache.ts` now holds the one builder, taking a single argument — every difference the three have lives in the CONSUMER (`tsconfigPaths`, the extension list), not the builder. The builder is shared, the memo deliberately is not: each adapter keeps its own `perFileSet`, hence its own index and its own `resolveCache`, because the three disagree about what a specifier resolves to and one shared cache would hand a language another language's answers. It buys no runtime reuse and the module says so — each provider pass builds its own `allFilePaths` Set, so the three are always different keys. C and C++'s `augmentedFilePaths` was a two-LEVEL memo, and needed no new abstraction: the outer memo's value is a function and a function is an object, so `perFileSet(perFileSet(...))` composes. The two instances stay one per file, and the reason is now in BOTH doc comments rather than only C++'s — cpp delegates to `resolveCImportTarget`, whose `suffixIndex` is keyed on the augmented set, so a shared memo would cross the two languages' indexes. Two sites are deliberately NOT migrated, each with the reason written at the declaration so the next sweep does not re-litigate them: - `configs/swift.ts` is a two-input memo keyed on one. `targets` is not derivable from the key; re-keying on `ctx` would force a banned non-null assertion or an unreachable fallback inside a memo builder. - `rust/qualified-call.ts` `MODULE_SCOPE_CACHE` is three inputs keyed on one, and sits ten lines below a `perFileSet` in the same file — the likeliest thing to be "fixed" by mistake. The other ten remaining `WeakMap`s are different concerns and stay: AST-node caches, worker-pool runtime state, graph metadata, mutable lazily-filled accumulators, and the C++ ADL / inline-namespace indexes, which are reassigned by explicit clear functions and epoch-stamped on read — validity rules beyond key identity that a closure over a private cache cannot express. Net −20 lines of code, +22 of the two "why not" notes. The primitive's own doc is where the cost sits: the Set-identity contract and the two design decisions are written once instead of being twenty-one implicit facts. Pure refactor: 1764 unit tests, 42 guard tests, all sixteen contract-test traversal counts unchanged (c 1, cobol 1, cpp 1, csharp 2, dart 1, go 1, java 2, javascript 2, kotlin 1, php 1, python 1, ruby 1, rust 0, swift 1, typescript 2, vue 2), 647 C/C++ tests, and every bench fingerprint unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Co58j4au9JLf8dmJpwdF9B * test(import-target): gate every registered language, not nine of sixteen The bench pinned output fingerprints and scaling for 9 of the 16 languages in `SCOPE_RESOLVERS`. The other seven — c, cpp, javascript, python, rust, swift, typescript, vue — resolve imports in production with nothing pinning their output or their cost. JavaScript was the sharpest case: the 25,972 us/import defect fixed earlier on this branch was gated by unit tests alone. All 16 are now gated, plus the `csharp_csproj` variant: 17 entries. **The nine existing languages are byte-identical** — 234 committed values (9 x 5 arms x 5 fields, plus 9 top-level fingerprints), 0 changed, and no pre-existing budget touched. Measured both before and after the memo consolidation in |