mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-03 02:21:44 +00:00
Merge upstream main into codex/atlascloud-wiki-provider
This commit is contained in:
commit
144d10b8ae
498 changed files with 61244 additions and 6312 deletions
|
|
@ -60,7 +60,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
|||
| Flag | Effect |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| `--force` | Force full regeneration |
|
||||
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
|
|
|
|||
|
|
@ -53,6 +53,14 @@ description: "Use when the user wants to know what will break if they change som
|
|||
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||
| >15 symbols or many processes | HIGH |
|
||||
| Critical path (auth, payments) | CRITICAL |
|
||||
| **Zero callers found** | **UNKNOWN** |
|
||||
|
||||
`UNKNOWN` is not a low rung on this scale — it means the walk could not answer.
|
||||
An empty caller set is equally consistent with "genuinely unused" and "the
|
||||
callers are not resolvable by the index" (plain-object property access, dynamic
|
||||
dispatch, cross-language calls), so few-callers ⇒ LOW does **not** apply. The
|
||||
result carries a `riskNote` saying so. Confirm with a text search before
|
||||
treating the symbol as safe to change or delete.
|
||||
|
||||
## Tools
|
||||
|
||||
|
|
@ -84,6 +92,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth, and reads like
|
||||
`UNKNOWN` above: a zero there means unseen, not unaffected. Re-run it rather
|
||||
than tick the pre-commit check.
|
||||
|
||||
## Example: "What breaks if I change validateUser?"
|
||||
|
||||
```
|
||||
|
|
|
|||
|
|
@ -124,12 +124,17 @@ phase that needs them.
|
|||
statement-level claims (never reconstructs fake edges).
|
||||
- No GitNexus at all → fallback mode: targeted grep/read exploration, findings
|
||||
labelled **source-derived**, with a recommendation to index.
|
||||
- Reading or publishing a plan requires Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
and `O_NOFOLLOW`; publication also requires a validated absolute Python 3
|
||||
PATH candidate with libc `renameat2(RENAME_NOREPLACE)` support, a
|
||||
writable target repository, and a shared filesystem for the plan and
|
||||
Git-admin vault. The writer fails closed when those guarantees are
|
||||
unavailable; it never redirects the plan elsewhere.
|
||||
- Reading or publishing a plan requires `O_DIRECTORY` and `O_NOFOLLOW`, plus
|
||||
`/proc/self/fd` on Linux; every other platform is refused. No interpreter is
|
||||
spawned and no native code is loaded. Publication is `link(2)`, which fails
|
||||
rather than replaces when the destination name is taken. Linux resolves every
|
||||
name against a held descriptor, so a parent swapped mid-write cannot redirect
|
||||
the operation; macOS has no equivalent path and instead pins each directory
|
||||
with an open descriptor and re-proves the chain either side of every step,
|
||||
which detects such a swap and aborts. Publishing also needs a writable target
|
||||
repository and a shared filesystem for the plan and Git-admin vault. The
|
||||
writer fails closed when those guarantees are unavailable; it never redirects
|
||||
the plan elsewhere.
|
||||
|
||||
## Limitations
|
||||
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -87,6 +87,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth: a short or empty
|
||||
list is not proof that only the expected files changed. Re-run it rather than
|
||||
treat the refactor as verified.
|
||||
|
||||
**cypher** — custom reference queries:
|
||||
|
||||
```cypher
|
||||
|
|
|
|||
|
|
@ -216,7 +216,10 @@ Work through plan §7 step by step, in order. For each step:
|
|||
`detect_changes` → commit as one unbroken sequence from the repository
|
||||
root — interleaving other work between the gate and the commit is how
|
||||
the gate gets skipped. Unexpected
|
||||
affected flows → investigate before committing, not after.
|
||||
affected flows → investigate before committing, not after. A result
|
||||
flagged `partial` (a graph query failed) or `truncated` (the symbol
|
||||
listing was capped) blocks the commit the same way: the gate did not
|
||||
see every changed symbol, so re-run it rather than read it as clean.
|
||||
|
||||
A relationship-affecting implementation edit or commit invalidates the
|
||||
procedure's prior proof. The next step must perform the required inter-step
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
2
.github/workflows/ci-e2e.yml
vendored
2
.github/workflows/ci-e2e.yml
vendored
|
|
@ -17,7 +17,7 @@ jobs:
|
|||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: dorny/paths-filter@7b450fff21473bca461d4b92ce414b9d0420d706 # v3
|
||||
- uses: dorny/paths-filter@ceb8a2b8f2d89434be7ff52d3de7ec3738c5cc9d # v3
|
||||
id: filter
|
||||
with:
|
||||
filters: |
|
||||
|
|
|
|||
139
.github/workflows/ci-tests.yml
vendored
139
.github/workflows/ci-tests.yml
vendored
|
|
@ -482,6 +482,14 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Cross-language scope-capture fingerprint + scaling guards
|
||||
# Runs even after an earlier guard fails (#2895). Every step here was
|
||||
# fail-fast, so the FIRST failing --check aborted the job and every guard
|
||||
# after it reported `skipped` — which reads identically to "nothing to do".
|
||||
# Audited across 13 benchmark runs on #2856: the job succeeded zero times
|
||||
# and the last two guards executed zero times for the life of the PR, while
|
||||
# two reviews read the checks summary and saw nothing wrong. `!cancelled()`
|
||||
# rather than `always()` so an explicit cancel still stops the job.
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts emit<Lang>ScopeCaptures output is unchanged
|
||||
# (fingerprint) and stays linear (scaling < 1.5) for go/csharp/rust/php/
|
||||
# ruby/cobol. Catches an O(n^2) re-regression without the worker pool.
|
||||
|
|
@ -489,6 +497,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Callable-value-flow target-index guards (#2693)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts buildGraphTargetIndex resolves an unchanged target
|
||||
# set (fingerprint), stays linear in def count, and that the #2693
|
||||
# widened gate — which now considers VALUE bindings, a population that
|
||||
|
|
@ -500,7 +509,19 @@ jobs:
|
|||
run: node --import tsx bench/callable-value-flow/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Re-export closure scaling guards (#2864)
|
||||
# Build-free: asserts buildReexportClosures stays linear in chain depth
|
||||
# and within an absolute ceiling on a wide package corpus. #2864 changed
|
||||
# this pass's input class from TypeScript barrels (a handful of shallow
|
||||
# edges) to every module-level Python `from m import x`, which is where
|
||||
# its two quadratic corners became reachable. The depth arm specifically
|
||||
# guards MAX_VIA_LENGTH — the bound that was removed once already, in
|
||||
# fc919ad6, and stayed invisible for as long as the input was shallow.
|
||||
run: node --import tsx bench/finalize-reexport/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: C++ qualified-namespace resolution guards (#2788)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts resolveCppQualifiedNamespaceMember resolves an
|
||||
# unchanged symbol set (fingerprint) and that per-call-site cost stays
|
||||
# independent of corpus size. Rationale and history: see the header of
|
||||
|
|
@ -508,7 +529,120 @@ jobs:
|
|||
run: node --import tsx bench/cpp-qualified-ns/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Import-target resolution guards (every registered language, #2877–#2909, PR #2911)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: runs EVERY import-target resolver registered in
|
||||
# SCOPE_RESOLVERS — plus C# a second time WITH csproj configs, over the
|
||||
# identical corpus, because the no-csproj arm returns before it can
|
||||
# reach the leg #2902 indexed. One arm per registered language over ONE
|
||||
# shared corpus, and no registered language ungated. That is enforced,
|
||||
# not enumerated: measure.mjs derives its list from a LANG_REGISTRY
|
||||
# table and its --check inventory arm reconciles that table against
|
||||
# SCOPE_RESOLVERS in both directions, so a language roster typed out
|
||||
# here would only be a second copy that can go stale — this one did.
|
||||
# A C/C++ #include is an import site for this purpose and is gated like
|
||||
# every other registered language (its headers arrive through
|
||||
# resolutionConfig rather than allFilePaths, which is the one structural
|
||||
# difference — see `newPass`).
|
||||
#
|
||||
# Asserts each returns an unchanged target set (a fingerprint per
|
||||
# language AND per arm), that per-import cost stays independent of
|
||||
# corpus size AND of path depth, that the absolute small-arm cost holds
|
||||
# — a constant-factor regression that grows both scale arms equally
|
||||
# passes every ratio — and that the per-pass index eight of them retain
|
||||
# stays within an absolute byte ceiling. The corpus SHAPE is asserted
|
||||
# too: a fingerprint alone cannot tell a legitimate resolution change
|
||||
# from a corpus quietly shrunk below the size the timing arms need.
|
||||
#
|
||||
# Several arms exist because an arm that stops MEASURING otherwise
|
||||
# passes. The heap arms drive real resolvers and carry a FLOOR as well
|
||||
# as a ceiling: when buildSuffixIndex's suffix maps went lazy, four arms
|
||||
# that called the builder directly read 0 B, and 0 B is under every
|
||||
# ceiling. EVERY budget is checked for PRESENCE first, timing and heap
|
||||
# alike, because `got > undefined` is false and `got < ceiling *
|
||||
# undefined` is false too, so deleting a budget key deleted its gate —
|
||||
# and the two heap scalars gate all eight heap arms at once. The heap
|
||||
# arm's own corpus shape (its two file counts, its path depth and the
|
||||
# probe it resolves) is asserted by the same loop as the timing arms,
|
||||
# because those four decide WHAT it measures. And an inventory arm
|
||||
# reconciles the bench's language table against SCOPE_RESOLVERS itself,
|
||||
# so a newly registered resolver cannot ship ungated the way JavaScript
|
||||
# did.
|
||||
#
|
||||
# The resolvers gated first were added as their own O(imports × files)
|
||||
# scans were indexed away (Ruby rebuilt a suffix index per `require`;
|
||||
# COBOL scanned twice per `COPY`), and the same corpus shape scores >3.3
|
||||
# against those pre-fix implementations. The rest were ungated until
|
||||
# this PR, which is not a theoretical gap: PR #2911 found JavaScript
|
||||
# reaching suffixResolve with no index at all — 25 972 µs per import at
|
||||
# 8000 files, protected only by unit tests. This step is what stops the
|
||||
# next one shipping.
|
||||
#
|
||||
# SCOPE: "independent of corpus size" holds for UNIQUE-LEAF layouts,
|
||||
# where no two directories share a last segment and no two files share a
|
||||
# basename — which is what the small/large/deep arms are, and where
|
||||
# every index bucket holds exactly one entry. The `collide` arm runs the
|
||||
# identical workload on the layout these languages are actually written
|
||||
# in (svcN/internal, SrcN/Models, a repeated basename per package, four
|
||||
# SPM modules instead of fifty); there the bucket grows with the file
|
||||
# count by construction and go, csharp, dart, java, swift and c/cpp
|
||||
# legitimately score 2.1–3.9, so that arm carries its own per-language
|
||||
# budget. It is a scope limit, not a regression — the indexed code is
|
||||
# still faster on that shape than the pre-change scan. Rust is the one
|
||||
# language whose collide arm is NOT a shared-leaf layout: it probes
|
||||
# candidate paths and is provably flat in the file count, so its arm is
|
||||
# a deep module tree that varies `::` segment count instead — the axis
|
||||
# its cost actually has.
|
||||
#
|
||||
# --expose-gc enables the retained-heap arm; --check REFUSES to run
|
||||
# without it rather than passing with the memory gate silently skipped.
|
||||
# ~44–45 s, which is essentially unchanged from the ~46 s it cost
|
||||
# before: the timing phase did fall from 39.8 s to 28.7 s when the
|
||||
# min-of-N estimator became per-language, but the inventory arm's one
|
||||
# dynamic import (pipeline/registry.ts pulls in every registered
|
||||
# provider) costs 6–10 s depending on the box and consumes almost all of
|
||||
# that. Report mode, which does not load the registry, is the mode that
|
||||
# got faster: ~33–35 s. Kept as-is because this job runs minutes clear
|
||||
# of the sharded coverage job that gates the merge, so the seconds buy
|
||||
# no merge latency — see COST in the bench header. The ts
|
||||
# family (javascript/typescript/vue) is still the largest block, 8.8 s,
|
||||
# because suffixResolve probes ~39 extensions per path part on a miss.
|
||||
# If this ever has to shrink, drop collide/collide_large for typescript
|
||||
# and vue (−3.9 s) — the only cut that removes near-duplicate work
|
||||
# rather than coverage. N is 15 (matching bench/cfg) for every language
|
||||
# whose cheapest arm is under 5 ms, because depth_ratio divides two
|
||||
# sub-3 ms numbers and at 5 or 7 it tripped its own budget roughly 1 run
|
||||
# in 20; the six languages whose cheapest arm is 20-28 ms drop to 7-8,
|
||||
# where the measured overshoot is at most 6.3%. The estimator was fixed
|
||||
# rather than the budget widened; distributions in _arms_note.
|
||||
# The Kotlin arm here is a second corpus, not a replacement for the
|
||||
# kotlin-import-target bench below, which carries tie-break probes (both
|
||||
# file-set iteration orders, the four-tier cascade) this one does not.
|
||||
# It sits with the other resolver-index guards rather than at the end of
|
||||
# the job: parking a new gate last is not safety, it is the slot least
|
||||
# likely to execute (#2895 measured the last two guards running zero
|
||||
# times in 13 runs). #2899 landed the `if: ${{ !cancelled() }}` below,
|
||||
# which is what makes position irrelevant — a failing step no longer
|
||||
# aborts the ones after it.
|
||||
# Rationale, budgets and the measured blind spot: see the header of
|
||||
# measure.mjs and _blind_spot in baselines.json.
|
||||
run: node --expose-gc --import tsx bench/import-target/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Kotlin import-resolution identity + scaling guards
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts resolveKotlinImportTarget resolves an unchanged
|
||||
# file set (fingerprint, in both file-set iteration orders — every
|
||||
# tie-break in that resolver is expressed only through iteration order)
|
||||
# and that per-import cost stays independent of workspace size. The
|
||||
# pre-index implementation scores 3.737 on this corpus against 0.99 for
|
||||
# the index, so the gate separates them by a wide margin. Rationale and
|
||||
# history: see the header of bench/kotlin-import-target/measure.mjs.
|
||||
run: node --import tsx bench/kotlin-import-target/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Receiver-resolution drop guards
|
||||
if: ${{ !cancelled() }}
|
||||
# NOT build-free: this one runs the real pipeline, so it needs dist/
|
||||
# (the setup action above builds). ~2m15s.
|
||||
#
|
||||
|
|
@ -534,6 +668,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Scope-emission guards (#2699)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts the JS/TS scope set is unchanged. Block scopes are
|
||||
# what make `let`/`const` in sibling blocks distinct bindings, but a
|
||||
# scope per `statement_block` triples the count and deepens every
|
||||
|
|
@ -546,6 +681,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: CFG construction time / disk / memory guards (#2081 M1)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts collectFunctionCfgs output is unchanged
|
||||
# (fingerprint) and that wall-time, cfgSideChannel disk bytes, AND
|
||||
# retained heap all stay sub-quadratic for the straight-line /
|
||||
|
|
@ -556,6 +692,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Emit-persistence throughput / byte-identity guards (#2203)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts streamAllCSVsToDisk output is byte-identical
|
||||
# (order-independent CSV-line fingerprint — the #2203 U2/U3 emit
|
||||
# optimisations must not change graph content) and that emit wall-time
|
||||
|
|
@ -565,6 +702,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Streaming PDG-emit byte-identity / bounded-RSS guards (#2202)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts the streaming PdgEmitSink emits a CSV row SET
|
||||
# byte-identical to the whole-graph streamAllCSVsToDisk emit, AND that
|
||||
# the in-memory graph retains zero BasicBlock nodes (the O(chunk) peak-RSS
|
||||
|
|
@ -574,6 +712,7 @@ jobs:
|
|||
working-directory: gitnexus
|
||||
|
||||
- name: Cross-language pipeline benchmarks (GITNEXUS_BENCH, serial)
|
||||
if: ${{ !cancelled() }}
|
||||
# cpp-adl-benchmark.test.ts is not a `*-pipeline-benchmark.test.ts` but
|
||||
# belongs here for the same reason: it is skipIf-gated on GITNEXUS_BENCH,
|
||||
# so it had never run in CI and the PR #1990 ADL emit-scaling guard it
|
||||
|
|
|
|||
4
.github/workflows/codeql.yml
vendored
4
.github/workflows/codeql.yml
vendored
|
|
@ -48,7 +48,7 @@ jobs:
|
|||
persist-credentials: false
|
||||
|
||||
- name: Initialize CodeQL
|
||||
uses: github/codeql-action/init@e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81 # v4.37.3
|
||||
uses: github/codeql-action/init@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
with:
|
||||
languages: ${{ matrix.language }}
|
||||
queries: security-and-quality
|
||||
|
|
@ -73,6 +73,6 @@ jobs:
|
|||
- '**/test/**/fixtures/**'
|
||||
|
||||
- name: Perform CodeQL Analysis
|
||||
uses: github/codeql-action/analyze@e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81 # v4.37.3
|
||||
uses: github/codeql-action/analyze@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
with:
|
||||
category: '/language:${{ matrix.language }}'
|
||||
|
|
|
|||
2
.github/workflows/scorecard.yml
vendored
2
.github/workflows/scorecard.yml
vendored
|
|
@ -53,6 +53,6 @@ jobs:
|
|||
retention-days: 5
|
||||
|
||||
- name: Upload to Security tab
|
||||
uses: github/codeql-action/upload-sarif@e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81 # v4.37.3
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
with:
|
||||
sarif_file: results.sarif
|
||||
|
|
|
|||
2
.github/workflows/trivy.yml
vendored
2
.github/workflows/trivy.yml
vendored
|
|
@ -76,7 +76,7 @@ jobs:
|
|||
exit-code: '0'
|
||||
|
||||
- name: Upload to Security tab
|
||||
uses: github/codeql-action/upload-sarif@e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81 # v4.37.3
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
with:
|
||||
sarif_file: trivy-${{ matrix.image.name }}.sarif
|
||||
category: trivy-${{ matrix.image.name }}
|
||||
|
|
|
|||
2
.github/workflows/workflow-lint.yml
vendored
2
.github/workflows/workflow-lint.yml
vendored
|
|
@ -76,7 +76,7 @@ jobs:
|
|||
continue-on-error: true
|
||||
|
||||
- name: Upload SARIF
|
||||
uses: github/codeql-action/upload-sarif@e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81 # v4.37.3
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
with:
|
||||
sarif_file: zizmor.sarif
|
||||
category: zizmor
|
||||
|
|
|
|||
|
|
@ -111,15 +111,16 @@ mirror. `gitnexus/test/unit/shipped-skills-sync.test.ts` guards the copies. Toke
|
|||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows). Use GitNexus graph tools to understand code, assess impact, and navigate safely.
|
||||
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows).
|
||||
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? Bootstrap with `npx`, `bunx`, or `pnpm dlx` — e.g. `bunx gitnexus@latest analyze` (npm 11 npx crash; #1939).
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze --index-only` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? Bootstrap with `npx`, `bunx`, or `pnpm dlx` — e.g. `bunx gitnexus@latest analyze` (npm 11 npx crash; #1939).
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). `partial: true` or `truncated: true` is not a clean check — a zero means unseen, not unaffected; re-run it. For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- **MUST treat `risk: UNKNOWN` as unresolved, not as low.** An empty caller set is not evidence the symbol is unused — it can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). `impact` pairs `UNKNOWN` with a `riskNote` saying so. Confirm with a text search before treating the symbol as safe to change or delete; do not proceed on the strength of a zero.
|
||||
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
|
||||
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
|
||||
|
|
@ -128,7 +129,7 @@ This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 rela
|
|||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis, and never read `UNKNOWN` as an all-clear — it means the walk could not answer, which is the one verdict that requires confirming by other means.
|
||||
- NEVER rename symbols with find-and-replace — use `rename` which understands the call graph.
|
||||
- NEVER commit before MCP/CLI graph change analysis.
|
||||
|
||||
|
|
|
|||
|
|
@ -98,7 +98,7 @@ scan → structure → [springConfig, markdown, cobol] → parse → [routes, to
|
|||
| `markdown` | `markdown.ts` | `structure` | Section nodes, cross-link edges from .md/.mdx |
|
||||
| `cobol` | `cobol.ts` | `structure` | COBOL program/paragraph/section nodes (regex, no tree-sitter) |
|
||||
| `parse` | `parse.ts` + `parse-impl.ts` | `structure`, `markdown`, `cobol` | Symbol nodes, IMPORTS/CALLS/EXTENDS edges, extracted routes/tools/ORM queries |
|
||||
| `routes` | `routes.ts` | `parse` | Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators) |
|
||||
| `routes` | `routes.ts` | `parse` | Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators, and JS/TS dispatch guards — see below) |
|
||||
| `tools` | `tools.ts` | `parse` | Tool nodes + HANDLES_TOOL edges |
|
||||
| `orm` | `orm.ts` | `parse` | QUERIES edges (Prisma, Supabase) |
|
||||
| `crossFile` | `cross-file.ts` + `cross-file-impl.ts` | `parse`, `routes`, `tools`, `orm` | Cross-file type propagation in topological import order |
|
||||
|
|
@ -164,6 +164,48 @@ export const myPhase: PipelinePhase<MyPhaseOutput> = {
|
|||
};
|
||||
```
|
||||
|
||||
### Where routes come from
|
||||
|
||||
`route-extractors/` holds four independent ways a route can be discovered, all
|
||||
converging on the routes phase's `(method, url)` registry:
|
||||
|
||||
| Source | Shape | Examples |
|
||||
| --- | --- | --- |
|
||||
| Filesystem convention | path → URL, no parsing | Next.js `app/`, Expo, PHP |
|
||||
| Single-file framework route | `isRouteFile` + worker extraction | Laravel `routes/*.php` |
|
||||
| Cross-file framework route | `discoverRootRouteFiles` + `extractRoutes` | Django `urlpatterns` |
|
||||
| AST-level route in a normal file | `extractDecoratorRoutes` | Spring, FastAPI, NestJS, **JS/TS dispatch guards** |
|
||||
|
||||
The last row is the one whose name undersells it. A route is DECLARED by a
|
||||
decorator, but it can also be **inferred** from a raw `node:http` server's own
|
||||
dispatch — `if (req.method === 'GET' && pathname === '/api/x')` is a route with
|
||||
a path, a verb and a handler, and nothing else in the pipeline could see it.
|
||||
`route-extractors/dispatch-guard.ts` reads that shape; the transport, dedup and
|
||||
handler resolution are shared with decorator routes, and
|
||||
`ExtractedDecoratorRoute.source` carries the provenance difference through to
|
||||
the `HANDLES_ROUTE` edge.
|
||||
|
||||
That extractor is deliberately **precision-weighted**: `route_map` presents its
|
||||
output as fact, so a `startsWith` namespace test, a bare `pathname === '/'`
|
||||
without a verb, and any regex it cannot translate exactly are all dropped rather
|
||||
than guessed at. A missing route is a coverage limit; an invented one is a lie.
|
||||
|
||||
Two rules there need more than one comparison to decide, and are worth knowing
|
||||
about before changing either:
|
||||
|
||||
- **Same-file constant folding.** `` pathname === `${basePath}/rules` `` is
|
||||
common enough that refusing it loses whole route modules — and loses them
|
||||
invisibly, since a module with unfoldable paths and a module with no routes
|
||||
produce the same empty answer. Folding is same-file, string literals only, one
|
||||
alias hop, and refuses on ambiguity (a name declared twice with different
|
||||
values is dropped, never guessed).
|
||||
- **Whole-repo reconciliation** (`reconcileDispatchGuardRoutes`, applied in the
|
||||
routes phase). A split route table — one module listing every path it
|
||||
recognises so the dispatcher can 404 early, handlers in others — otherwise
|
||||
lists every route twice, once verb-less with the table as its "handler". It
|
||||
applies to dispatch-guard routes only: a framework route with no verb is
|
||||
method-agnostic *by declaration*, which is a fact, not a weaker observation.
|
||||
|
||||
---
|
||||
|
||||
## Semantic model
|
||||
|
|
@ -214,6 +256,9 @@ Language-agnostic scope-resolution resolver. This is the resolution path for eve
|
|||
│ emitReferencesViaLookup ── uses handledSites + deferred-site skip set
|
||||
│ emitPropertyDispatchCalls ── registration USES + conservative CALLS
|
||||
│ emitCallableValueFlow ── assigned/passed callable invocation CALLS
|
||||
│ emitImportedValueReferences ── cross-file value reads via finalized imports
|
||||
│ emitUniqueNamePropertyAccesses ── LAST-RESORT property reads by name,
|
||||
│ narrowed same-file → direct-import, refusing to choose otherwise
|
||||
│ emitImportEdges
|
||||
▼
|
||||
KnowledgeGraph (IMPORTS / CALLS / ACCESSES / INHERITS / USES)
|
||||
|
|
@ -232,6 +277,12 @@ The solver is flow-insensitive but bounded: dependency-indexed work items rerun
|
|||
|
||||
Property-key dispatch remains a separate conservative fallback. Its per-key fan-out cap is 32; capped keys synthesize no partial calls and are reported at warning level with language, skipped-key count, dropped key names (bounded), and cap; the count also travels in `RunScopeResolutionStats.propertyDispatchSkippedKeys`.
|
||||
|
||||
Interface-dispatch fan-out walks the subtype closure of the receiver's interface and is **generic-instantiation aware** (#2912): a call through `IValidator<string>` must not reach an implementor of `IValidator<int>`, which shares its declaration and therefore its subtype list. Each heritage clause's arguments reach resolution by one of three routes — read off the `@reference.inherits` anchor's own spelling where that anchor spans the whole base (most languages, no query change), through the `@reference.type-arguments` sub-tag where the anchor is the bare name and moving it would renumber inheritance edge ids (Rust `impl T<A> for S`, Dart `extends`), or on a heritage MARKER payload for clauses that never become reference sites (Dart `implements`/`with`). Whichever pass emits the edge records the pair through one sink: `preEmitInheritanceEdges` for heritage clauses, `ScopeResolver.emitHeritageEdges` for the rest.
|
||||
|
||||
The walk then carries a substitution: a subtype's own type parameters bind to the receiver's arguments, so `class Wrapper<T> : IValidator<T>` stays reachable from every instantiation while `class IntValidator : IValidator<int>` is pruned from the `string` one. Receiver arguments come from the declared type (Case 4), a class-level field's declared type (Case 6), or — for a compound receiver such as `this._repo` — the spelling the compound fold typed that position from, reported back through `recordReceiverType` and accepted only when it names the class the fold returned.
|
||||
|
||||
The filter prunes only on positive evidence: an unknown instantiation on either side, an argument list whose arity does not line up, a name that may be a type variable the language's captures never recorded, or an unresolved spelling whose simple name matches all keep the target. A type parameter of the declaration ENCLOSING either side is recognised as such and never compared — `void Run<T>(IValidator<T> v)` writes a receiver with no known instantiation, so it keeps the unfiltered fan-out. That recognition is what generic METHODS now carry `@declaration.type-parameters` for in C#, Java and Kotlin (TypeScript already did): without it an unbounded `T` grounds to nothing and a bounded one grounds to its BOUND, and both compare unequal to an implementor's concrete argument. Languages that capture neither type arguments nor type parameters therefore emit exactly the pre-#2912 fan-out. The fan-out cap (32, `GITNEXUS_MAX_INTERFACE_DISPATCH_FANOUT`) and its skipped-target reporting are unchanged and apply after filtering. Note the fan-out itself still fires only for a receiver whose folded type is an `Interface` symbol, so a Rust `Trait` or a Dart abstract `Class` receiver emits no secondary targets to filter in the first place.
|
||||
|
||||
Standalone (regex-based) providers such as COBOL participate via `ScopeResolver.scopeResolutionEdgeMode: 'callable-flow-only'`: `runScopeResolution` runs for them, but every ordinary emission path — heritage, interface implementations, receiver-bound, free-call fallback, reference/import edges, post-resolution hooks — is gated off, so their legacy phase (e.g. `cobolPhase`) remains the sole owner of structural edges and the callable solver's `CALLS` are purely additive. A callable-flow-only provider whose files emitted no callable facts exits early, before finalize, keeping the opt-in proportional to source scanning.
|
||||
|
||||
### Receiver chains and the drop census (#2766)
|
||||
|
|
|
|||
|
|
@ -62,15 +62,16 @@ See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.m
|
|||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows). Use GitNexus graph tools to understand code, assess impact, and navigate safely.
|
||||
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows).
|
||||
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? Bootstrap with `npx`, `bunx`, or `pnpm dlx` — e.g. `bunx gitnexus@latest analyze` (npm 11 npx crash; #1939).
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze --index-only` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? Bootstrap with `npx`, `bunx`, or `pnpm dlx` — e.g. `bunx gitnexus@latest analyze` (npm 11 npx crash; #1939).
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). `partial: true` or `truncated: true` is not a clean check — a zero means unseen, not unaffected; re-run it. For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- **MUST treat `risk: UNKNOWN` as unresolved, not as low.** An empty caller set is not evidence the symbol is unused — it can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). `impact` pairs `UNKNOWN` with a `riskNote` saying so. Confirm with a text search before treating the symbol as safe to change or delete; do not proceed on the strength of a zero.
|
||||
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
|
||||
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
|
||||
|
|
@ -79,7 +80,7 @@ This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 rela
|
|||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis, and never read `UNKNOWN` as an all-clear — it means the walk could not answer, which is the one verdict that requires confirming by other means.
|
||||
- NEVER rename symbols with find-and-replace — use `rename` which understands the call graph.
|
||||
- NEVER commit before MCP/CLI graph change analysis.
|
||||
|
||||
|
|
|
|||
|
|
@ -31,7 +31,7 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
### Stale graph after edits
|
||||
|
||||
- **Trigger:** MCP warns index is behind `HEAD`, or search doesn't match latest commit.
|
||||
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used). Runs incrementally by default — the pipeline parses every file every run (cross-file resolution requires it), but tree-sitter dispatch is skipped for unchanged file chunks via the content-addressed cache, and only changed-file rows (plus their importers, transitively) are rewritten in LadybugDB. When the effective write set exceeds ~50% of the repo's files (minimum 50 files), the run transparently switches to the full wipe + bulk-COPY write plan and logs "switching to a full DB write" — expected behavior, not a bug, and file-level bookkeeping stays incremental.
|
||||
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used). Runs incrementally by default — the pipeline parses every file every run (cross-file resolution requires it), but tree-sitter dispatch is skipped for unchanged file chunks via the content-addressed cache, and only changed-file rows (plus their importers, transitively) are rewritten in LadybugDB. When the effective write set exceeds ~50% of the repo's files (minimum 50 files), the run transparently switches to the full wipe + bulk-COPY write plan and logs "switching to a full DB write" — expected behavior, not a bug, and file-level bookkeeping stays incremental. That same line also appears — regardless of write-set size, even for a one-file change — when a LadybugDB extension the existing index depends on cannot load on this machine (VECTOR, #2623; FTS, #2841), because a DB carrying those indexes refuses all row-level DML until the extension is loaded; run `gitnexus doctor` for live extension status and re-run with `GITNEXUS_LBUG_EXTENSION_INSTALL=auto` (with network access) to allow one bounded install attempt. The rebuild is one-shot: it clears the indexes, so the next run goes back to the incremental plan.
|
||||
- **Why:** Tools query LadybugDB from last analyze; git changes are invisible until re-indexed.
|
||||
|
||||
### Index seems corrupt or "incremental" is misbehaving
|
||||
|
|
@ -52,6 +52,12 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
- **Do:** Re-run plain `npx gitnexus analyze` — no `--embeddings` flag needed. A retained `embeddingCheckpoint` in the index metadata forces embedding generation for exactly the pending nodes regardless of flags, and clears once they succeed. `--drop-embeddings` abandons the pending nodes instead of retrying them; `--force` also discards the checkpoint (with a warning) and rebuilds without resuming it.
|
||||
- **Why:** A long analyze run against a flaky HTTP embedding endpoint tolerates bounded sub-batch failures instead of aborting the whole run: it deletes the affected nodes' embedding rows (so they hold zero rows, never a partial set) and records those nodes as pending in `embeddingCheckpoint`. `stats.embeddings` stays an honest, non-zero count of everything that did succeed, so this state never trips the "Embeddings vanished" Sign above — `embedding-checkpoint-pending` is the only reliable signal.
|
||||
|
||||
### Analyze reports INCOMPLETE with a collapsed graph write
|
||||
|
||||
- **Trigger:** `npx gitnexus status` reports `incompleteReasons: ["graph-write-collapsed"]`; the analyze summary printed `Repository indexed INCOMPLETELY` naming an expected and a persisted relationship count, and the CLI exited non-zero.
|
||||
- **Do:** Re-run `npx gitnexus analyze --force`. If it recurs, check free disk space on the volume holding `.gitnexus/`, confirm no second `analyze` is running against the same repo (both stage through `.gitnexus/csv`), then run `npx gitnexus doctor`.
|
||||
- **Why:** The run finished and wrote metadata, but far fewer relationships are readable back than the pipeline produced. Nothing throws: the DB holds rows and the metadata is valid, so every query answers with missing edges rather than an error — a confident empty answer, which is worse than a failure because it looks like a result. Unlike `incremental-in-progress` and `embedding-checkpoint-pending`, which describe a run that did what it said and left work for next time, this one means most of your edges are gone, so it is the one incomplete reason that also fails the exit code. The check compares in-memory totals (including rows streamed out of the heap) against the post-write count, refuses to answer when the count cannot be read, and is skipped on incremental runs where whole-scope counts are not comparable.
|
||||
|
||||
### MCP lists no repos
|
||||
|
||||
- **Trigger:** MCP stderr says no indexed repos.
|
||||
|
|
|
|||
13
MIGRATION.md
13
MIGRATION.md
|
|
@ -106,6 +106,19 @@ Running `npx gitnexus analyze` writes both `gitnexus.json` and `meta.json`
|
|||
with identical content. A pre-existing repo that only has `meta.json` gets
|
||||
`gitnexus.json` bootstrapped from it on the first run.
|
||||
|
||||
### Process ids are not stable across this release
|
||||
|
||||
`Process` ids are positional (`proc_<idx>_<entry>`), and this release changes
|
||||
both which execution flows are detected and the order they are selected in:
|
||||
tracing is depth-first, sibling branches follow source order, and selection
|
||||
round-robins across terminals so one flow cannot take every slot. A given
|
||||
`proc_7_handle` before the upgrade is not the same flow afterwards.
|
||||
|
||||
Nothing in GitNexus persists or joins on a raw process id across a re-index —
|
||||
the MCP resource keys by label — so this is one-time index churn rather than a
|
||||
broken consumer. If you have external tooling that stored a process id, re-
|
||||
resolve it by label after the next analyze.
|
||||
|
||||
### What about rollback?
|
||||
|
||||
Downgrading to an older GitNexus version is safe: `meta.json` is always
|
||||
|
|
|
|||
10
RUNBOOK.md
10
RUNBOOK.md
|
|
@ -66,6 +66,16 @@ npx gitnexus analyze
|
|||
|
||||
No `--embeddings` flag needed — a retained checkpoint forces embedding generation for the pending nodes regardless of flags, and clears once they succeed. `--drop-embeddings` abandons the pending nodes instead of retrying them; `--force` also discards the checkpoint (with a warning) and rebuilds without resuming it.
|
||||
|
||||
**Collapsed graph write (analyze exits NON-ZERO and says INCOMPLETE):** A run can finish writing metadata while only a fraction of the relationships it produced are readable back from the index — edges collapsing to a small share of what was built, or a `CodeRelation` table that never materialized (which reads as a persisted count of zero). Because the metadata IS written and the DB does hold rows, nothing looks broken: queries answer with missing edges rather than an error, which is a confident empty answer rather than a failure. `npx gitnexus status` reports `incompleteReasons: ["graph-write-collapsed"]`, the analyze summary prints `Repository indexed INCOMPLETELY` with the expected and persisted counts, and the CLI exits non-zero so automation is not told an unusable index is fine.
|
||||
|
||||
Recovery is a full rebuild:
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze --force
|
||||
```
|
||||
|
||||
If it recurs, the cause is almost always environmental rather than a code defect: check free disk space on the volume holding `.gitnexus/`, make sure no second `analyze` is running against the same repo (both use `.gitnexus/csv` for staging), then run `npx gitnexus doctor`. The check compares in-memory relationship totals (including streamed rows) against what the DB hands back, and is deliberately skipped on incremental runs, where the two counts are not comparable.
|
||||
|
||||
**Large repos:** Analyze may skip or limit embedding work when node counts are very high; watch CLI output.
|
||||
|
||||
---
|
||||
|
|
|
|||
|
|
@ -543,7 +543,7 @@ function handlePostToolUse(input) {
|
|||
// If HEAD matches last indexed commit, no reindex needed
|
||||
if (currentHead && currentHead === lastCommit) return;
|
||||
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings, indexOnly: true });
|
||||
sendHookResponse(
|
||||
'PostToolUse',
|
||||
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
|
|
|
|||
|
|
@ -276,7 +276,13 @@ function formatBunxCommand(gitnexusArgs) {
|
|||
}
|
||||
|
||||
function formatAnalyzeCommand(options = {}, deps = {}) {
|
||||
const suffix = options.embeddings ? ' --embeddings' : '';
|
||||
// `--index-only` is what a routine "your index is stale" nudge wants: it
|
||||
// reindexes without rewriting AGENTS.md / CLAUDE.md / skills, so an agent
|
||||
// following the nudge on every commit cannot churn the tracked agent guides
|
||||
// (#2907). Callers that actually want the docs refreshed omit it.
|
||||
const suffix = `${options.indexOnly ? ' --index-only' : ''}${
|
||||
options.embeddings ? ' --embeddings' : ''
|
||||
}`;
|
||||
// Keep the stale-index hook budget tight by querying each tool at most once.
|
||||
// The memoized `probe` is a spawn-free PATH scan (resolveOnPath) shared with
|
||||
// resolveInvocationMode, so `gitnexus` is scanned only once and no subprocess
|
||||
|
|
|
|||
|
|
@ -60,7 +60,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
|||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
|
||||
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
|
|
|
|||
|
|
@ -53,6 +53,14 @@ description: "Use when the user wants to know what will break if they change som
|
|||
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||
| >15 symbols or many processes | HIGH |
|
||||
| Critical path (auth, payments) | CRITICAL |
|
||||
| **Zero callers found** | **UNKNOWN** |
|
||||
|
||||
`UNKNOWN` is not a low rung on this scale — it means the walk could not answer.
|
||||
An empty caller set is equally consistent with "genuinely unused" and "the
|
||||
callers are not resolvable by the index" (plain-object property access, dynamic
|
||||
dispatch, cross-language calls), so few-callers ⇒ LOW does **not** apply. The
|
||||
result carries a `riskNote` saying so. Confirm with a text search before
|
||||
treating the symbol as safe to change or delete.
|
||||
|
||||
## Tools
|
||||
|
||||
|
|
@ -84,6 +92,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth, and reads like
|
||||
`UNKNOWN` above: a zero there means unseen, not unaffected. Re-run it rather
|
||||
than tick the pre-commit check.
|
||||
|
||||
## Example: "What breaks if I change validateUser?"
|
||||
|
||||
```
|
||||
|
|
|
|||
|
|
@ -124,12 +124,17 @@ phase that needs them.
|
|||
statement-level claims (never reconstructs fake edges).
|
||||
- No GitNexus at all → fallback mode: targeted grep/read exploration, findings
|
||||
labelled **source-derived**, with a recommendation to index.
|
||||
- Reading or publishing a plan requires Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
and `O_NOFOLLOW`; publication also requires a validated absolute Python 3
|
||||
PATH candidate with libc `renameat2(RENAME_NOREPLACE)` support, a
|
||||
writable target repository, and a shared filesystem for the plan and
|
||||
Git-admin vault. The writer fails closed when those guarantees are
|
||||
unavailable; it never redirects the plan elsewhere.
|
||||
- Reading or publishing a plan requires `O_DIRECTORY` and `O_NOFOLLOW`, plus
|
||||
`/proc/self/fd` on Linux; every other platform is refused. No interpreter is
|
||||
spawned and no native code is loaded. Publication is `link(2)`, which fails
|
||||
rather than replaces when the destination name is taken. Linux resolves every
|
||||
name against a held descriptor, so a parent swapped mid-write cannot redirect
|
||||
the operation; macOS has no equivalent path and instead pins each directory
|
||||
with an open descriptor and re-proves the chain either side of every step,
|
||||
which detects such a swap and aborts. Publishing also needs a writable target
|
||||
repository and a shared filesystem for the plan and Git-admin vault. The
|
||||
writer fails closed when those guarantees are unavailable; it never redirects
|
||||
the plan elsewhere.
|
||||
|
||||
## Limitations
|
||||
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -87,6 +87,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth: a short or empty
|
||||
list is not proof that only the expected files changed. Re-run it rather than
|
||||
treat the refactor as verified.
|
||||
|
||||
**cypher** — custom reference queries:
|
||||
|
||||
```cypher
|
||||
|
|
|
|||
|
|
@ -216,7 +216,10 @@ Work through plan §7 step by step, in order. For each step:
|
|||
`detect_changes` → commit as one unbroken sequence from the repository
|
||||
root — interleaving other work between the gate and the commit is how
|
||||
the gate gets skipped. Unexpected
|
||||
affected flows → investigate before committing, not after.
|
||||
affected flows → investigate before committing, not after. A result
|
||||
flagged `partial` (a graph query failed) or `truncated` (the symbol
|
||||
listing was capped) blocks the commit the same way: the gate did not
|
||||
see every changed symbol, so re-run it rather than read it as clean.
|
||||
|
||||
A relationship-affecting implementation edit or commit invalidates the
|
||||
procedure's prior proof. The next step must perform the required inter-step
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -36,6 +36,8 @@ description: Analyze blast radius before making code changes
|
|||
- [ ] Assess risk level and report to user
|
||||
```
|
||||
|
||||
> `partial: true` (a graph query failed) or `truncated: true` (the changed-symbol listing was capped) means the result is short of the truth: a zero there means unseen, not unaffected. Re-run it rather than tick the pre-commit check.
|
||||
|
||||
## Understanding Output
|
||||
|
||||
| Depth | Risk Level | Meaning |
|
||||
|
|
@ -52,6 +54,14 @@ description: Analyze blast radius before making code changes
|
|||
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||
| >15 symbols or many processes | HIGH |
|
||||
| Critical path (auth, payments) | CRITICAL |
|
||||
| **Zero callers found** | **UNKNOWN** |
|
||||
|
||||
`UNKNOWN` is not a low rung on this scale — it means the walk could not answer.
|
||||
An empty caller set is equally consistent with "genuinely unused" and "the
|
||||
callers are not resolvable by the index" (plain-object property access, dynamic
|
||||
dispatch, cross-language calls), so few-callers ⇒ LOW does **not** apply. The
|
||||
result carries a `riskNote` saying so. Confirm with a text search before
|
||||
treating the symbol as safe to change or delete.
|
||||
|
||||
## Tools
|
||||
|
||||
|
|
|
|||
|
|
@ -23,6 +23,8 @@ description: Plan safe refactors using blast radius and dependency mapping
|
|||
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
> Every `detect_changes()` below: `partial: true` (a graph query failed) or `truncated: true` (the changed-symbol listing was capped) means the result is short of the truth — a short or empty list is not proof that only the expected files changed. Re-run it rather than treat the refactor as verified.
|
||||
|
||||
## Checklists
|
||||
|
||||
### Rename Symbol
|
||||
|
|
|
|||
375
gitnexus-shared/package-lock.json
generated
375
gitnexus-shared/package-lock.json
generated
|
|
@ -8,21 +8,382 @@
|
|||
"name": "gitnexus-shared",
|
||||
"version": "1.0.0",
|
||||
"devDependencies": {
|
||||
"typescript": "^6.0.3"
|
||||
"typescript": "^7.0.2"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-aix-ppc64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-aix-ppc64/-/typescript-aix-ppc64-7.0.2.tgz",
|
||||
"integrity": "sha512-MTKKkWB7p/0E9xi1d1tHtZ5PiLkGEMIq88pK2CubZjOsLtYTLqhgIgi6zepFa+9GHZ6h05NMCkQxGKiPXMxXtQ==",
|
||||
"cpu": [
|
||||
"ppc64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"aix"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-darwin-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-darwin-arm64/-/typescript-darwin-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-gowzar9MwS/aRWp6f3a4KUqzRjAZjOsmGNCM6LcTgXum+dBfgsBVMN+AgvOCCbguXyick6LJhpBszxMebJ8syA==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"darwin"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-darwin-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-darwin-x64/-/typescript-darwin-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-SZ9xZInqApNlNGc9s0W1VSsktYSOe9cFqNOIqmN1Gs8SmkjKZYFt017G4VwPxASInODuAdbTW7sXiFUf893RgA==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"darwin"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-freebsd-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-freebsd-arm64/-/typescript-freebsd-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-W5NH4y/J0plIIS5b2xvTEkU7JFxyqdMAOgf+Ilhl0vHQXKO5dZoxd+C/jEtq56c4F3wk71RB4BMRQ2XdI+bwYQ==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"freebsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-freebsd-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-freebsd-x64/-/typescript-freebsd-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-UMGDx5sTpzNw3WiPebH7l90IWfJggEd+egHt/q6p7/Cm3zqoV7VxkGXt+3DxPIw8CcmvAB0j3sVVfbhX+M4Tpw==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"freebsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-arm": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-arm/-/typescript-linux-arm-7.0.2.tgz",
|
||||
"integrity": "sha512-gffT3xPz9sR7j/YJExkyPntrI0P2EP9XbOyWzth2/Gs0RstK+90RBcO0ncXoXy/beYll1SXw846Nf2zdnEz0QQ==",
|
||||
"cpu": [
|
||||
"arm"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-arm64/-/typescript-linux-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-Qh4eU4/y3yDjnfjjyPYihMj5/ODIlmt+Bzu17OI+fiSRDW57QmU5SiN63exPRNJPKUzcc1INa1NXdrJ+MqHjUQ==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-loong64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-loong64/-/typescript-linux-loong64-7.0.2.tgz",
|
||||
"integrity": "sha512-uEHck9i8hoAzXPiYRib1O7miOnz23SxIeVl6F4LXox+qov1K35jHcEW6VHKvZI+pyvl7fZEP4MCU5LYvIq1GuQ==",
|
||||
"cpu": [
|
||||
"loong64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-mips64el": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-mips64el/-/typescript-linux-mips64el-7.0.2.tgz",
|
||||
"integrity": "sha512-R4KvAMnE43W5Qeqb0Ly56O3mWMWIAgsMyz36DCaycd5nbg/9kzm0liw3JocfRqyJY0KPmzFjbswozXyW0DnIYA==",
|
||||
"cpu": [
|
||||
"mips64el"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-ppc64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-ppc64/-/typescript-linux-ppc64-7.0.2.tgz",
|
||||
"integrity": "sha512-DORx5b3sd/4S7eayxm4FQv+A7CrkUIGRaHiwI8oiHTAI1fAPWhF4J0vAlkC8biAlHSVVwxMQ3tjZ2/DVbnQiiA==",
|
||||
"cpu": [
|
||||
"ppc64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-riscv64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-riscv64/-/typescript-linux-riscv64-7.0.2.tgz",
|
||||
"integrity": "sha512-wf0jqEDOjrPRnKwYRyyJDRo11KMbvMFrU+q4zqKyChODBzvlkbhNQfKvLxQCcwTpdDaXSHZTVuh0JoCrKCUMHQ==",
|
||||
"cpu": [
|
||||
"riscv64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-s390x": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-s390x/-/typescript-linux-s390x-7.0.2.tgz",
|
||||
"integrity": "sha512-IkwJc3L7yhytWd/ewjyxNDfOmswCm9GWMJT/ue/dU4aZNbwZeYAetq42VyLmsmSjvoX7z74X6ZaYCtzAr0EuGw==",
|
||||
"cpu": [
|
||||
"s390x"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-linux-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-linux-x64/-/typescript-linux-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-EYdf2cNg7rgCWJnxCdJ+F3V39O8ihb37eHAu1LK8oAFizgTQbPOK7zHHXbPt8rX24COqODXeI3sIf0fCXG7H/A==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"linux"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-netbsd-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-netbsd-arm64/-/typescript-netbsd-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-+polYF4MF04aPpO5FTkHran9yUQDSXqy5GiSDKpsll5jy3l3+g9QLhpf39T+ePtefhXLOGrLl0QIjkQP6VnelA==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"netbsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-netbsd-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-netbsd-x64/-/typescript-netbsd-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-8YIT0EHM/3dq10ZOVF/A7pc/YSMtbcecct4rWtexrnSCHOPcpC2KTLXfTCR6vDpnSiY12heNb1GiN/wu+T/FyA==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"netbsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-openbsd-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-openbsd-arm64/-/typescript-openbsd-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-APT8+ClYnuYm1u9+kgGXoMj2VzWzcymwh2gNSQVySHfkRDGOTVkoWLjCmOQSaO+PoqQ57B0flRp9SA+7GnnkzQ==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"openbsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-openbsd-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-openbsd-x64/-/typescript-openbsd-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-yX7s+Q0Dln0Dt9tEzZsAjXXR/+ytBM7AlglaqyeMPxQszJ1JhlJdZ6jLA+IzldHtflX81em7lDao1xXu+aRRkg==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"openbsd"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-sunos-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-sunos-x64/-/typescript-sunos-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-dLJDGaLZ1D4HPQn62u1n8mBDkJREwMsAkCdkwd4Ieqw+x3TUyTsqY0YiBCtE6H6OzzgGk3iuZ3vFWRS+E8/d1g==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"sunos"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-win32-arm64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-win32-arm64/-/typescript-win32-arm64-7.0.2.tgz",
|
||||
"integrity": "sha512-Gyl1Vy6OsWesLzmq+EP0Fb7b4Nid5232AvcA2SFcdYreldpNtYFFofPjnt62y9hQy7VTaZp65ICJjuAQRaVcIQ==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"win32"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@typescript/typescript-win32-x64": {
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/@typescript/typescript-win32-x64/-/typescript-win32-x64-7.0.2.tgz",
|
||||
"integrity": "sha512-0BQ3HkAHHlKLSp1qRvf3SUhGpGsDuhB/jgFw75guyqbxJqEaS0Cw/VFO8i2nHglJUzQCRtMMR/IBAKE3ETMC4g==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"win32"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=16.20.0"
|
||||
}
|
||||
},
|
||||
"node_modules/typescript": {
|
||||
"version": "6.0.3",
|
||||
"resolved": "https://registry.npmjs.org/typescript/-/typescript-6.0.3.tgz",
|
||||
"integrity": "sha512-y2TvuxSZPDyQakkFRPZHKFm+KKVqIisdg9/CZwm9ftvKXLP8NRWj38/ODjNbr43SsoXqNuAisEf1GdCxqWcdBw==",
|
||||
"version": "7.0.2",
|
||||
"resolved": "https://registry.npmjs.org/typescript/-/typescript-7.0.2.tgz",
|
||||
"integrity": "sha512-8FYau96o3NKOhbjKi/qNvG/W5jhzxkbdm5sj9AbZ/5T5sWqn3hJgLfGx27sRKZWTvyzCP8dLRBTf5tBTSRVUNA==",
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"bin": {
|
||||
"tsc": "bin/tsc",
|
||||
"tsserver": "bin/tsserver"
|
||||
"tsc": "bin/tsc"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=14.17"
|
||||
"node": ">=16.20.0"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@typescript/typescript-aix-ppc64": "7.0.2",
|
||||
"@typescript/typescript-darwin-arm64": "7.0.2",
|
||||
"@typescript/typescript-darwin-x64": "7.0.2",
|
||||
"@typescript/typescript-freebsd-arm64": "7.0.2",
|
||||
"@typescript/typescript-freebsd-x64": "7.0.2",
|
||||
"@typescript/typescript-linux-arm": "7.0.2",
|
||||
"@typescript/typescript-linux-arm64": "7.0.2",
|
||||
"@typescript/typescript-linux-loong64": "7.0.2",
|
||||
"@typescript/typescript-linux-mips64el": "7.0.2",
|
||||
"@typescript/typescript-linux-ppc64": "7.0.2",
|
||||
"@typescript/typescript-linux-riscv64": "7.0.2",
|
||||
"@typescript/typescript-linux-s390x": "7.0.2",
|
||||
"@typescript/typescript-linux-x64": "7.0.2",
|
||||
"@typescript/typescript-netbsd-arm64": "7.0.2",
|
||||
"@typescript/typescript-netbsd-x64": "7.0.2",
|
||||
"@typescript/typescript-openbsd-arm64": "7.0.2",
|
||||
"@typescript/typescript-openbsd-x64": "7.0.2",
|
||||
"@typescript/typescript-sunos-x64": "7.0.2",
|
||||
"@typescript/typescript-win32-arm64": "7.0.2",
|
||||
"@typescript/typescript-win32-x64": "7.0.2"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -24,6 +24,6 @@
|
|||
"src"
|
||||
],
|
||||
"devDependencies": {
|
||||
"typescript": "^6.0.3"
|
||||
"typescript": "^7.0.2"
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -30,7 +30,11 @@ export type { PipelinePhase, PipelineProgress } from './pipeline.js';
|
|||
|
||||
// ─── Scope-based resolution — RFC #909 (Ring 1 #910) ────────────────────────
|
||||
// Data model (RFC §2)
|
||||
export type { ParameterTypeClass, SymbolDefinition } from './scope-resolution/symbol-definition.js';
|
||||
export type {
|
||||
ParameterTypeClass,
|
||||
SymbolDefinition,
|
||||
TypeParameter,
|
||||
} from './scope-resolution/symbol-definition.js';
|
||||
export type {
|
||||
ScopeId,
|
||||
DefId,
|
||||
|
|
|
|||
|
|
@ -373,6 +373,8 @@ function makeEdgeDrafts(
|
|||
targetFile: null,
|
||||
targetExportedName: extractExportedName(parsed),
|
||||
kind: edgeKindFor(parsed),
|
||||
...typeOnlyFor(parsed),
|
||||
...runsOnlyWhenCalledFor(parsed),
|
||||
linkStatus: 'unresolved',
|
||||
};
|
||||
return [
|
||||
|
|
@ -392,7 +394,13 @@ function makeEdgeDrafts(
|
|||
// and resolved-dynamic imports are terminal at the file level — no
|
||||
// `targetDefId` needed since they materialize no `BindingRef`. Pre-
|
||||
// finalize them here so the fixpoint loop skips them entirely.
|
||||
const targetFiles = Array.isArray(targetFile) ? targetFile : [targetFile];
|
||||
// Annotated rather than inferred: `isArray`'s `arg is any[]` predicate widens
|
||||
// the true branch to a MUTABLE array, and a resolver may hand back a cached,
|
||||
// frozen candidate list (Kotlin's `dirChildren` buckets do). Only `.map` is
|
||||
// wanted here, so pinning `readonly` makes an in-place `.sort()`/`.push()` —
|
||||
// which would reorder that resolver's index for the rest of the run — a
|
||||
// compile error rather than a runtime TypeError.
|
||||
const targetFiles: readonly string[] = Array.isArray(targetFile) ? targetFile : [targetFile];
|
||||
const isFileLevelTerminal = parsed.kind === 'side-effect' || parsed.kind === 'dynamic-resolved';
|
||||
return targetFiles.map((tf) => {
|
||||
const base: ImportEdge = {
|
||||
|
|
@ -403,6 +411,8 @@ function makeEdgeDrafts(
|
|||
hooks.isNamespaceImport?.(parsed, tf, file.filePath) === true
|
||||
? 'namespace'
|
||||
: edgeKindFor(parsed),
|
||||
...typeOnlyFor(parsed),
|
||||
...runsOnlyWhenCalledFor(parsed),
|
||||
};
|
||||
return {
|
||||
source: parsed,
|
||||
|
|
@ -420,6 +430,73 @@ function edgeKindFor(parsed: ParsedImport): ImportEdge['kind'] {
|
|||
return parsed.kind;
|
||||
}
|
||||
|
||||
/**
|
||||
* Carry `ParsedImport.typeOnly` onto the edge — the erasure fact `check
|
||||
* --cycles` needs and cannot re-derive, because `kind` is identical for the
|
||||
* erased and the runtime spelling of the same import (`import type D` and
|
||||
* `import D` both arrive as `alias`).
|
||||
*
|
||||
* `'typeOnly' in parsed` rather than a switch over the erasable kinds: only
|
||||
* four variants declare the property, so `parsed.typeOnly` does not compile
|
||||
* against the whole union, and `in` narrows it without naming them. That is
|
||||
* also the safer shape — an enumeration has to be updated when a variant gains
|
||||
* the property or the fact silently stops reaching the edge, while this form
|
||||
* handles a new variant correctly whether or not it declares one.
|
||||
*
|
||||
* Returns a spreadable object rather than a `boolean` so an edge that is not
|
||||
* type-only keeps the exact property set it had before this field existed.
|
||||
* Every `finalized` edge is built by spreading `base`, so setting it here is
|
||||
* enough for all of them.
|
||||
*/
|
||||
function typeOnlyFor(parsed: ParsedImport): { typeOnly?: true } {
|
||||
return 'typeOnly' in parsed && parsed.typeOnly === true ? { typeOnly: true } : {};
|
||||
}
|
||||
|
||||
/**
|
||||
* Re-carry both runtime-presence flags from an existing edge onto a derived
|
||||
* one.
|
||||
*
|
||||
* `expandWildcard` builds each `wildcard-expanded` edge from scratch rather
|
||||
* than spreading the source (three fields differ per exported name), so every
|
||||
* field it does not name is dropped. That is exactly how both flags were lost
|
||||
* once already. Naming the pair here keeps "these two travel together" in one
|
||||
* place, so a third presence flag is added in one place too.
|
||||
*/
|
||||
function carriedPresenceFlags(edge: Pick<ImportEdge, 'typeOnly' | 'runsOnlyWhenCalled'>): {
|
||||
typeOnly?: true;
|
||||
runsOnlyWhenCalled?: true;
|
||||
} {
|
||||
return {
|
||||
...(edge.typeOnly === true ? { typeOnly: true } : {}),
|
||||
...(edge.runsOnlyWhenCalled === true ? { runsOnlyWhenCalled: true } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Carry `ParsedImport.runsOnlyWhenCalled` onto the edge — the position fact
|
||||
* `check --cycles` needs and, unlike every other property of an import, cannot
|
||||
* look up for itself.
|
||||
*
|
||||
* The scope an import was written in does not survive to here:
|
||||
* `FinalizeFile.parsedImports` is a flat per-file list, and Phase 4 publishes
|
||||
* the finalized edges under `file.moduleScope` (see `linkedByScope.set` above),
|
||||
* so the consumer's map is keyed by the Module scope for every file. Walking
|
||||
* that map's key to look for an enclosing `Function` therefore always starts —
|
||||
* and ends — at a `Module`. Only the extractor still knows, so the edge has to
|
||||
* carry what it decided.
|
||||
*
|
||||
* No `in` guard, unlike {@link typeOnlyFor}: position is a property of where
|
||||
* the statement sits, so every variant declares `runsOnlyWhenCalled` and
|
||||
* `parsed.runsOnlyWhenCalled` compiles against the whole union. A new variant
|
||||
* that omits it is a build break here, which is the right outcome.
|
||||
*
|
||||
* Returns a spreadable object rather than a `boolean` so an edge that is not
|
||||
* deferred keeps the exact property set it had before this field existed.
|
||||
*/
|
||||
function runsOnlyWhenCalledFor(parsed: ParsedImport): { runsOnlyWhenCalled?: true } {
|
||||
return parsed.runsOnlyWhenCalled === true ? { runsOnlyWhenCalled: true } : {};
|
||||
}
|
||||
|
||||
function extractLocalName(parsed: ParsedImport): string {
|
||||
switch (parsed.kind) {
|
||||
case 'wildcard':
|
||||
|
|
@ -515,9 +592,11 @@ function tryFinalize(
|
|||
return null;
|
||||
}
|
||||
|
||||
const viaFiles = [targetFile, ...followed.via];
|
||||
// Capped here too, not just inside the closure: this is the last hop, the
|
||||
// one the emitted edge carries.
|
||||
const viaFiles = extendVia(targetFile, followed.via);
|
||||
const transitiveVia =
|
||||
draft.source.kind === 'reexport' || viaFiles.length > 1 ? Object.freeze(viaFiles) : undefined;
|
||||
draft.source.kind === 'reexport' || viaFiles.length > 1 ? viaFiles : undefined;
|
||||
|
||||
return {
|
||||
...draft.base,
|
||||
|
|
@ -549,11 +628,19 @@ type FileReexportClosure = ReadonlyMap<string, ReexportClosureEntry>;
|
|||
* level import graph. Replaces the legacy recursive
|
||||
* `followReexportChain` crawl with a bounded, stack-safe pass:
|
||||
*
|
||||
* 1. **Sub-graph.** Build a directed graph whose edges are
|
||||
* `reexport` and `wildcard` drafts only (regular imports do not
|
||||
* contribute to the export surface, and `namespace`/
|
||||
* `reexport-namespace` are terminal — their target def lives in
|
||||
* `localDefs`).
|
||||
* 1. **Sub-graph.** Build a directed graph whose edges are `wildcard`
|
||||
* drafts, `reexport` drafts, and `named`/`alias` drafts flagged
|
||||
* `reexportsName` by their provider. `namespace`/`reexport-namespace`
|
||||
* are terminal — their target def lives in `localDefs` — and are
|
||||
* excluded on `base.kind`, after any `isNamespaceImport`
|
||||
* reclassification.
|
||||
*
|
||||
* The flagged-named case is what languages with no dedicated
|
||||
* re-export form need (today: Python, whose module-level
|
||||
* `from m import x` both binds and republishes). For those providers
|
||||
* the sub-graph is close to the file-level named-import graph, NOT a
|
||||
* sparse barrel graph — measured ~20× more edges on the CPython
|
||||
* stdlib — so read every bound below with that input class in mind.
|
||||
* 2. **SCC condensation.** Run the same iterative `tarjanSccs` over
|
||||
* the sub-graph. Output is in reverse-topological order (leaves
|
||||
* first), so when we process an SCC every out-of-SCC neighbor
|
||||
|
|
@ -567,21 +654,34 @@ type FileReexportClosure = ReadonlyMap<string, ReexportClosureEntry>;
|
|||
* the cycle; first-wins precedence keeps the map monotone
|
||||
* so the fixpoint converges in at most |SCC| hops).
|
||||
*
|
||||
* **Precedence semantics — preserved from the recursive crawl.**
|
||||
* **Precedence semantics.**
|
||||
* * Named re-exports take precedence over wildcards.
|
||||
* * Within each kind, declaration order wins (first match for a
|
||||
* given exported name is kept; later drafts skip).
|
||||
* given exported name is kept; later drafts skip). This is only sound
|
||||
* where the language makes a duplicate export illegal — true for TS
|
||||
* and Rust `kind: 'reexport'`, false for the flagged-named form, where
|
||||
* the module namespace rebinds (last write wins) and `if`/`try` pairs
|
||||
* execute exactly one branch. For those, an in-file collision on the
|
||||
* same published name with two different in-workspace targets is
|
||||
* genuinely ambiguous and is dropped instead of guessed — see
|
||||
* `collectAmbiguousReexports`.
|
||||
*
|
||||
* **Complexity.**
|
||||
* * Pre-pass: O(V + E_re) for SCC, plus O(|SCC| × Σ drafts) per cyclic
|
||||
* SCC. For tree-shaped barrel graphs (the common case) it
|
||||
* collapses to O(E_re) total.
|
||||
* * Per-edge lookup at finalize time: O(1).
|
||||
* SCC. Tree-shaped barrel graphs collapse to O(E_re) total; the
|
||||
* flagged-named input class does not — the CPython stdlib produces 10
|
||||
* cyclic SCCs here where TypeScript-shaped input produced none.
|
||||
* * Per-edge lookup at finalize time: O(1). Target `localDefs` are
|
||||
* indexed by simple name on first use (`findExportByName`), so the
|
||||
* per-hop cost is O(1) rather than a linear scan of the target file.
|
||||
* * `transitiveVia` preserves the exact file path chain for diagnostics
|
||||
* and graph provenance. Building those arrays copies the inherited path,
|
||||
* which is O(depth²) in a pathological single-name barrel chain; practical
|
||||
* TypeScript barrel chains are shallow enough that we keep exact paths
|
||||
* instead of capping or summarizing them.
|
||||
* which is Θ(depth²) in a single-name chain, and Θ(|SCC|²) for a cyclic
|
||||
* SCC whose chain tracks the cycle. `MAX_REEXPORT_DEPTH = 100` bounded
|
||||
* this until it was removed in `fc919ad6` for shallow TypeScript
|
||||
* barrels; **nothing bounds it now**, and the flagged-named class feeds
|
||||
* it far deeper input. Real `__init__.py` chains measure ≤ ~6, so this
|
||||
* is a known unenforced assumption, not a live regression.
|
||||
* * Pathological deep chains that previously needed
|
||||
* `MAX_REEXPORT_DEPTH=100` to bound stack growth now resolve
|
||||
* in full and are bounded only by available memory — the
|
||||
|
|
@ -595,19 +695,22 @@ function buildReexportClosures(
|
|||
const closures = new Map<string, Map<string, ReexportClosureEntry>>();
|
||||
for (const file of files) closures.set(file.filePath, new Map());
|
||||
|
||||
// ── Step 1: build the re-export sub-graph (only resolvable
|
||||
// reexport/wildcard targets contribute edges).
|
||||
// ── Step 1: build the re-export sub-graph (only resolvable wildcard /
|
||||
// reexport / flagged-named targets contribute edges), and collect the
|
||||
// per-file ambiguous names in the same walk.
|
||||
const subGraph = new Map<string, Set<string>>();
|
||||
const ambiguous = new Map<string, ReadonlySet<string>>();
|
||||
for (const file of files) {
|
||||
const targets = new Set<string>();
|
||||
const drafts = edgeIndex.get(file.filePath);
|
||||
if (drafts !== undefined) {
|
||||
for (const d of drafts) {
|
||||
if (d.source.kind !== 'reexport' && d.source.kind !== 'wildcard') continue;
|
||||
if (!contributesReexportEdge(d)) continue;
|
||||
if (d.targetFile === null) continue;
|
||||
if (!byFilePath.has(d.targetFile)) continue;
|
||||
targets.add(d.targetFile);
|
||||
}
|
||||
ambiguous.set(file.filePath, collectAmbiguousReexports(drafts, byFilePath));
|
||||
}
|
||||
subGraph.set(file.filePath, targets);
|
||||
}
|
||||
|
|
@ -623,7 +726,7 @@ function buildReexportClosures(
|
|||
if (!scc.isCycle) {
|
||||
const filePath = scc.files[0];
|
||||
if (filePath !== undefined) {
|
||||
populateFileClosure(filePath, byFilePath, edgeIndex, closures);
|
||||
populateFileClosure(filePath, byFilePath, edgeIndex, closures, ambiguous);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
|
@ -637,7 +740,7 @@ function buildReexportClosures(
|
|||
progressed = false;
|
||||
iter++;
|
||||
for (const filePath of scc.files) {
|
||||
if (populateFileClosure(filePath, byFilePath, edgeIndex, closures)) {
|
||||
if (populateFileClosure(filePath, byFilePath, edgeIndex, closures, ambiguous)) {
|
||||
progressed = true;
|
||||
}
|
||||
}
|
||||
|
|
@ -647,6 +750,95 @@ function buildReexportClosures(
|
|||
return closures;
|
||||
}
|
||||
|
||||
/**
|
||||
* Does this import republish names from its target under the *importing* file,
|
||||
* making it an edge in the re-export sub-graph?
|
||||
*
|
||||
* `reexport` and `wildcard` are the explicit forms; `named`/`alias` drafts
|
||||
* flagged `reexportsName` cover providers whose ordinary import syntax also
|
||||
* republishes (see that field on `ParsedImport` for the contract).
|
||||
*
|
||||
* Tested on `base.kind`, not `source.kind`: `isNamespaceImport` can reclassify
|
||||
* a `named` draft to `namespace` (Python's `from . import submodule`), and a
|
||||
* namespace import aliases the target *module* — it publishes no name, so
|
||||
* admitting it would republish whatever def happens to share the module's
|
||||
* simple name.
|
||||
*/
|
||||
function contributesReexportEdge(draft: ImportEdgeDraft): boolean {
|
||||
if (draft.base.kind === 'namespace') return false;
|
||||
if (draft.source.kind === 'wildcard') return true;
|
||||
return isNamedReexport(draft);
|
||||
}
|
||||
|
||||
/**
|
||||
* Named (non-wildcard) re-export. The narrowed type lets `populateFileClosure`
|
||||
* read `localName` (the name this file publishes) and `importedName` (the name
|
||||
* the target exports) without re-discriminating on `kind`.
|
||||
*/
|
||||
function isNamedReexport(draft: ImportEdgeDraft): draft is ImportEdgeDraft & {
|
||||
readonly source: Extract<ParsedImport, { kind: 'named' | 'alias' | 'reexport' }>;
|
||||
} {
|
||||
if (draft.base.kind === 'namespace') return false;
|
||||
const source = draft.source;
|
||||
if (source.kind === 'reexport') return true;
|
||||
return (source.kind === 'named' || source.kind === 'alias') && source.reexportsName === true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Names this file publishes ambiguously, which the closure must decline to
|
||||
* answer for rather than guess at.
|
||||
*
|
||||
* Declaration-order first-wins is sound only where a duplicate export is
|
||||
* illegal — two `export { X } from …` is a TypeScript compile error, so the
|
||||
* rule never fires. The flagged-named form has no such guarantee: CPython's
|
||||
* module namespace rebinds, so
|
||||
*
|
||||
* from .v1 import Client # legacy, left behind
|
||||
* from .v2 import Client # the actual public Client
|
||||
*
|
||||
* binds `v2`, and first-wins would attribute every `from pkg import Client` in
|
||||
* the repo to the dead implementation. Last-wins is not the answer either —
|
||||
* for the equally common `try:`/`except ImportError:` and `if
|
||||
* sys.version_info` pairs exactly one branch runs, and which one is not
|
||||
* decidable here. So both directions are wrong on real code and the entry is
|
||||
* dropped: the importer stays unresolved, which is exactly the pre-#2864
|
||||
* answer, and the file-level IMPORTS edge is unaffected.
|
||||
*
|
||||
* Computed once per file from data phase 0 froze (`edgeIndex`, `targetFile`)
|
||||
* and never revised, so the closure map stays monotone and the `|SCC| + 1`
|
||||
* fixpoint cap keeps the meaning it has above. A set that could grow mid-
|
||||
* fixpoint would need retraction to propagate to files that already inherited
|
||||
* the name, and would break both.
|
||||
*
|
||||
* Only two flagged drafts resolving to two *different in-workspace files*
|
||||
* count. Duplicates of the same target are harmless, and an unresolvable
|
||||
* target (`null` — the `try: import ujson / except: import json` shape, both
|
||||
* external) never entered the closure to begin with.
|
||||
*
|
||||
* ponytail: named-vs-named only. Wildcard-vs-wildcard collisions are also
|
||||
* first-wins today, but their inherited half depends on target closures that
|
||||
* are still filling in, so detecting them needs a set that grows during the
|
||||
* fixpoint — the thing this pre-pass exists to avoid.
|
||||
*/
|
||||
function collectAmbiguousReexports(
|
||||
drafts: readonly ImportEdgeDraft[],
|
||||
byFilePath: ReadonlyMap<string, FinalizeFile>,
|
||||
): ReadonlySet<string> {
|
||||
const firstTarget = new Map<string, string>();
|
||||
const conflicting = new Set<string>();
|
||||
for (const draft of drafts) {
|
||||
if (!isNamedReexport(draft)) continue;
|
||||
if (draft.source.kind === 'reexport') continue; // explicit form: duplicates are illegal upstream
|
||||
const targetFile = draft.targetFile;
|
||||
if (targetFile === null || !byFilePath.has(targetFile)) continue;
|
||||
const localName = draft.source.localName;
|
||||
const seen = firstTarget.get(localName);
|
||||
if (seen === undefined) firstTarget.set(localName, targetFile);
|
||||
else if (seen !== targetFile) conflicting.add(localName);
|
||||
}
|
||||
return conflicting;
|
||||
}
|
||||
|
||||
/**
|
||||
* Populate one file's re-export closure for one pass. Returns `true`
|
||||
* iff the closure grew (signalling fixpoint progress to the caller).
|
||||
|
|
@ -666,24 +858,29 @@ function populateFileClosure(
|
|||
byFilePath: ReadonlyMap<string, FinalizeFile>,
|
||||
edgeIndex: ReadonlyMap<string, ImportEdgeDraft[]>,
|
||||
closures: Map<string, Map<string, ReexportClosureEntry>>,
|
||||
ambiguousByFile: ReadonlyMap<string, ReadonlySet<string>>,
|
||||
): boolean {
|
||||
const myClosure = closures.get(filePath);
|
||||
if (myClosure === undefined) return false;
|
||||
const before = myClosure.size;
|
||||
const drafts = edgeIndex.get(filePath);
|
||||
if (drafts === undefined) return false;
|
||||
// Fixed for the whole run — see `collectAmbiguousReexports`. Consulted in
|
||||
// both loops below: suppressing only the named one would let a later
|
||||
// `import *` refill the name and reinstate an arbitrary winner.
|
||||
const ambiguous = ambiguousByFile.get(filePath) ?? EMPTY_NAME_SET;
|
||||
|
||||
// Named re-exports — precedence over wildcards, declaration order
|
||||
// first-wins for duplicates of the same exported name.
|
||||
for (const draft of drafts) {
|
||||
if (draft.source.kind !== 'reexport') continue;
|
||||
if (!isNamedReexport(draft)) continue;
|
||||
const targetFile = draft.targetFile;
|
||||
if (targetFile === null) continue;
|
||||
const targetModule = byFilePath.get(targetFile);
|
||||
if (targetModule === undefined) continue;
|
||||
|
||||
const localName = draft.source.localName;
|
||||
if (myClosure.has(localName)) continue;
|
||||
if (ambiguous.has(localName) || myClosure.has(localName)) continue;
|
||||
|
||||
const importedName = draft.source.importedName;
|
||||
const direct = findExportByName(targetModule.localDefs, importedName);
|
||||
|
|
@ -695,7 +892,7 @@ function populateFileClosure(
|
|||
if (inherited !== undefined) {
|
||||
myClosure.set(localName, {
|
||||
def: inherited.def,
|
||||
via: Object.freeze([targetFile, ...inherited.via]),
|
||||
via: extendVia(targetFile, inherited.via),
|
||||
});
|
||||
}
|
||||
// Else: target's closure is still empty (in-SCC, awaiting next
|
||||
|
|
@ -714,16 +911,16 @@ function populateFileClosure(
|
|||
|
||||
for (const def of targetModule.localDefs) {
|
||||
const name = deriveSimpleName(def);
|
||||
if (name === null || myClosure.has(name)) continue;
|
||||
if (name === null || ambiguous.has(name) || myClosure.has(name)) continue;
|
||||
myClosure.set(name, { def, via: Object.freeze([targetFile]) });
|
||||
}
|
||||
const targetClosure = closures.get(targetFile);
|
||||
if (targetClosure !== undefined) {
|
||||
for (const [name, entry] of targetClosure) {
|
||||
if (myClosure.has(name)) continue;
|
||||
if (ambiguous.has(name) || myClosure.has(name)) continue;
|
||||
myClosure.set(name, {
|
||||
def: entry.def,
|
||||
via: Object.freeze([targetFile, ...entry.via]),
|
||||
via: extendVia(targetFile, entry.via),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
|
@ -732,6 +929,35 @@ function populateFileClosure(
|
|||
return myClosure.size > before;
|
||||
}
|
||||
|
||||
/**
|
||||
* Longest `transitiveVia` chain kept intact. Beyond this the tail is replaced
|
||||
* by {@link VIA_TRUNCATED}, so the entry still says "this came through a long
|
||||
* chain" without carrying it.
|
||||
*
|
||||
* Reinstates a bound the algorithm lost. Each hop copies the inherited path,
|
||||
* so an uncapped chain is Θ(depth²) in both time and retained memory, and
|
||||
* Θ(|SCC|²) for a cycle whose chain tracks it. `MAX_REEXPORT_DEPTH = 100`
|
||||
* covered this until `fc919ad6` removed it — correctly, for the TypeScript
|
||||
* barrels that were then the only input, which are shallow. Admitting
|
||||
* flagged-named imports changes the input class, so the bound comes back.
|
||||
*
|
||||
* 32 against a measured real-world worst case of ~6 for `__init__.py` chains:
|
||||
* five times the deepest chain anyone has, and it turns the quadratic into
|
||||
* O(depth × 32). Safe to truncate because `ImportEdge.transitiveVia` has no
|
||||
* production reader — it is diagnostic provenance, emitted and typed but not
|
||||
* consumed by graph emission (`emitImportEdges` dedups on source→target and
|
||||
* drops it).
|
||||
*/
|
||||
const MAX_VIA_LENGTH = 32;
|
||||
const VIA_TRUNCATED = '…';
|
||||
|
||||
function extendVia(head: string, inherited: readonly string[]): readonly string[] {
|
||||
if (inherited.length + 1 <= MAX_VIA_LENGTH) return Object.freeze([head, ...inherited]);
|
||||
// Already truncated one hop down: re-truncating keeps the array at the cap
|
||||
// rather than growing it by one per hop, which is the whole point.
|
||||
return Object.freeze([head, ...inherited.slice(0, MAX_VIA_LENGTH - 2), VIA_TRUNCATED]);
|
||||
}
|
||||
|
||||
/**
|
||||
* O(1) lookup into a precomputed re-export closure. Replaces the legacy
|
||||
* recursive `followReexportChain` traversal with a single map indexing.
|
||||
|
|
@ -792,15 +1018,52 @@ function findExportByName(
|
|||
//
|
||||
// See `gitnexus/test/integration/resolvers/typescript-hof-callbacks.test.ts`
|
||||
// for the cross-file regression this rule prevents.
|
||||
let fallback: SymbolDefinition | undefined;
|
||||
for (const d of defs) {
|
||||
if (deriveSimpleName(d) !== name) continue;
|
||||
if (isCallableOrTypeLike(d.type)) return d;
|
||||
if (fallback === undefined) fallback = d;
|
||||
}
|
||||
return fallback;
|
||||
return indexExportsByName(defs).get(name);
|
||||
}
|
||||
|
||||
/**
|
||||
* `simple name → winning def` for one file's `localDefs`, built once and
|
||||
* memoized on the array itself.
|
||||
*
|
||||
* Every caller of `findExportByName` sits in a loop that revisits the same
|
||||
* target files: the phase-3 fixpoint rescans a target once per iteration, and
|
||||
* `populateFileClosure` scans once per admitted re-export — which for a
|
||||
* provider setting `reexportsName` is every named import in the file, where it
|
||||
* used to be zero. Keeping the scan turned that into O(edges × defs).
|
||||
*
|
||||
* Safe to key on identity because `FinalizeFile.localDefs` is documented static
|
||||
* input that the fixpoint never mutates; a `WeakMap` ties each index to its
|
||||
* array's lifetime with no cross-pass state to invalidate. Same shape as the
|
||||
* `defById` map `materializeBindings` already builds for the same reason.
|
||||
*/
|
||||
const EXPORTS_BY_NAME = new WeakMap<
|
||||
readonly SymbolDefinition[],
|
||||
ReadonlyMap<string, SymbolDefinition>
|
||||
>();
|
||||
|
||||
function indexExportsByName(
|
||||
defs: readonly SymbolDefinition[],
|
||||
): ReadonlyMap<string, SymbolDefinition> {
|
||||
const cached = EXPORTS_BY_NAME.get(defs);
|
||||
if (cached !== undefined) return cached;
|
||||
const index = new Map<string, SymbolDefinition>();
|
||||
for (const d of defs) {
|
||||
const name = deriveSimpleName(d);
|
||||
if (name === null) continue;
|
||||
const existing = index.get(name);
|
||||
// First match wins within a tier; a callable displaces a stored value
|
||||
// shadow but never another callable — identical to the linear scan's
|
||||
// "first callable if any, else first match".
|
||||
if (existing === undefined) index.set(name, d);
|
||||
else if (!isCallableOrTypeLike(existing.type) && isCallableOrTypeLike(d.type))
|
||||
index.set(name, d);
|
||||
}
|
||||
EXPORTS_BY_NAME.set(defs, index);
|
||||
return index;
|
||||
}
|
||||
|
||||
const EMPTY_NAME_SET: ReadonlySet<string> = new Set();
|
||||
|
||||
const CALLABLE_OR_TYPE_LIKE: ReadonlySet<string> = new Set([
|
||||
'Function',
|
||||
'Method',
|
||||
|
|
@ -874,6 +1137,35 @@ function expandWildcard(
|
|||
kind: 'wildcard-expanded',
|
||||
targetModuleScope: edge.targetModuleScope,
|
||||
targetDefId: def.nodeId,
|
||||
// Every expanded edge inherits the presence facts of the ONE statement it
|
||||
// came from. They are built fresh rather than spread from `edge` because
|
||||
// `localName`, `targetExportedName` and `targetDefId` all differ per name
|
||||
// — which is exactly how a property added to the wildcard edge upstream
|
||||
// gets silently dropped here, and how `runsOnlyWhenCalled` was.
|
||||
//
|
||||
// `runsOnlyWhenCalled`: Ruby's `def f; require './m'; end` is one
|
||||
// statement inside one method body — and every Ruby `require` is a
|
||||
// wildcard, since the required file's whole surface becomes visible — so
|
||||
// each name it brings in is bound only when `f` runs. Losing the flag
|
||||
// here re-reports the pair as an initialization dependency and
|
||||
// suppresses nothing — it INVENTS a cycle (`check --cycles`), which is
|
||||
// why this is carried and not derived.
|
||||
//
|
||||
// Ruby is the reachable spelling. Python has no function-local
|
||||
// `from x import *` — it is a SyntaxError — and Rust's `fn f() { use
|
||||
// m::*; }`, which IS legal, is not deferred at all: `use` is a
|
||||
// compile-time path alias, so the Rust provider opts out of the position
|
||||
// rule (`LanguageProvider.importsExecuteWhereWritten`).
|
||||
//
|
||||
// `typeOnly`: unreachable today and deliberately kept. No provider emits
|
||||
// a type-only wildcard — `reexport-wildcard` returns `kind: 'wildcard'`
|
||||
// with no `typeOnly` because `export type *` is unparseable by the
|
||||
// vendored grammar (documented on `ParsedImport`'s `wildcard` variant).
|
||||
// It is propagated so the day that gap closes does not silently
|
||||
// reintroduce this same defect for erasure. Do not delete it as dead
|
||||
// code; `typeOnlyFor` is the gate that decides whether it can ever be
|
||||
// set, and it is where the correspondence is enforced.
|
||||
...carriedPresenceFlags(edge),
|
||||
});
|
||||
}
|
||||
return expanded;
|
||||
|
|
|
|||
|
|
@ -82,6 +82,28 @@ export interface ReferenceSite {
|
|||
* otherwise, in which case resolution is unchanged.
|
||||
*/
|
||||
readonly rawQualifiedName?: string;
|
||||
/**
|
||||
* Top-level generic/template arguments the source wrote ON this reference —
|
||||
* `class UserValidator : IValidator<string>` yields `['string']` on the
|
||||
* `inherits` site whose `name` is `IValidator`.
|
||||
*
|
||||
* `name` is the BASE name and stays that way: every lookup in resolution is
|
||||
* keyed by it, and one declaration answers for every instantiation of itself.
|
||||
* This records what the erasure threw away, so a consumer that needs the
|
||||
* INSTANTIATION — receiver-bound interface dispatch, which must not fan a
|
||||
* `IValidator<string>` receiver out to an `IValidator<int>` implementor
|
||||
* (#2912) — can ask for it without re-parsing the source.
|
||||
*
|
||||
* Derived generically from the anchor capture's own text (see
|
||||
* `collectReferenceSites`), so no language query change is needed: an emitter
|
||||
* whose `@reference.inherits` anchor spans the whole base gets this for free,
|
||||
* and one whose anchor is the bare name simply leaves it absent.
|
||||
*
|
||||
* ABSENT MEANS UNKNOWN, never "not generic" — the two are indistinguishable
|
||||
* here, and only the first is safe to act on. Consumers must fail OPEN on
|
||||
* absence (keep the target), matching `SymbolDefinition.typeParameters`.
|
||||
*/
|
||||
readonly typeArguments?: readonly string[];
|
||||
/** Source-text range of this reference. */
|
||||
readonly atRange: Range;
|
||||
/**
|
||||
|
|
|
|||
|
|
@ -24,6 +24,38 @@ export interface ParameterTypeClass {
|
|||
templateArguments?: string[];
|
||||
}
|
||||
|
||||
/**
|
||||
* One declared generic/template TYPE PARAMETER — `T` in `class Box<T extends
|
||||
* Repo>`, `template <class T> struct Vec`, `interface Repo<T>`.
|
||||
*
|
||||
* NOT the same axis as `SymbolDefinition.templateArguments`, and conflating the
|
||||
* two is the defect this shape exists to end. `templateArguments` records the
|
||||
* arguments a declaration was written AGAINST (`template <> struct Vec<bool>` →
|
||||
* `['bool']`); `typeParameters` records the parameters it was written IN TERMS
|
||||
* OF. A declaration can carry both — a C++ partial specialization
|
||||
* `template <class T> struct Vec<T*>` has `templateArguments: ['T*']` AND
|
||||
* `typeParameters: [{name: 'T'}]` — and that pairing is precisely what tells a
|
||||
* partial specialization apart from the full specialization `template <> struct
|
||||
* Vec<T*>`, which carries the identical `templateArguments` and NO parameters.
|
||||
*/
|
||||
export interface TypeParameter {
|
||||
/** The parameter's declared name, exactly as written (`T`, `Ts`, `TKey`). */
|
||||
name: string;
|
||||
/**
|
||||
* The declared upper bound / constraint, verbatim and un-split, when the
|
||||
* declaration states one inline: `T extends Repo` → `Repo`, `T : Repo` →
|
||||
* `Repo`, `T extends Repo & Closeable` → `Repo & Closeable`.
|
||||
*
|
||||
* VERBATIM because the intersection/compound spellings differ per language
|
||||
* and a shared consumer that wants the first bound can take the first token
|
||||
* itself, while one that wants to round-trip the source cannot recover what a
|
||||
* split threw away. Absent when the parameter is unbounded, and absent when
|
||||
* the bound is declared OUT OF LINE (C# `where T : IRepo`, Kotlin/Rust
|
||||
* `where` clauses) — see `parseTypeParameterList`.
|
||||
*/
|
||||
bound?: string;
|
||||
}
|
||||
|
||||
export interface SymbolDefinition {
|
||||
nodeId: string;
|
||||
filePath: string;
|
||||
|
|
@ -48,6 +80,18 @@ export interface SymbolDefinition {
|
|||
declaredType?: string;
|
||||
/** Generic/template specialization arguments for class-like symbols (e.g. ['User'], ['T*']). */
|
||||
templateArguments?: string[];
|
||||
/**
|
||||
* Declared generic/template TYPE PARAMETERS, in DECLARATION ORDER — see
|
||||
* {@link TypeParameter} for how this differs from `templateArguments`.
|
||||
*
|
||||
* ORDER IS LOAD-BEARING: substitution is positional (`Repo<User>` binds the
|
||||
* FIRST parameter), so a set or a name-keyed map would discard exactly the
|
||||
* information this carries. Absent for a non-generic declaration and for every
|
||||
* language whose captures do not populate it, so a reader MUST treat absence
|
||||
* as "unknown", never as "not generic" — the two are indistinguishable here
|
||||
* and only the first is safe to act on.
|
||||
*/
|
||||
typeParameters?: TypeParameter[];
|
||||
/** Per-language constraint payload for template / generic overloads
|
||||
* (e.g. C++ `enable_if_t<P, T>` predicate trees, C++20 `requires` clauses).
|
||||
* Opaque to shared code — the producing language adapter owns the shape
|
||||
|
|
@ -63,6 +107,10 @@ export interface SymbolDefinition {
|
|||
* Unavailable callables still participate in overload selection, but a
|
||||
* selected unavailable target must suppress edge emission. */
|
||||
isDeleted?: boolean;
|
||||
/** True when the declaration identity was synthesized rather than written in
|
||||
* source (for example an anonymous class). Consumers may use this only as a
|
||||
* conservative priority hint; it does not change graph-node identity. */
|
||||
isSynthetic?: boolean;
|
||||
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
|
||||
ownerId?: string;
|
||||
/** #1982/#1993: bridge-held enclosing-namespace path (e.g. `NS1`, `Outer.Inner`)
|
||||
|
|
|
|||
|
|
@ -119,8 +119,96 @@ export type ParsedImport =
|
|||
readonly importedName: string;
|
||||
readonly targetRaw: string;
|
||||
/** Provider-specific imported symbol category when module and symbol
|
||||
* namespaces have distinct resolution rules (for example PHP). */
|
||||
* namespaces have distinct resolution rules (for example PHP).
|
||||
*
|
||||
* **Not** the same fact as {@link ParsedImport.typeOnly} — see the note
|
||||
* on `typeOnly` below, which is documented on this variant. */
|
||||
readonly importedSymbolKind?: 'type' | 'function' | 'const';
|
||||
/**
|
||||
* Is this import ERASED before the module ever runs?
|
||||
*
|
||||
* TypeScript `import type { X } from './m'` and `import { type X }` are
|
||||
* deleted by `tsc`: no `require`/`import` for `./m` survives in the
|
||||
* emitted JavaScript, so the pair cannot force a module-INITIALIZATION
|
||||
* order and cannot participate in an init cycle. That is the one thing
|
||||
* `check --cycles` exists to find, so the fact has to survive from the
|
||||
* syntax down to the emitted `IMPORTS` edge — see `ImportEdge.typeOnly`
|
||||
* and `graph-bridge/imports-to-edges.ts`.
|
||||
*
|
||||
* **Distinct from `importedSymbolKind: 'type'`, which is NOT a substitute.**
|
||||
* That field is a resolution-NAMESPACE category (PHP's `use function` /
|
||||
* `use const` split), it exists only on this variant, and it says "the
|
||||
* thing imported is a type". A symbol being a type says nothing about
|
||||
* whether the import STATEMENT is erased, and PHP erases nothing at all.
|
||||
* This field is about the statement's runtime existence, not the symbol's
|
||||
* category.
|
||||
*
|
||||
* Set only by providers whose syntax marks it. Absent everywhere else,
|
||||
* which reads as "not erased" — the fail-safe direction, since it only
|
||||
* makes `check --cycles` over-report.
|
||||
*
|
||||
* That fail-safe matters more than it first looks, because an explicit
|
||||
* `type` is a SUFFICIENT signal of erasure and not a necessary one. With
|
||||
* neither `verbatimModuleSyntax` nor `importsNotUsedAsValues: preserve`
|
||||
* set — this repo sets neither — `tsc` also elides a plain
|
||||
* `import { SomeInterface }` whose bindings are every one of them used in
|
||||
* type position. Those statements are erased at run time and carry no
|
||||
* marker, so they stay tagged as initializing and `check --cycles` can
|
||||
* still report a cycle that cannot exist. Closing that gap needs
|
||||
* whole-program binding USE information, not import syntax, which is why
|
||||
* this field stops at what the syntax states.
|
||||
*/
|
||||
readonly typeOnly?: boolean;
|
||||
/**
|
||||
* Was this import written inside a function body — so that it runs only
|
||||
* when something CALLS that function, never while the module itself is
|
||||
* initializing?
|
||||
*
|
||||
* Python's `def f(): from x import Y` and a CommonJS
|
||||
* `function f() { const { Y } = require('./x'); }` are the spellings.
|
||||
* Both are syntactically ordinary imports — no `kind` tells them apart
|
||||
* from a top-level one, and nothing about the target does either. Only
|
||||
* their POSITION defers them.
|
||||
*
|
||||
* Not every language's imports are like that, and the rule is wrong for
|
||||
* the ones that are not: Rust's `use` and C/C++'s `#include` are legal
|
||||
* in a function body and are deferred by NOTHING, because neither is an
|
||||
* executed statement. Those providers opt out — see
|
||||
* `LanguageProvider.importsExecuteWhereWritten`, below.
|
||||
*
|
||||
* **Why this cannot be re-derived downstream — the whole reason the
|
||||
* field exists.** The natural place to decide it looks like the graph
|
||||
* bridge, by walking the scope the finalized edges hang off; that is
|
||||
* exactly what `graph-bridge/imports-to-edges.ts` once attempted, and it
|
||||
* is dead code by construction. `finalize-algorithm.ts:295` publishes
|
||||
* every file's finalized edges as
|
||||
* `linkedByScope.set(file.moduleScope, …)`, so the map the bridge
|
||||
* receives is keyed by the file's **Module** scope and by nothing else:
|
||||
* the walk starts at a `Module` every time and answers `false` for every
|
||||
* import in the tree. Finalize cannot recover the position either —
|
||||
* `FinalizeFile.parsedImports` is a flat per-file `ParsedImport[]` with
|
||||
* no scope attached. The extractor is the last stage that still knows
|
||||
* where the statement sat (`scope-extractor.ts`, Pass 3), so it marks the
|
||||
* fact here and it rides the edge from there — see
|
||||
* {@link ImportEdge.runsOnlyWhenCalled}.
|
||||
*
|
||||
* Consumed by `check --cycles`, which asks "can these modules be
|
||||
* initialized in any order?". A deferred import carries no
|
||||
* initialization order, and deferring one is the standard way to BREAK
|
||||
* an init cycle, so counting it reports the fix as the bug.
|
||||
*
|
||||
* Set by the central extractor for every language, not by providers —
|
||||
* except that a provider may declare that its imports do not execute
|
||||
* where they are written (`LanguageProvider.importsExecuteWhereWritten:
|
||||
* false`) and be skipped entirely. C, C++, Rust and COBOL do. A `#include`
|
||||
* or a `use` inside a function body is not deferred: the header is
|
||||
* spliced and the path alias is resolved before anything runs, so the
|
||||
* pair really is a dependency and the cycle it can form is real.
|
||||
*
|
||||
* Absent reads as "runs at initialization" — the fail-safe direction,
|
||||
* since it only makes `check --cycles` over-report.
|
||||
*/
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
/**
|
||||
* Set by providers when `targetRaw` already names the imported symbol
|
||||
* rather than only its containing module. Consumers that compose
|
||||
|
|
@ -128,6 +216,40 @@ export type ParsedImport =
|
|||
* duplicating `importedName`.
|
||||
*/
|
||||
readonly targetIncludesImportedName?: boolean;
|
||||
/**
|
||||
* Set by providers whose import syntax *also* republishes the name from
|
||||
* the importing module, so a third file can import it from there.
|
||||
*
|
||||
* Python has no dedicated re-export form: a module-level
|
||||
* `from pkg.impl import X` binds `X` locally **and** publishes it as
|
||||
* `pkg.X`, which is the standard way a package `__init__.py` declares
|
||||
* its public surface. Languages with an explicit form (TS `export … from`,
|
||||
* Rust `pub use`) emit `kind: 'reexport'` instead and leave this unset.
|
||||
*
|
||||
* **The flag must track actual republication, not syntax.** Only a
|
||||
* module-level statement publishes: the same `from m import X` inside a
|
||||
* `def` or `class` body binds locally and puts nothing in the module
|
||||
* namespace, so flagging it fabricates a re-export of a name no importer
|
||||
* can reach. `if` / `try` / `for` / `with` do not suppress it — Python
|
||||
* has no block scope. A provider that cannot tell these apart at
|
||||
* interpret time must carry the fact down from its capture emitter,
|
||||
* where the syntax node is still available.
|
||||
*
|
||||
* **Why not `kind: 'reexport'`.** Not because that form drops the local
|
||||
* binding — `materializeBindings` creates a module-scope `BindingRef`
|
||||
* for every linked edge, re-export included. It is that `reexport`
|
||||
* changes what the binding *is*: `origin` flips to `'reexport'`, which
|
||||
* carries different evidence weight and `ORIGIN_PRIORITY`, and it
|
||||
* misreports the parse-time syntax Python actually wrote. A flag adds
|
||||
* the export-surface fact without restating the import as something the
|
||||
* source does not say.
|
||||
*
|
||||
* Consumed by `buildReexportClosures` (`finalize-algorithm.ts`), which
|
||||
* also documents how ambiguous duplicates of one published name are
|
||||
* handled — the precedence rules that hold for an explicit re-export do
|
||||
* not carry over.
|
||||
*/
|
||||
readonly reexportsName?: boolean;
|
||||
}
|
||||
/**
|
||||
* Per-name import with rename.
|
||||
|
|
@ -144,8 +266,18 @@ export type ParsedImport =
|
|||
readonly targetRaw: string;
|
||||
/** See the same field on the `named` variant. */
|
||||
readonly importedSymbolKind?: 'type' | 'function' | 'const';
|
||||
/** See the same field on the `named` variant — including why it is not
|
||||
* interchangeable with `importedSymbolKind`. Reaches this variant from
|
||||
* `import type D from './m'` and `import { type X as Y } from './m'`. */
|
||||
readonly typeOnly?: boolean;
|
||||
/** See the same field on the `named` variant. Reaches this variant from
|
||||
* Python's `def f(): from x import Y as Z` and a CommonJS
|
||||
* `function f() { const { Y: Z } = require('./x'); }`. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
/** See the same field on the `named` variant. */
|
||||
readonly targetIncludesImportedName?: boolean;
|
||||
/** See the same field on the `named` variant. */
|
||||
readonly reexportsName?: boolean;
|
||||
}
|
||||
/**
|
||||
* Qualified module handle, with or without rename. `importedName` is the
|
||||
|
|
@ -165,6 +297,12 @@ export type ParsedImport =
|
|||
/** Module being aliased (e.g. `numpy` in `import numpy as np`). */
|
||||
readonly importedName: string;
|
||||
readonly targetRaw: string;
|
||||
/** See the same field on the `named` variant. Reaches this variant from
|
||||
* TypeScript `import type * as N from './m'`. */
|
||||
readonly typeOnly?: boolean;
|
||||
/** See the same field on the `named` variant. Reaches this variant from
|
||||
* Python's `def f(): import numpy as np`. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
}
|
||||
/**
|
||||
* Syntactically-detectable parse-time re-export. Finalize may still produce
|
||||
|
|
@ -186,6 +324,19 @@ export type ParsedImport =
|
|||
readonly targetRaw: string;
|
||||
/** Set when the re-export renames the symbol (e.g. `export { X as Y } from './y'`). */
|
||||
readonly alias?: string;
|
||||
/** See the same field on the `named` variant. Reaches this variant from
|
||||
* TypeScript `export type { X } from './y'` and `export { type X } from './y'`. */
|
||||
readonly typeOnly?: boolean;
|
||||
/** See the same field on the `named` variant. NO spelling reaches this
|
||||
* variant today: the two providers that emit `reexport` are TypeScript
|
||||
* / JavaScript, whose `export … from` is a module-top-level-only
|
||||
* declaration, and Rust, whose `pub use` is a compile-time path alias
|
||||
* that its provider exempts from the position rule outright
|
||||
* (`LanguageProvider.importsExecuteWhereWritten`). Kept because the
|
||||
* extractor sets the field with no `switch` on `kind`, so a re-export
|
||||
* form that IS an executed statement would be tagged the moment one
|
||||
* appears — not because anything sets it now. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
}
|
||||
/**
|
||||
* Wildcard import — brings every exported name from the target module into
|
||||
|
|
@ -197,10 +348,26 @@ export type ParsedImport =
|
|||
* - Python `from foo import *` → `{ kind: 'wildcard', targetRaw: 'foo' }`
|
||||
* - JS `export * from './foo'` → `{ kind: 'wildcard', targetRaw: './foo' }`
|
||||
* - Rust `pub use foo::*` → `{ kind: 'wildcard', targetRaw: 'foo' }`
|
||||
*
|
||||
* No `typeOnly` here on purpose. The one syntax that would set it,
|
||||
* TypeScript 5.0's `export type * from './m'`, is not parsed by the
|
||||
* vendored tree-sitter-typescript grammar — it yields an `ERROR` node
|
||||
* holding the bare `type` token, so the fact is not readable at the
|
||||
* statement level (see `typescript/import-decomposer.ts`). Add the field
|
||||
* with the grammar that can express it, not before.
|
||||
*/
|
||||
| {
|
||||
readonly kind: 'wildcard';
|
||||
readonly targetRaw: string;
|
||||
/** See the same field on the `named` variant. Present here although
|
||||
* `typeOnly` is not: erasure is a syntactic fact this spelling cannot
|
||||
* express, but POSITION is not — Ruby's `def f; require './m'; end` is
|
||||
* a wildcard (everything in the required file becomes visible) and IS
|
||||
* deferred. Python cannot reach it: `from x import *` inside a `def` is
|
||||
* a SyntaxError. Rust's fn-local `use foo::*` is legal but not
|
||||
* deferred — `use` does not execute
|
||||
* (`LanguageProvider.importsExecuteWhereWritten`). */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
}
|
||||
/**
|
||||
* Runtime-computed target — the import path is not a static literal at
|
||||
|
|
@ -217,6 +384,9 @@ export type ParsedImport =
|
|||
readonly localName: string;
|
||||
/** Source text of the unresolved expression when available; `null` otherwise. */
|
||||
readonly targetRaw: string | null;
|
||||
/** See the same field on the `named` variant. Set by position like every
|
||||
* other variant; this kind links no target, so nothing reads it here. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
}
|
||||
/**
|
||||
* Lazy / dynamic import whose target IS a static string literal at parse
|
||||
|
|
@ -238,6 +408,10 @@ export type ParsedImport =
|
|||
| {
|
||||
readonly kind: 'dynamic-resolved';
|
||||
readonly targetRaw: string;
|
||||
/** See the same field on the `named` variant. Redundant on this kind —
|
||||
* `import()` is already deferred wherever it is written — but set
|
||||
* uniformly, because position is decided without consulting `kind`. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
}
|
||||
/**
|
||||
* Bare-source / side-effect import that introduces no local name binding
|
||||
|
|
@ -253,6 +427,10 @@ export type ParsedImport =
|
|||
| {
|
||||
readonly kind: 'side-effect';
|
||||
readonly targetRaw: string;
|
||||
/** See the same field on the `named` variant. Reaches this variant from
|
||||
* a bare CommonJS `function f() { require('./polyfill'); }` — the ESM
|
||||
* spelling `import './polyfill'` cannot, being top-level only. */
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
};
|
||||
|
||||
/**
|
||||
|
|
@ -348,6 +526,37 @@ export interface ImportEdge {
|
|||
| 'side-effect';
|
||||
/** Re-export chain, for provenance (e.g., `['./y']` when re-exported via `./y`). */
|
||||
readonly transitiveVia?: readonly string[];
|
||||
/**
|
||||
* The import is erased before the module runs — see `ParsedImport`'s
|
||||
* `typeOnly` on the `named` variant for the full note, including why
|
||||
* `importedSymbolKind: 'type'` is a different fact and not a substitute.
|
||||
*
|
||||
* Carried straight from the `ParsedImport` by `makeEdgeDrafts`. The edge is
|
||||
* still emitted: a type-only import is a real source-level dependency that
|
||||
* `impact` and `trace` must see, and editing the target still breaks the
|
||||
* importer's typecheck. What the flag removes is the claim that the pair
|
||||
* forces an INITIALIZATION order.
|
||||
*/
|
||||
readonly typeOnly?: boolean;
|
||||
/**
|
||||
* The import was written inside a function body, so it runs only when that
|
||||
* function is called — never during module initialization. See
|
||||
* `ParsedImport`'s `runsOnlyWhenCalled` on the `named` variant for the full
|
||||
* note, including why the consumer cannot re-derive this from the scope tree
|
||||
* and therefore has to be told (`finalize-algorithm.ts:295`).
|
||||
*
|
||||
* Carried straight from the `ParsedImport` by `makeEdgeDrafts`, for the same
|
||||
* reason `typeOnly` is: the edge is where `graph-bridge/imports-to-edges.ts`
|
||||
* can still see it. The edge is still emitted either way — a deferred import
|
||||
* is a real dependency. What the flag removes is the claim that the pair
|
||||
* forces an INITIALIZATION order.
|
||||
*
|
||||
* Distinct from `kind === 'dynamic-resolved'`, which records the OTHER way an
|
||||
* import can be deferred (`import('./m')`). Neither implies the other: a
|
||||
* top-level `import()` is deferred with this flag unset, and a function-local
|
||||
* `from x import Y` is deferred with an ordinary `named` kind.
|
||||
*/
|
||||
readonly runsOnlyWhenCalled?: boolean;
|
||||
/** Set to `'unresolved'` when the SCC fixpoint could not link this edge. */
|
||||
readonly linkStatus?: 'unresolved';
|
||||
}
|
||||
|
|
|
|||
136
gitnexus-web/package-lock.json
generated
136
gitnexus-web/package-lock.json
generated
|
|
@ -11,14 +11,14 @@
|
|||
"@langchain/anthropic": "^1.5.1",
|
||||
"@langchain/core": "^1.2.3",
|
||||
"@langchain/google-genai": "^2.2.0",
|
||||
"@langchain/langgraph": "^1.4.8",
|
||||
"@langchain/langgraph": "^1.4.9",
|
||||
"@langchain/ollama": "^1.3.0",
|
||||
"@langchain/openai": "^1.5.3",
|
||||
"@sigma/edge-curve": "^3.1.0",
|
||||
"@tailwindcss/vite": "^4.3.3",
|
||||
"axios": "^1.18.1",
|
||||
"d3": "^7.9.0",
|
||||
"dompurify": "^3.4.12",
|
||||
"dompurify": "^3.4.13",
|
||||
"gitnexus-shared": "file:../gitnexus-shared",
|
||||
"graphology": "^0.26.0",
|
||||
"graphology-indices": "^0.17.0",
|
||||
|
|
@ -28,14 +28,14 @@
|
|||
"graphology-utils": "^2.3.0",
|
||||
"i18next": "^26.3.6",
|
||||
"i18next-browser-languagedetector": "^8.2.1",
|
||||
"langchain": "^1.4.6",
|
||||
"langchain": "^1.5.4",
|
||||
"lru-cache": "^11.5.2",
|
||||
"lucide-react": "^1.23.0",
|
||||
"mermaid": "^11.15.0",
|
||||
"lucide-react": "^1.28.0",
|
||||
"mermaid": "^11.16.1",
|
||||
"mnemonist": "^0.40.4",
|
||||
"pandemonium": "^2.4.0",
|
||||
"react": "^19.2.5",
|
||||
"react-dom": "^19.2.7",
|
||||
"react-dom": "^19.2.8",
|
||||
"react-i18next": "^17.0.11",
|
||||
"react-markdown": "^10.1.0",
|
||||
"react-syntax-highlighter": "^16.1.1",
|
||||
|
|
@ -55,10 +55,10 @@
|
|||
"@types/dompurify": "^3.2.0",
|
||||
"@types/node": "^26.0.1",
|
||||
"@types/react": "^19.2.14",
|
||||
"@types/react-dom": "^19.2.3",
|
||||
"@types/react-dom": "^19.2.4",
|
||||
"@types/react-syntax-highlighter": "^15.5.13",
|
||||
"@vercel/node": "^5.8.23",
|
||||
"@vitejs/plugin-react": "^6.0.4",
|
||||
"@vitejs/plugin-react": "^6.0.5",
|
||||
"@vitest/coverage-v8": "^4.1.9",
|
||||
"jsdom": "^29.1.1",
|
||||
"tree-sitter-wasms": "^0.1.13",
|
||||
|
|
@ -289,9 +289,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@braintree/sanitize-url": {
|
||||
"version": "7.1.1",
|
||||
"resolved": "https://registry.npmjs.org/@braintree/sanitize-url/-/sanitize-url-7.1.1.tgz",
|
||||
"integrity": "sha512-i1L7noDNxtFyL5DmZafWy1wRVhGehQmzZaz1HiN5e7iylJMSZR7ekOV7NsIqa5qBldlLrsKv4HbgFUVlQrz8Mw==",
|
||||
"version": "7.1.2",
|
||||
"resolved": "https://registry.npmjs.org/@braintree/sanitize-url/-/sanitize-url-7.1.2.tgz",
|
||||
"integrity": "sha512-jigsZK+sMF/cuiB7sERuo9V7N9jx+dhmHHnQyDSVdpZwVutaBu7WvNYqMDLSgFgfB30n452TP3vjDAvFC973mA==",
|
||||
"license": "MIT"
|
||||
},
|
||||
"node_modules/@bramus/specificity": {
|
||||
|
|
@ -1172,13 +1172,13 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@langchain/langgraph": {
|
||||
"version": "1.4.8",
|
||||
"resolved": "https://registry.npmjs.org/@langchain/langgraph/-/langgraph-1.4.8.tgz",
|
||||
"integrity": "sha512-DN1Np1XefdBEbp1qBKlt39cwoL743AAGpR5Ipja0gY2YbWvsoQnOTIrjnj/orSAhaUYsdTKS8VSWdFzsHZo6Ig==",
|
||||
"version": "1.4.9",
|
||||
"resolved": "https://registry.npmjs.org/@langchain/langgraph/-/langgraph-1.4.9.tgz",
|
||||
"integrity": "sha512-EvD9rS66Cya09y6rbMgD3Ir8miAkJQFo7FyJOPRPO736Kz3y5TeyeBDOS8ctff/jRc788bPijHx2NVFM79Qqig==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@langchain/langgraph-checkpoint": "^1.1.3",
|
||||
"@langchain/langgraph-sdk": "~1.9.26",
|
||||
"@langchain/langgraph-sdk": "~1.9.28",
|
||||
"@langchain/protocol": "^0.0.18",
|
||||
"@standard-schema/spec": "1.1.0"
|
||||
},
|
||||
|
|
@ -1330,12 +1330,12 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@mermaid-js/parser": {
|
||||
"version": "1.1.1",
|
||||
"resolved": "https://registry.npmjs.org/@mermaid-js/parser/-/parser-1.1.1.tgz",
|
||||
"integrity": "sha512-VuHdsYMK1bT6X2JbcAaWAhugTRvRBRyuZgd+c22swUeI9g/ntaxF7CY7dYarhZovofCbUNO0G7JesfmNtjYOCw==",
|
||||
"version": "1.2.0",
|
||||
"resolved": "https://registry.npmjs.org/@mermaid-js/parser/-/parser-1.2.0.tgz",
|
||||
"integrity": "sha512-oYPyv8A4As1yH5Bx+04iQEQxXuIQDe0GKCNSRgao6z8AM9jixXIfP0vsppRLvGf+nKIOb9/LdpWA4YuJiVvESA==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@chevrotain/types": "~11.1.1"
|
||||
"@chevrotain/types": "~11.1.2"
|
||||
}
|
||||
},
|
||||
"node_modules/@napi-rs/wasm-runtime": {
|
||||
|
|
@ -1755,12 +1755,6 @@
|
|||
"tailwindcss": "4.3.3"
|
||||
}
|
||||
},
|
||||
"node_modules/@tailwindcss/node/node_modules/tailwindcss": {
|
||||
"version": "4.3.2",
|
||||
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.2.tgz",
|
||||
"integrity": "sha512-WtctNNSH8A9jlMIqxzuYumOHU5uGZyRv0Q5svQl+oEPy5w84YpBxdb7MdqyiSPQge5jTJ6zFQLq0PFygdccSBA==",
|
||||
"license": "MIT"
|
||||
},
|
||||
"node_modules/@tailwindcss/oxide": {
|
||||
"version": "4.3.3",
|
||||
"resolved": "https://registry.npmjs.org/@tailwindcss/oxide/-/oxide-4.3.3.tgz",
|
||||
|
|
@ -2075,12 +2069,6 @@
|
|||
"vite": "^5.2.0 || ^6 || ^7 || ^8"
|
||||
}
|
||||
},
|
||||
"node_modules/@tailwindcss/vite/node_modules/tailwindcss": {
|
||||
"version": "4.3.2",
|
||||
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.2.tgz",
|
||||
"integrity": "sha512-WtctNNSH8A9jlMIqxzuYumOHU5uGZyRv0Q5svQl+oEPy5w84YpBxdb7MdqyiSPQge5jTJ6zFQLq0PFygdccSBA==",
|
||||
"license": "MIT"
|
||||
},
|
||||
"node_modules/@testing-library/dom": {
|
||||
"version": "10.4.1",
|
||||
"resolved": "https://registry.npmjs.org/@testing-library/dom/-/dom-10.4.1.tgz",
|
||||
|
|
@ -2601,9 +2589,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@types/react-dom": {
|
||||
"version": "19.2.3",
|
||||
"resolved": "https://registry.npmjs.org/@types/react-dom/-/react-dom-19.2.3.tgz",
|
||||
"integrity": "sha512-jp2L/eY6fn+KgVVQAOqYItbF0VY/YApe5Mz2F0aykSO8gx31bYCZyvSeYxCHKvzHG5eZjc+zyaS5BrBWya2+kQ==",
|
||||
"version": "19.2.4",
|
||||
"resolved": "https://registry.npmjs.org/@types/react-dom/-/react-dom-19.2.4.tgz",
|
||||
"integrity": "sha512-Bsc+QHgp+P/F02XDzNCY9jnZNCUuLki36KT7VKrTXXLdHf+vHMNZnW1rVu5DNW/rCK+fya3DATySbLM4yhtKUw==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"peerDependencies": {
|
||||
|
|
@ -2800,9 +2788,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@vitejs/plugin-react": {
|
||||
"version": "6.0.4",
|
||||
"resolved": "https://registry.npmjs.org/@vitejs/plugin-react/-/plugin-react-6.0.4.tgz",
|
||||
"integrity": "sha512-XcCQz0TBpBgljhj0gMuuDj49i6Ytqh5q1osT/Gp5uAVJUCTWxyskk/l1jwYYiu2xcNHHipdMz40EGfM1VdamVg==",
|
||||
"version": "6.0.5",
|
||||
"resolved": "https://registry.npmjs.org/@vitejs/plugin-react/-/plugin-react-6.0.5.tgz",
|
||||
"integrity": "sha512-BOVzne/NL162sMdResB25mUv+vWMF5NoAjNf09TeGlE7ZpszZWSD3winycicLJw72yeVsoCn/2kOhEuCvEShMA==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
|
|
@ -3470,9 +3458,9 @@
|
|||
"license": "MIT"
|
||||
},
|
||||
"node_modules/cytoscape": {
|
||||
"version": "3.33.1",
|
||||
"resolved": "https://registry.npmjs.org/cytoscape/-/cytoscape-3.33.1.tgz",
|
||||
"integrity": "sha512-iJc4TwyANnOGR1OmWhsS9ayRS3s+XQ185FmuHObThD+5AeJCakAAbWv8KimMTt08xCCLNgneQwFp+JRJOr9qGQ==",
|
||||
"version": "3.34.0",
|
||||
"resolved": "https://registry.npmjs.org/cytoscape/-/cytoscape-3.34.0.tgz",
|
||||
"integrity": "sha512-62rNSrioXw93uliKFBwjukeQyeWwH2PqDrTac31r2P6464u3AUvTk0xS4LVvT251g7IgkFunrI48ZEZGjywSOg==",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": ">=0.10"
|
||||
|
|
@ -4021,9 +4009,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/dayjs": {
|
||||
"version": "1.11.19",
|
||||
"resolved": "https://registry.npmjs.org/dayjs/-/dayjs-1.11.19.tgz",
|
||||
"integrity": "sha512-t5EcLVS6QPBNqM2z8fakk/NKel+Xzshgt8FFKAn+qwlD1pzZWxh0nVCrvFK7ZDb6XucZeF9z8C7CBWTRIVApAw==",
|
||||
"version": "1.11.21",
|
||||
"resolved": "https://registry.npmjs.org/dayjs/-/dayjs-1.11.21.tgz",
|
||||
"integrity": "sha512-98IT+HOahAisibz/yjKbzuOBwYcjJ7BCLPzARyHiyEBmRz4fatF+KPJszEHXsGYjUG234aH/cOjW1wwTbKUZlA==",
|
||||
"license": "MIT"
|
||||
},
|
||||
"node_modules/debug": {
|
||||
|
|
@ -4121,9 +4109,9 @@
|
|||
"peer": true
|
||||
},
|
||||
"node_modules/dompurify": {
|
||||
"version": "3.4.12",
|
||||
"resolved": "https://registry.npmjs.org/dompurify/-/dompurify-3.4.12.tgz",
|
||||
"integrity": "sha512-zQvGet8Z2sWbQhCmfFz/T5QWH2oBmjnqK3qvOjaqaNLrLEF912WamU+ohnTp0TCep/MFVHpdJuCZEdFOdTnEFg==",
|
||||
"version": "3.4.13",
|
||||
"resolved": "https://registry.npmjs.org/dompurify/-/dompurify-3.4.13.tgz",
|
||||
"integrity": "sha512-2vmYIoqjze2d+kakP8S/nS5shfsl587kzwEjcGlTdiksUVgFHnFCsLYDVj/JNqJVOQZGSYBTmuycv0PodwmnMQ==",
|
||||
"license": "(MPL-2.0 OR Apache-2.0)",
|
||||
"optionalDependencies": {
|
||||
"@types/trusted-types": "^2.0.7"
|
||||
|
|
@ -5345,9 +5333,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/katex": {
|
||||
"version": "0.16.27",
|
||||
"resolved": "https://registry.npmjs.org/katex/-/katex-0.16.27.tgz",
|
||||
"integrity": "sha512-aeQoDkuRWSqQN6nSvVCEFvfXdqo1OQiCmmW1kc9xSdjutPv7BGO7pqY9sQRJpMOGrEdfDgF2TfRXe5eUAD2Waw==",
|
||||
"version": "0.16.47",
|
||||
"resolved": "https://registry.npmjs.org/katex/-/katex-0.16.47.tgz",
|
||||
"integrity": "sha512-Eeo8Ys1doU1z+x8AZsPpQu+p/QcZBI5PeOo7QGQdy2x2m0MU/hYagBbGOmXwr5KVbEfVuWv9LpnQWeehogurjg==",
|
||||
"funding": [
|
||||
"https://opencollective.com/katex",
|
||||
"https://github.com/sponsors/katex"
|
||||
|
|
@ -5375,13 +5363,13 @@
|
|||
"integrity": "sha512-Ls993zuzfayK269Svk9hzpeGUKob/sIgZzyHYdjQoAdQetRKpOLj+k/QQQ/6Qi0Yz65mlROrfd+Ev+1+7dz9Kw=="
|
||||
},
|
||||
"node_modules/langchain": {
|
||||
"version": "1.4.6",
|
||||
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.4.6.tgz",
|
||||
"integrity": "sha512-pwuFmGOyiMezptLVLrpb5jILirvYPGHI5uJCFHL5K5WPxMy2XuPLI5QNMKtoHkdiL6a2dLebqugKw87cneaESw==",
|
||||
"version": "1.5.4",
|
||||
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.5.4.tgz",
|
||||
"integrity": "sha512-9Rq6Ih77UOy3+7bCbxMJS16MRUJwfxuljU0yW2KOXDgEKWE8cmaZJE6ONEy4HdWGMsbj3qyv3vD5UvV7fvNksg==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@langchain/langgraph": "^1.3.4",
|
||||
"@langchain/langgraph-checkpoint": "^1.0.4",
|
||||
"@langchain/langgraph": "^1.4.7",
|
||||
"@langchain/langgraph-checkpoint": "^1.1.3",
|
||||
"langsmith": ">=0.5.0 <1.0.0",
|
||||
"zod": "^3.25.76 || ^4"
|
||||
},
|
||||
|
|
@ -5389,7 +5377,7 @@
|
|||
"node": ">=20"
|
||||
},
|
||||
"peerDependencies": {
|
||||
"@langchain/core": "^1.2.0"
|
||||
"@langchain/core": "^1.2.3"
|
||||
}
|
||||
},
|
||||
"node_modules/langsmith": {
|
||||
|
|
@ -5727,9 +5715,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/lucide-react": {
|
||||
"version": "1.23.0",
|
||||
"resolved": "https://registry.npmjs.org/lucide-react/-/lucide-react-1.23.0.tgz",
|
||||
"integrity": "sha512-38BpJcD0JhFosxHApP/BYsBetLpQFRoTRzEzstM/XCc3jsAG7wqaY1lgVwxiUe3xqYE+lNxo2PkCmYwXWrwwIw==",
|
||||
"version": "1.28.0",
|
||||
"resolved": "https://registry.npmjs.org/lucide-react/-/lucide-react-1.28.0.tgz",
|
||||
"integrity": "sha512-fARAFJULsGuDDydjp6+6blekG/sBIM29TerzLjc9bQUKAcEfrSc4ZQKb25KRz4OMKd87cZTb5dgq0w/T6KufVg==",
|
||||
"license": "ISC",
|
||||
"peerDependencies": {
|
||||
"react": "^16.5.1 || ^17.0.0 || ^18.0.0 || ^19.0.0"
|
||||
|
|
@ -6138,26 +6126,26 @@
|
|||
}
|
||||
},
|
||||
"node_modules/mermaid": {
|
||||
"version": "11.15.0",
|
||||
"resolved": "https://registry.npmjs.org/mermaid/-/mermaid-11.15.0.tgz",
|
||||
"integrity": "sha512-pTMbcf3rWdtLiYGpmoTjHEpeY8seiy6sR+9nD7LOs8KfUbHE4lOUAprTRqRAcWSQ6MQpdX+YEsxShtGsINtPtw==",
|
||||
"version": "11.16.1",
|
||||
"resolved": "https://registry.npmjs.org/mermaid/-/mermaid-11.16.1.tgz",
|
||||
"integrity": "sha512-TQsq6u22fAn3rek5VOubrhKPo1g5hwC3FXUN9hiyupTckcYiGuuKGkNQrKYwGJkXUxZdojwRG46gsSCFZMDp4g==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@braintree/sanitize-url": "^7.1.1",
|
||||
"@braintree/sanitize-url": "^7.1.2",
|
||||
"@iconify/utils": "^3.0.2",
|
||||
"@mermaid-js/parser": "^1.1.1",
|
||||
"@mermaid-js/parser": "^1.2.0",
|
||||
"@types/d3": "^7.4.3",
|
||||
"@upsetjs/venn.js": "^2.0.0",
|
||||
"cytoscape": "^3.33.1",
|
||||
"cytoscape": "^3.33.3",
|
||||
"cytoscape-cose-bilkent": "^4.1.0",
|
||||
"cytoscape-fcose": "^2.2.0",
|
||||
"d3": "^7.9.0",
|
||||
"d3-sankey": "^0.12.3",
|
||||
"dagre-d3-es": "7.0.14",
|
||||
"dayjs": "^1.11.19",
|
||||
"dompurify": "^3.3.1",
|
||||
"dayjs": "^1.11.20",
|
||||
"dompurify": "^3.3.3",
|
||||
"es-toolkit": "^1.45.1",
|
||||
"katex": "^0.16.25",
|
||||
"katex": "^0.16.45",
|
||||
"khroma": "^2.1.0",
|
||||
"marked": "^16.3.0",
|
||||
"roughjs": "^4.6.6",
|
||||
|
|
@ -7405,24 +7393,24 @@
|
|||
"license": "MIT"
|
||||
},
|
||||
"node_modules/react": {
|
||||
"version": "19.2.7",
|
||||
"resolved": "https://registry.npmjs.org/react/-/react-19.2.7.tgz",
|
||||
"integrity": "sha512-HNe9WslTbXmFK8o8cmwgAeJFSBvt1bPdHCVKtaaV+WlAN36mpT4hcRpwbf3fY56ar2oIXzsBpOAiIRHAdY0OlQ==",
|
||||
"version": "19.2.8",
|
||||
"resolved": "https://registry.npmjs.org/react/-/react-19.2.8.tgz",
|
||||
"integrity": "sha512-PWaYA1L/q9u2u7xYQi+Y3L3Yfnie7XyLeaJICV1MGD6LprsBxcAqGjYyr0eY3p+QdsA+x/Irkt4Qif8D63+Sbw==",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": ">=0.10.0"
|
||||
}
|
||||
},
|
||||
"node_modules/react-dom": {
|
||||
"version": "19.2.7",
|
||||
"resolved": "https://registry.npmjs.org/react-dom/-/react-dom-19.2.7.tgz",
|
||||
"integrity": "sha512-t0BRVXvbiE/o20Hfw669rLbMCDWtYZLvmJigy2f0MxsXF+71pxhR3xOkspmsO8h3ZlNzyibAmtCa3l4lYKk6gQ==",
|
||||
"version": "19.2.8",
|
||||
"resolved": "https://registry.npmjs.org/react-dom/-/react-dom-19.2.8.tgz",
|
||||
"integrity": "sha512-rVprimfGBG3DR+Tq0IQG2DT5PxKth1WIGDmj5yPmlzr4YBe7uyE+Du4oVqTDXZSHGGGXRtTJEGSSePyQCMBglQ==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"scheduler": "^0.27.0"
|
||||
},
|
||||
"peerDependencies": {
|
||||
"react": "^19.2.7"
|
||||
"react": "^19.2.8"
|
||||
}
|
||||
},
|
||||
"node_modules/react-i18next": {
|
||||
|
|
|
|||
|
|
@ -21,14 +21,14 @@
|
|||
"@langchain/anthropic": "^1.5.1",
|
||||
"@langchain/core": "^1.2.3",
|
||||
"@langchain/google-genai": "^2.2.0",
|
||||
"@langchain/langgraph": "^1.4.8",
|
||||
"@langchain/langgraph": "^1.4.9",
|
||||
"@langchain/ollama": "^1.3.0",
|
||||
"@langchain/openai": "^1.5.3",
|
||||
"@sigma/edge-curve": "^3.1.0",
|
||||
"@tailwindcss/vite": "^4.3.3",
|
||||
"axios": "^1.18.1",
|
||||
"d3": "^7.9.0",
|
||||
"dompurify": "^3.4.12",
|
||||
"dompurify": "^3.4.13",
|
||||
"gitnexus-shared": "file:../gitnexus-shared",
|
||||
"graphology": "^0.26.0",
|
||||
"graphology-indices": "^0.17.0",
|
||||
|
|
@ -38,14 +38,14 @@
|
|||
"graphology-utils": "^2.3.0",
|
||||
"i18next": "^26.3.6",
|
||||
"i18next-browser-languagedetector": "^8.2.1",
|
||||
"langchain": "^1.4.6",
|
||||
"langchain": "^1.5.4",
|
||||
"lru-cache": "^11.5.2",
|
||||
"lucide-react": "^1.23.0",
|
||||
"mermaid": "^11.15.0",
|
||||
"lucide-react": "^1.28.0",
|
||||
"mermaid": "^11.16.1",
|
||||
"mnemonist": "^0.40.4",
|
||||
"pandemonium": "^2.4.0",
|
||||
"react": "^19.2.5",
|
||||
"react-dom": "^19.2.7",
|
||||
"react-dom": "^19.2.8",
|
||||
"react-i18next": "^17.0.11",
|
||||
"react-markdown": "^10.1.0",
|
||||
"react-syntax-highlighter": "^16.1.1",
|
||||
|
|
@ -65,10 +65,10 @@
|
|||
"@types/dompurify": "^3.2.0",
|
||||
"@types/node": "^26.0.1",
|
||||
"@types/react": "^19.2.14",
|
||||
"@types/react-dom": "^19.2.3",
|
||||
"@types/react-dom": "^19.2.4",
|
||||
"@types/react-syntax-highlighter": "^15.5.13",
|
||||
"@vercel/node": "^5.8.23",
|
||||
"@vitejs/plugin-react": "^6.0.4",
|
||||
"@vitejs/plugin-react": "^6.0.5",
|
||||
"@vitest/coverage-v8": "^4.1.9",
|
||||
"jsdom": "^29.1.1",
|
||||
"tree-sitter-wasms": "^0.1.13",
|
||||
|
|
|
|||
|
|
@ -21,7 +21,13 @@ import {
|
|||
fetchOpenRouterModels,
|
||||
} from '../core/llm/settings-service';
|
||||
import { getAuthToken, setAuthToken } from '../services/backend-client';
|
||||
import type { LLMSettings, LLMProvider } from '../core/llm/types';
|
||||
import type { LLMSettings, LLMProvider, MiniMaxThinkingMode } from '../core/llm/types';
|
||||
import {
|
||||
getMiniMaxModelCapabilities,
|
||||
MINIMAX_ANTHROPIC_BASE_URLS,
|
||||
MINIMAX_DOCS_ROOTS,
|
||||
MINIMAX_MODEL_IDS,
|
||||
} from '../core/llm/types';
|
||||
import { DEFAULT_OLLAMA_BASE_URL } from '../config/ui-constants';
|
||||
import { ProviderConfigCard } from './settings/ProviderConfigCard';
|
||||
import { SecretInput } from './settings/SecretInput';
|
||||
|
|
@ -341,6 +347,20 @@ export const SettingsPanel = ({
|
|||
|
||||
if (!isOpen) return null;
|
||||
|
||||
const miniMaxModel = settings.minimax?.model ?? MINIMAX_MODEL_IDS[0];
|
||||
const miniMaxCapabilities = getMiniMaxModelCapabilities(miniMaxModel);
|
||||
const configuredMiniMaxThinkingMode = settings.minimax?.thinkingMode;
|
||||
const miniMaxThinkingMode =
|
||||
configuredMiniMaxThinkingMode &&
|
||||
miniMaxCapabilities?.thinkingModes.includes(configuredMiniMaxThinkingMode)
|
||||
? configuredMiniMaxThinkingMode
|
||||
: (miniMaxCapabilities?.thinkingModes[0] ?? configuredMiniMaxThinkingMode ?? 'adaptive');
|
||||
const miniMaxBaseUrl = settings.minimax?.baseUrl ?? MINIMAX_ANTHROPIC_BASE_URLS.global_en;
|
||||
const miniMaxDocsRoot =
|
||||
miniMaxBaseUrl === MINIMAX_ANTHROPIC_BASE_URLS.cn_zh
|
||||
? MINIMAX_DOCS_ROOTS.cn_zh
|
||||
: MINIMAX_DOCS_ROOTS.global_en;
|
||||
|
||||
const providers: LLMProvider[] = [
|
||||
'openai',
|
||||
'gemini',
|
||||
|
|
@ -864,7 +884,7 @@ export const SettingsPanel = ({
|
|||
value: settings.minimax?.apiKey ?? '',
|
||||
placeholder: t('settings:providers.minimax.apiKeyPlaceholder'),
|
||||
helperText: t('settings:providers.minimax.helperText'),
|
||||
helperLink: 'https://platform.minimax.io',
|
||||
helperLink: miniMaxDocsRoot,
|
||||
helperLinkLabel: t('settings:providers.minimax.helperLinkLabel'),
|
||||
isVisible: !!showApiKey['minimax'],
|
||||
onChange: (value) =>
|
||||
|
|
@ -875,16 +895,79 @@ export const SettingsPanel = ({
|
|||
onToggleVisibility: () => toggleApiKeyVisibility('minimax'),
|
||||
}}
|
||||
model={{
|
||||
value: settings.minimax?.model ?? 'MiniMax-M2.5',
|
||||
value: miniMaxModel,
|
||||
placeholder: t('settings:providers.minimax.modelPlaceholder'),
|
||||
onChange: (value) =>
|
||||
setSettings((prev) => ({
|
||||
...prev,
|
||||
minimax: { ...prev.minimax!, model: value },
|
||||
minimax: {
|
||||
...prev.minimax!,
|
||||
model: value,
|
||||
thinkingMode:
|
||||
getMiniMaxModelCapabilities(value)?.thinkingModes[0] ??
|
||||
prev.minimax?.thinkingMode,
|
||||
},
|
||||
})),
|
||||
helperText: t('settings:providers.minimax.helperModel'),
|
||||
}}
|
||||
/>
|
||||
>
|
||||
<div className="space-y-2">
|
||||
<label className="text-sm font-medium text-text-secondary">
|
||||
{t('settings:providers.minimax.endpoint')}
|
||||
</label>
|
||||
<select
|
||||
value={miniMaxBaseUrl}
|
||||
onChange={(event) =>
|
||||
setSettings((prev) => ({
|
||||
...prev,
|
||||
minimax: { ...prev.minimax!, baseUrl: event.target.value },
|
||||
}))
|
||||
}
|
||||
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 font-mono text-sm text-text-primary transition-all outline-none focus:border-accent focus:ring-2 focus:ring-accent/20"
|
||||
>
|
||||
<option value={MINIMAX_ANTHROPIC_BASE_URLS.global_en}>
|
||||
{t('settings:providers.minimax.endpoints.global')}
|
||||
</option>
|
||||
<option value={MINIMAX_ANTHROPIC_BASE_URLS.cn_zh}>
|
||||
{t('settings:providers.minimax.endpoints.china')}
|
||||
</option>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div className="space-y-2">
|
||||
<label className="text-sm font-medium text-text-secondary">
|
||||
{t('settings:providers.minimax.thinking')}
|
||||
</label>
|
||||
<select
|
||||
value={miniMaxThinkingMode}
|
||||
disabled={miniMaxCapabilities?.thinkingModes.length === 1}
|
||||
onChange={(event) =>
|
||||
setSettings((prev) => ({
|
||||
...prev,
|
||||
minimax: {
|
||||
...prev.minimax!,
|
||||
thinkingMode: event.target.value as MiniMaxThinkingMode,
|
||||
},
|
||||
}))
|
||||
}
|
||||
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 text-sm text-text-primary transition-all outline-none focus:border-accent focus:ring-2 focus:ring-accent/20 disabled:cursor-not-allowed disabled:opacity-60"
|
||||
>
|
||||
{(miniMaxCapabilities?.thinkingModes ?? ['adaptive', 'disabled']).map((mode) => (
|
||||
<option key={mode} value={mode}>
|
||||
{t(`settings:providers.minimax.thinkingModes.${mode}`)}
|
||||
</option>
|
||||
))}
|
||||
</select>
|
||||
{miniMaxCapabilities && (
|
||||
<p className="text-xs text-text-muted">
|
||||
{t('settings:providers.minimax.capabilities', {
|
||||
contextWindow: miniMaxCapabilities.contextWindow.toLocaleString(),
|
||||
modalities: miniMaxCapabilities.inputModalities.join(', '),
|
||||
})}
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
</ProviderConfigCard>
|
||||
)}
|
||||
|
||||
{/* DeepSeek Settings */}
|
||||
|
|
|
|||
|
|
@ -20,6 +20,7 @@ import { ChatOllama } from '@langchain/ollama';
|
|||
import type { BaseChatModel } from '@langchain/core/language_models/chat_models';
|
||||
import { createGraphRAGTools, type GraphRAGBackend } from './tools';
|
||||
import type {
|
||||
AgentUserContent,
|
||||
ProviderConfig,
|
||||
OpenAIConfig,
|
||||
AzureOpenAIConfig,
|
||||
|
|
@ -32,7 +33,9 @@ import type {
|
|||
DeepSeekConfig,
|
||||
AgentStreamChunk,
|
||||
AgentHistoryMessage,
|
||||
MiniMaxThinkingMode,
|
||||
} from './types';
|
||||
import { getMiniMaxModelCapabilities, MINIMAX_ANTHROPIC_BASE_URLS } from './types';
|
||||
import {
|
||||
type CodebaseContext,
|
||||
buildDynamicSystemPrompt,
|
||||
|
|
@ -275,14 +278,28 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
|
|||
throw new Error('MiniMax API key is required but was not provided');
|
||||
}
|
||||
|
||||
const capabilities = getMiniMaxModelCapabilities(minimaxConfig.model);
|
||||
const requestedThinkingMode = minimaxConfig.thinkingMode;
|
||||
const thinkingMode: MiniMaxThinkingMode | undefined =
|
||||
requestedThinkingMode && capabilities?.thinkingModes.includes(requestedThinkingMode)
|
||||
? requestedThinkingMode
|
||||
: (capabilities?.thinkingModes[0] ?? requestedThinkingMode);
|
||||
const thinking =
|
||||
thinkingMode && thinkingMode !== 'always_on' ? { type: thinkingMode } : undefined;
|
||||
const temperature =
|
||||
thinkingMode === 'adaptive' || thinkingMode === 'always_on'
|
||||
? undefined
|
||||
: (minimaxConfig.temperature ?? 0.1);
|
||||
|
||||
return new ChatAnthropic({
|
||||
anthropicApiKey: minimaxConfig.apiKey,
|
||||
model: minimaxConfig.model,
|
||||
temperature: minimaxConfig.temperature ?? 0.1,
|
||||
...(temperature !== undefined ? { temperature } : {}),
|
||||
maxTokens: minimaxConfig.maxTokens ?? 8192,
|
||||
streaming: true,
|
||||
...(thinking ? { thinking } : {}),
|
||||
clientOptions: {
|
||||
baseURL: 'https://api.minimax.io/anthropic',
|
||||
baseURL: minimaxConfig.baseUrl ?? MINIMAX_ANTHROPIC_BASE_URLS.global_en,
|
||||
},
|
||||
});
|
||||
}
|
||||
|
|
@ -393,7 +410,7 @@ export const createGraphRAGAgent = (
|
|||
/**
|
||||
* Message type for agent conversation
|
||||
*/
|
||||
export type AgentMessage = { role: 'user'; content: string } | AgentHistoryMessage;
|
||||
export type AgentMessage = { role: 'user'; content: AgentUserContent } | AgentHistoryMessage;
|
||||
|
||||
export interface AgentRuntimeOptions {
|
||||
/** Capture assistant/tool messages for providers that require exact transcript replay. */
|
||||
|
|
@ -412,7 +429,9 @@ const isAbortError = (error: unknown, signal?: AbortSignal): boolean => {
|
|||
export const buildLangChainMessages = (messages: AgentMessage[]): BaseMessage[] =>
|
||||
messages.map((message) => {
|
||||
if (message.role === 'user') {
|
||||
return new HumanMessage(message.content);
|
||||
return typeof message.content === 'string'
|
||||
? new HumanMessage(message.content)
|
||||
: new HumanMessage({ content: message.content as any });
|
||||
}
|
||||
if (message.role === 'tool') {
|
||||
return new ToolMessage({
|
||||
|
|
@ -542,6 +561,7 @@ export async function* streamAgentResponse(
|
|||
|
||||
// Handle content that can be string or array of content blocks
|
||||
let content: string = '';
|
||||
let thinkingContent: string = '';
|
||||
if (typeof rawContent === 'string') {
|
||||
content = rawContent;
|
||||
} else if (Array.isArray(rawContent)) {
|
||||
|
|
@ -550,6 +570,14 @@ export async function* streamAgentResponse(
|
|||
.filter((block: any) => block.type === 'text' || typeof block === 'string')
|
||||
.map((block: any) => (typeof block === 'string' ? block : block.text || ''))
|
||||
.join('');
|
||||
thinkingContent = rawContent
|
||||
.filter((block: any) => block?.type === 'thinking')
|
||||
.map((block: any) => block.thinking || '')
|
||||
.join('');
|
||||
}
|
||||
|
||||
if (thinkingContent) {
|
||||
yield { type: 'reasoning', reasoning: thinkingContent };
|
||||
}
|
||||
|
||||
// If chunk has content, stream it
|
||||
|
|
|
|||
|
|
@ -19,12 +19,32 @@ import {
|
|||
GLMConfig,
|
||||
DeepSeekConfig,
|
||||
ProviderConfig,
|
||||
MINIMAX_MODEL_IDS,
|
||||
} from './types';
|
||||
import { DEFAULT_OPENROUTER_BASE_URL, DEFAULT_OLLAMA_BASE_URL } from '../../config/ui-constants';
|
||||
import { resilientFetch } from 'gitnexus-shared';
|
||||
|
||||
const STORAGE_KEY = 'gitnexus-llm-settings';
|
||||
|
||||
const mergeMiniMaxSettings = (
|
||||
stored?: LLMSettings['minimax'],
|
||||
): NonNullable<LLMSettings['minimax']> => {
|
||||
const merged = {
|
||||
...DEFAULT_LLM_SETTINGS.minimax,
|
||||
...stored,
|
||||
};
|
||||
|
||||
if (!(MINIMAX_MODEL_IDS as readonly string[]).includes(merged.model ?? '')) {
|
||||
return {
|
||||
...merged,
|
||||
model: DEFAULT_LLM_SETTINGS.minimax?.model,
|
||||
thinkingMode: DEFAULT_LLM_SETTINGS.minimax?.thinkingMode,
|
||||
};
|
||||
}
|
||||
|
||||
return merged;
|
||||
};
|
||||
|
||||
const mergeWithDefaults = (parsed?: Partial<LLMSettings> | null): LLMSettings => ({
|
||||
...DEFAULT_LLM_SETTINGS,
|
||||
...parsed,
|
||||
|
|
@ -52,10 +72,7 @@ const mergeWithDefaults = (parsed?: Partial<LLMSettings> | null): LLMSettings =>
|
|||
...DEFAULT_LLM_SETTINGS.openrouter,
|
||||
...parsed?.openrouter,
|
||||
},
|
||||
minimax: {
|
||||
...DEFAULT_LLM_SETTINGS.minimax,
|
||||
...parsed?.minimax,
|
||||
},
|
||||
minimax: mergeMiniMaxSettings(parsed?.minimax),
|
||||
glm: {
|
||||
...DEFAULT_LLM_SETTINGS.glm,
|
||||
...parsed?.glm,
|
||||
|
|
@ -437,7 +454,7 @@ export const getAvailableModels = (provider: LLMProvider): string[] => {
|
|||
case 'ollama':
|
||||
return ['llama3.2', 'llama3.1', 'mistral', 'codellama', 'deepseek-coder'];
|
||||
case 'minimax':
|
||||
return ['MiniMax-M2.5', 'MiniMax-M2.5-highspeed'];
|
||||
return [...MINIMAX_MODEL_IDS];
|
||||
case 'glm':
|
||||
return ['GLM-5', 'GLM-5-Turbo', 'GLM-4.7', 'GLM-4.5'];
|
||||
case 'deepseek':
|
||||
|
|
|
|||
|
|
@ -1233,7 +1233,20 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`,
|
|||
}
|
||||
}
|
||||
|
||||
return `No ${direction} dependencies found for "${target}" (types: ${activeRelTypes.join(', ')}). This code appears to be ${direction === 'upstream' ? 'unused (not called by anything)' : 'self-contained (no outgoing dependencies)'}.${multipleMatchWarning}`;
|
||||
// An empty UPSTREAM walk is not evidence of disuse — it is the absence
|
||||
// of evidence. The symbol may be reached only through a reference class
|
||||
// the index does not record (a property access on a plain object, a
|
||||
// dynamic dispatch, a call from a language whose resolver is weaker
|
||||
// here). The Node/MCP path reports `risk: UNKNOWN` with a `riskNote`
|
||||
// for exactly this case; this surface answers in prose rather than an
|
||||
// enum, so it carries the same MEANING rather than the same field —
|
||||
// saying "appears to be unused" here is the identical false certainty.
|
||||
//
|
||||
// Downstream keeps its wording: no outgoing dependencies really does
|
||||
// describe the symbol itself, not a claim about the rest of the repo.
|
||||
return direction === 'upstream'
|
||||
? `No ${direction} dependencies found for "${target}" (types: ${activeRelTypes.join(', ')}). This does NOT establish the symbol is unused — an empty caller set can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). Confirm with a text search before treating it as dead code.${multipleMatchWarning}`
|
||||
: `No ${direction} dependencies found for "${target}" (types: ${activeRelTypes.join(', ')}). This code appears to be self-contained (no outgoing dependencies).${multipleMatchWarning}`;
|
||||
}
|
||||
|
||||
const depth1 = byDepth.get(1) || [];
|
||||
|
|
|
|||
|
|
@ -20,6 +20,71 @@ export type LLMProvider =
|
|||
| 'glm'
|
||||
| 'deepseek';
|
||||
|
||||
export const MINIMAX_ANTHROPIC_BASE_URLS = {
|
||||
global_en: 'https://api.minimax.io/anthropic',
|
||||
cn_zh: 'https://api.minimaxi.com/anthropic',
|
||||
} as const;
|
||||
|
||||
export const MINIMAX_DOCS_ROOTS = {
|
||||
global_en: 'https://platform.minimax.io/docs',
|
||||
cn_zh: 'https://platform.minimaxi.com/docs',
|
||||
} as const;
|
||||
|
||||
export const MINIMAX_MODEL_IDS = ['MiniMax-M3', 'MiniMax-M2.7'] as const;
|
||||
|
||||
export type MiniMaxModelId = (typeof MINIMAX_MODEL_IDS)[number];
|
||||
export type MiniMaxThinkingMode = 'adaptive' | 'disabled' | 'always_on';
|
||||
export type MiniMaxInputModality = 'text' | 'image' | 'video';
|
||||
|
||||
export interface MiniMaxModelCapabilities {
|
||||
contextWindow: number;
|
||||
inputModalities: readonly MiniMaxInputModality[];
|
||||
thinkingModes: readonly MiniMaxThinkingMode[];
|
||||
}
|
||||
|
||||
export const MINIMAX_MODEL_CAPABILITIES: Record<MiniMaxModelId, MiniMaxModelCapabilities> = {
|
||||
'MiniMax-M3': {
|
||||
contextWindow: 1_000_000,
|
||||
inputModalities: ['text', 'image', 'video'],
|
||||
thinkingModes: ['adaptive', 'disabled'],
|
||||
},
|
||||
'MiniMax-M2.7': {
|
||||
contextWindow: 204_800,
|
||||
inputModalities: ['text'],
|
||||
thinkingModes: ['always_on'],
|
||||
},
|
||||
};
|
||||
|
||||
export const getMiniMaxModelCapabilities = (model: string): MiniMaxModelCapabilities | undefined =>
|
||||
MINIMAX_MODEL_CAPABILITIES[model as MiniMaxModelId];
|
||||
|
||||
export type MiniMaxMediaDetail = 'low' | 'default' | 'high';
|
||||
|
||||
export type MiniMaxMediaSource =
|
||||
| {
|
||||
type: 'url';
|
||||
url: string;
|
||||
detail?: MiniMaxMediaDetail;
|
||||
fps?: number;
|
||||
max_long_side_pixel?: number;
|
||||
}
|
||||
| {
|
||||
type: 'base64';
|
||||
media_type: string;
|
||||
data: string;
|
||||
detail?: MiniMaxMediaDetail;
|
||||
fps?: number;
|
||||
max_long_side_pixel?: number;
|
||||
};
|
||||
|
||||
export type AgentUserContent =
|
||||
| string
|
||||
| Array<
|
||||
| { type: 'text'; text: string }
|
||||
| { type: 'image'; source: MiniMaxMediaSource }
|
||||
| { type: 'video'; source: MiniMaxMediaSource }
|
||||
>;
|
||||
|
||||
/**
|
||||
* Base configuration shared by all providers
|
||||
*/
|
||||
|
|
@ -94,7 +159,9 @@ export interface OpenRouterConfig extends BaseProviderConfig {
|
|||
export interface MiniMaxConfig extends BaseProviderConfig {
|
||||
provider: 'minimax';
|
||||
apiKey: string;
|
||||
model: string; // e.g., 'MiniMax-M2.5', 'MiniMax-M2.5-highspeed'
|
||||
model: string;
|
||||
baseUrl?: string;
|
||||
thinkingMode?: MiniMaxThinkingMode;
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -200,7 +267,9 @@ export const DEFAULT_LLM_SETTINGS: LLMSettings = {
|
|||
},
|
||||
minimax: {
|
||||
apiKey: '',
|
||||
model: 'MiniMax-M2.5',
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
baseUrl: MINIMAX_ANTHROPIC_BASE_URLS.global_en,
|
||||
thinkingMode: 'adaptive',
|
||||
temperature: 0.1,
|
||||
},
|
||||
glm: {
|
||||
|
|
|
|||
|
|
@ -76,8 +76,20 @@
|
|||
"apiKeyPlaceholder": "Enter your MiniMax API key",
|
||||
"helperText": "Get your API key from",
|
||||
"helperLinkLabel": "MiniMax Platform",
|
||||
"modelPlaceholder": "e.g., MiniMax-M2.5, MiniMax-M2.5-highspeed",
|
||||
"helperModel": "Available: MiniMax-M2.5 (default), MiniMax-M2.5-highspeed (faster)"
|
||||
"modelPlaceholder": "e.g., MiniMax-M3 or MiniMax-M2.7",
|
||||
"helperModel": "Available: MiniMax-M3 (default) and MiniMax-M2.7",
|
||||
"endpoint": "Regional endpoint",
|
||||
"endpoints": {
|
||||
"global": "Global (api.minimax.io)",
|
||||
"china": "China (api.minimaxi.com)"
|
||||
},
|
||||
"thinking": "Thinking mode",
|
||||
"thinkingModes": {
|
||||
"adaptive": "Adaptive",
|
||||
"disabled": "Disabled",
|
||||
"always_on": "Always on"
|
||||
},
|
||||
"capabilities": "{{contextWindow}} token context | Inputs: {{modalities}}"
|
||||
},
|
||||
"glm": {
|
||||
"apiKeyPlaceholder": "Enter your Z.AI API key"
|
||||
|
|
|
|||
|
|
@ -76,8 +76,20 @@
|
|||
"apiKeyPlaceholder": "输入 MiniMax API Key",
|
||||
"helperText": "从这里获取 API Key:",
|
||||
"helperLinkLabel": "MiniMax Platform",
|
||||
"modelPlaceholder": "例如:MiniMax-M2.5、MiniMax-M2.5-highspeed",
|
||||
"helperModel": "可用:MiniMax-M2.5(默认)、MiniMax-M2.5-highspeed(更快)"
|
||||
"modelPlaceholder": "例如:MiniMax-M3 或 MiniMax-M2.7",
|
||||
"helperModel": "可用:MiniMax-M3(默认)和 MiniMax-M2.7",
|
||||
"endpoint": "区域端点",
|
||||
"endpoints": {
|
||||
"global": "全球(api.minimax.io)",
|
||||
"china": "中国(api.minimaxi.com)"
|
||||
},
|
||||
"thinking": "思考模式",
|
||||
"thinkingModes": {
|
||||
"adaptive": "自适应",
|
||||
"disabled": "关闭",
|
||||
"always_on": "始终开启"
|
||||
},
|
||||
"capabilities": "{{contextWindow}} token 上下文 | 输入:{{modalities}}"
|
||||
},
|
||||
"glm": {
|
||||
"apiKeyPlaceholder": "输入 Z.AI API Key"
|
||||
|
|
|
|||
|
|
@ -95,3 +95,34 @@ describe('streamAgentResponse abort', () => {
|
|||
expect(chunks).toEqual([{ type: 'error', error: 'Cannot abort the current transaction' }]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('streamAgentResponse content blocks', () => {
|
||||
const userMessage: AgentMessage[] = [{ role: 'user', content: 'hello' }];
|
||||
|
||||
it('emits thinking blocks as reasoning', async () => {
|
||||
const agent = {
|
||||
stream: async function* () {
|
||||
yield [
|
||||
'messages',
|
||||
[
|
||||
{
|
||||
_getType: () => 'ai',
|
||||
content: [{ type: 'thinking', thinking: 'Reviewing the repository context.' }],
|
||||
tool_calls: [],
|
||||
},
|
||||
],
|
||||
];
|
||||
},
|
||||
};
|
||||
|
||||
const chunks = [];
|
||||
for await (const chunk of streamAgentResponse(agent as any, userMessage)) {
|
||||
chunks.push(chunk);
|
||||
}
|
||||
|
||||
expect(chunks).toEqual([
|
||||
{ type: 'reasoning', reasoning: 'Reviewing the repository context.' },
|
||||
{ type: 'done', historyMessages: undefined },
|
||||
]);
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -10,6 +10,7 @@ import {
|
|||
DeepSeekChatOpenAI,
|
||||
DeepSeekChatOpenAICompletions,
|
||||
} from '../../src/core/llm/deepseek-chat-model';
|
||||
import { MINIMAX_ANTHROPIC_BASE_URLS, MINIMAX_MODEL_IDS } from '../../src/core/llm/types';
|
||||
|
||||
describe('buildLangChainMessages', () => {
|
||||
it('reconstructs assistant tool-call turns for replay', () => {
|
||||
|
|
@ -50,6 +51,24 @@ describe('buildLangChainMessages', () => {
|
|||
]);
|
||||
expect((langChainMessages[2] as any).tool_call_id).toBe('call_weather');
|
||||
});
|
||||
|
||||
it('preserves MiniMax image and video content blocks', () => {
|
||||
const content = [
|
||||
{ type: 'text' as const, text: 'Compare these inputs.' },
|
||||
{
|
||||
type: 'image' as const,
|
||||
source: { type: 'url' as const, url: 'https://example.com/image.png' },
|
||||
},
|
||||
{
|
||||
type: 'video' as const,
|
||||
source: { type: 'url' as const, url: 'https://example.com/video.mp4', fps: 1 },
|
||||
},
|
||||
];
|
||||
|
||||
const [message] = buildLangChainMessages([{ role: 'user', content }]);
|
||||
|
||||
expect((message as any).content).toEqual(content);
|
||||
});
|
||||
});
|
||||
|
||||
describe('serializeAgentHistoryMessages', () => {
|
||||
|
|
@ -206,6 +225,48 @@ it('drops reasoningContent from serialized assistant messages without tool calls
|
|||
});
|
||||
|
||||
describe('createChatModel', () => {
|
||||
it('configures MiniMax-M3 adaptive thinking on the China endpoint', () => {
|
||||
const model = createChatModel({
|
||||
provider: 'minimax',
|
||||
apiKey: 'minimax-test-key',
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
baseUrl: MINIMAX_ANTHROPIC_BASE_URLS.cn_zh,
|
||||
thinkingMode: 'adaptive',
|
||||
temperature: 0.1,
|
||||
} as any) as any;
|
||||
|
||||
expect(model.model).toBe(MINIMAX_MODEL_IDS[0]);
|
||||
expect(model.clientOptions.baseURL).toBe(MINIMAX_ANTHROPIC_BASE_URLS.cn_zh);
|
||||
expect(model.thinking).toEqual({ type: 'adaptive' });
|
||||
expect(model.temperature).toBeUndefined();
|
||||
});
|
||||
|
||||
it('supports disabled thinking for MiniMax-M3', () => {
|
||||
const model = createChatModel({
|
||||
provider: 'minimax',
|
||||
apiKey: 'minimax-test-key',
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
thinkingMode: 'disabled',
|
||||
temperature: 0.1,
|
||||
} as any) as any;
|
||||
|
||||
expect(model.thinking).toEqual({ type: 'disabled' });
|
||||
expect(model.temperature).toBe(0.1);
|
||||
});
|
||||
|
||||
it('keeps MiniMax-M2.7 thinking always on', () => {
|
||||
const model = createChatModel({
|
||||
provider: 'minimax',
|
||||
apiKey: 'minimax-test-key',
|
||||
model: MINIMAX_MODEL_IDS[1],
|
||||
thinkingMode: 'disabled',
|
||||
temperature: 0.1,
|
||||
} as any) as any;
|
||||
|
||||
expect(model.invocationParams({}).thinking).toBeUndefined();
|
||||
expect(model.temperature).toBeUndefined();
|
||||
});
|
||||
|
||||
it('keeps DeepSeek model subclasses on withConfig clones used for tool binding', () => {
|
||||
const model = createChatModel({
|
||||
provider: 'deepseek',
|
||||
|
|
|
|||
|
|
@ -10,6 +10,12 @@ import {
|
|||
getAvailableModels,
|
||||
getProviderCapabilities,
|
||||
} from '../../src/core/llm/settings-service';
|
||||
import {
|
||||
getMiniMaxModelCapabilities,
|
||||
MINIMAX_ANTHROPIC_BASE_URLS,
|
||||
MINIMAX_MODEL_IDS,
|
||||
} from '../../src/core/llm/types';
|
||||
import { createChatModel } from '../../src/core/llm/agent';
|
||||
|
||||
describe('loadSettings', () => {
|
||||
it('returns defaults when nothing is stored', () => {
|
||||
|
|
@ -17,6 +23,11 @@ describe('loadSettings', () => {
|
|||
expect(settings.activeProvider).toBeDefined();
|
||||
expect(settings.openai).toBeDefined();
|
||||
expect(settings.ollama).toBeDefined();
|
||||
expect(settings.minimax).toMatchObject({
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
baseUrl: MINIMAX_ANTHROPIC_BASE_URLS.global_en,
|
||||
thinkingMode: 'adaptive',
|
||||
});
|
||||
});
|
||||
|
||||
it('merges stored values with defaults', () => {
|
||||
|
|
@ -35,6 +46,30 @@ describe('loadSettings', () => {
|
|||
expect(settings.openai).toBeDefined();
|
||||
});
|
||||
|
||||
it('migrates unsupported legacy MiniMax models to the current default', () => {
|
||||
sessionStorage.setItem(
|
||||
'gitnexus-llm-settings',
|
||||
JSON.stringify({
|
||||
activeProvider: 'minimax',
|
||||
minimax: {
|
||||
apiKey: 'minimax-test-key',
|
||||
model: 'MiniMax-M2.5',
|
||||
temperature: 0.1,
|
||||
},
|
||||
}),
|
||||
);
|
||||
|
||||
const settings = loadSettings();
|
||||
expect(settings.minimax).toMatchObject({
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
thinkingMode: 'adaptive',
|
||||
});
|
||||
|
||||
const model = createChatModel(getActiveProviderConfig()!) as any;
|
||||
expect(model.model).toBe(MINIMAX_MODEL_IDS[0]);
|
||||
expect(model.thinking).toEqual({ type: 'adaptive' });
|
||||
});
|
||||
|
||||
it('returns defaults on corrupted JSON', () => {
|
||||
sessionStorage.setItem('gitnexus-llm-settings', 'not-json{{{');
|
||||
const settings = loadSettings();
|
||||
|
|
@ -116,6 +151,26 @@ describe('getActiveProviderConfig', () => {
|
|||
expect(config!.provider).toBe('deepseek');
|
||||
});
|
||||
|
||||
it('returns the regional endpoint and thinking mode for MiniMax', () => {
|
||||
const settings = loadSettings();
|
||||
settings.activeProvider = 'minimax';
|
||||
settings.minimax = {
|
||||
...settings.minimax,
|
||||
apiKey: 'minimax-test-key',
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
baseUrl: MINIMAX_ANTHROPIC_BASE_URLS.cn_zh,
|
||||
thinkingMode: 'disabled',
|
||||
};
|
||||
saveSettings(settings);
|
||||
|
||||
expect(getActiveProviderConfig()).toMatchObject({
|
||||
provider: 'minimax',
|
||||
model: MINIMAX_MODEL_IDS[0],
|
||||
baseUrl: MINIMAX_ANTHROPIC_BASE_URLS.cn_zh,
|
||||
thinkingMode: 'disabled',
|
||||
});
|
||||
});
|
||||
|
||||
it('returns null for openrouter with empty API key', () => {
|
||||
const settings = loadSettings();
|
||||
settings.activeProvider = 'openrouter';
|
||||
|
|
@ -161,6 +216,20 @@ describe('getAvailableModels', () => {
|
|||
expect(getAvailableModels('ollama').length).toBeGreaterThan(0);
|
||||
expect(getAvailableModels('anthropic')).toContain('claude-sonnet-4-20250514');
|
||||
expect(getAvailableModels('deepseek')).toContain('deepseek-v4-flash');
|
||||
expect(getAvailableModels('minimax')).toEqual([...MINIMAX_MODEL_IDS]);
|
||||
});
|
||||
|
||||
it('describes MiniMax model input and thinking capabilities', () => {
|
||||
expect(getMiniMaxModelCapabilities(MINIMAX_MODEL_IDS[0])).toEqual({
|
||||
contextWindow: 1_000_000,
|
||||
inputModalities: ['text', 'image', 'video'],
|
||||
thinkingModes: ['adaptive', 'disabled'],
|
||||
});
|
||||
expect(getMiniMaxModelCapabilities(MINIMAX_MODEL_IDS[1])).toEqual({
|
||||
contextWindow: 204_800,
|
||||
inputModalities: ['text'],
|
||||
thinkingModes: ['always_on'],
|
||||
});
|
||||
});
|
||||
|
||||
it('returns empty array for unknown provider', () => {
|
||||
|
|
|
|||
|
|
@ -1,6 +1,7 @@
|
|||
{
|
||||
"fingerprint": "69e9182ae205183ade24c3d8ad5d7292aea677144b1cbe443dd631bc25b0cafe",
|
||||
"fingerprint": "4ee15e742a9839671a900df4f57c1c91196c64256c8cab2ac445bec605a092d5",
|
||||
"scaling_budget": 1.8,
|
||||
"max_ms_large": 1000,
|
||||
"_note": "fingerprint = sha256 over per-file digests (filename + sha256(file bytes)), entry list sorted — binds each emitted line to its file so a row routed to the WRONG pair file changes the hash, AND catches within-file row reordering (file bytes hashed as-written). Byte-identity gate for #2203 U2/U3. NOTE: a future change that legitimately reorders emit (without changing the node/edge SET) will trip --check; regenerate then. scaling_budget bounds (t_large/t_small)/(LARGE/SMALL): observed ~0.95-1.05 (linear); 1.8 tolerates disk-I/O timing noise on CI while still catching an O(n^2) re-regression (~4x). max_ms_large=1000ms is a coarse absolute backstop (observed ~200ms) that catches a gross uniform slowdown the ratio gate misses; generous so CI host noise won't flake it. Regenerate via `node --import tsx bench/emit-persistence/measure.mjs`."
|
||||
"_rebaselined_2856_property_is_detail": "Third and last of the bench guards this branch left red. The Property node table gained an `isDetail` BOOLEAN column (see PROPERTY_SCHEMA in src/core/lbug/schema.ts), so `streamAllCSVsToDisk` writes one more header field and one more cell per Property row — csv-generator.ts `propertyHeader` and the `node.label === 'Property'` tail. Verified to be header-only drift rather than a change in what is emitted: dumping every CSV this bench produces on `origin/main` and on this branch and diffing per-file (filename, byte length, sha256) shows the file SET is identical at 35 CSVs on both sides, 34 of the 35 are byte-identical, and the sole difference is `property.csv` growing 68 -> 77 bytes, `id,name,filePath,startLine,endLine,content,description,declaredType` -> `...,declaredType,isDetail`. The synthetic graph has no Property nodes, so no ROW moved at all. That is the check that matters here: a row routed to the wrong pair file, or a within-file reordering, is what this fingerprint exists to catch, and neither happened. Prior 69e9182ae205183ade24c3d8ad5d7292aea677144b1cbe443dd631bc25b0cafe -> 4ee15e742a9839671a900df4f57c1c91196c64256c8cab2ac445bec605a092d5. Both timing gates passed unchanged while this was red (scaling_ratio 0.783 vs budget 1.8, elapsed_ms_large 229ms vs the 1000ms backstop), so no throughput claim is being rebaselined away.",
|
||||
"_note": "fingerprint = sha256 over per-file digests (filename + sha256(file bytes)), entry list sorted — binds each emitted line to its file so a row routed to the WRONG pair file changes the hash, AND catches within-file row reordering (file bytes hashed as-written). Byte-identity gate for #2203 U2/U3. NOTE: a future change that legitimately reorders emit (without changing the node/edge SET) will trip --check; regenerate then, and record WHY in a `_rebaselined_<reason>` key alongside — bench/scope-capture/baselines.json sets that convention and it is what makes a regenerated hash reviewable. scaling_budget bounds (t_large/t_small)/(LARGE/SMALL): observed ~0.95-1.05 (linear); 1.8 tolerates disk-I/O timing noise on CI while still catching an O(n^2) re-regression (~4x). max_ms_large=1000ms is a coarse absolute backstop (observed ~200ms) that catches a gross uniform slowdown the ratio gate misses; generous so CI host noise won't flake it. Regenerate via `node --import tsx bench/emit-persistence/measure.mjs`."
|
||||
}
|
||||
|
|
|
|||
246
gitnexus/bench/finalize-reexport/measure.mjs
Normal file
246
gitnexus/bench/finalize-reexport/measure.mjs
Normal file
|
|
@ -0,0 +1,246 @@
|
|||
/**
|
||||
* Build-free scaling bench for `buildReexportClosures`, the re-export closure
|
||||
* pass inside `finalize`.
|
||||
*
|
||||
* WHY THIS EXISTS. Until #2864 the closure sub-graph admitted only `reexport`
|
||||
* and `wildcard` drafts, so its input was TypeScript barrel files: a handful
|
||||
* of edges, shallow chains. #2864 admits `named`/`alias` drafts flagged
|
||||
* `reexportsName`, which for Python is every module-level `from m import x` —
|
||||
* measured ~20x more edges on the CPython stdlib, and cyclic SCCs where there
|
||||
* were none. The pass went from "rarely runs" to "runs over the whole named
|
||||
* import graph", and nothing measured it.
|
||||
*
|
||||
* The specific regression this guards is a QUADRATIC, and it has already
|
||||
* happened once. `populateFileClosure` copies the inherited `via` array at
|
||||
* every hop, so an unbounded chain is Theta(depth^2) in time AND retained
|
||||
* memory. `MAX_REEXPORT_DEPTH = 100` bounded it until commit `fc919ad6`
|
||||
* removed it — a correct call for shallow TS barrels, invisible for years,
|
||||
* and wrong the moment the input class changed. `MAX_VIA_LENGTH` restores the
|
||||
* bound; this bench is what notices if it goes away again. Measured at
|
||||
* depth 400: 67 ms / 145 MB uncapped vs 25 ms / 40 MB capped.
|
||||
*
|
||||
* TWO ARMS, deliberately not one, and only one of them is a timing arm:
|
||||
*
|
||||
* - `max_via_len` — EXACT and deterministic. Builds a chain far deeper than
|
||||
* the cap and asserts the longest emitted `transitiveVia` is exactly
|
||||
* `MAX_VIA_LENGTH`. Removing the cap is directly observable as a longer
|
||||
* array, so this catches it with zero flake.
|
||||
*
|
||||
* This started life as a `depth_ratio` timing arm and that was a BAD GATE.
|
||||
* Sampled five times capped it scored 2.71-3.52, and three times uncapped
|
||||
* it scored 5.87-7.65 — the ranges nearly touch, and one uncapped run came
|
||||
* in UNDER the budget. A gate that passes a third of the time on a broken
|
||||
* build is worse than no gate, because it is read as evidence. The
|
||||
* quadratic is real, but at these depths the pass's linear work dilutes it
|
||||
* enough that wall-clock cannot separate the two cleanly. The structural
|
||||
* assertion can, so it is the one that gates.
|
||||
*
|
||||
* - `width_ms` — an absolute ceiling on a wide, shallow, realistic package
|
||||
* corpus (the shape a real Python repo actually has). Structural checks
|
||||
* cannot see a constant factor: reintroducing a per-lookup linear scan of
|
||||
* a target's `localDefs` leaves every array length untouched while making
|
||||
* every real analyze slower. This arm IS timing-sensitive — re-run on an
|
||||
* idle machine before investigating. Its budget is deliberately loose; it
|
||||
* is here to catch a doubling, not to police drift.
|
||||
*
|
||||
* Both arms feed `finalize` through INDEXED hooks. The obvious mistake is to
|
||||
* reuse the unit tests' `defaultHooks`, whose `resolveImportTarget` does
|
||||
* `files.some(...)` per import — that is O(imports x files) in the FIXTURE,
|
||||
* and it swamps the pass under test so completely that removing the cap
|
||||
* measures as no change at all.
|
||||
*
|
||||
* Usage:
|
||||
* node --import tsx bench/finalize-reexport/measure.mjs # report
|
||||
* node --import tsx bench/finalize-reexport/measure.mjs --check # CI gate
|
||||
*/
|
||||
import { performance } from 'node:perf_hooks';
|
||||
import { finalize } from 'gitnexus-shared';
|
||||
|
||||
/** Must equal `MAX_VIA_LENGTH` in `gitnexus-shared`'s finalize-algorithm.ts. */
|
||||
const EXPECTED_MAX_VIA = 32;
|
||||
// Generous absolute ceiling — this arm exists to catch a restored O(n^2)
|
||||
// scan (which more than doubles it), not to police small drift.
|
||||
const WIDTH_MS_BUDGET = 1200;
|
||||
|
||||
const PROBE_DEPTH = 400;
|
||||
|
||||
const deriveSimple = (d) => {
|
||||
const q = d.qualifiedName;
|
||||
if (q === undefined || q.length === 0) return null;
|
||||
const dot = q.lastIndexOf('.');
|
||||
return dot === -1 ? q : q.slice(dot + 1);
|
||||
};
|
||||
|
||||
function hooksFor(files) {
|
||||
const byPath = new Map(files.map((f) => [f.filePath, f]));
|
||||
const byScope = new Map(files.map((f) => [f.moduleScope, f]));
|
||||
return {
|
||||
resolveImportTarget: (raw) => (raw !== null && byPath.has(raw) ? raw : null),
|
||||
expandsWildcardTo: (scope) => {
|
||||
const t = byScope.get(scope);
|
||||
return t === undefined ? [] : t.localDefs.map(deriveSimple).filter((n) => n !== null);
|
||||
},
|
||||
mergeBindings: (existing, incoming) => [...existing, ...incoming],
|
||||
};
|
||||
}
|
||||
|
||||
const mkFile = (filePath, localDefs, parsedImports) => ({
|
||||
filePath,
|
||||
moduleScope: `scope:${filePath}#1:0-9999:0:Module`,
|
||||
localDefs,
|
||||
parsedImports,
|
||||
});
|
||||
const mkDef = (qn) => ({ nodeId: `def:${qn}`, filePath: 'x', type: 'Function', qualifiedName: qn });
|
||||
const reexporting = (name, targetRaw) => ({
|
||||
kind: 'named',
|
||||
localName: name,
|
||||
importedName: name,
|
||||
targetRaw,
|
||||
reexportsName: true,
|
||||
});
|
||||
|
||||
/** A `__init__.py` chain N deep, each hop republishing the same names. */
|
||||
function chainCorpus(depth, names = 20) {
|
||||
const files = [
|
||||
mkFile(
|
||||
'leaf.py',
|
||||
Array.from({ length: names }, (_, j) => mkDef(`leaf.fn${j}`)),
|
||||
[],
|
||||
),
|
||||
];
|
||||
let prev = 'leaf.py';
|
||||
for (let d = 0; d < depth; d++) {
|
||||
const p = `hop${d}.py`;
|
||||
files.push(
|
||||
mkFile(
|
||||
p,
|
||||
[],
|
||||
Array.from({ length: names }, (_, j) => reexporting(`fn${j}`, prev)),
|
||||
),
|
||||
);
|
||||
prev = p;
|
||||
}
|
||||
files.push(
|
||||
mkFile(
|
||||
'app.py',
|
||||
[],
|
||||
Array.from({ length: names }, (_, j) => ({
|
||||
kind: 'named',
|
||||
localName: `fn${j}`,
|
||||
importedName: `fn${j}`,
|
||||
targetRaw: prev,
|
||||
})),
|
||||
),
|
||||
);
|
||||
return files;
|
||||
}
|
||||
|
||||
/** Wide and shallow: the layout a real Python repo has. */
|
||||
function packageCorpus({ leaves, defsPerLeaf, pkgSize, consumers, importsPerConsumer }) {
|
||||
const files = [];
|
||||
const leafPaths = [];
|
||||
for (let i = 0; i < leaves; i++) {
|
||||
const p = `pkg${Math.floor(i / pkgSize)}/mod${i}.py`;
|
||||
leafPaths.push(p);
|
||||
files.push(
|
||||
mkFile(
|
||||
p,
|
||||
Array.from({ length: defsPerLeaf }, (_, j) => mkDef(`mod${i}.fn${j}`)),
|
||||
[],
|
||||
),
|
||||
);
|
||||
}
|
||||
const initPaths = [];
|
||||
for (let g = 0; g < Math.ceil(leaves / pkgSize); g++) {
|
||||
const p = `pkg${g}/__init__.py`;
|
||||
initPaths.push(p);
|
||||
const imports = [];
|
||||
for (let i = g * pkgSize; i < Math.min((g + 1) * pkgSize, leaves); i++) {
|
||||
for (let j = 0; j < defsPerLeaf; j++) imports.push(reexporting(`fn${j}_${i}`, leafPaths[i]));
|
||||
}
|
||||
files.push(mkFile(p, [], imports));
|
||||
}
|
||||
for (let c = 0; c < consumers; c++) {
|
||||
const imports = [];
|
||||
for (let k = 0; k < importsPerConsumer; k++) {
|
||||
const g = (c * 7 + k) % initPaths.length;
|
||||
imports.push({
|
||||
kind: 'named',
|
||||
localName: `fn0_${g * pkgSize}`,
|
||||
importedName: `fn0_${g * pkgSize}`,
|
||||
targetRaw: initPaths[g],
|
||||
});
|
||||
}
|
||||
files.push(mkFile(`app/consumer${c}.py`, [], imports));
|
||||
}
|
||||
return files;
|
||||
}
|
||||
|
||||
function timeMedian(files, reps = 5) {
|
||||
const hooks = hooksFor(files);
|
||||
finalize({ files, workspaceIndex: undefined }, hooks); // warm
|
||||
const times = [];
|
||||
for (let r = 0; r < reps; r++) {
|
||||
const t0 = performance.now();
|
||||
finalize({ files, workspaceIndex: undefined }, hooks);
|
||||
times.push(performance.now() - t0);
|
||||
}
|
||||
times.sort((a, b) => a - b);
|
||||
return times[Math.floor(times.length / 2)];
|
||||
}
|
||||
|
||||
/** Longest `transitiveVia` any edge in this graph carries. */
|
||||
function maxViaLength(files) {
|
||||
const out = finalize({ files, workspaceIndex: undefined }, hooksFor(files));
|
||||
let max = 0;
|
||||
for (const edges of out.imports.values()) {
|
||||
for (const e of edges) {
|
||||
if (e.transitiveVia !== undefined) max = Math.max(max, e.transitiveVia.length);
|
||||
}
|
||||
}
|
||||
return max;
|
||||
}
|
||||
|
||||
const deepChain = chainCorpus(PROBE_DEPTH);
|
||||
const maxVia = maxViaLength(deepChain);
|
||||
const chainMs = timeMedian(deepChain);
|
||||
const widthMs = timeMedian(
|
||||
packageCorpus({
|
||||
leaves: 6000,
|
||||
defsPerLeaf: 8,
|
||||
pkgSize: 12,
|
||||
consumers: 3000,
|
||||
importsPerConsumer: 15,
|
||||
}),
|
||||
3,
|
||||
);
|
||||
|
||||
console.log(`chain depth ${PROBE_DEPTH} : ${chainMs.toFixed(1)} ms`);
|
||||
console.log(`max_via_len : ${maxVia} (must equal ${EXPECTED_MAX_VIA})`);
|
||||
console.log(`width_ms : ${widthMs.toFixed(1)} (budget <= ${WIDTH_MS_BUDGET})`);
|
||||
|
||||
if (process.argv.includes('--check')) {
|
||||
let failed = false;
|
||||
if (maxVia !== EXPECTED_MAX_VIA) {
|
||||
failed = true;
|
||||
console.error(
|
||||
`\nFAIL max_via_len: ${maxVia}, expected exactly ${EXPECTED_MAX_VIA}.\n` +
|
||||
`A LARGER value means the \`via\` chain copy lost its bound — see ` +
|
||||
`MAX_VIA_LENGTH in gitnexus-shared/src/scope-resolution/finalize-algorithm.ts. ` +
|
||||
`Each hop copies the inherited path, so an unbounded chain is O(depth^2) ` +
|
||||
`in time and retained memory (measured 67 ms / 145 MB vs 25 ms / 40 MB at ` +
|
||||
`depth ${PROBE_DEPTH}).\nA SMALLER value means the cap moved; update ` +
|
||||
`EXPECTED_MAX_VIA here and the two finalize-algorithm tests that pin it.`,
|
||||
);
|
||||
}
|
||||
if (widthMs > WIDTH_MS_BUDGET) {
|
||||
failed = true;
|
||||
console.error(
|
||||
`\nFAIL width_ms: ${widthMs.toFixed(1)} exceeds budget ${WIDTH_MS_BUDGET}. ` +
|
||||
`With max_via_len healthy this points at a per-lookup linear scan coming ` +
|
||||
`back (see indexExportsByName). Re-run on an idle machine first.`,
|
||||
);
|
||||
}
|
||||
if (failed) process.exit(1);
|
||||
console.log('\nOK — within budget.');
|
||||
}
|
||||
1012
gitnexus/bench/import-target/baselines.json
Normal file
1012
gitnexus/bench/import-target/baselines.json
Normal file
File diff suppressed because one or more lines are too long
2923
gitnexus/bench/import-target/measure.mjs
Normal file
2923
gitnexus/bench/import-target/measure.mjs
Normal file
File diff suppressed because it is too large
Load diff
15
gitnexus/bench/kotlin-import-target/baselines.json
Normal file
15
gitnexus/bench/kotlin-import-target/baselines.json
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
{
|
||||
"_comment": "Baselines for bench/kotlin-import-target/measure.mjs --check. `fingerprint` is a sha256 over every `fileSet | fromFile | targetRaw -> result` record the correctness corpus resolves, in BOTH file-set iteration orders; it is a CORRECTNESS gate, so drift means Kotlin import resolution started returning a different file set and IMPORTS/CALLS edges moved in every Kotlin repository. Explain it, never re-baseline to make CI green. `cases` and `non_null` are asserted beside it because a shrunken or hollowed corpus produces a perfectly valid fingerprint over a smaller surface — all three are one re-baseline, never separate ones. `scaling_budget`, `depth_budget` and `small_ms_ceiling` are timing gates and carry deliberate headroom for shared CI runners.",
|
||||
"_provenance": "RE-BASELINED ONCE, DELIBERATELY, IN #2881. The previous value ebf1790bf1d42dad483a51f2cbdeb2351e493b9e8236e4eedeef592dd81e2c5c (13256 non-null) was the PRE-INDEX implementation's, and the index that replaced it in #2872 reproduced it byte for byte — that is what made #2872 a performance change. #2881 changes resolution on purpose: `getKotlinFileIndex` no longer requires the parent directory to be the FIRST occurrence of that name in the path, so a file whose package directory name repeats higher in its own path is now a child of that package (`data/src/main/kotlin/com/example/data/Repo.kt` IS a child of `data`, and `import data.helper` resolves instead of returning null). The drift was not read off the new code and accepted; the corpus was dumped from both implementations and diffed record by record. That census was RE-RUN with a shape classifier after review found its taxonomy — 54 NULL -> resolved, 149 reselections, 32 wider fan-outs — was entirely SHAPE-PRESERVING and so had no bucket for a class this change introduces. Both ends of the re-run are validated against numbers this file already publishes, so the census is provably over the surface they describe: driven over this bench's own corpus, the BASE resolver reproduces ebf1790bf1… at 13256 non_null and the HEAD resolver reproduces d91110bee3… at 13310, across 20106 records of which 19968 are distinct. 235 distinct records moved, classified by SHAPE rather than by null-ness: null -> string 38 and null -> array 16, which together are the first census's '54 NULL -> resolved' and exactly the +54 in non_null; string -> string (a different member of a now-wider bucket) 149; array -> array (the fan-out grew) 32; string -> array 0; and ZERO of every other transition — nothing went string -> null, array -> null or array -> string, and no array shrank or reordered. So the old three buckets reappear inside the shape taxonomy exactly, and its two structural claims hold when checked directly instead of inferred: all 32 growths are order-preserving SUPERSETS of the base answer, no record lost a member, and in all 149 reselections the new answer's parent directory is named by a segment of the import and carries the same directory NAME the base answer's parent did. WHAT THE OLD TAXONOMY HAD NO BUCKET FOR is `string -> array`, and it is the one class here that is not shape-preserving: it is a RESOLVED -> UNRESOLVED transition. Tier 3 (`findKotlinPackageFiles`) runs before tier 4 (`findByProgressivePrefixStrip`), so a bucket the removed guards left empty returned null and let tier 4 answer with a single BOUND file; a now-populated bucket stops tier 4 running at all and hands back a fan-out array that need not contain the imported name at all. Two files reproduce it, in both iteration orders: ['data/src/main/kotlin/com/example/data/Repo.kt', 'common/helper.kt'] with `import data.helper` answers 'common/helper.kt' at base and ['data/src/main/kotlin/com/example/data/Repo.kt'] at head. ITS COUNT OVER THIS CORPUS IS 0, AND THAT IS A FACT ABOUT THE CORPUS RATHER THAN ABOUT THE CLASS. This file's own fuzz generator, run at ten times the repositories (4000, ~198600 distinct records), hits the class 12, 4, 10 and 10 times over four seeds — ~5e-5 per record, an expectation of about ONE over the 19968 records here — so 0 is this corpus being an order of magnitude too small to reach it, not the shape being unreachable. The consequence is worth stating plainly: the fingerprint below is blind to a resolved -> unresolved class this change introduces, by corpus SIZE and not by construction, and no arm in this bench gates it today. Adding a hand-written case for it is a deliberate fingerprint move and a fourth re-baseline of this file; it is worth doing and it is not this change. The corpus itself is untouched, which is why `cases` is unchanged at 20106 — the fingerprint is over the same surface as the value it replaces.",
|
||||
"_gate_controls": "The gate is only worth its baseline if a plausible regression moves it, so each arm was checked against the mutation it exists to catch, with the resolver otherwise untouched. All values below are against the CURRENT baseline (#2881, guards removed + per-directory key memo + bucket compaction). Caught: skipping the dirChildren component walk above depth 8 (fingerprint 41bb550b76d4…, non_null 13310 -> 12800); capping a dirChildren bucket at 17 entries (a7681945b752…, non_null UNCHANGED — the fingerprint is the only arm that sees it, and note the compaction pass now rewrites those same buckets, so this control was re-run after it); capping suffixByStem key depth at 7 (d24b8a2bd822…, non_null unchanged); and a HALF fix that drops only the `startsWith` guard while keeping the `indexOf` first-occurrence check (836977b83bf0…, non_null 13310 -> 13282), which leaves every mid-path repeat such as `top/data/mid/data/Repo.kt` broken and is the mutation #2881 itself makes plausible. Added with the memo: keying `dirKeys` on the directory's LAST SEGMENT instead of its full path (36a4e9dad313…, non_null 13310 -> 13305) — the memo's whole safety argument is that its key determines the key SET a directory contributes, so a coarser key silently hands one directory another's bucket list, and that is the one way this optimization can move an answer. Also caught, with the RESOLVER untouched and only the corpus edited: dropping the competing file from the exact-beats-earlier-suffix case and emptying the repeated-directory case (44df5093ee…). All of these passed silently before this corpus carried deep paths, packages above 16 files, queries against suffix keys deeper than 7, and the file set inside the hashed record. Re-check them after any corpus edit — a corpus that stops spanning an axis takes the gate with it. NOTE what no fingerprint control here can catch: the memo and the compaction are both invisible to this bench by design (identical output), so no arm in this file gates either one, and the honest version of where they ARE gated is narrower than a claim about comparing the three maps would suggest. The memo's gate is test/unit/scope-resolution/kotlin/kotlin-index-internals.test.ts, which drives the resolver's OBSERVABLE SURFACE rather than the built index — the index is module-private — and reconstructs what it needs from the tiers. It pins: bucket CONTENTS and ORDER, read back from the fan-out tier, which hands out the bucket array itself; that the first-child tier reads position 0 of that SAME array; bucket IDENTITY across two calls on one Set, which is what proves the memo's hit path ran at all, since only a second file in the same directory reaches it; the frozen state of the array actually handed out, on the multi-child path, on the `length === 1` skip path, and once per key of a multi-key directory; that the memo keys on the NORMALIZED directory while storing the raw path; and the one mutation that can move an answer — keying `dirKeys` on the directory's last segment instead of the whole `dir` — which fails three of its arms. KEY INSERTION ORDER is unasserted there BY DESIGN and not by omission: `dirChildren` is only ever read by `.get(key)`, so key order has no consumer, and that file says so. The COMPACTION is unasserted there too and cannot be asserted there at all — a JS array's backing-store capacity has no reflective surface, so deleting `bucket.slice()` and freezing the grown bucket in place leaves every arm in that file green, `Object.isFrozen` included. Its only instrument is the retained-heap arm in bench/import-target, whose kotlin ceiling was tightened to 1.0747x its recorded reading precisely so that the +12.57% the slice reclaims fails `--check`; see `_heap_compaction_gate` in bench/import-target/baselines.json for the measurement and for how to tell that failure apart from a runner's heapUsed accounting moving under the whole file.",
|
||||
"fingerprint": "d91110bee389891c313811c5b4bae61d909561156e1458d38d487be969f0059c",
|
||||
"cases": 20106,
|
||||
"non_null": 13310,
|
||||
"scaling_budget": 1.6,
|
||||
"depth_budget": 2.0,
|
||||
"small_ms_ceiling": 40,
|
||||
"_scaling_note": "(t_large/t_small)/(1600/400). ~1.0 is linear. OBSERVED BAND: 0.99-1.04 on a 12-core dev box, small arm ~6 ms. Read that band as a floor, not a spec — independent runs on other hardware during review came out 0.954-1.014, 0.965-1.036 and ~0.95-1.08, so a 1.2 reading is noise and should be re-run, not investigated. IMPORTS_PER_FILE is sized so the small arm lands in the ms rather than the ~2 ms a first revision measured, where timer granularity and JIT warm-up, not scaling, set the number; bench/cpp-qualified-ns documents the same artifact. TRIAGE: every timing arm here is a TIMING signal — RE-RUN IT on an idle machine before investigating; runner contention dominates. The fingerprint arm is the opposite: deterministic, a re-run never changes it, and it must never be wished away. FLOOR CHECK: the pre-index implementation — i.e. exactly the regression this gate exists to catch — measures ratio 3.737 on this corpus (2207.8 ms small, 33003.5 ms large, one cold run) against ~1.0 for the index. Independent review runs measured its floor at 3.905-4.297. Treat the absolute times as an order of magnitude only: the floor arm is one cold run because best-of-seven against a quadratic implementation costs minutes, while the index arm is best-of-seven after two warmups.",
|
||||
"_depth_note": "deep_ms/shallow_ms at a FIXED file count, paths 24 components against 8. scaling_ratio divides the file count out, so it is scale-invariant and structurally cannot see a cost that grows with path depth instead — and the two loops the index is built from are depth loops (one suffixByStem entry per '/' in a stem, one dirChildren pass per component of dir). OBSERVED BAND, five runs each on one box: 1.44-1.51 before #2881; 1.27-1.40 after its guard removal, which deleted two string comparisons per component of every dir; 1.20-1.26 after the same issue's per-directory key memo, which turns that whole component walk from once-per-FILE into once-per-DIRECTORY. Both movements are per-depth work, which is why this arm sees them and the file-count arm does not. The BUDGET moved with the band both times — 2.4 -> 2.2 -> 2.0 — holding the ~1.6x headroom over the band's top that 2.4 expressed against the original; left at 2.4 it would quietly have become 1.9x, which is how a gate goes slack without anyone deciding to loosen it. Note what this budget is NOT for: a revert of #2881 scores ~1.5 and passes at any of those numbers, and that is correct — reverting it restores a resolution bug, which is the FINGERPRINT's job to catch, not a timing arm's. It sits above 1.0 legitimately: 3x the depth is 3x the suffix keys per file, so the build genuinely does more work; what the budget forbids is that growing faster than the depth ratio itself.",
|
||||
"_ceiling_note": "small_ms_ceiling is an ABSOLUTE bound, because scaling_ratio is a ratio and a constant-factor regression that grows both arms equally passes it. Measured during review: a full workspace scan reintroduced on 1-in-16 imports is caught by the ratio (1.814), but at 1-in-32 it passes at 1.490 while running 2.8x slower in absolute terms. 40 ms against an observed 5.9-6.1 ms leaves ~6x of headroom for a loaded shared runner while still catching that shape.",
|
||||
"_blind_spot": "WHAT THIS BENCH CANNOT SEE, measured rather than guessed. Its scaling corpus gives every module a UNIQUE package leaf (`com/example/mod{N}`), so a `dirChildren` query matches exactly one directory. That makes it blind to any cost that grows with the number of DIRECTORIES sharing a queried segment — the shape a real Kotlin monorepo has, where 200 modules each hold `data`, `ui` and `domain`. Established by building the reuse this file's memo argues against: swapping `dirChildren` for the shared `import-resolvers/package-dir-index.ts` (with its first-occurrence rule off) is OUTPUT-IDENTICAL — same fingerprint, same cases, same non_null, 0 divergences over 107948 answers — and on THIS corpus it costs only 1.37x-1.50x and passes every arm here. On a repeated-leaf corpus the same swap measures 13.5x per first-child query, 409x per fan-out, and 8114x on `import data.*` at 200 matching directories (it merges and SORTS every candidate, per import), for 3.1x-5.7x end to end and a bench-style scaling_ratio of 3.465 against this file's 1.6 budget — i.e. back to the pre-index quadratic floor of 3.737. A change that regresses this resolver to the very shape the bench exists to catch would go GREEN here. The trade it buys is real and also measured: 26.2% less retained memory, 12.18 MiB at 32000 files. If that memory is ever wanted, the shape to build is per-suffix keys -> DIRECTORY lists plus files-per-directory (8.29 MiB against 15.87 measured, single-directory query still one hash lookup) — and the repeated-leaf arm to measure it against already exists one directory over: bench/import-target's kotlin `collide` layout puts `com/example/models` under 200 modules at the 1600-file scale, with `collide_scaling_budget` 1.8 against a measured 1.081. The swap scores 3.465 there. So the gate for this decision is that arm, not a new one here; what this file lacks is only a repeated-leaf arm of its own, which would be duplicated coverage."
|
||||
}
|
||||
539
gitnexus/bench/kotlin-import-target/measure.mjs
Normal file
539
gitnexus/bench/kotlin-import-target/measure.mjs
Normal file
|
|
@ -0,0 +1,539 @@
|
|||
/**
|
||||
* Build-free identity + scaling bench for `resolveKotlinImportTarget`, the
|
||||
* Kotlin import resolver.
|
||||
*
|
||||
* Before this bench's companion change the resolver walked the ENTIRE
|
||||
* `allFilePaths` Set on every import. Its four tiers — exact/suffix,
|
||||
* directory child, package fan-out, progressive prefix strip — each ran
|
||||
* `for (const raw of allFilePaths)` with a `replace(/\\/g, '/')` and several
|
||||
* string comparisons per entry, and they are tried in cascade, so one
|
||||
* unresolved import cost two to four full passes. Resolution was therefore
|
||||
* O(imports x files). Once a repository reaches tens of thousands of Kotlin
|
||||
* files that is on the order of 10^10 string operations on one thread:
|
||||
* `analyze` sits at exactly 1.00 core with a flat heap and emits nothing for
|
||||
* hours, because every allocation is a short-lived string and nothing
|
||||
* accumulates to hint at progress.
|
||||
*
|
||||
* This is the same shape #1918 fixed for Python and #2788 for C++, and it
|
||||
* returns the same way: someone adds a tier, reaches for `allFilePaths`, and
|
||||
* writes a loop. Neither existing gate can catch it here —
|
||||
* `bench/python-scope/import-target-fingerprint.mjs` drives the Python
|
||||
* resolver only, and `bench/scope-capture/measure.mjs` fingerprints
|
||||
* `emit<Lang>ScopeCaptures`, a different function that never calls import
|
||||
* resolution. Hence this bench, in an always-on CI step.
|
||||
*
|
||||
* TWO ARMS, and they fail for opposite reasons:
|
||||
*
|
||||
* - `fingerprint` — a sha256 over every `fromFile | targetRaw -> result`
|
||||
* triple the correctness corpus resolves (an exhaustive branch matrix plus
|
||||
* a deterministic fuzz). This is a CORRECTNESS gate. Drift means Kotlin
|
||||
* imports started resolving a DIFFERENT file set, i.e. CALLS/IMPORTS edges
|
||||
* moved in every Kotlin repository. It is deterministic: a re-run never
|
||||
* changes it, and it must never be re-baselined to make CI green. This
|
||||
* value is the one the pre-index implementation produced — see
|
||||
* `_provenance` in baselines.json.
|
||||
*
|
||||
* - `scaling_ratio` — `(t_large/t_small)/(LARGE/SMALL)` over a synthetic
|
||||
* Kotlin monorepo at two scales, timing the index build TOGETHER with
|
||||
* resolving every import. ~1.0 is linear; a reintroduced per-import scan
|
||||
* measures ~4 at this scale gap. This is a TIMING gate: re-run it on an
|
||||
* idle machine before investigating.
|
||||
*
|
||||
* A ratio cannot see a constant factor and a file-count ratio cannot see a
|
||||
* depth cost, so `--check` also asserts a DEPTH ratio (file count fixed, paths
|
||||
* ~3x deeper) and an absolute ceiling on the small arm. A full workspace scan
|
||||
* reintroduced on 1-in-32 imports scores 1.490 — inside the scaling budget —
|
||||
* while running 2.8x slower; the ceiling is what catches that shape.
|
||||
*
|
||||
* One honest limit: at a very small import count the index loses. Building it
|
||||
* is one workspace pass, so a single import into a 100k-file workspace costs
|
||||
* ~0.8 s against ~0 for a scan that returns on its first hit. It inverts at
|
||||
* roughly 15 imports, and in the polyglot case that motivates the worry —
|
||||
* 100k files, 5% Kotlin, a couple of imports — the index already wins, because
|
||||
* the build skips non-`.kt` entries as cheaply as the scan did.
|
||||
*
|
||||
* Five properties of the corpora are load-bearing and must not be
|
||||
* "simplified" away:
|
||||
*
|
||||
* 1. **The correctness corpus fuzzes each file set in BOTH iteration
|
||||
* orders.** Every tie-break in this resolver is expressed only through
|
||||
* Set-iteration order — "first suffix match wins", and the two stem maps
|
||||
* keeping the FIRST path inserted per key. A single-order corpus scores an
|
||||
* implementation that keeps the LAST match identically.
|
||||
* 2. **The correctness corpus contains repeated directory names at BOTH the
|
||||
* leading and the mid-path position** (`data/src/main/kotlin/com/example/
|
||||
* data/Repo.kt` and `top/data/mid/data/Repo.kt`). Until #2881 the resolver
|
||||
* required a file's package directory to be the FIRST occurrence of that
|
||||
* name in its own path, so neither file was a child of `data`; both are
|
||||
* now, and that is what the fingerprint pins. Two positions, not one,
|
||||
* because the old rule was two guards and a half fix that drops only the
|
||||
* leading-position one still leaves the mid-path shape broken — see
|
||||
* `_gate_controls` in baselines.json. Without these shapes the fingerprint
|
||||
* cannot tell the current rule from either predecessor.
|
||||
* 3. **~40% of the scaling corpus's imports are unresolvable.** The old cost
|
||||
* was worst when nothing matched, because only then did all four tiers
|
||||
* run. A corpus where every import hits tier 1 exits after one pass and
|
||||
* scores a per-import scan far closer to linear.
|
||||
* 4. **The hashed record includes the FILE SET, not just the query and the
|
||||
* result.** Otherwise a corpus edit that swaps the workspace under a case
|
||||
* while leaving its result string alone is invisible: dropping the
|
||||
* competing file from the "exact beats an earlier suffix" case, or
|
||||
* emptying the repeated-directory negative case, each leaves `cases`,
|
||||
* `non_null` and the fingerprint byte-identical and the gate green.
|
||||
* 5. **Path depth and package size are spanned, not pinned.** Both loops this
|
||||
* change added are driven by depth — one `suffixByStem` entry per '/' in a
|
||||
* stem, one `dirChildren` pass per component of `dir` — and the fan-out
|
||||
* tier returns a bucket whose length is the package size. While the corpus
|
||||
* capped depth at 8 components and packages at 16 files, three plausible
|
||||
* follow-up guards (cap suffix depth at 7, skip the `dirChildren` suffix
|
||||
* loop above depth 8, cap a bucket at 17) all passed `--check` with a
|
||||
* byte-identical fingerprint — while on a standard Gradle layout the depth
|
||||
* skip resolved EVERY package import to null and the bucket cap truncated
|
||||
* fan-out by 58%. Import ARITY, by contrast, was never blind: a tier-4 cap
|
||||
* at 4 dotted segments already failed the gate, because the branch matrix
|
||||
* carries 6- and 8-segment cases.
|
||||
*
|
||||
* Run:
|
||||
* node --import tsx bench/kotlin-import-target/measure.mjs # report
|
||||
* node --import tsx bench/kotlin-import-target/measure.mjs --check # CI gate
|
||||
*/
|
||||
import crypto from 'node:crypto';
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { resolveKotlinImportTarget } from '../../src/core/ingestion/languages/kotlin/import-target.ts';
|
||||
|
||||
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
const BASELINE_PATH = path.resolve(__dirname, 'baselines.json');
|
||||
|
||||
const SMALL = 400;
|
||||
const LARGE = 1600;
|
||||
/** Imports per file. Keeps the import count proportional to the file count, so
|
||||
* a per-import workspace scan shows up as a quadratic ratio rather than being
|
||||
* amortized away by a fixed import budget. Sized so the SMALL arm measures in
|
||||
* the tens of ms: at ~2 ms timer granularity and JIT warm-up, not scaling, set
|
||||
* the ratio — the same artifact bench/cpp-qualified-ns documents. */
|
||||
const IMPORTS_PER_FILE = 32;
|
||||
/** Depth arm: same file count either side, ~3x the path depth on one side. */
|
||||
const DEPTH_FILES = 800;
|
||||
const DEPTH_PAD = 16;
|
||||
const WARMUP = 2;
|
||||
const REPS = 7;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Correctness arm
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const lines = [];
|
||||
let nonNull = 0;
|
||||
|
||||
function resolve(files, targetRaw, fromFile) {
|
||||
return resolveKotlinImportTarget(
|
||||
{ kind: 'named', localName: 'X', importedName: 'X', targetRaw },
|
||||
{ fromFile, allFilePaths: new Set(files) },
|
||||
);
|
||||
}
|
||||
|
||||
/** Record one case in BOTH file-set iteration orders — see header property 1. */
|
||||
function record(files, targetRaw, fromFile = 'App.kt') {
|
||||
for (const [order, list] of [
|
||||
['fwd', files],
|
||||
['rev', [...files].reverse()],
|
||||
]) {
|
||||
const r = resolve(list, targetRaw, fromFile);
|
||||
if (r !== null) nonNull++;
|
||||
const rendered = r === null ? 'NULL' : Array.isArray(r) ? `[${r.join(',')}]` : r;
|
||||
// The FILE SET is part of the hashed record, not just the query and the
|
||||
// result — see header property 4.
|
||||
lines.push(`${order}\t${list.join('|')}\t${fromFile}\t${targetRaw}\t${rendered}`);
|
||||
}
|
||||
}
|
||||
|
||||
// ---- 1. Exhaustive branch matrix ------------------------------------------
|
||||
|
||||
// Tier 1, exact.
|
||||
record(['util/User.kt', 'util/Repo.kt'], 'util.User');
|
||||
// Tier 1, suffix (import is not workspace-rooted).
|
||||
record(['src/main/kotlin/util/User.kt'], 'util.User');
|
||||
// Exact anywhere beats a suffix found earlier.
|
||||
record(['deep/util/User.kt', 'util/User.kt'], 'util.User');
|
||||
// No exact match: first suffix in iteration order wins.
|
||||
record(['a/util/User.kt', 'b/util/User.kt'], 'util.User');
|
||||
// .kt / .kts sharing a stem.
|
||||
record(['dup/Thing.kt', 'dup/Thing.kts'], 'dup.Thing');
|
||||
// Multi-segment suffix query.
|
||||
record(['src/main/com/example/User.kt'], 'com.example.User');
|
||||
record(['a/b/com/example/User.kt', 'com/example/User.kt'], 'com.example.User');
|
||||
// Tier 2: stripped path matches a file (class-or-object holding the member).
|
||||
record(['util/OneArg.kt'], 'util.OneArg.writeAudit');
|
||||
record(['src/main/kotlin/util/OneArg.kt'], 'util.OneArg.writeAudit');
|
||||
// Tier 3: package fan-out to every direct child, in order.
|
||||
record(['models/User.kt', 'models/Repo.kt', 'models/sub/Deep.kt'], 'models.getRepo');
|
||||
record(['models/User.kt', 'models/sub/Deep.kt', 'models/Repo.kt'], 'models.getRepo');
|
||||
// Fan-out where the package directory is reached by suffix, not at the root.
|
||||
record(['app/src/main/kotlin/models/User.kt', 'app/src/main/kotlin/models/Repo.kt'], 'models.get');
|
||||
// Tier 4: progressive prefix strip, one and several skip levels.
|
||||
record(['x/y/z/Deep.kt'], 'com.example.z.Deep');
|
||||
record(['z/Deep.kt'], 'a.b.c.d.z.Deep');
|
||||
record(['q/Deep.kt'], 'a.b.c.d.e.f.q.Deep');
|
||||
// Tier 4 reaching the fan-out tier after stripping.
|
||||
record(['pkg/A.kt', 'pkg/B.kt'], 'com.example.pkg.someFunction');
|
||||
// Backslash normalization.
|
||||
record(['win\\pkg\\A.kt'], 'win.pkg.A');
|
||||
record(['win\\pkg\\A.kt', 'win\\pkg\\B.kt'], 'win.pkg.someFunction');
|
||||
// Non-Kotlin files never resolve.
|
||||
record(['pkg/A.java', 'pkg/A.md', 'pkg/A.kt.txt'], 'pkg.A');
|
||||
// Kotlin file alongside non-Kotlin noise of the same stem.
|
||||
record(['pkg/A.java', 'pkg/A.kt'], 'pkg.A');
|
||||
// Header property 2: repeated directory name — a child of the repeated package
|
||||
// since #2881, at the leading position here and mid-path below.
|
||||
record(['data/src/main/kotlin/com/example/data/Repo.kt'], 'data.something');
|
||||
record(['data/src/main/kotlin/com/example/data/Repo.kt'], 'data.Repo');
|
||||
record(['a/c/b/c/File.kt'], 'c.X');
|
||||
record(['c/b/c/File.kt'], 'c.X');
|
||||
// Doubly nested same-name directory, both below the root.
|
||||
record(['top/data/mid/data/Repo.kt'], 'data.something');
|
||||
// A path starting with the directory name is not its child unless direct.
|
||||
record(['data/sub/Repo.kt'], 'data.something');
|
||||
record(['data/Repo.kt'], 'data.something');
|
||||
// Repo-root file has no package directory.
|
||||
record(['Root.kt'], 'Root');
|
||||
record(['Root.kt', 'pkg/Root.kt'], 'Root');
|
||||
// Wildcard: `.*` is stripped and lands on the single-file tier, not fan-out.
|
||||
record(['models/User.kt', 'models/Repo.kt'], 'models.*');
|
||||
record(['models/Repo.kt', 'models/User.kt'], 'models.*');
|
||||
record(['util/User.kt'], 'util.User.*');
|
||||
// Unknown target.
|
||||
record(['pkg/A.kt'], 'nowhere.Thing');
|
||||
// Single-segment target with no directory anywhere.
|
||||
record(['pkg/A.kt'], 'A');
|
||||
// Empty-ish and degenerate targets.
|
||||
record(['pkg/A.kt'], '*');
|
||||
record(['pkg/A.kt'], 'pkg.');
|
||||
// fromFile variation must not change the outcome (this resolver ignores it) —
|
||||
// pinned so a future change that starts consulting it is visible here.
|
||||
record(['util/User.kt'], 'util.User', 'deep/nested/Caller.kt');
|
||||
|
||||
// ---- 1b. Depth and package size, the two axes the loops scale on ----------
|
||||
//
|
||||
// Header property 5. The index writes one `suffixByStem` entry per '/' in a
|
||||
// stem and walks `dir` once per component, so DEPTH is what those two loops
|
||||
// cost, and `dirChildren` bucket length is what the fan-out tier returns. A
|
||||
// corpus that pins either as a constant cannot see a guard on it: capping
|
||||
// suffix-key depth at 7, skipping the `dirChildren` suffix loop above depth 8,
|
||||
// or capping a bucket at 17 entries all left the fingerprint, `cases` and
|
||||
// `non_null` byte-identical before these cases existed — while, on a standard
|
||||
// Gradle layout, the depth skip resolved EVERY package import to null and the
|
||||
// bucket cap silently truncated fan-out by 58%.
|
||||
const DEEP = 'core/data/src/main/kotlin/com/example/core/data/repository';
|
||||
// 11 components — ordinary for Android/Gradle source, which runs 9-12.
|
||||
record([`${DEEP}/UserRepository.kt`], 'com.example.core.data.repository.UserRepository');
|
||||
record([`${DEEP}/UserRepository.kt`], 'repository.UserRepository');
|
||||
record([`${DEEP}/UserRepository.kt`], 'core.data.repository.UserRepository');
|
||||
record([`${DEEP}/UserRepository.kt`, `${DEEP}/PostRepository.kt`], 'repository.findAll');
|
||||
record([`${DEEP}/UserRepository.kt`, `${DEEP}/PostRepository.kt`], 'core.data.repository.findAll');
|
||||
// Deeper still, and with the repeated-name shape at depth.
|
||||
const DEEPER = 'feature/home/src/main/kotlin/com/example/feature/home/data/local/dao';
|
||||
record([`${DEEPER}/UserDao.kt`], 'dao.UserDao');
|
||||
record([`${DEEPER}/UserDao.kt`, `${DEEPER}/PostDao.kt`], 'dao.insertAll');
|
||||
record([`${DEEPER}/UserDao.kt`], 'home.data.local.dao.UserDao');
|
||||
// Suffix keys deeper than 7 components. Depth in the FILE is not enough on its
|
||||
// own: a cap on how many component-suffixes a stem contributes stays invisible
|
||||
// unless something QUERIES one of the deep keys, and every Gradle-shaped import
|
||||
// above is 6 segments or fewer. These reach the top of the stem.
|
||||
record(
|
||||
[`${DEEP}/UserRepository.kt`],
|
||||
'src.main.kotlin.com.example.core.data.repository.UserRepository',
|
||||
);
|
||||
record(
|
||||
[`${DEEP}/UserRepository.kt`],
|
||||
'data.src.main.kotlin.com.example.core.data.repository.UserRepository',
|
||||
);
|
||||
record([`${DEEPER}/UserDao.kt`], 'src.main.kotlin.com.example.feature.home.data.local.dao.UserDao');
|
||||
record(
|
||||
[`${DEEPER}/UserDao.kt`],
|
||||
'home.src.main.kotlin.com.example.feature.home.data.local.dao.UserDao',
|
||||
);
|
||||
record(
|
||||
[`${DEEP}/UserRepository.kt`, `${DEEP}/PostRepository.kt`],
|
||||
'src.main.kotlin.com.example.core.data.repository.findAll',
|
||||
);
|
||||
|
||||
// A package larger than any plausible bucket cap. 40 files in one package is
|
||||
// ordinary; a silent sibling cap is exactly what #2732 shipped on the JVM side.
|
||||
const BIG_PACKAGE = Array.from({ length: 40 }, (_, i) => `${DEEP}/Item${i}.kt`);
|
||||
record(BIG_PACKAGE, 'repository.someTopLevelFun');
|
||||
record(BIG_PACKAGE, 'com.example.core.data.repository.someTopLevelFun');
|
||||
record([...BIG_PACKAGE, `${DEEP}/sub/Nested.kt`], 'repository.someTopLevelFun');
|
||||
|
||||
// ---- 2. Deterministic fuzz -------------------------------------------------
|
||||
|
||||
/** xorshift32 — seeded, so the corpus is identical on every machine. */
|
||||
let seed = 0x9e3779b9;
|
||||
function rnd() {
|
||||
seed ^= seed << 13;
|
||||
seed ^= seed >>> 17;
|
||||
seed ^= seed << 5;
|
||||
seed >>>= 0;
|
||||
return seed / 0x100000000;
|
||||
}
|
||||
function pick(arr) {
|
||||
return arr[Math.floor(rnd() * arr.length)];
|
||||
}
|
||||
|
||||
const DIRS = [
|
||||
'',
|
||||
'app',
|
||||
'core',
|
||||
'data',
|
||||
'feature/home',
|
||||
'lib/data',
|
||||
'src/main/kotlin',
|
||||
'src/main/kotlin/com/example',
|
||||
'module/src/main/kotlin/com/example/data',
|
||||
'data/src/main/kotlin/com/example/data',
|
||||
'top/data/mid/data',
|
||||
'win\\pkg',
|
||||
// Depth beyond the Gradle norm, so the fuzz spans the axis too rather than
|
||||
// leaving it to the hand-written cases above.
|
||||
'core/data/src/main/kotlin/com/example/core/data/repository',
|
||||
'feature/home/src/main/kotlin/com/example/feature/home/data/local/dao',
|
||||
'a/b/c/d/e/f/g/h/i/j/k/l',
|
||||
];
|
||||
// Segment alphabet overlaps the DIRS entries on purpose: a random dotted target
|
||||
// only exercises a deep suffix key if its segments can actually align with a
|
||||
// deep path.
|
||||
const SEGS = [
|
||||
'User',
|
||||
'Repo',
|
||||
'Util',
|
||||
'Service',
|
||||
'Model',
|
||||
'data',
|
||||
'core',
|
||||
'api',
|
||||
'store',
|
||||
'sub',
|
||||
'src',
|
||||
'main',
|
||||
'kotlin',
|
||||
'com',
|
||||
'example',
|
||||
'repository',
|
||||
'dao',
|
||||
];
|
||||
const EXTS = ['.kt', '.kt', '.kt', '.kts', '.java', '.md'];
|
||||
|
||||
function randPath() {
|
||||
const dir = pick(DIRS);
|
||||
const base = pick(SEGS);
|
||||
const file = `${base}${pick(EXTS)}`;
|
||||
if (dir === '') return file;
|
||||
return dir.includes('\\') ? `${dir}\\${file}` : `${dir}/${file}`;
|
||||
}
|
||||
function randDotted() {
|
||||
// Up to 9 segments, not 4: import arity is the one axis the branch matrix
|
||||
// already spanned, but the fuzz should cover it too now that the corpus
|
||||
// carries paths deep enough for a long target to align with one.
|
||||
const n = 1 + Math.floor(rnd() * 9);
|
||||
const parts = [];
|
||||
for (let i = 0; i < n; i++) parts.push(pick(SEGS));
|
||||
return rnd() < 0.12 ? `${parts.join('.')}.*` : parts.join('.');
|
||||
}
|
||||
|
||||
// File counts run to 45, not 16: a package that never exceeds 16 direct
|
||||
// children cannot distinguish an uncapped `dirChildren` bucket from one capped
|
||||
// at 17 (header property 5).
|
||||
for (let repo = 0; repo < 400; repo++) {
|
||||
const fileCount = 3 + Math.floor(rnd() * 43);
|
||||
const files = [];
|
||||
for (let i = 0; i < fileCount; i++) files.push(randPath());
|
||||
const fromFile = randPath();
|
||||
for (let imp = 0; imp < 25; imp++) record(files, randDotted(), fromFile);
|
||||
}
|
||||
|
||||
const correctnessFingerprint = crypto
|
||||
.createHash('sha256')
|
||||
.update([...lines].sort().join('\n'))
|
||||
.digest('hex');
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Scaling arm
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** A synthetic Kotlin monorepo: Gradle-module roots over a shared package
|
||||
* namespace, at the path depth real Kotlin source has (the index stores one
|
||||
* suffix entry per '/' in a stem and walks `dir` once per component, so depth
|
||||
* is a cost driver and a flat corpus would understate the build).
|
||||
*
|
||||
* `padDepth` inserts filler segments so the depth arm below can hold the file
|
||||
* count fixed and vary only depth — the scaling ratio is scale-invariant in
|
||||
* FILE COUNT and would otherwise never see a depth-driven cost regression. */
|
||||
function buildCorpus(fileCount, padDepth = 0) {
|
||||
const pad = Array.from({ length: padDepth }, (_, d) => `p${d}`).join('/');
|
||||
const files = [];
|
||||
for (let i = 0; i < fileCount; i++) {
|
||||
const mod = i % 16;
|
||||
const root = pad === '' ? `lib${mod}` : `lib${mod}/${pad}`;
|
||||
files.push(`${root}/src/main/kotlin/com/example/mod${mod}/Class${i}.kt`);
|
||||
}
|
||||
return files;
|
||||
}
|
||||
|
||||
/** Import targets for the corpus, ~40% of them unresolvable — see header
|
||||
* property 3: only a miss drives all four tiers, which is where the
|
||||
* per-import scan was worst. */
|
||||
function buildImports(fileCount) {
|
||||
const imports = [];
|
||||
for (let i = 0; i < fileCount * IMPORTS_PER_FILE; i++) {
|
||||
const kind = i % 5;
|
||||
const mod = i % 16;
|
||||
if (kind === 0)
|
||||
imports.push(`com.example.mod${mod}.Class${i % fileCount}`); // tier 1 hit
|
||||
else if (kind === 1)
|
||||
imports.push(`com.example.mod${mod}.someFunction`); // fan-out
|
||||
else if (kind === 2)
|
||||
imports.push(`mod${mod}.Class${i % fileCount}`); // suffix
|
||||
else imports.push(`org.absent.pkg${mod}.Missing${i}`); // full cascade, no hit
|
||||
}
|
||||
return imports;
|
||||
}
|
||||
|
||||
function fastest(values) {
|
||||
return Math.min(...values);
|
||||
}
|
||||
|
||||
/**
|
||||
* Time one full pass: the index build PLUS resolving every import. The build is
|
||||
* the work the per-import scan was traded for, so hiding it would let an index
|
||||
* that is itself quadratic pass. Each pass gets its own Set object, because the
|
||||
* index is memoized on Set identity and a shared Set would build once and make
|
||||
* every later pass free. The Sets are constructed OUTSIDE the timer so their
|
||||
* own O(files) cost never lands in the measurement.
|
||||
*/
|
||||
function timeResolution(files, imports) {
|
||||
const sets = [];
|
||||
for (let i = 0; i < WARMUP + REPS; i++) sets.push(new Set(files));
|
||||
const fromFile = files[0];
|
||||
|
||||
for (let w = 0; w < WARMUP; w++) {
|
||||
for (const t of imports) {
|
||||
resolveKotlinImportTarget(
|
||||
{ kind: 'named', localName: 'X', importedName: 'X', targetRaw: t },
|
||||
{ fromFile, allFilePaths: sets[w] },
|
||||
);
|
||||
}
|
||||
}
|
||||
const samples = [];
|
||||
for (let r = 0; r < REPS; r++) {
|
||||
const set = sets[WARMUP + r];
|
||||
const t0 = performance.now();
|
||||
for (const t of imports) {
|
||||
resolveKotlinImportTarget(
|
||||
{ kind: 'named', localName: 'X', importedName: 'X', targetRaw: t },
|
||||
{ fromFile, allFilePaths: set },
|
||||
);
|
||||
}
|
||||
samples.push(performance.now() - t0);
|
||||
}
|
||||
return fastest(samples);
|
||||
}
|
||||
|
||||
const scales = {};
|
||||
for (const [name, fileCount] of [
|
||||
['small', SMALL],
|
||||
['large', LARGE],
|
||||
]) {
|
||||
const files = buildCorpus(fileCount);
|
||||
const imports = buildImports(fileCount);
|
||||
scales[name] = {
|
||||
files: fileCount,
|
||||
imports: imports.length,
|
||||
ms: Number(timeResolution(files, imports).toFixed(3)),
|
||||
};
|
||||
}
|
||||
|
||||
const scalingRatio = scales.large.ms / scales.small.ms / (LARGE / SMALL);
|
||||
|
||||
// Depth arm: file count fixed, depth roughly tripled. `scaling_ratio` divides
|
||||
// out the file count, so it is scale-INVARIANT and structurally cannot see a
|
||||
// cost that grows with path depth instead — and both loops this PR added are
|
||||
// depth loops. Same corpus size, same imports, only the paths get longer.
|
||||
const depthFiles = buildCorpus(DEPTH_FILES, 0);
|
||||
const depthFilesPadded = buildCorpus(DEPTH_FILES, DEPTH_PAD);
|
||||
const depthImports = buildImports(DEPTH_FILES);
|
||||
const shallowMs = timeResolution(depthFiles, depthImports);
|
||||
const deepMs = timeResolution(depthFilesPadded, depthImports);
|
||||
const depthRatio = deepMs / shallowMs;
|
||||
|
||||
const report = {
|
||||
small: scales.small,
|
||||
large: scales.large,
|
||||
scaling_ratio: Number(scalingRatio.toFixed(3)),
|
||||
depth: {
|
||||
files: DEPTH_FILES,
|
||||
shallow_components: 8,
|
||||
deep_components: 8 + DEPTH_PAD,
|
||||
shallow_ms: Number(shallowMs.toFixed(3)),
|
||||
deep_ms: Number(deepMs.toFixed(3)),
|
||||
},
|
||||
depth_ratio: Number(depthRatio.toFixed(3)),
|
||||
cases: lines.length,
|
||||
non_null: nonNull,
|
||||
fingerprint: correctnessFingerprint,
|
||||
};
|
||||
|
||||
if (!process.argv.includes('--check')) {
|
||||
console.log(JSON.stringify(report, null, 2));
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const baseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||||
const failures = [];
|
||||
if (report.fingerprint !== baseline.fingerprint) {
|
||||
failures.push(
|
||||
`fingerprint drift: ${report.fingerprint} != ${baseline.fingerprint} — Kotlin import ` +
|
||||
`resolution returned a DIFFERENT file set. That is a behaviour change, not a perf one: ` +
|
||||
`IMPORTS/CALLS edges move in every Kotlin repository. Explain it, never re-baseline to ` +
|
||||
`make CI green.`,
|
||||
);
|
||||
}
|
||||
for (const field of ['cases', 'non_null']) {
|
||||
if (report[field] !== baseline[field]) {
|
||||
failures.push(
|
||||
`${field} ${report[field]} != ${baseline[field]} — the corpus itself changed, so the ` +
|
||||
`fingerprint above is computed over a different surface and proves nothing about the ` +
|
||||
`resolver. Re-baseline every corpus field together, deliberately.`,
|
||||
);
|
||||
}
|
||||
}
|
||||
if (report.scaling_ratio > baseline.scaling_budget) {
|
||||
failures.push(
|
||||
`scaling ${report.scaling_ratio} > budget ${baseline.scaling_budget} — per-import cost grows ` +
|
||||
`with workspace size again, i.e. a tier went back to walking allFilePaths. Timing arm: ` +
|
||||
`re-run on an idle machine before investigating (see _scaling_note in baselines.json); the ` +
|
||||
`fingerprint arm is deterministic and never warrants a re-run.`,
|
||||
);
|
||||
}
|
||||
if (report.depth_ratio > baseline.depth_budget) {
|
||||
failures.push(
|
||||
`depth ratio ${report.depth_ratio} > budget ${baseline.depth_budget} — cost now grows with ` +
|
||||
`PATH DEPTH at a fixed file count. scaling_ratio divides the file count out and cannot ` +
|
||||
`see this. Timing arm: re-run on an idle machine first.`,
|
||||
);
|
||||
}
|
||||
if (report.small.ms > baseline.small_ms_ceiling) {
|
||||
failures.push(
|
||||
`small arm ${report.small.ms} ms > ceiling ${baseline.small_ms_ceiling} ms — scaling_ratio is ` +
|
||||
`a RATIO, so a constant-factor regression that grows both arms equally passes it (a full ` +
|
||||
`scan reintroduced on 1-in-32 imports measured 1.490, inside the budget, while running ` +
|
||||
`2.8x slower). This ceiling is what catches that. Timing arm: re-run on an idle machine.`,
|
||||
);
|
||||
}
|
||||
|
||||
console.log(JSON.stringify(report, null, 2));
|
||||
if (failures.length > 0) {
|
||||
console.error(`[kotlin-import-target --check] FAIL\n - ${failures.join('\n - ')}`);
|
||||
process.exit(1);
|
||||
}
|
||||
console.log('[kotlin-import-target --check] PASS');
|
||||
|
|
@ -1 +1 @@
|
|||
a0da3e7c00f603e4bdad91a376b3fc181577a73c2ca1719ab7449d3463c671e0
|
||||
2600a1f6f8a042eb4f520a7870c34d9ca292765824537c3bc861b40dac8769a8
|
||||
|
|
|
|||
|
|
@ -364,8 +364,13 @@ also resolves, so PHP nullable field types already work.
|
|||
|
||||
**C++ — the base already resolves, but `this->` field receivers do not.**
|
||||
`pointerArrowChain` and `valueDotChain` both RESOLVE, so a decorated C++ base is
|
||||
not a gap. But `this->repo.save()` and `this->repo->save()` are both
|
||||
INVISIBLE-GAP — a distinct defect, not a decoration one.
|
||||
not a gap. `this->repo.save()` and `this->repo->save()` were both INVISIBLE-GAP
|
||||
when this was written — a distinct defect, not a decoration one — and #2833
|
||||
closed it: a language that declares `this` IS the enclosing class
|
||||
(`resolveThisViaEnclosingClass`) synthesizes no `this` typeBinding anywhere, so
|
||||
a chain whose BASE is `this` could never seed its head. It was never a generics
|
||||
gap; the NON-generic control failed identically. C++'s `fieldReceiverCall` and
|
||||
`decoratedFieldType` cells moved INVISIBLE-GAP -> RESOLVES with it.
|
||||
|
||||
**Rust — the decorated receiver is NOT a gap.** `&mut self` resolves, so Go is
|
||||
the only language whose method receiver decoration defeats the lookup. Rust's
|
||||
|
|
|
|||
|
|
@ -61,9 +61,9 @@
|
|||
"awaitParen": "N/A",
|
||||
"explicitTypeArgs": "VISIBLE-GAP",
|
||||
"indexElement": "RESOLVES",
|
||||
"fieldReceiverCall": "INVISIBLE-GAP",
|
||||
"fieldReceiverCall": "RESOLVES",
|
||||
"decoratedReceiverBase": "N/A",
|
||||
"decoratedFieldType": "INVISIBLE-GAP"
|
||||
"decoratedFieldType": "RESOLVES"
|
||||
},
|
||||
"go": {
|
||||
"plainChain": "RESOLVES",
|
||||
|
|
@ -200,10 +200,11 @@
|
|||
},
|
||||
"countArm": {
|
||||
"callDrops": 102,
|
||||
"totalDropsAllKinds": 129,
|
||||
"totalDropsAllKinds": 148,
|
||||
"bySiteKind": {
|
||||
"call": 102,
|
||||
"read": 27
|
||||
"read": 27,
|
||||
"write": 19
|
||||
},
|
||||
"callDropsByExtension": {
|
||||
".java": 49,
|
||||
|
|
|
|||
|
|
@ -13,9 +13,10 @@ node --import tsx bench/schema-pairs/measure.mjs --check # gate vs baselines.
|
|||
|
||||
`src/core/lbug/schema.ts` generates its relation pairs from two cross products,
|
||||
and declines to add a third one **on the strength of a number** — roughly 1.04×
|
||||
at 450 declared pairs, 1.6× at 786, 2.1× at 1024. That measurement used to live
|
||||
in a scratch directory, so nobody proposing a third rule could re-run it. This
|
||||
harness is that measurement, committed — and it reproduces those figures.
|
||||
near production's pair count, 1.6× at 786, 2.1× at 1024. That measurement used
|
||||
to live in a scratch directory, so nobody proposing a third rule could re-run
|
||||
it. This harness is that measurement, committed — and it reproduces those
|
||||
figures.
|
||||
|
||||
Run it before widening a rule, and quote the new ratio in the review.
|
||||
|
||||
|
|
@ -29,8 +30,27 @@ Observed on the reference box, **four runs** (ratios vs the 332-pair list):
|
|||
| 786 | 1.52–1.75× | 1.19–1.31× |
|
||||
| 1024 | 2.03–2.34× | 1.31–1.57× |
|
||||
|
||||
Production's 450 came out _faster_ than 332 on three of the four runs, so at this
|
||||
size the pair count is inside run-to-run noise. Everything past ~640 is not.
|
||||
Production's former 450-pair surface came out _faster_ than 332 on three of the
|
||||
four runs, so at this size the pair count is inside run-to-run noise. Everything
|
||||
past ~640 is not.
|
||||
|
||||
#2801 remeasured the new 461-pair production surface on Windows six times:
|
||||
|
||||
| run | untyped ratio | typed ratio | interpretation |
|
||||
| --- | ------------- | ----------- | ----------------------------------- |
|
||||
| 1 | 1.101× | 1.157× | noise-dominated (`typed > untyped`) |
|
||||
| 2 | 1.324× | 1.122× | below the operational budget |
|
||||
| 3 | 2.705× | 1.089× | exceeds the operational budget |
|
||||
| 4 | 1.417× | 1.065× | below the operational budget |
|
||||
| 5 | 1.196× | 1.050× | below the operational budget |
|
||||
| 6 | 1.585× | 2.577× | noise-dominated (`typed > untyped`) |
|
||||
|
||||
The three comparable Windows runs below the 1.5× operational ceiling span
|
||||
**1.20–1.42× untyped / 1.05–1.12× typed**. Run 3 is published rather than
|
||||
silently discarded: no pre-registered rule excludes it, and `--check` would
|
||||
correctly reject it. These Windows measurements are not combined with the
|
||||
historical reference-box rows to infer cross-size ordering.
|
||||
|
||||
**Quote the range, not a single run** — one run is not evidence here.
|
||||
|
||||
## What it measures
|
||||
|
|
@ -52,8 +72,8 @@ data**, then times two query shapes over 40 anchors × 15 reps (median):
|
|||
control; the real cost of widening sits between it and `ratio_*`. A run where
|
||||
`typed_ratio` moves _more_ than `ratio` is noise-dominated and should be
|
||||
rerun.
|
||||
- **`ratio_<size>`** — `untyped_ms_<size> / untyped_ms_332`. `ratio_450` is the
|
||||
figure `schema.ts` quotes.
|
||||
- **`ratio_<size>`** — `untyped_ms_<size> / untyped_ms_332`. The
|
||||
production-size ratio is the figure `schema.ts` quotes.
|
||||
|
||||
### Sizes
|
||||
|
||||
|
|
@ -66,7 +86,8 @@ harness fails if the row counts ever differ across sizes.
|
|||
| size | what it is |
|
||||
| ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| 332 | the pre-#2792 hand-written list — the reference for every ratio |
|
||||
| 450 | production today (two cross products + 72 hand-declared pairs) |
|
||||
| 450 | production before Record became linkable (#2801) |
|
||||
| 461 | production today (two cross products + 69 hand-declared pairs) |
|
||||
| 641 | the third cross product `schema.ts` defers (`DEFINITION_ANCHOR_LABELS × {CodeElement, Section, Typedef, Union, Namespace, Impl, TypeAlias, Static, Template}`), which would leave ~29 hand-declared lines |
|
||||
| 786 | the size an earlier revision of that comment attributed to the third rule — it is 641; kept as a measured waypoint |
|
||||
| 1024 | the full cross product, the ceiling |
|
||||
|
|
@ -77,11 +98,13 @@ Before timing anything, the harness round-trips the **real** `SCHEMA_QUERIES`
|
|||
through a real database and asserts that `CALL SHOW_CONNECTION('CodeRelation')`
|
||||
reports exactly the pairs `parseRelationSchemaPairs` finds in `RELATION_SCHEMA`.
|
||||
|
||||
No magic number is baked in: the invariant is that the DDL LadybugDB _accepted_
|
||||
carries the pair set our own parser believes it declares. The absolute count is
|
||||
reported as `declared_pairs`. A pair declared twice would not reach this check at
|
||||
all — LadybugDB rejects the `CREATE REL TABLE` outright, which is why a duplicate
|
||||
kills every `analyze` rather than one repository's.
|
||||
No production-size magic number is baked in: the measured production size and
|
||||
budget key are derived from that parsed DDL count. A missing `ratio_<size>_budget`
|
||||
entry makes `--check` fail closed. The invariant is that the DDL LadybugDB
|
||||
_accepted_ carries the pair set our own parser believes it declares. The
|
||||
absolute count is reported as `declared_pairs`. A pair declared twice would not
|
||||
reach this check at all — LadybugDB rejects the `CREATE REL TABLE` outright,
|
||||
which is why a duplicate kills every `analyze` rather than one repository's.
|
||||
|
||||
## What it does NOT measure
|
||||
|
||||
|
|
@ -93,8 +116,11 @@ kills every `analyze` rather than one repository's.
|
|||
|
||||
## Regenerating the baseline
|
||||
|
||||
`baselines.json` holds one budget, `ratio_450_budget` — the ceiling on what
|
||||
`baselines.json` holds one production-size budget — the ceiling on what
|
||||
production's own pair count may cost relative to the 332-pair hand-list it
|
||||
replaced. Re-run without `--check` **several times** and copy the top of the
|
||||
observed `ratio_450` range plus headroom — the spread between runs on this box
|
||||
is wider than the effect being measured at 450, so a single run cannot set it.
|
||||
observed production-size ratio range plus headroom — the spread between runs on
|
||||
this box is wider than the effect being measured near production, so a single
|
||||
run cannot set it. Publish the raw ratios and apply only the pre-registered
|
||||
`typed_ratio > ratio` noise rule; do not silently discard another run to make a
|
||||
budget pass.
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
{
|
||||
"_comment": "ratio_450_budget — ceiling on what production's 450-pair set may cost on untyped-endpoint anchored queries, relative to the 332-pair hand-list it replaced. Observed 0.94x and 1.05x across two runs on the reference box (i.e. inside run-to-run noise; it came out faster than 332 once). The budget carries headroom for that spread — compare typed_ratio_450 (1.10-1.17x) for this box's floor. Raise it only with a measured range, never a single run.",
|
||||
"ratio_450_budget": 1.3
|
||||
"_comment": "ratio_461_budget — operational ceiling on what production's 461-pair set may cost on untyped-endpoint anchored queries, relative to the 332-pair hand-list it replaced. #2801 measured 1.20-1.42x across three comparable Windows runs (typed floor 1.05-1.12x), so 1.5x leaves explicit host headroom. All six raw runs are published in README.md; two meet the pre-registered typed_ratio > ratio noise rule, while one additional 2.705x run is not silently discarded and would fail this gate. Raise the budget only with a published measured range, never a single run.",
|
||||
"ratio_461_budget": 1.5
|
||||
}
|
||||
|
|
|
|||
|
|
@ -3,10 +3,10 @@
|
|||
*
|
||||
* `src/core/lbug/schema.ts` declares its relation pairs from two cross products
|
||||
* plus a small hand-written remainder, and it justifies NOT adding a third cross
|
||||
* product with a number: anchored queries cost ~1.04× at 450 declared pairs but
|
||||
* 1.6× at 786 and 2.1× at 1024. That measurement previously lived in a scratch
|
||||
* directory, so the claim could not be re-checked when someone proposed
|
||||
* widening a rule. This is it, committed.
|
||||
* product with a number: anchored queries cost ~1.04× near production's pair
|
||||
* count but 1.6× at 786 and 2.1× at 1024. That measurement previously lived in
|
||||
* a scratch directory, so the claim could not be re-checked when someone
|
||||
* proposed widening a rule. This is it, committed.
|
||||
*
|
||||
* WHAT IT MEASURES. Against a real `@ladybugdb/core` database, with byte-identical
|
||||
* DATA at every size, it times the query shape whose plan actually depends on the
|
||||
|
|
@ -34,7 +34,8 @@
|
|||
* same query at every size, and the only variable is how many UNUSED pairs the
|
||||
* table declares:
|
||||
* - 332 — the pre-#2792 hand-written list (the historical baseline);
|
||||
* - 450 — production today (two cross products + 72 hand-declared);
|
||||
* - 450 — production before Record became linkable (#2801);
|
||||
* - 461 — production today (two cross products + 69 hand-declared);
|
||||
* - 641 — the third cross product schema.ts defers
|
||||
* (`DEFINITION_ANCHOR_LABELS × {CodeElement, Section, Typedef, Union,
|
||||
* Namespace, Impl, TypeAlias, Static, Template}`), which would leave
|
||||
|
|
@ -43,8 +44,8 @@
|
|||
* third rule (it is 641; 786 is kept as a measured waypoint);
|
||||
* - 1024 — the full cross product, the ceiling.
|
||||
*
|
||||
* Ratios are reported against 332, the smallest size — `ratio_450` is the
|
||||
* number schema.ts quotes.
|
||||
* Ratios are reported against 332, the smallest size — the production-size
|
||||
* ratio is the number schema.ts quotes.
|
||||
*
|
||||
* CORRECTNESS GATE. Before timing anything it round-trips the REAL
|
||||
* `SCHEMA_QUERIES` through a real database and asserts that
|
||||
|
|
@ -61,9 +62,9 @@
|
|||
* node --import tsx bench/schema-pairs/measure.mjs # print JSON lines
|
||||
* node --import tsx bench/schema-pairs/measure.mjs --check # gate vs baselines.json
|
||||
*
|
||||
* `--check` fails if the correctness gate breaks, or if `ratio_450` exceeds its
|
||||
* budget — i.e. if production's own pair count starts costing materially more
|
||||
* than the hand-written list it replaced.
|
||||
* `--check` fails if the correctness gate breaks, or if the production-size
|
||||
* ratio exceeds its budget — i.e. if production's own pair count starts
|
||||
* costing materially more than the hand-written list it replaced.
|
||||
*/
|
||||
import fs from 'node:fs';
|
||||
import os from 'node:os';
|
||||
|
|
@ -85,9 +86,14 @@ const lbug = (await import('@ladybugdb/core')).default;
|
|||
|
||||
// ---- sizes + the pair enumeration every size is a prefix of ----
|
||||
|
||||
const SIZES = [332, 450, 641, 786, 1024];
|
||||
const REFERENCE_SIZE = 332; // ratios are relative to this
|
||||
const PRODUCTION_SIZE = 450; // the size schema.ts ships
|
||||
// Derive production from the same executable DDL the correctness gate
|
||||
// round-trips. A LINKABLE_LABELS widening must not require a second copied
|
||||
// count here — and cannot silently select a stale/missing budget key.
|
||||
const PRODUCTION_SIZE = parseRelationSchemaPairs(RELATION_SCHEMA).size;
|
||||
const SIZES = [...new Set([REFERENCE_SIZE, 450, PRODUCTION_SIZE, 641, 786, 1024])].sort(
|
||||
(a, b) => a - b,
|
||||
);
|
||||
|
||||
// The four pairs the synthetic data uses. Pinned to the FRONT of the
|
||||
// enumeration so they are declared at every size — otherwise a smaller pair set
|
||||
|
|
@ -338,8 +344,13 @@ if (!CHECK) {
|
|||
process.stdout.write(JSON.stringify(summary) + '\n');
|
||||
} else {
|
||||
const baselines = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf8'));
|
||||
const budget = baselines[`ratio_${PRODUCTION_SIZE}_budget`];
|
||||
if (budget !== undefined && summary[`ratio_${PRODUCTION_SIZE}`] >= budget) {
|
||||
const budgetKey = `ratio_${PRODUCTION_SIZE}_budget`;
|
||||
const budget = baselines[budgetKey];
|
||||
if (budget === undefined) {
|
||||
failures.push(`no ${budgetKey} in baselines.json — the production gate is disarmed`);
|
||||
} else if (typeof budget !== 'number' || !Number.isFinite(budget)) {
|
||||
failures.push(`${budgetKey} must be a finite number (got ${JSON.stringify(budget)})`);
|
||||
} else if (summary[`ratio_${PRODUCTION_SIZE}`] >= budget) {
|
||||
failures.push(
|
||||
`production pair set (${PRODUCTION_SIZE}) costs ${summary[`ratio_${PRODUCTION_SIZE}`]}× vs ` +
|
||||
`${REFERENCE_SIZE} pairs, >= budget ${budget} (untyped ${reference.untyped_ms}ms -> ` +
|
||||
|
|
|
|||
|
|
@ -1,27 +1,28 @@
|
|||
{
|
||||
"_comment": "Per-language baselines for bench/scope-capture/measure.mjs --check. fingerprint = order-independent sha256 over the lang-resolution/<lang>-* fixture corpus + a 20-entity synthetic source (correctness gate; re-baseline intentionally on a legitimate capture change). scaling_budget = max allowed (t800/t250)/(800/250); ~1.0 is linear, ~3.2 is quadratic. The synthetic source is now HERITAGE-BEARING for every language (each Entity extends/implements/embeds/uses-trait/conforms-to a shared base) so the #1951 @reference.inherits synth is gated at scale, not just the base capture loop. All languages thread the tree-sitter captured node instead of re-deriving it with findNodeAtRange(tree.rootNode,...) per match, so all are linear (go #1915, python #1918, ruby/php/rust/csharp #1951, java #1956).",
|
||||
"go": {
|
||||
"fingerprint": "e386598526e502d131e52a17d219635b3a4196d94f1ebdd25922a2582c985d18",
|
||||
"fingerprint": "9c554a9d698a2b79fb419852daadca87b8aae88180cceabf9c8d82f3e3300f2e",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior 3d4e32e7490c830516126e28931827949baa3594cb521f7a3d8dcfed95b6018a -> 57b3c55135af8d2af33b9a7c4bf89796a7bee5b5822b402a2dea91af7232cf4a; scaling 1.058 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: provider-owned callable assignment/copy/formal/argument/invoke facts with invocation/constructor-result suppression. Prior 09ecd94911b830f52fa8807560abcbd79f163d02a2072870c1a59297e9a326e1 -> 3d4e32e7490c830516126e28931827949baa3594cb521f7a3d8dcfed95b6018a; scaling 1.039 < 1.5.",
|
||||
"_rebaselined": "#1976: F33 generic composite literal constructor inference adds generic_type captures in composite_literal patterns; fingerprint drift expected.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 57b3c55135af8d2af33b9a7c4bf89796a7bee5b5822b402a2dea91af7232cf4a -> 5d6c59c2f2c0dd937c53bf5d736e0f8376b2899a381e488a33aec23524823efb.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 57b3c55135af8d2af33b9a7c4bf89796a7bee5b5822b402a2dea91af7232cf4a -> 5d6c59c2f2c0dd937c53bf5d736e0f8376b2899a381e488a33aec23524823efb.",
|
||||
"_rebaselined_2766_go_pointer_receiver_fixture": "#2766: added test/fixtures/lang-resolution/go-pointer-receiver-field-chain/ (2 Go files) as the committed regression fixture for pointer-receiver base resolution. Go fixture_count 100 -> 102. Prior 5d6c59c2f2c0dd937c53bf5d736e0f8376b2899a381e488a33aec23524823efb -> 8cba537ff211fab3bac5fb4456cd1ffba14d6a2db75c40acae28ab8bf29f3d2e. FIXTURE-CORPUS GROWTH, NOT A CAPTURE CHANGE: the accompanying fix is a resolution-time lookup fallback (stripTypePreservingDecoration) and cannot move capture output; go was the ONLY language whose fingerprint drifted, and every other language matched its baseline on the same run.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 8cba537ff211fab3bac5fb4456cd1ffba14d6a2db75c40acae28ab8bf29f3d2e -> 8162272bb897b0b89472c406321cf8d88a5ae4ea83ea9e3c45f8e817041bff9f.",
|
||||
"_rebaselined_2766_await_subscript_emission": "#2766: extractMixedChain now walks THROUGH await and subscript nodes and peels transparent wrappers at loop entry, so sites whose receiver is `repos[0]` or `(await f())` mint a receiver chain where they previously minted none. EMISSION CHANGE: more sites carry `@reference.receiver-chain`; no existing chain changed shape. Only go and kotlin drifted of 15 — the two whose fixture corpora contain such receivers. Prior 8162272bb897b0b89472c406321cf8d88a5ae4ea83ea9e3c45f8e817041bff9f -> c9c908f441e3be12fad2448120ed3ea35dc235a12b3f63b0ec532ffdae11d9e9.",
|
||||
"_rebaselined_2766_phantom_callee_read_site": "#2766: Go's `@reference.read` pattern matches EVERY selector_expression, so a member call `h.dep.Work()` minted THREE sites — the call, the genuine `h.dep` field read, and a PHANTOM read on the callee `h.dep.Work`. The phantom resolved through findOwnedMember (which prefers methods over fields) and emitted an ACCESSES edge to the METHOD duplicating the CALLS edge at the same position; visible today on any receiver the text cascade can type (`RunFromValueReceiver -> DoWork`). The emitter now drops a read match whose selector is in FUNCTION position. FEWER capture matches for Go, no other language affected — go was the only fingerprint of 15 that moved. A method VALUE (`f := h.dep.Work`) is not in function position and is untouched. Prior c9c908f441e3be12fad2448120ed3ea35dc235a12b3f63b0ec532ffdae11d9e9 -> 7bb524a32a2eed57a15b454e3a33480e92a496c683e6856ef02179693c0e02e3.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 8cba537ff211fab3bac5fb4456cd1ffba14d6a2db75c40acae28ab8bf29f3d2e -> 8162272bb897b0b89472c406321cf8d88a5ae4ea83ea9e3c45f8e817041bff9f.",
|
||||
"_rebaselined_2766_await_subscript_emission": "#2766: extractMixedChain now walks THROUGH await and subscript nodes and peels transparent wrappers at loop entry, so sites whose receiver is `repos[0]` or `(await f())` mint a receiver chain where they previously minted none. EMISSION CHANGE: more sites carry `@reference.receiver-chain`; no existing chain changed shape. Only go and kotlin drifted of 15 \u2014 the two whose fixture corpora contain such receivers. Prior 8162272bb897b0b89472c406321cf8d88a5ae4ea83ea9e3c45f8e817041bff9f -> c9c908f441e3be12fad2448120ed3ea35dc235a12b3f63b0ec532ffdae11d9e9.",
|
||||
"_rebaselined_2766_phantom_callee_read_site": "#2766: Go's `@reference.read` pattern matches EVERY selector_expression, so a member call `h.dep.Work()` minted THREE sites \u2014 the call, the genuine `h.dep` field read, and a PHANTOM read on the callee `h.dep.Work`. The phantom resolved through findOwnedMember (which prefers methods over fields) and emitted an ACCESSES edge to the METHOD duplicating the CALLS edge at the same position; visible today on any receiver the text cascade can type (`RunFromValueReceiver -> DoWork`). The emitter now drops a read match whose selector is in FUNCTION position. FEWER capture matches for Go, no other language affected \u2014 go was the only fingerprint of 15 that moved. A method VALUE (`f := h.dep.Work`) is not in function position and is untouched. Prior c9c908f441e3be12fad2448120ed3ea35dc235a12b3f63b0ec532ffdae11d9e9 -> 7bb524a32a2eed57a15b454e3a33480e92a496c683e6856ef02179693c0e02e3.",
|
||||
"_rebaselined_2766_callee_position_marker": "#2766 review fix: a call's callee selector is no longer DROPPED at capture. An earlier commit on this branch dropped it outright, which also deleted the genuine field read on a func-typed struct field (`h.dep.Work()` where `Work func() error`) - callback/hook/mock structs lost their only ACCESSES evidence. The match is now emitted carrying `@reference.callee-position`, and the phantom is suppressed at EMIT by the resolved target's kind instead. Go only: the other 14 languages' fingerprints are byte-identical, which is the check that this is not a cross-language capture change. Prior 7bb524a32a2eed57a15b454e3a33480e92a496c683e6856ef02179693c0e02e3 -> e47302079e17a5e73711bbed5416557b49327cb67e4932008700ec6b8fb468b3; scaling 1.001 < 1.5; fixtures 102 (unchanged), capture_groups_fp 2103.",
|
||||
"_rebaselined_2813_interface_field_dispatch_fixture": "#2813: added test/fixtures/lang-resolution/go-interface-field-dispatch/ (8 Go files) as the committed regression fixture for calls through an interface-typed struct field. Go fixture_count 102 -> 110. FIXTURE-CORPUS GROWTH, NOT A CAPTURE CHANGE: the accompanying fixes are a detection-time method-set change (interface-impls.ts) and a resolution-time fan-out in the shared receiver pass, neither of which emits captures; go/query.ts and go/captures.ts are untouched. Go was the ONLY language whose fingerprint drifted, and every other language matched its baseline on the same run - the same check used for the #2766 fixture growth above. Prior e47302079e17a5e73711bbed5416557b49327cb67e4932008700ec6b8fb468b3 -> cffee41cadbf350855d99bd5aee7c015b1e8b31d1c343d02f113540abe86c765; scaling 1.074 < 1.5, capture_groups_fp 2303.",
|
||||
"_rebaselined_2837": "#2837: Go struct/interface captures re-anchored from the type_declaration onto the type_spec (@scope.class/@declaration.struct/@declaration.interface in languages/go/query.ts, @definition.struct/@definition.interface in GO_QUERIES). A grouped `type (...)` block used to yield ONE scope and ONE node for every type in it, so each type after the first lost its field typeBindings and every field-receiver call in the file emitted nothing. Capture COUNT is unchanged; only ranges moved, plus the new go-grouped-type-decl fixture. Prior c27fb803598581fa4eb7ddf5ef6f8369b9e3a150082d11362e7aa3ec8faaa832 -> e386598526e502d131e52a17d219635b3a4196d94f1ebdd25922a2582c985d18; scaling 1.054 < 1.5."
|
||||
"_rebaselined_2837": "#2837: Go struct/interface captures re-anchored from the type_declaration onto the type_spec (@scope.class/@declaration.struct/@declaration.interface in languages/go/query.ts, @definition.struct/@definition.interface in GO_QUERIES). A grouped `type (...)` block used to yield ONE scope and ONE node for every type in it, so each type after the first lost its field typeBindings and every field-receiver call in the file emitted nothing. Capture COUNT is unchanged; only ranges moved, plus the new go-grouped-type-decl fixture. Prior c27fb803598581fa4eb7ddf5ef6f8369b9e3a150082d11362e7aa3ec8faaa832 -> e386598526e502d131e52a17d219635b3a4196d94f1ebdd25922a2582c985d18; scaling 1.054 < 1.5.",
|
||||
"_rebaselined_2873_undecided_satisfaction_fixtures": "#2873: added test/fixtures/lang-resolution/go-extern-qualified-signatures/ (5 Go files) and go-undecided-satisfaction/ (1 Go file) as the committed regression fixtures for out-of-repo package qualifiers in method signatures and for a satisfaction check that cannot be decided. Go fixture_count 116 -> 122. Prior e386598526e502d131e52a17d219635b3a4196d94f1ebdd25922a2582c985d18 -> 9c554a9d698a2b79fb419852daadca87b8aae88180cceabf9c8d82f3e3300f2e. FIXTURE-CORPUS GROWTH, NOT A CAPTURE CHANGE: the accompanying fix is resolution-time (signatureContextForFile recovers an identity for unresolvable imports) plus a tri-state verdict, neither of which runs during capture; go was the ONLY language whose fingerprint drifted and every other language matched its baseline on the same run."
|
||||
},
|
||||
"cobol": {
|
||||
"fingerprint": "c8c00b56a7da24e04080eb885714fbbf45e3903324f0cf9df0754f5b5a92e3aa",
|
||||
"_rebaselined_2813_exact_method_sets": "#2813: Go embedded fields now emit `@reference.embedded-pointer` when spelled `*T` rather than `T`. A CAPTURE-EMISSION CHANGE, not fixture growth: fixture_count is unchanged at 110 and capture_groups_fp moves 2303 -> 2339 (+36), which is the new marker plus the WrongSigRepo/Recount rows added to two existing fixture files. The marker is required for exactness — Go gives `struct{ Base }` and `struct{ *Base }` different method sets, so structural interface satisfaction cannot be correct without knowing which was written (go.dev/ref/spec#Struct_types). Go was the ONLY language of 15 whose fingerprint moved, which is the check that this is a Go capture change and not a cross-language regression. Accompanied by SCHEMA_BUMP 39 -> 43 (skipping 40/41/42, taken by origin/main during review) so a warm cache cannot replay the pre-marker capture set. Prior cffee41cadbf350855d99bd5aee7c015b1e8b31d1c343d02f113540abe86c765 -> c27fb803598581fa4eb7ddf5ef6f8369b9e3a150082d11362e7aa3ec8faaa832; scaling 0.987 < 1.5.",
|
||||
"_rebaselined_2813_exact_method_sets": "#2813: Go embedded fields now emit `@reference.embedded-pointer` when spelled `*T` rather than `T`. A CAPTURE-EMISSION CHANGE, not fixture growth: fixture_count is unchanged at 110 and capture_groups_fp moves 2303 -> 2339 (+36), which is the new marker plus the WrongSigRepo/Recount rows added to two existing fixture files. The marker is required for exactness \u2014 Go gives `struct{ Base }` and `struct{ *Base }` different method sets, so structural interface satisfaction cannot be correct without knowing which was written (go.dev/ref/spec#Struct_types). Go was the ONLY language of 15 whose fingerprint moved, which is the check that this is a Go capture change and not a cross-language regression. Accompanied by SCHEMA_BUMP 39 -> 43 (skipping 40/41/42, taken by origin/main during review) so a warm cache cannot replay the pre-marker capture set. Prior cffee41cadbf350855d99bd5aee7c015b1e8b31d1c343d02f113540abe86c765 -> c27fb803598581fa4eb7ddf5ef6f8369b9e3a150082d11362e7aa3ec8faaa832; scaling 0.987 < 1.5.",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: COBOL procedure-pointer callable flow facts; multi-topic extraction now consumes each grouped scope/declaration match once instead of requiring a duplicate declaration-only match. Prior 68ee0e95eb9f86f2d92ca35f730f4c2d4d83abc1b5241ae767ff3437780ec8d1 -> d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e; scaling 0.853 < 1.5.",
|
||||
"_note": "Updated for F17-F23 fixes (P2: TIMES guard, ADD GIVING, SQL AS alias). See PR #1959.",
|
||||
"_rebaselined_2793_declaratives": "PR #2793: corpus-only re-baseline. `cobol-declaratives` was added to test/fixtures/lang-resolution to reproduce the `Namespace→Record` analyze abort (DECLARATIVES / USE AFTER STANDARD ERROR ON <file>), and this bench globs `lang-resolution/cobol-*`, so the corpus grew 14 -> 15 files. Verified capture-neutral: with that one fixture moved aside the fingerprint is byte-identical to the prior d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e. No COBOL capture code changed in that PR. Scaling 0.677 < 1.5."
|
||||
"_rebaselined_2793_declaratives": "PR #2793: corpus-only re-baseline. `cobol-declaratives` was added to test/fixtures/lang-resolution to reproduce the `Namespace\u2192Record` analyze abort (DECLARATIVES / USE AFTER STANDARD ERROR ON <file>), and this bench globs `lang-resolution/cobol-*`, so the corpus grew 14 -> 15 files. Verified capture-neutral: with that one fixture moved aside the fingerprint is byte-identical to the prior d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e. No COBOL capture code changed in that PR. Scaling 0.677 < 1.5."
|
||||
},
|
||||
"c": {
|
||||
"fingerprint": "3418cded9f7072152f68992f0a426f43ae7d9d553579a47075fc0cab185848a5",
|
||||
|
|
@ -29,13 +30,15 @@
|
|||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior 57fee292147ae6d2db7062da1e07d17122cf355207c8967fa85fd2ec9ca398a4 -> 3418cded9f7072152f68992f0a426f43ae7d9d553579a47075fc0cab185848a5; scaling 1.073 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: C function-pointer signatures plus direct-callee argument metadata and invocation-result suppression. Prior 75bcdbbf006bf9bd263c0f5857461b118f39b164e9f821cb0651ad0ec46ef6ae -> 57fee292147ae6d2db7062da1e07d17122cf355207c8967fa85fd2ec9ca398a4; scaling 1.035 < 1.5.",
|
||||
"_rebaselined_callable_flow": "Callable-value-flow facts for C function pointers, copies, pointer-to-pointer cells, arguments, and indirect invokes. Prior 12a196b2d6249c8d86a931b12ecebc2a0cdf8d6f47683acdd0d8e9d8bc7657f5 -> 75bcdbbf006bf9bd263c0f5857461b118f39b164e9f821cb0651ad0ec46ef6ae; measured scaling ratio 0.980 < 1.5.",
|
||||
"_added": "#1956: c added to the scope-capture bench (was UNBENCHED). C has no inheritance — flat scale source. Adding it exposed + fixed a pre-existing O(n^2) findNodeAtRange root-walk in c/captures.ts (threaded c.node, byte-identical over c-* fixtures); scaling 3.475 -> 0.96.",
|
||||
"_note": "#1983: + c-static-linkage-worker fixture (caller.c/lib.c/lib.h/local.c — worker-path static-linkage side-channel test). Pure fixture-corpus drift: no c/captures.ts or query change branch-vs-main, existing fixtures' captures byte-identical (c-captures.test.ts 45/45), scaling stays linear (~0.97). The baseline was missed when the fixture landed; regenerated here. fingerprint 0de009b->39f3a83.",
|
||||
"_added": "#1956: c added to the scope-capture bench (was UNBENCHED). C has no inheritance \u2014 flat scale source. Adding it exposed + fixed a pre-existing O(n^2) findNodeAtRange root-walk in c/captures.ts (threaded c.node, byte-identical over c-* fixtures); scaling 3.475 -> 0.96.",
|
||||
"_note": "#1983: + c-static-linkage-worker fixture (caller.c/lib.c/lib.h/local.c \u2014 worker-path static-linkage side-channel test). Pure fixture-corpus drift: no c/captures.ts or query change branch-vs-main, existing fixtures' captures byte-identical (c-captures.test.ts 45/45), scaling stays linear (~0.97). The baseline was missed when the fixture landed; regenerated here. fingerprint 0de009b->39f3a83.",
|
||||
"_rebaselined": "#1919 open-language coverage: new lang-resolution fixtures + intended capture additions (F5/F9 c-cpp, F26/F28/F29 dart, F47/F48/F49/F51/F52 kotlin, F75/F79 swift). Fingerprint-only drift; scaling_ratio ~1.0 (linear, no perf regression)."
|
||||
},
|
||||
"cpp": {
|
||||
"fingerprint": "856d02f3f9d22cb973877211100aee8e052d4bc545922f78704b1a21ce49ddcc",
|
||||
"fingerprint": "bf3587674267be1759e7c45abef143c3b81fe8629cfd17da5f8af40e83cc39ec",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_2833_qualified_member_fields": "#2833 follow-up: the six per-qualifier-depth `field_declaration` type-binding rules for a QUALIFIED generic member are replaced by three depth-agnostic ones that match the outer `qualified_identifier` itself, with the qualifier reduced to its top-level tail in `interpret.ts` (`cppQualifiedTail`). This is a CAPTURE-LOGIC change and it moves the fingerprint in two places at once. (1) A qualified NON-generic member (`ns::Address addr;`, `std::string name;`) was captured by nothing at all and now binds \u2014 that is the whole +24 on the fixture corpus, every one of them a `std::string` member. (2) Qualifier depth is no longer enumerated, so `a::b::c::Repo<User>` (depth 3+) is captured where the old rules stopped at 2. Capture-name histogram, cpp-* corpus (278 files): `@type-binding.field` 8 -> 32, `@type-binding.name` and `@type-binding.type` 401 -> 425; synthetic DAO-20: `@type-binding.field` 40 -> 60, `@type-binding.name` and `@type-binding.type` 61 -> 81 (= 20 entities x the one `std::string name;` member the DAO unit already declared). NO OTHER TAG MOVED in either set \u2014 not one `@declaration.*`, `@scope.*` or `@reference.*` count \u2014 which is the property that says three rules replaced six without widening what a field_declaration matches. Measured over the 13 cpp-* fixture repos whose sources gained a binding, the distinct CALLS edge set is byte-identical before and after (32 edges): a reduced tail that names no workspace class binds nothing. Prior bd47c82d09a83cbf0ac857f41876fa31d22304043735582e913bccde06cf2c1a -> db1156d81b3e3341faf5e938a4a34417f4fd246588b6150b4686481823262529; scaling 1.04 < 1.5.",
|
||||
"_rebaselined_2833_generic_member_fields": "#2833 review follow-up: the cpp DAO generator's unit gains two GENERIC member fields \u2014 `Repo<Entity_n> repo;` (bare template_type) and `std::vector<Entity_n> items;` (qualified_identifier wrapping a template_type) \u2014 plus the header declaring `template <typename T> class Repo`. CORPUS CHANGE, NOT A CAPTURE-LOGIC CHANGE: no extractor edit accompanies it. It exists because the corpus had ZERO template-typed member fields and, across 279 cpp-* fixtures, not one qualified generic member either, so BOTH rounds of new `field_declaration` type-binding rules landed with a byte-identical cpp fingerprint \u2014 the gate was structurally blind to the exact thing being changed. Measured under the new corpus, the three states now differ: pre-#2833 query 0e7cbda71360b7ff35dd76091c77f288d6af6a5cfa9185ad85a372aae8c85191 (4521 groups) -> the three template_type field rules de07d8b5300ed867b460918e16b4d80259c7eb6efc1034d32bebe9ff7cab126d (4541) -> the six qualified rules bd47c82d09a83cbf0ac857f41876fa31d22304043735582e913bccde06cf2c1a (4561); under the OLD corpus all three were 856d02f3f9d22cb973877211100aee8e052d4bc545922f78704b1a21ce49ddcc. Capture-name histogram over the synthetic DAO-20: `@type-binding.field` 0 -> 40, `@declaration.field` 40 -> 80, `@type-binding.type`/`@type-binding.name` 20 -> 61, `@declaration.name` 104 -> 147 \u2014 40 = 20 entities x 2 fields, with the residual +1/+2/+3 attributable to the one-off header declaration; every `@reference.*` count is unchanged. Prior 856d02f3f9d22cb973877211100aee8e052d4bc545922f78704b1a21ce49ddcc -> bd47c82d09a83cbf0ac857f41876fa31d22304043735582e913bccde06cf2c1a; scaling 1.058 < 1.5. `c` is unaffected (3418cded..., unchanged).",
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature/cv metadata. Prior dde874d2c30bda9f634f9799281a66de800cad9f76cf65e7c31839e2ae9da9ff -> 57860dd2a8d4b06c6d2dd0d854c08b781faee3da8f2b6c42ba0c68a9f70e5ccb; scaling 1.090 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: C++ overload-aware function/reference/member-pointer flow facts with invocation/constructor-result suppression. Prior 3a503a1513e7eede3f7a223dcce0896c06d15bdfa920445224c9025848c0d710 -> dde874d2c30bda9f634f9799281a66de800cad9f76cf65e7c31839e2ae9da9ff; scaling 1.034 < 1.5.",
|
||||
"_rebaselined_callable_flow": "Callable-value-flow facts for C++ function pointers/references, reference aliases, contextual arity, arguments, and member-pointer syntax. Prior 6ab657c8f9bfe988a3759098c2cffdcc0443def75ff263f1282b82c21d96e931 -> 3a503a1513e7eede3f7a223dcce0896c06d15bdfa920445224c9025848c0d710; measured scaling ratio 1.069 < 1.5.",
|
||||
|
|
@ -43,37 +46,50 @@
|
|||
"_note_1899_followup": "#1899 follow-up: braced-init metadata now carries element count, intentionally changing C++ capture output; CI benchmark scaling remains linear (1.129 < 1.5).",
|
||||
"_added": "#1956: cpp added to the scope-capture bench (was UNBENCHED). Heritage-bearing scale source (: public Base, public Mixin) drives emitCppInheritanceCaptures at scale. Adding it exposed + fixed a pre-existing O(n^2) findNodeAtRange root-walk in cpp/captures.ts (~12 sites, threaded c.node, byte-identical over 263 cpp-* fixtures); scaling 2.30 -> 1.12.",
|
||||
"_rebaselined": "#1919 open-language coverage: new lang-resolution fixtures + intended capture additions (F5/F9 c-cpp, F26/F28/F29 dart, F47/F48/F49/F51/F52 kotlin, F75/F79 swift). Fingerprint-only drift; scaling_ratio ~1.0 (linear, no perf regression). #2094: deleted C++ declarations retain @declaration.is-deleted metadata; deleted operator and pointer-return shapes plus the expanded deleted-overload fixture are included. Intended capture drift; scaling remains linear (1.139 < 1.5).",
|
||||
"_note": "#1975: + cpp-out-of-line-class fixture, fixture_count 263->265. #1990: + cpp-adl-ns-plus-hidden-friend-same-name fixture (ADL hidden-friend + namespace-callable merge parity test). Pure fixture-corpus drift — no scope-extractor change; existing fixtures' captures byte-identical. fixture_count 265->267. #1995: + cpp-union-nested-tail-collision and cpp-anon-ns-tail-collision fixtures — pure fixture-corpus drift; fixture_count 270->272, fingerprint 538e8be->d63ded6. #1993: + cpp-cross-namespace-same-tail fixture — pure fixture-corpus drift; fixture_count 272->273, fingerprint d63ded6->6d6207ae. #2077 review follow-up: cpp-member-lattice adds cross-file, qualified-base, nested-template, inherited-using, this-receiver, and non-virtual-override regressions; fixture_count 274->275. Capture scaling remains linear (1.134 < 1.5). #1899: braced-init call arguments emit a conservative parameter-type capture; fixture_count 277, scaling remains linear (1.141 < 1.5).",
|
||||
"_note": "#1975: + cpp-out-of-line-class fixture, fixture_count 263->265. #1990: + cpp-adl-ns-plus-hidden-friend-same-name fixture (ADL hidden-friend + namespace-callable merge parity test). Pure fixture-corpus drift \u2014 no scope-extractor change; existing fixtures' captures byte-identical. fixture_count 265->267. #1995: + cpp-union-nested-tail-collision and cpp-anon-ns-tail-collision fixtures \u2014 pure fixture-corpus drift; fixture_count 270->272, fingerprint 538e8be->d63ded6. #1993: + cpp-cross-namespace-same-tail fixture \u2014 pure fixture-corpus drift; fixture_count 272->273, fingerprint d63ded6->6d6207ae. #2077 review follow-up: cpp-member-lattice adds cross-file, qualified-base, nested-template, inherited-using, this-receiver, and non-virtual-override regressions; fixture_count 274->275. Capture scaling remains linear (1.134 < 1.5). #1899: braced-init call arguments emit a conservative parameter-type capture; fixture_count 277, scaling remains linear (1.141 < 1.5).",
|
||||
"_rebaselined_2522_review_fixes": "PR #2522 review fixes: outermost-chain passing modes; ->* ERROR-recovery role order; member-store visibility. Prior 57860dd2a8d4b06c6d2dd0d854c08b781faee3da8f2b6c42ba0c68a9f70e5ccb -> f29bc3f7b1622954d6f6b7647bc9cf6c7a2629ffcc0fe00ac7918e4925876b65; scaling ratio re-verified within budget.",
|
||||
"_rebaselined_2522_prototype_value_cells": "Plain function/method prototypes no longer index as callable value cells (only pointer/parenthesized variable declarators do) — removes the spurious indirect-invoke facts that leaked phantom CALLS past two-phase suppression. Prior f29bc3f7b1622954d6f6b7647bc9cf6c7a2629ffcc0fe00ac7918e4925876b65 -> a70625bb0a9ef74e760d9d79cc5557485d0f0d3fb935e8a22a0c9556c65b5bb1; scaling re-verified within budget.",
|
||||
"_rebaselined_2522_prototype_value_cells": "Plain function/method prototypes no longer index as callable value cells (only pointer/parenthesized variable declarators do) \u2014 removes the spurious indirect-invoke facts that leaked phantom CALLS past two-phase suppression. Prior f29bc3f7b1622954d6f6b7647bc9cf6c7a2629ffcc0fe00ac7918e4925876b65 -> a70625bb0a9ef74e760d9d79cc5557485d0f0d3fb935e8a22a0c9556c65b5bb1; scaling re-verified within budget.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747: additionally adds the `cpp-receiver-chain-arrow` fixture, the behavioural proof for a `->` BASE receiver (`svc->getUser()->save()`) that the rollout fixed and that `cpp-chain-call/` could never catch because it uses the value `.` form. Prior a70625bb0a9ef74e760d9d79cc5557485d0f0d3fb935e8a22a0c9556c65b5bb1 -> 7e27aea46f3e17f33c41babbe0ddd982d1ab5920f143864763e0a1c6aef882a5.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 7e27aea46f3e17f33c41babbe0ddd982d1ab5920f143864763e0a1c6aef882a5 -> 856d02f3f9d22cb973877211100aee8e052d4bc545922f78704b1a21ce49ddcc."
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 7e27aea46f3e17f33c41babbe0ddd982d1ab5920f143864763e0a1c6aef882a5 -> 856d02f3f9d22cb973877211100aee8e052d4bc545922f78704b1a21ce49ddcc.",
|
||||
"capture_groups_small": 5021,
|
||||
"capture_groups_large": 16021,
|
||||
"capture_groups_fp": 4605,
|
||||
"fixture_count": 279
|
||||
},
|
||||
"csharp": {
|
||||
"_rebaselined": "#1956 synth-widening: + csharp-qualified-base fixture; the synth now walks record_declaration + struct_declaration base_lists and handles alias_qualified_name (matching the #1940 legacy leg), so record/struct heritage now emits. csharp-record-base gains a record inherits capture. (record->record SAME-namespace EXTENDS is a separate registry resolution gap, tracked as follow-up.) Linear (~1.00). (Earlier #1956: heritage-bearing scale source.) | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged. | #1924 F16: record primary-constructor base bindings now exclude constructor arguments; capture fingerprint changes, scaling remains linear. | #2036 review follow-up: csharp-record-base now exercises primary-constructor base dispatch end to end; +2 capture groups, scaling remains linear.",
|
||||
"fingerprint": "476d98a7cc659951c315d63319c8077bbcf0e5f3ec12d32ed773992a1f3a2adc",
|
||||
"fingerprint": "2930ef49fdce984a4c051409880bddfe8445e30e1c6bf802bd90a0a0f8f6b094",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior f31544530924748f9aa37d11cec570bc10c3ddf9d9b237e6df7a17623fd2bb3a -> 75cf380209fa7d1a8a3ec873be1a9424b4e5173be0b08234c2291e8521a9b3c1; scaling 1.061 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: C# method-group/delegate callable flow facts with invocation-result suppression. Prior 2bb5bc8c19cb8eb08c9590545ad8a1968a7152951f7e12746e2d7901d542fed9 -> f31544530924748f9aa37d11cec570bc10c3ddf9d9b237e6df7a17623fd2bb3a; scaling 1.115 < 1.5.",
|
||||
"_note": "#2046: F35 qualified-constructor captures now emit @reference.qualified-name + a simple-name @reference.name on `new Ns.Foo()`/`new A.B.Foo()`; namespace_declaration/file_scoped_namespace_declaration now emit @declaration.namespace name captures (feeding the non-destructive namespacePrefix sidecar for `new B.Foo()` same-tail disambiguation). + csharp-interface-only-base and csharp-namespace-qualified-ctor fixtures. Pure capture-additive + fixture-corpus drift; scaling stays linear (~1.11).",
|
||||
"_rebaselined_2563_instance_ownership": "#2563: csharp-using-static adds same-file ownership, local-function, overload, partial-class, and cross-namespace same-name coverage. Prior 75cf380209fa7d1a8a3ec873be1a9424b4e5173be0b08234c2291e8521a9b3c1 -> e05dc27456bde8175948586c9e7689033a378fa40e9ca4ce78cce41fbea0f2f8; scaling 1.058 < 1.5.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 05a85bae70cf9c94f42459c843cfc36e3e81c872e5dcc7d77bc42fbc390f4bfe -> 8a282254b93b3ef2ff34c2fdba819ebc95c53c4fcb09942cbad99f96d3687855.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 8a282254b93b3ef2ff34c2fdba819ebc95c53c4fcb09942cbad99f96d3687855 -> 476d98a7cc659951c315d63319c8077bbcf0e5f3ec12d32ed773992a1f3a2adc."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 05a85bae70cf9c94f42459c843cfc36e3e81c872e5dcc7d77bc42fbc390f4bfe -> 8a282254b93b3ef2ff34c2fdba819ebc95c53c4fcb09942cbad99f96d3687855.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 8a282254b93b3ef2ff34c2fdba819ebc95c53c4fcb09942cbad99f96d3687855 -> 476d98a7cc659951c315d63319c8077bbcf0e5f3ec12d32ed773992a1f3a2adc.",
|
||||
"capture_groups_small": 4259,
|
||||
"capture_groups_large": 13609,
|
||||
"capture_groups_fp": 2657,
|
||||
"fixture_count": 178
|
||||
},
|
||||
"rust": {
|
||||
"fingerprint": "6174889b8c98e0af430fa54c268dc781989ca9a8172d690eebae37a95f77e809",
|
||||
"fingerprint": "e61653008ff2de506cfd47f905fa9eb22d82fbbfe94d2a1d8190c358211b57b7",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_mod_node_identity_2745_review": "#2745 review: added rust-2742-mod-members, rust-2742-nested-mods and rust-2742-type-vs-module under lang-resolution for the container/owner-edge fix, nested inline modules, and the imported-type-vs-module precedence. emitRustScopeCaptures is unchanged — verified by removing ONLY those three fixture dirs and re-running, which reproduces the prior fingerprint exactly, so the shift is purely corpus growth (fixture_count 196 -> 202, capture_groups_fp 3432 -> 3556). Prior 90fda086a4e13aa069a5981f63ed58ab1c71f1ed3da5e1480a080e1992b0d3e5 -> 05acbaca48427e0d9e0793bcd0ce4057712d3716b5e7868189c12e05ef8dd300; scaling 1.022 local / 1.057 CI < 1.5. NOTE for the next fixture author: a new rust-* fixture drifts BOTH this bench baseline and the rust-captures-golden snapshot. Updating only the golden is how this reached CI red.",
|
||||
"_rebaselined_generic_instantiation_2912": "#2912: RUST_SCOPE_QUERY tags trait-impl heritage with the instantiation the impl was written with (`impl Validator<String> for V`), so interface dispatch can prune implementors of an instantiation the receiver cannot hold. Additive capture text on existing impl matches \u2014 the same matches are minted, carrying one more field \u2014 so this is digest drift, not a capture-set change: capture_groups_fp (3556) and fixture_count (202) are both unchanged, which is the check that no match appeared or vanished. Prior 116a971fee0004f340477aff69fa110a1d92bd8ba882d7c926483c6b1e8ca2b9 -> e61653008ff2de506cfd47f905fa9eb22d82fbbfe94d2a1d8190c358211b57b7; scaling 1.018 < 1.5. Only rust and dart move; the other 13 languages are byte-identical.",
|
||||
"_rebaselined_mod_node_identity_2745_review": "#2745 review: added rust-2742-mod-members, rust-2742-nested-mods and rust-2742-type-vs-module under lang-resolution for the container/owner-edge fix, nested inline modules, and the imported-type-vs-module precedence. emitRustScopeCaptures is unchanged \u2014 verified by removing ONLY those three fixture dirs and re-running, which reproduces the prior fingerprint exactly, so the shift is purely corpus growth (fixture_count 196 -> 202, capture_groups_fp 3432 -> 3556). Prior 90fda086a4e13aa069a5981f63ed58ab1c71f1ed3da5e1480a080e1992b0d3e5 -> 05acbaca48427e0d9e0793bcd0ce4057712d3716b5e7868189c12e05ef8dd300; scaling 1.022 local / 1.057 CI < 1.5. NOTE for the next fixture author: a new rust-* fixture drifts BOTH this bench baseline and the rust-captures-golden snapshot. Updating only the golden is how this reached CI red.",
|
||||
"_rebaselined_dyn_trait_object_2604": "#2604: RUST_SCOPE_QUERY now captures function_signature_item (abstract trait methods, no body) as a scope + declaration, so a &dyn Trait receiver can dispatch a CALLS edge to the trait's own method. Additive capture shift across every bench fixture with a required trait method. Prior df369c5a5f8de7753fc8bab8b4108ef5081750974ea5085ba9a867675ac9eb29 -> f7742f65f14d7d6590df7f16303fc3cc9dc0c233cd80bf90c98b084933cd3846; scaling 1.033 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior 65e5bca66bb1ca117949409e8fb5c80ee69d6f1b5318908eaaecf08da0482e5c -> df369c5a5f8de7753fc8bab8b4108ef5081750974ea5085ba9a867675ac9eb29; scaling 1.065 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: Rust fn-value callable flow facts with invocation/constructor-result suppression. Prior ac610bbe97666bf285923479dd7b43a2fe4c5354aae8df1bcbafdc04fb220f82 -> 65e5bca66bb1ca117949409e8fb5c80ee69d6f1b5318908eaaecf08da0482e5c; scaling 1.024 < 1.5.",
|
||||
"_rebaselined": "#1956 tri-review U1: rust-qualified-trait fixture (scoped + generic-of-scoped impl trait paths); bareTypeIdentifier now resolves scoped_type_identifier bases by their name: tail (additive, no existing-fixture drift); linear (~1.04). #1975: + rust-scoped-impl fixture (impl a::Inner / b::Inner inherent scoped impls) — legacy @definition.impl scoped arm + findEnclosingClassInfo inherent-impl scoped target; rust scope-extractor captures byte-identical. | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_note": "PR #1934: F66/F68 let-binding pattern narrowing; F71 union (Struct-labeled, now materialized via legacy @definition.struct + resolvable); F72 macro FULLY WIRED — @declaration.macro/@reference.macro + MacroRegistry → USES edges to Macro nodes (never a same-named fn). + rust-macro / rust-union fixtures and merged with origin/main #1975 rust-scoped-impl; fingerprint re-baselined (scaling ~0.99, fixture_count 126). #1992: + rust-nested-tail-collision-generic and rust-generic-impl-same-method-name (F3) fixtures — pure fixture-corpus drift, no scope-extractor change; fixture_count 127->129, fingerprint 56ffc1c0->b00aea0f.",
|
||||
"_rebaselined": "#1956 tri-review U1: rust-qualified-trait fixture (scoped + generic-of-scoped impl trait paths); bareTypeIdentifier now resolves scoped_type_identifier bases by their name: tail (additive, no existing-fixture drift); linear (~1.04). #1975: + rust-scoped-impl fixture (impl a::Inner / b::Inner inherent scoped impls) \u2014 legacy @definition.impl scoped arm + findEnclosingClassInfo inherent-impl scoped target; rust scope-extractor captures byte-identical. | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_note": "PR #1934: F66/F68 let-binding pattern narrowing; F71 union (Struct-labeled, now materialized via legacy @definition.struct + resolvable); F72 macro FULLY WIRED \u2014 @declaration.macro/@reference.macro + MacroRegistry \u2192 USES edges to Macro nodes (never a same-named fn). + rust-macro / rust-union fixtures and merged with origin/main #1975 rust-scoped-impl; fingerprint re-baselined (scaling ~0.99, fixture_count 126). #1992: + rust-nested-tail-collision-generic and rust-generic-impl-same-method-name (F3) fixtures \u2014 pure fixture-corpus drift, no scope-extractor change; fixture_count 127->129, fingerprint 56ffc1c0->b00aea0f.",
|
||||
"_rebaselined_import_disambiguation_2514": "#2514: added rust-import-* and rust-dup-* fixtures under lang-resolution for the range-binding ambiguity latch + import-disambiguated resolution (for-loops / struct destructuring across explicit/aliased/glob use imports). emitRustScopeCaptures is unchanged; the corpus fingerprint shifts purely because the fixture set grew (130 -> 174). Prior f7742f65f14d7d6590df7f16303fc3cc9dc0c233cd80bf90c98b084933cd3846 -> 655aed01cf1b6b84fa0c64d48dfb2526ecb67f47d90f0a91edabacd269a212db; scaling 1.06 < 1.5.",
|
||||
"_rebaselined_self_type_binding_2714": "#2714: a Rust `Self` type binding now records the enclosing impl's type instead of the literal 'Self'. `let fresh = Self { .. }` inside `impl User` binds `fresh: User`; recorded verbatim it bound `fresh: Self`, which resolves to nothing. The type-env channel already substituted this (type-extractors/rust.ts findEnclosingImplType); the scope-resolution channel did not, so the two disagreed. The gap was invisible while lookupCore Step 1 still walked the lexical chain for NAMED receivers — the impl scope binds the method by name, so fresh.validate() resolved by accident — and became a lost CALLS edge when #2714 stopped that walk. Only the rust fingerprint moves; the other 14 languages are byte-identical.",
|
||||
"_rebaselined_self_type_binding_2714": "#2714: a Rust `Self` type binding now records the enclosing impl's type instead of the literal 'Self'. `let fresh = Self { .. }` inside `impl User` binds `fresh: User`; recorded verbatim it bound `fresh: Self`, which resolves to nothing. The type-env channel already substituted this (type-extractors/rust.ts findEnclosingImplType); the scope-resolution channel did not, so the two disagreed. The gap was invisible while lookupCore Step 1 still walked the lexical chain for NAMED receivers \u2014 the impl scope binds the method by name, so fresh.validate() resolved by accident \u2014 and became a lost CALLS edge when #2714 stopped that walk. Only the rust fingerprint moves; the other 14 languages are byte-identical.",
|
||||
"_rebaselined_module_tree_2730": "#2730 + #2741 review: RUST_SCOPE_QUERY captures mod_item as @declaration.namespace (a Rust module is an item, mirroring the C++ namespace_definition capture) and tags scoped call sites with @reference.qualified-name so the written path survives to resolution. Both are additive captures: every bench fixture holding a mod block or a Foo::bar() call gains groups, and the corpus also grew by the rust-2730-* fixtures added for the fix and its review (workspace-crates, type-qualified, gaps, samename-wrapper, crate-layout). Prior 7f1240b38457468f06b7931e0c2c578f218f922774d0dc7e2ee6ef3b08d4d689 -> 90fda086a4e13aa069a5981f63ed58ab1c71f1ed3da5e1480a080e1992b0d3e5; scaling 1.061 < 1.5; fixture_count 196. Only the rust fingerprint moves; the other 14 languages are byte-identical. The earlier revision of this note cited 655aed01... as the prior value, which was two rebaselines stale (it predates #2604 and #2714); the CI gate compares live fingerprints, not this prose, so nothing caught it.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 05acbaca48427e0d9e0793bcd0ce4057712d3716b5e7868189c12e05ef8dd300 -> 83812d82f0e2c3eb552f3246381ca3dd5ccd6783d63aba3325f1343e7772280c.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 83812d82f0e2c3eb552f3246381ca3dd5ccd6783d63aba3325f1343e7772280c -> 6174889b8c98e0af430fa54c268dc781989ca9a8172d690eebae37a95f77e809."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 05acbaca48427e0d9e0793bcd0ce4057712d3716b5e7868189c12e05ef8dd300 -> 83812d82f0e2c3eb552f3246381ca3dd5ccd6783d63aba3325f1343e7772280c.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 83812d82f0e2c3eb552f3246381ca3dd5ccd6783d63aba3325f1343e7772280c -> 6174889b8c98e0af430fa54c268dc781989ca9a8172d690eebae37a95f77e809.",
|
||||
"capture_groups_small": 5507,
|
||||
"capture_groups_large": 17607,
|
||||
"capture_groups_fp": 3556,
|
||||
"fixture_count": 202
|
||||
},
|
||||
"php": {
|
||||
"fingerprint": "b213a872342da2d866b04681dede988770e4d3dfdc0d6e9f62212ec5b59cdc2c",
|
||||
|
|
@ -81,9 +97,9 @@
|
|||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior df7b1565f9115d66b1ae32e4a408d651afb2521b14e5ca615f3be426c29af618 -> 4a688fa5a7016546f7f3c6d44de023608ae80c5b0e3670c16f6e61b3632608fd; scaling 1.078 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: PHP first-class callable and variable-invocation flow facts with invocation-result suppression. Prior 31c9e3f3cb7094a2bf9021cf9db859036e002f8b44605cd993b470fc600e97cb -> df7b1565f9115d66b1ae32e4a408d651afb2521b14e5ca615f3be426c29af618; scaling 1.074 < 1.5.",
|
||||
"_rebaselined": "#1956: heritage-bearing scale source (class extends Base + use trait); both forms gated at scale; linear (~1.04). | #2481/#2482: PHP imports carry a symbol-kind capture so function/constant imports resolve by declaring file; capture shape changes, scaling remains linear (~1.04).",
|
||||
"_note": "PR #1931: F53 import multi-clause, F54 enum_case, F55 anonymous_class — fixture count 138→140, fingerprint drift expected.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 4a688fa5a7016546f7f3c6d44de023608ae80c5b0e3670c16f6e61b3632608fd -> 3745662053c76b6ae0a84a29aad319626ed5ccb88f7b9376c2680d3dc6502e28.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 3745662053c76b6ae0a84a29aad319626ed5ccb88f7b9376c2680d3dc6502e28 -> b213a872342da2d866b04681dede988770e4d3dfdc0d6e9f62212ec5b59cdc2c."
|
||||
"_note": "PR #1931: F53 import multi-clause, F54 enum_case, F55 anonymous_class \u2014 fixture count 138\u2192140, fingerprint drift expected.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 4a688fa5a7016546f7f3c6d44de023608ae80c5b0e3670c16f6e61b3632608fd -> 3745662053c76b6ae0a84a29aad319626ed5ccb88f7b9376c2680d3dc6502e28.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 3745662053c76b6ae0a84a29aad319626ed5ccb88f7b9376c2680d3dc6502e28 -> b213a872342da2d866b04681dede988770e4d3dfdc0d6e9f62212ec5b59cdc2c."
|
||||
},
|
||||
"ruby": {
|
||||
"fingerprint": "1c8c9c4b54036fa24c2a81e39ea530e938645c856d369075e5f437da78218c57",
|
||||
|
|
@ -91,10 +107,10 @@
|
|||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior cff273ae6cb7232c977d9241581834a2a2fa8bcf6369f7bd8f2471cd4419a6ef -> bf50ec6a53c8c91680dc6feac63a8956e78b1059249232dc25a0cfed25f31236; scaling 1.103 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: Ruby Method/Proc callable flow facts with invocation/constructor-result suppression. Prior b5ea93bb3d0469c3821a8c70f5d5991c6f326e41097c119ad691154301dcc753 -> cff273ae6cb7232c977d9241581834a2a2fa8bcf6369f7bd8f2471cd4419a6ef; scaling 1.086 < 1.5.",
|
||||
"_rebaselined": "#1956 synth-widening: + ruby-qualified-base fixture; synth now reduces a scope_resolution superclass (class C < Mod::Super) to its trailing constant (matching the #1940 legacy leg), at parity. Linear (~1.03). (Earlier #1956: heritage-bearing scale source.) | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_note": "F62: + scope_resolution class/module declaration captures — fixture count 78→81, fingerprint drift expected. #1975: + ruby-tail-collision fixture (Foo::Bar vs Baz::Bar stay distinct nodes) — pure fixture-corpus drift, scope-extractor captures unchanged; 81→82. #1991: + ruby-nested-mixin-tail-collision fixture (85→86). Recomputed on the #942 merge (fixture-comment rewording shifts capture byte-positions, capture LOGIC unchanged): bf6b13a -> b5ea93bb.",
|
||||
"_note": "F62: + scope_resolution class/module declaration captures \u2014 fixture count 78\u219281, fingerprint drift expected. #1975: + ruby-tail-collision fixture (Foo::Bar vs Baz::Bar stay distinct nodes) \u2014 pure fixture-corpus drift, scope-extractor captures unchanged; 81\u219282. #1991: + ruby-nested-mixin-tail-collision fixture (85\u219286). Recomputed on the #942 merge (fixture-comment rewording shifts capture byte-positions, capture LOGIC unchanged): bf6b13a -> b5ea93bb.",
|
||||
"_rebaselined_2522_review_fixes": "PR #2522 review fixes: bare identifiers are calls, not callable references (bareNamesAreCalls). Prior bf50ec6a53c8c91680dc6feac63a8956e78b1059249232dc25a0cfed25f31236 -> 070e4e11502442998ddf4048c2981cf1b2b735a87362ff854c5d14d71f98f4e2; scaling ratio re-verified within budget.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior fea3edf82f521995147874b7f6c5f9e2eb88efdebf6365668f3260e913f0b558 -> fc81941b0a921074fa80dc448284de9a23bd07358ddc84d4894797cc08c3fe83.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior fc81941b0a921074fa80dc448284de9a23bd07358ddc84d4894797cc08c3fe83 -> 1c8c9c4b54036fa24c2a81e39ea530e938645c856d369075e5f437da78218c57."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior fea3edf82f521995147874b7f6c5f9e2eb88efdebf6365668f3260e913f0b558 -> fc81941b0a921074fa80dc448284de9a23bd07358ddc84d4894797cc08c3fe83.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior fc81941b0a921074fa80dc448284de9a23bd07358ddc84d4894797cc08c3fe83 -> 1c8c9c4b54036fa24c2a81e39ea530e938645c856d369075e5f437da78218c57."
|
||||
},
|
||||
"swift": {
|
||||
"fingerprint": "adef9284feaecd39cb490aebce83876e15b9150c7a04b00a396feb78b7e1e0a9",
|
||||
|
|
@ -103,13 +119,14 @@
|
|||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: Swift function-value callable flow facts with invocation-result suppression. Prior 180ac68e780bdf6f9089d53f51cbb9a66aed3e7774631cc3fcbaae5020213998 -> 5f923c6604d825d12b249f31c155b0f4d13a8379d532e5dde64a0f9b15cf4725; scaling 1.043 < 1.5.",
|
||||
"_rebaselined": "#1919 open-language coverage: new lang-resolution fixtures + intended capture additions (F5/F9 c-cpp, F26/F28/F29 dart, F47/F48/F49/F51/F52 kotlin, F75/F79 swift). Fingerprint-only drift; scaling_ratio ~1.0 (linear, no perf regression).",
|
||||
"_rebaselined_2522_review_fixes": "PR #2522 review fixes: assignment target:/result: fields join the shared fallback. Prior 7687ee2466e16020a12440a03fbda53e63aa05f94b4481f6133c09867a0d560d -> 115c5da807e36bb12fdeba28e44f2b6484ef322ff26c19fa0f191febaf774248; scaling ratio re-verified within budget.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 115c5da807e36bb12fdeba28e44f2b6484ef322ff26c19fa0f191febaf774248 -> a6fca5f052ae5ec635b56051e28a168c864a988b2221a3279ddd69807378ba0b.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior a6fca5f052ae5ec635b56051e28a168c864a988b2221a3279ddd69807378ba0b -> 2f04ae960123cf50138a49fabdc5a146c2963170cecf5755c552b23c9055a9e7.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 115c5da807e36bb12fdeba28e44f2b6484ef322ff26c19fa0f191febaf774248 -> a6fca5f052ae5ec635b56051e28a168c864a988b2221a3279ddd69807378ba0b.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior a6fca5f052ae5ec635b56051e28a168c864a988b2221a3279ddd69807378ba0b -> 2f04ae960123cf50138a49fabdc5a146c2963170cecf5755c552b23c9055a9e7.",
|
||||
"_rebaselined_inferred_field_receiver_2807": "#2807: optional property annotations (`var a: Outer?`) now emit a type binding. The prior pattern required the `user_type` to be a DIRECT child of the annotation, so an `optional_type` wrapper meant an optional field was never typed at all and its receiver could not resolve. ADDS @type-binding.annotation captures on the optional form only; no capture is removed. Prior 2f04ae960123cf50138a49fabdc5a146c2963170cecf5755c552b23c9055a9e7 -> adef9284feaecd39cb490aebce83876e15b9150c7a04b00a396feb78b7e1e0a9; scaling 1.023 < 1.5."
|
||||
},
|
||||
"dart": {
|
||||
"fingerprint": "ba93c90dcd341259e8e088816bc8c76ad27882419f665e35c056dc22fa54cf73",
|
||||
"fingerprint": "3a8ddabbeb1cba47a4757451d4f79d726ca230fd15e860772b11526fbb1c6687",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_generic_instantiation_2912": "#2912: the Dart heritage marker carries a fourth field \u2014 the type arguments the clause was written with (`implements Validator<String>`) \u2014 so interface dispatch can prune implementors of a mismatched instantiation. Additive marker text on existing heritage matches rather than a new match, so this is digest drift only; a marker from a pre-#2912 cache simply has no fourth field and reads as unknown. Prior ba93c90dcd341259e8e088816bc8c76ad27882419f665e35c056dc22fa54cf73 -> 3a8ddabbeb1cba47a4757451d4f79d726ca230fd15e860772b11526fbb1c6687; scaling 1.027 < 1.5.",
|
||||
"_rebaselined_2538": "#2538: Dart extension type headers are preprocessed into normal extension declarations before scope capture, so extension type symbols and their methods are now emitted. Intentional Dart-only capture fingerprint drift; CI measured scaling 1.042 < 1.5.",
|
||||
"_rebaselined_2538_implements": "#2538 tri-review follow-up: Dart extension type implements clauses now emit heritage markers and fixture coverage asserts IMPLEMENTS edges, including multi-arg generic interfaces. Prior committed baseline 66a46d5ff09f3d11b2771db0f48596fe7057e95c5bc8f56241fdb911137298c3 -> ba93c90dcd341259e8e088816bc8c76ad27882419f665e35c056dc22fa54cf73; scaling 0.945 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior 29ce2bfe70b246b1c9d5e99c0ec11e850c22e9672737592207242b7f4cc824b8 -> 66a46d5ff09f3d11b2771db0f48596fe7057e95c5bc8f56241fdb911137298c3; scaling 1.054 < 1.5.",
|
||||
|
|
@ -118,11 +135,14 @@
|
|||
"_rebaselined": "#1919 review CF3 fix: extended kotlin-local-property-owner (init/accessor destructuring) + new dart-accessor-owner fixture (getter/setter ownership). Fingerprint-only corpus drift; scaling ~1.0."
|
||||
},
|
||||
"java": {
|
||||
"fingerprint": "a9943355e945e03ddb87c800f4cc1f62b3d04feefb3ec64c258d8e0bb3b3fcd9",
|
||||
"fingerprint": "2bf47cc19b595a9889ac21ec0154c6ce6786271d68551f21d1bc14c626bcd4ff",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_2935_synthetic_declarations": "PR #2935 review follow-up: synthesized Java anonymous classes and bodied enum constants now carry the presence-only @declaration.is-synthetic sidecar used to preserve source-written dispatch targets at the fanout cap. DIGEST DRIFT ONLY, NOT A CAPTURE-SET CHANGE: the tag is attached to existing synthetic declaration matches; capture groups and fixture count remain 5755/18405, 3512, and 206. Prior 36d689c58526c4482fbd701d1d9ca156623a3970734ead145717858712271ab5 -> 2e2150b4f4d64519e3f4c6d7a2c12259178d3117872203c904fab8cba96a694a; CI scaling 0.971 < 1.5.",
|
||||
"_rebaselined_2917_record_component_accessors": "#2917: every implicit Java record-component accessor now emits a component-bounded @scope.function plus @declaration.method/name/zero-arity/return-type metadata. The scope boundary prevents subsequent record-body references from being attributed to the accessor. Java was the only general language fingerprint to move; capture groups scale by exactly two per generated record component (small 5755 -> 6255, large 18405 -> 20005). Prior 36d689c58526c4482fbd701d1d9ca156623a3970734ead145717858712271ab5 -> 901a66c7dc0f071eeef9e4864b2519e5b58a1a141a1f9a7817ea42f7ff70eafb; scaling 0.961 < 1.5. Re-measured after merging origin/main, which carries #2935's is-synthetic sidecar on top of the same corpus: 2e2150b4f4d64519e3f4c6d7a2c12259178d3117872203c904fab8cba96a694a -> 79dafc369eaeb7183ee8cc1149b1a6c21ad672c7e5b806fe8b0060e5a952c79a; scaling 1.085 < 1.5, capture groups 6255/20005, capture_groups_fp 3560, fixture_count 206 (unchanged by the merge).",
|
||||
"_rebaselined_2900_record_heritage": "#2900 review follow-up: the Java scale unit now includes a record implementing Marker, so the record-declaration @reference.inherits path is fingerprinted and exercised at scale. Prior b29e263524f55151dcb7cfc4c929d3d1d7bb360355cee4e832158f927857f663 -> 36d689c58526c4482fbd701d1d9ca156623a3970734ead145717858712271ab5; scaling 1.042 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata; same-name lexical regions use an O(ancestor-depth) ID-set lookup. Prior d5c59d7dc9e206637515d5aea1163f7c1cdd76410c38c5fe6143d13d19677d6a -> 004a3592998dca1193bd1429a8284513725de7764f2a3eceedaaa984cfd763b4; scaling 0.992 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: Java method-reference/SAM callable flow facts with invocation-result suppression. Prior 062d754764aaa8a6772fb90875c710502a63e3e7a300e633942381ed914faada -> d5c59d7dc9e206637515d5aea1163f7c1cdd76410c38c5fe6143d13d19677d6a; scaling 1.074 < 1.5.",
|
||||
"_rebaselined": "#2357 (supersedes #2353): + java-cast-receiver, java-this-field-chain, java-this-dispatch fixtures (cast-wrapped receivers, this.field chains incl. initializer contexts, bare-this dispatch pinning). Drift is purely fixture-additive: with the three new dirs parked, the fingerprint reproduces the prior baseline byte-identically — no emit/capture change. #1956 synth-widening: + java-iface-extends fixture; synthesizeJavaInheritanceReferences now ALSO walks interface_declaration extends_interfaces (interface IA extends IB, IC<T>), matching the #1940 legacy leg. (Earlier U2+review: java-qualified-base fixture covers 2- AND 3-segment qualified bases guarding the legacy end-anchor; synth tail-resolves scoped bases.) Linear (~1.03). (Earliest: java added to bench, exposed+fixed the O(n^2) findNodeAtRange root-walk; 3.09 -> ~0.99.) | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_rebaselined": "#2357 (supersedes #2353): + java-cast-receiver, java-this-field-chain, java-this-dispatch fixtures (cast-wrapped receivers, this.field chains incl. initializer contexts, bare-this dispatch pinning). Drift is purely fixture-additive: with the three new dirs parked, the fingerprint reproduces the prior baseline byte-identically \u2014 no emit/capture change. #1956 synth-widening: + java-iface-extends fixture; synthesizeJavaInheritanceReferences now ALSO walks interface_declaration extends_interfaces (interface IA extends IB, IC<T>), matching the #1940 legacy leg. (Earlier U2+review: java-qualified-base fixture covers 2- AND 3-segment qualified bases guarding the legacy end-anchor; synth tail-resolves scoped bases.) Linear (~1.03). (Earliest: java added to bench, exposed+fixed the O(n^2) findNodeAtRange root-walk; 3.09 -> ~0.99.) | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_note": "#1928 / #2045: F35 adds qualified + qualified-generic constructor query captures (`new pkg.Foo()`, `new a.b.Foo()`, `new pkg.Box<T>()`); F38 synthesizes `@reference.call.constructor` on `super(...)`/`this(...)` explicit_constructor_invocation nodes; F41 generic-aware stripQualifier in interpret (type-binding normalization). + java-qualified-constructor and java-explicit-constructor fixtures. Pure capture-additive + fixture-corpus drift; scaling stays linear (~1.06).",
|
||||
"_rebaselined_2522_review_fixes": "PR #2522 review fixes: get/test dropped from callableProtocolMethods. Prior 004a3592998dca1193bd1429a8284513725de7764f2a3eceedaaa984cfd763b4 -> f3b4f4b6610e07c3ac90deb1c53d3572b6ad55a36e5d7134984876d30031ff67; scaling ratio re-verified within budget.",
|
||||
"_rebaselined_2550_instance_model": "PR #2549 (#2550): anonymous class bodies emit synthesized @declaration.class/@declaration.name (Worker$N), an @reference.inherits to the constructed type, and receiver @type-binding.* captures; six new java-* fixtures joined the corpus. Prior f3b4f4b6610e07c3ac90deb1c53d3572b6ad55a36e5d7134984876d30031ff67 -> d79c3b92acfc866094981499b977388ca14f90839bca0c040342ab1cec00aa90; scaling 1.058 < 1.5.",
|
||||
|
|
@ -130,34 +150,51 @@
|
|||
"_rebaselined_2564_record_capture": "PR for #2564: JAVA_QUERIES gained a (record_declaration name: (identifier) @name) @definition.record capture, previously entirely missing (record_declaration had no structure-phase capture at all, unlike class/interface/enum) - a record's methods existed as ownerless Method nodes with no HAS_METHOD edge. Two new java-* fixtures (java-record-methods, java-new-expr-chain-call) joined the corpus. Prior 975b68aaac6d06094260fb0c67f9b1bc03692ba7220669d192aca9dccd5fc0ca -> 85fc7af9c3c1bceac76cb4f27214410b04967682a2eaa7e468e26efd1f4e2537; scaling 1.059 < 1.5.",
|
||||
"_rebaselined_2561_enum_constant_receiver": "PR for #2561: synthesizeJavaAnonymousClassDeclarations now emits a class-scope @type-binding.annotation/name/type per enum constant (constant simple name -> its E$N synthesized class when bodied, else the host enum) so E.CONST.method() resolves through the existing compound-receiver chain walk. Two drivers of the drift, both in the java-enum-constant-body fixture (this bench's corpus IS test/fixtures/lang-resolution): (1) one extra type-binding match per enum_constant from the capture change; (2) review follow-up added a body-less Plain.java enum + EnumConst.dispatchToConstant/dispatchInherited methods (bodied-override, inherited-via-MRO, and body-less dispatch call sites). The review's fail-safe hardening (bodied constant binds ONLY to E$N, never the host enum, when name synthesis fails on a malformed tree) is output-neutral on this well-formed corpus (verified: fingerprint identical with and without it). Prior 85fc7af9c3c1bceac76cb4f27214410b04967682a2eaa7e468e26efd1f4e2537 -> d04298a91beec76d0fa7099b3d71265723be60c1df688969aa954f135dd49686; scaling < 1.5.",
|
||||
"_rebaselined_2562_local_classes": "#2562: Java block-local classes, enums, records, and interfaces use source-type-relative JLS 13.1 Host$NLocal identities with javac-compatible per-(host, simple-name) numbering; anonymous numbering remains separate. Lexical aliases begin at each declaration and end with its immediate block. Expanded java-local-class-naming fixtures cover declaration order, disjoint blocks, initializers, lambdas, local type kinds, and recursive local/member/anonymous host chains. Prior d04298a91beec76d0fa7099b3d71265723be60c1df688969aa954f135dd49686 -> 6dd5913a58400a191ff54abf9b852b03d5add657d16c11e60a7c4608ba186197; scaling 1.204 < 1.5.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 6dd5913a58400a191ff54abf9b852b03d5add657d16c11e60a7c4608ba186197 -> 310adbc2e0827b5ac749acaa981cd12d256fc5b7cbc5592c5bee219e92abf9ee.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 310adbc2e0827b5ac749acaa981cd12d256fc5b7cbc5592c5bee219e92abf9ee -> a9943355e945e03ddb87c800f4cc1f62b3d04feefb3ec64c258d8e0bb3b3fcd9."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 6dd5913a58400a191ff54abf9b852b03d5add657d16c11e60a7c4608ba186197 -> 310adbc2e0827b5ac749acaa981cd12d256fc5b7cbc5592c5bee219e92abf9ee.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 310adbc2e0827b5ac749acaa981cd12d256fc5b7cbc5592c5bee219e92abf9ee -> a9943355e945e03ddb87c800f4cc1f62b3d04feefb3ec64c258d8e0bb3b3fcd9.",
|
||||
"capture_groups_small": 6255,
|
||||
"capture_groups_large": 20005,
|
||||
"capture_groups_fp": 3586,
|
||||
"fixture_count": 209,
|
||||
"_rebaselined_2910_declared_package_fixtures": "#2910 adds three Java resolver fixture files covering an external JDK lookalike, a path/package mismatch, and wildcard package membership. Fixture-corpus growth only: Java query rules and synthetic scaling sources are unchanged; capture_groups_small/large remain 6255/20005. capture_groups_fp 3560 -> 3586 and fixture_count 206 -> 209."
|
||||
},
|
||||
"java-local-types": {
|
||||
"fingerprint": "8c50bbc83dff4f7f5abd06078aa6abc6b64af05fddb17ee826b5f3df3d346633",
|
||||
"fingerprint": "bdde823fa725e636e257940efb4c8655aa23124c1727cbaa8856d1ad8f71729e",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_2935_synthetic_declarations": "PR #2935 review follow-up: the local-type stress corpus includes synthesized anonymous declarations, which now carry the presence-only @declaration.is-synthetic sidecar. DIGEST DRIFT ONLY, NOT A CAPTURE-SET CHANGE. Prior 8c50bbc83dff4f7f5abd06078aa6abc6b64af05fddb17ee826b5f3df3d346633 -> 560734cd053fb4f4b23aa04bc7870c22089a8deedb0217fa9c1b4db689e02a97; CI scaling 1.002 < 1.5.",
|
||||
"_rebaselined_2917_record_component_accessors": "#2917: the focused local-type fixture corpus contains local records, so their implicit component accessors add the same bounded scope/declaration captures as the general Java corpus. No local-type naming logic changed. Prior 8c50bbc83dff4f7f5abd06078aa6abc6b64af05fddb17ee826b5f3df3d346633 -> 3e22f368a4ee139be7cb91ff4fb77ddadf60c55efe8d66955ec81f366a46e460; scaling 1.032 < 1.5, capture_groups_fp 680. Re-measured on top of #2935's is-synthetic sidecar after merging origin/main: 560734cd053fb4f4b23aa04bc7870c22089a8deedb0217fa9c1b4db689e02a97 -> bdde823fa725e636e257940efb4c8655aa23124c1727cbaa8856d1ad8f71729e; scaling 0.997 < 1.5, capture_groups_fp 680.",
|
||||
"_added": "#2562 performance follow-up: co-scales same-host, same-name local classes and anonymous classes to gate JLS binary-name ordinal allocation. Precomputed per-sequence ordinals reduce the focused 100->800 workload from 176->6655ms to 141->752ms; normalized 250->800 scaling is 1.054.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior a9ad88de21ca6747a923260dbdf677fb74a004abbf9d57781f745e3a9027530b -> 3ca67847ea2b9a71b0a41e09f943767e5a2d3a113d3e203499ee364e37f40236.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 3ca67847ea2b9a71b0a41e09f943767e5a2d3a113d3e203499ee364e37f40236 -> 8c50bbc83dff4f7f5abd06078aa6abc6b64af05fddb17ee826b5f3df3d346633."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior a9ad88de21ca6747a923260dbdf677fb74a004abbf9d57781f745e3a9027530b -> 3ca67847ea2b9a71b0a41e09f943767e5a2d3a113d3e203499ee364e37f40236.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 3ca67847ea2b9a71b0a41e09f943767e5a2d3a113d3e203499ee364e37f40236 -> 8c50bbc83dff4f7f5abd06078aa6abc6b64af05fddb17ee826b5f3df3d346633.",
|
||||
"capture_groups_fp": 680
|
||||
},
|
||||
"typescript": {
|
||||
"fingerprint": "7a960908031331360ce582f5b55b7681e1cd7f8a2eabfd73c00982cb17f2a949",
|
||||
"fingerprint": "05d1dadd6c9ef35c74079fa50f341b1b36e4fb02c9a89dd1b59f32b7cfd5e633",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_2934_import_type_only": "#2934: `import-decomposer.ts` attaches a presence-only `@import.type-only` synthetic capture to specifiers `tsc` erases, so `check --cycles` can stop counting type-only edges as initialization cycles. DIGEST DRIFT ONLY, NOT A CAPTURE-SET CHANGE \u2014 the tag is added to import matches that already existed, never a new match, the same shape as the #2747 receiver-chain rebaseline. Every count is unchanged: capture_groups_fp 2414, fixture_count 155, capture_groups_small/large 4503/14403 (those measure the SYNTHETIC scaling source, which has no imports at all). The fingerprint moves because `canonicalizeMatch` in measure.mjs hashes every TAG on every match, synthetics included, so one extra presence-only tag on an existing match rewrites that match's canonical string. Attribution is exact, not inferred: neutralizing ONLY the `m['@import.type-only'] = \u2026` assignment in import-decomposer.ts and re-running returns the fingerprint to c2fbf8a89e5686dd\u2026 byte-for-byte, so nothing else in the TypeScript capture stream moved. All 14 other languages report ok. Scaling 0.997 < 1.5. NOTE ON THE CONTROL: javascript did not move (2026993b\u2026, 43 fixtures), but it is a WEAK control here \u2014 `import type` is TypeScript-only syntax, so a JS corpus cannot express the construct and could not have drifted either way. It evidences no collateral damage, not the correctness of the TS change; the exact-attribution check above is what does that. Prior c2fbf8a89e5686dd1ff3659b20d41d8b05ebcc9790356e3653ee0c8ca5d365c8 -> f719163eb03a447c9e40ca316a905dd76cee82192a75a403df478ebbdc13e98f.",
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior 27f937bfb47d4bded316ea3c785ff659c8cd88a5761d928f113477a08c802c78 -> e05446620c5b80b7aae291cfdf32f693580fada2ae687124769b04a0c03bfe63; scaling 0.983 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: lexical callable bindings, direct-callee argument metadata, and invocation-result suppression. Prior db5933cc6760234ed7d495123410feba6de243646d583f20d43032b9459f81fd -> 27f937bfb47d4bded316ea3c785ff659c8cd88a5761d928f113477a08c802c78; scaling 0.975 < 1.5.",
|
||||
"_rebaselined_callable_flow": "Callable assignment/copy/formal/argument/invoke facts (also consumed by Vue script blocks). Prior 25de86fd3377132c4e35d3d98f4f94a58e0cfeb7c22948a8ea3be4e793be74fd -> db5933cc6760234ed7d495123410feba6de243646d583f20d43032b9459f81fd; measured scaling ratio 0.951 < 1.5.",
|
||||
"_rebaselined": "#1962: F44 (class scope@), F85 (enum member declarations), F87 (optional_parameter type annotations) add new captures — fingerprint drift expected.",
|
||||
"_note": "#1968: F44, F85, F87 — fingerprint drift expected.",
|
||||
"_rebaselined": "#1962: F44 (class scope@), F85 (enum member declarations), F87 (optional_parameter type annotations) add new captures \u2014 fingerprint drift expected.",
|
||||
"_note": "#1968: F44, F85, F87 \u2014 fingerprint drift expected.",
|
||||
"_rebaselined_2522": "#2522 intentional @reference.value-ref/property-key capture additions. GitHub Actions run 29553361660 job 87800394279: prior 3f44a4a6892698df2d145c8ff2812c3b318807648983c88aca28fbd694f172f9 -> 25de86fd3377132c4e35d3d98f4f94a58e0cfeb7c22948a8ea3be4e793be74fd; scaling ratio 0.987 < 1.5.",
|
||||
"_rebaselined_2550_instance_model": "PR #2549 (#2545/#2551): object literals emit @scope.object (was unscoped, then @scope.block during development). Prior e05446620c5b80b7aae291cfdf32f693580fada2ae687124769b04a0c03bfe63 -> 3280b13d3f9378ab23eee31c2edc779b5a9ae1e7bb510c23a24855b44406d2f4; scaling 0.981 < 1.5.",
|
||||
"_rebaselined_receiver_owner_2701": "#2701: every non-arrow function form now carries a `@receiver-owner.this` marker on the same node as `@scope.function`, so a scope that BINDS its own `this` can stop the receiver walk (`Scope.ownsReceivers`). Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus against 1d3088173f6f93827641b476d614d5d15cd4f3ea: the ONLY delta is @receiver-owner.this (typescript +143, javascript +32) — every other capture count is byte-identical, so no existing capture moved. Prior 3280b13d3f9378ab23eee31c2edc779b5a9ae1e7bb510c23a24855b44406d2f4 -> 281e95484203b481094729ca249ef0423c41273eac35e424cdfd032a0dac7699.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior cad25be9f81d6e021ebae8dcb166bc0af3a1ba8021f1506f6ca93fd4c2649000 -> 9e112415f1169f08576826c12ea1d137d1994e34b44c45986c9ffee83b8b4edc.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 9e112415f1169f08576826c12ea1d137d1994e34b44c45986c9ffee83b8b4edc -> cdefe88d3c275f31953216c676ef32c7bf5727d56b9c3840b81ee6bf85749dff.",
|
||||
"_rebaselined_inferred_field_receiver_2807": "#2807: inference-typed class fields now emit a type binding — `public_field_definition` with a `new_expression` value, and `this.<field> = new ...` carrying a @type-binding.this-field marker. ADDS @type-binding.constructor captures only; no capture is removed, and the annotated form is unchanged because annotation outranks constructor-inferred in typeBindingStrength. Prior cdefe88d3c275f31953216c676ef32c7bf5727d56b9c3840b81ee6bf85749dff -> 248b56f0d7a0a6fc7a949dc7afb8611e135ed642bccc2631b96ebb9d686bb965; scaling 0.994 < 1.5.",
|
||||
"_rebaselined_ts_heritage_2842": "#2842 review: TypeScript heritage capture now emits `@reference.inherits` for `interface_declaration` (bases on `extends_type_clause`) and `abstract_class_declaration` (bases on `class_heritage`), which were both silently skipped — so `interface B extends A` and `abstract class X implements I` produced no edge and every interface-dispatch walk dead-ended on a bodiless declaration. Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus (145 files) with and without the change: the ONLY deltas are @reference.inherits 17 -> 20 (+3) and its paired @reference.name 245 -> 248 (+3), emitted together by emitTsInheritanceBase. Every other capture count is byte-identical, so no existing capture moved. The +3 is the three `interface X extends BasePayload` declarations in typescript-generic-calls/src/{auth,admin,guest}.ts. javascript is unchanged (no interfaces in the language). Prior 248b56f0d7a0a6fc7a949dc7afb8611e135ed642bccc2631b96ebb9d686bb965 -> 7a960908031331360ce582f5b55b7681e1cd7f8a2eabfd73c00982cb17f2a949."
|
||||
"_rebaselined_receiver_owner_2701": "#2701: every non-arrow function form now carries a `@receiver-owner.this` marker on the same node as `@scope.function`, so a scope that BINDS its own `this` can stop the receiver walk (`Scope.ownsReceivers`). Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus against 1d3088173f6f93827641b476d614d5d15cd4f3ea: the ONLY delta is @receiver-owner.this (typescript +143, javascript +32) \u2014 every other capture count is byte-identical, so no existing capture moved. Prior 3280b13d3f9378ab23eee31c2edc779b5a9ae1e7bb510c23a24855b44406d2f4 -> 281e95484203b481094729ca249ef0423c41273eac35e424cdfd032a0dac7699.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior cad25be9f81d6e021ebae8dcb166bc0af3a1ba8021f1506f6ca93fd4c2649000 -> 9e112415f1169f08576826c12ea1d137d1994e34b44c45986c9ffee83b8b4edc.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 9e112415f1169f08576826c12ea1d137d1994e34b44c45986c9ffee83b8b4edc -> cdefe88d3c275f31953216c676ef32c7bf5727d56b9c3840b81ee6bf85749dff.",
|
||||
"_rebaselined_inferred_field_receiver_2807": "#2807: inference-typed class fields now emit a type binding \u2014 `public_field_definition` with a `new_expression` value, and `this.<field> = new ...` carrying a @type-binding.this-field marker. ADDS @type-binding.constructor captures only; no capture is removed, and the annotated form is unchanged because annotation outranks constructor-inferred in typeBindingStrength. Prior cdefe88d3c275f31953216c676ef32c7bf5727d56b9c3840b81ee6bf85749dff -> 248b56f0d7a0a6fc7a949dc7afb8611e135ed642bccc2631b96ebb9d686bb965; scaling 0.994 < 1.5.",
|
||||
"_rebaselined_ts_heritage_2842": "#2842 review: TypeScript heritage capture now emits `@reference.inherits` for `interface_declaration` (bases on `extends_type_clause`) and `abstract_class_declaration` (bases on `class_heritage`), which were both silently skipped \u2014 so `interface B extends A` and `abstract class X implements I` produced no edge and every interface-dispatch walk dead-ended on a bodiless declaration. Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus (145 files) with and without the change: the ONLY deltas are @reference.inherits 17 -> 20 (+3) and its paired @reference.name 245 -> 248 (+3), emitted together by emitTsInheritanceBase. Every other capture count is byte-identical, so no existing capture moved. The +3 is the three `interface X extends BasePayload` declarations in typescript-generic-calls/src/{auth,admin,guest}.ts. javascript is unchanged (no interfaces in the language). Prior 248b56f0d7a0a6fc7a949dc7afb8611e135ed642bccc2631b96ebb9d686bb965 -> 7a960908031331360ce582f5b55b7681e1cd7f8a2eabfd73c00982cb17f2a949.",
|
||||
"capture_groups_small": 4503,
|
||||
"capture_groups_large": 14403,
|
||||
"capture_groups_fp": 2465,
|
||||
"fixture_count": 167,
|
||||
"_rebaselined_blind_spots_2856": "#2856 blind-spots series: the JS/TS SCOPE queries gained capture rules, so fingerprint drift is expected and additive. Verified before re-baselining by diffing the capture-name sets in both scope queries against origin/main: TypeScript gained exactly @reference.read.identifier (A2 bare-identifier reads in value positions) and @reference.type (R2-2 type references, so a declared contract stops reporting incoming:{}); JavaScript gained exactly @reference.read.identifier, @reference.read.destructured (R2-1c) and @reference.write.property-key (R2-1b record-construction writes). NOTHING was removed on either side \u2014 the delta is a pure superset, which is the check that no existing capture moved. capture_groups_small/large are unchanged (4503/14403) because those measure the SYNTHETIC scaling source, which this branch does not touch; only the fixture-corpus count moves. capture_groups_fp 2097 -> 2338 and fixture_count 146 -> 151 from 21 new lang-resolution fixtures. Scaling stayed linear and inside budget: typescript 1.116 < 1.5, javascript 1.010 < 1.5. Prior typescript ed92588e0fc7b28b3a0174339ac378b4dd85965fe007db1208dea97a65ce0571 -> f66a3e6f1e096431e7046505129a627deaa00ca0de5bc846b080591b397248f7; prior javascript 806f70ad3cce5fc849f6d06a08ace8a95f92a1ea84a2418fddabb1eef5846594 -> 2026993b81b873839dd2ef8797d9c14d9c48516b2b57b05ac17d8d43f2f4eba3.",
|
||||
"_rebaselined_type_parameter_shadowing_w2_8": "W2-8: `@declaration.type-parameters` is now captured on generic FUNCTIONS, generator functions and type ALIASES, not only on class/interface declarations. NO NEW CAPTURE NAME \u2014 verified by diffing the capture-name sets against the wave-1 branch, which returns empty; the tag already existed and simply fires on more declarations. That is the whole delta: capture_groups_fp 2338 -> 2371 (+33 occurrences of an existing tag) and fixture_count 151 -> 152 (one new fixture, typescript-type-parameters). capture_groups_small/large unchanged at 4503/14403, since those measure the synthetic scaling source this does not touch. Scaling 1.06 < 1.5. JavaScript is untouched \u2014 it has no type parameters \u2014 and its fingerprint does not move, which is the check that this is the TS declaration rules and not something broader. Prior f66a3e6f1e096431e7046505129a627deaa00ca0de5bc846b080591b397248f7 -> 62c7f1bfbe568eed927fb78f00061ed5e49d12511fd8260648b876df386f3b4c.",
|
||||
"_rebaselined_2899_review_type_parameter_scope_fixtures": "PR #2899 review follow-up: FIXTURE-CORPUS GROWTH ONLY \u2014 no query rule changed and no capture name was added or removed. `typescript/query.ts` is byte-identical to the previous baseline; the type-parameter shadowing defect was fixed on the RESOLUTION side (`walkers.ts` gains a `declarationOpenedScope` gate so a declaration's `typeParameters` bind only inside the scope that declaration opened, and the `USES` guard moved from `graph-bridge/references-to-edges.ts` to `resolve-references.ts` where the spelled `site.name` is in hand). The fingerprint moves because measure.mjs fingerprints the whole `lang-resolution/typescript-*` fixture corpus and the regression tests add three files to `typescript-type-parameters/src/` (values.ts, aliased.ts, namespaced.ts) plus two scope-less generic aliases in shapes.ts. Per-file accounting sums exactly to the delta: shapes.ts 33->35 (+2), values.ts +11, aliased.ts +10, namespaced.ts +20 = +43. capture_groups_fp 2371 -> 2414; fixture_count 152 -> 155. capture_groups_small/large unchanged at 4503/14403 (they measure the SYNTHETIC scaling source, untouched). JAVASCRIPT IS THE CONTROL AND DID NOT MOVE (fingerprint 2026993b..., 43 fixtures) \u2014 which is the check that this is corpus growth and not a capture regression; all 14 other languages report `ok`. Scaling 0.976 < 1.5. Prior 62c7f1bfbe568eed927fb78f00061ed5e49d12511fd8260648b876df386f3b4c -> c2fbf8a89e5686dd1ff3659b20d41d8b05ebcc9790356e3653ee0c8ca5d365c8.",
|
||||
"_rebaselined_2953_workspace_fixture": "#2953 adds test/fixtures/lang-resolution/typescript-pnpm-workspace-imports, a pnpm monorepo of 12 .ts files, and the TypeScript capture corpus is collected from test/fixtures. CORPUS GROWTH ONLY, NOT A CAPTURE CHANGE: fixture_count 155 -> 167 and capture_groups_fp 2414 -> 2465 are the 12 new files' own matches; capture_groups_small/large are unchanged at 4503/14403 because those measure the SYNTHETIC scaling source, which the fixture corpus does not feed. Attribution is exact rather than inferred: moving that one fixture directory aside and re-running returns typescript to f719163eb03a447c9e40ca316a905dd76cee82192a75a403df478ebbdc13e98f byte-for-byte with fixture_count back at 155, and [scope-capture --check] PASSES for all 15 languages - so nothing in the TypeScript capture stream moved. #2953 changes import RESOLUTION, which runs after capture and feeds no capture tag. Prior f719163eb03a447c9e40ca316a905dd76cee82192a75a403df478ebbdc13e98f -> 05d1dadd6c9ef35c74079fa50f341b1b36e4fb02c9a89dd1b59f32b7cfd5e633."
|
||||
},
|
||||
"javascript": {
|
||||
"fingerprint": "806f70ad3cce5fc849f6d06a08ace8a95f92a1ea84a2418fddabb1eef5846594",
|
||||
"fingerprint": "2026993b81b873839dd2ef8797d9c14d9c48516b2b57b05ac17d8d43f2f4eba3",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior b59fe8135b6a31a12bc3f872b224054b16592588153ae3661d03958d787c76f3 -> 479927409bbdd9852a36172c8260aa56df260e99129a7a9c20a0d1903dd5538b; scaling 1.050 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: lexical callable bindings, direct-callee argument metadata, and invocation-result suppression. Prior 917a9cd975ba035bdad71fdb70cd72eeddec58c25797e5a1addfa6172808a55c -> b59fe8135b6a31a12bc3f872b224054b16592588153ae3661d03958d787c76f3; scaling 1.093 < 1.5.",
|
||||
|
|
@ -166,23 +203,28 @@
|
|||
"_rebaselined": "#1956 synth-widening: + javascript-qualified-base fixture; synthesizeJsInheritanceReferences now handles a member_expression base (class S extends ns.Base -> Base), matching the #1940 legacy leg + the TS terminalTsTypeNameNode property_identifier case, at parity. Linear (~1.05). | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged.",
|
||||
"_rebaselined_2522": "#2522 intentional @reference.value-ref/property-key capture additions. GitHub Actions run 29553361660 job 87800394279: prior d72f03c6c502235d2d4b74d66baa5c7d361f040d7a1b72e84acad61210d05ae8 -> 5567dd47e7ba29821a518c4a9852adc3b774e25ef3e7a6e2b3ecb7b59ddab73c; scaling ratio 1.031 < 1.5.",
|
||||
"_rebaselined_2550_instance_model": "PR #2549 (#2545/#2551): object literals emit @scope.object. Prior 479927409bbdd9852a36172c8260aa56df260e99129a7a9c20a0d1903dd5538b -> f1ccf42a36895c8e34dcb724286f247d469835f2dcbb23ad3347190adc7fde1c; scaling 1.096 < 1.5.",
|
||||
"_rebaselined_receiver_owner_2701": "#2701: every non-arrow function form now carries a `@receiver-owner.this` marker on the same node as `@scope.function`, so a scope that BINDS its own `this` can stop the receiver walk (`Scope.ownsReceivers`). Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus against 1d3088173f6f93827641b476d614d5d15cd4f3ea: the ONLY delta is @receiver-owner.this (typescript +143, javascript +32) — every other capture count is byte-identical, so no existing capture moved. Prior f1ccf42a36895c8e34dcb724286f247d469835f2dcbb23ad3347190adc7fde1c -> 90601494695b834d3a9af7ac4844eac603f4f432809a05554cc59de0674a4354.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 1c71ef628eb75a3b111afa8c2a7c351c16a7f5aab9fac2f098f82b2866312aa8 -> 83344b7cba093702f4528eeee44e438809c229d43b12e69ed288812ce7ffc7bc.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 83344b7cba093702f4528eeee44e438809c229d43b12e69ed288812ce7ffc7bc -> 806f70ad3cce5fc849f6d06a08ace8a95f92a1ea84a2418fddabb1eef5846594."
|
||||
"_rebaselined_receiver_owner_2701": "#2701: every non-arrow function form now carries a `@receiver-owner.this` marker on the same node as `@scope.function`, so a scope that BINDS its own `this` can stop the receiver walk (`Scope.ownsReceivers`). Verified before re-baselining by diffing the capture-name histogram over this same fixture corpus against 1d3088173f6f93827641b476d614d5d15cd4f3ea: the ONLY delta is @receiver-owner.this (typescript +143, javascript +32) \u2014 every other capture count is byte-identical, so no existing capture moved. Prior f1ccf42a36895c8e34dcb724286f247d469835f2dcbb23ad3347190adc7fde1c -> 90601494695b834d3a9af7ac4844eac603f4f432809a05554cc59de0674a4354.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 1c71ef628eb75a3b111afa8c2a7c351c16a7f5aab9fac2f098f82b2866312aa8 -> 83344b7cba093702f4528eeee44e438809c229d43b12e69ed288812ce7ffc7bc.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior 83344b7cba093702f4528eeee44e438809c229d43b12e69ed288812ce7ffc7bc -> 806f70ad3cce5fc849f6d06a08ace8a95f92a1ea84a2418fddabb1eef5846594.",
|
||||
"_rebaselined_blind_spots_2856": "#2856 blind-spots series: the JS/TS SCOPE queries gained capture rules, so fingerprint drift is expected and additive. Verified before re-baselining by diffing the capture-name sets in both scope queries against origin/main: TypeScript gained exactly @reference.read.identifier (A2 bare-identifier reads in value positions) and @reference.type (R2-2 type references, so a declared contract stops reporting incoming:{}); JavaScript gained exactly @reference.read.identifier, @reference.read.destructured (R2-1c) and @reference.write.property-key (R2-1b record-construction writes). NOTHING was removed on either side \u2014 the delta is a pure superset, which is the check that no existing capture moved. capture_groups_small/large are unchanged (4503/14403) because those measure the SYNTHETIC scaling source, which this branch does not touch; only the fixture-corpus count moves. capture_groups_fp 2097 -> 2338 and fixture_count 146 -> 151 from 21 new lang-resolution fixtures. Scaling stayed linear and inside budget: typescript 1.116 < 1.5, javascript 1.010 < 1.5. Prior typescript ed92588e0fc7b28b3a0174339ac378b4dd85965fe007db1208dea97a65ce0571 -> f66a3e6f1e096431e7046505129a627deaa00ca0de5bc846b080591b397248f7; prior javascript 806f70ad3cce5fc849f6d06a08ace8a95f92a1ea84a2418fddabb1eef5846594 -> 2026993b81b873839dd2ef8797d9c14d9c48516b2b57b05ac17d8d43f2f4eba3."
|
||||
},
|
||||
"kotlin": {
|
||||
"fingerprint": "efd5dbf80ffcd3bab2834d1010f6fe2b239dcc5d58229938dea9cff8d0f380f2",
|
||||
"fingerprint": "a184f8ff0ae40d246db855b63f7ff26bda3afac03e5f4c76e4593c7e2cefce54",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined_callable_flow_2522_review": "PR #2522 review hardening: callable operands retain expression/qualified identity and formals retain signature metadata. Prior bddba25d5a88152bbbee8d70e82c944b5302accb4b625df782adb1d4f7a7ac12 -> e856951c2a779163d555dadc8e1bf59304a86caed78ac1f450d9caa2b50f63d1; scaling 1.090 < 1.5.",
|
||||
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: Kotlin callable-reference flow facts with invocation-result suppression. Prior 4900431791f2b9280009deb2b82659c26ead8aa6fb8731190a7c505dec5a9041 -> bddba25d5a88152bbbee8d70e82c944b5302accb4b625df782adb1d4f7a7ac12; scaling 0.880 < 1.5.",
|
||||
"_added": "#1951: bench coverage added (was ungated); scale source heritage-bearing (: Base()); js/kotlin O(n^2) findNodeAtRange-per-match fixed to threaded captured node, now linear.",
|
||||
"_rebaselined": "#1919 review CF3 fix: extended kotlin-local-property-owner (init/accessor destructuring) + new dart-accessor-owner fixture (getter/setter ownership). Fingerprint-only corpus drift; scaling ~1.0.",
|
||||
"_rebaselined_2271": "PR #2271: re-vendored tree-sitter-kotlin 0.3.8 -> unreleased fwcd main c8ac3d26 for `fun interface` support + new kotlin-fun-interface fixture in the corpus. Drift is both corpus-additive (the fixture) and grammar-driven (the new grammar parses `fun interface` as a class_declaration, not an ERROR node). Baselined to the NEW grammar's fingerprint, so this --check passes only once the regenerated prebuilds land — until then CI loads the committed 0.3.8 binary and the bench is red, same as the kotlin fun-interface integration tests. scaling ~0.83 (linear).",
|
||||
"_rebaselined_2271": "PR #2271: re-vendored tree-sitter-kotlin 0.3.8 -> unreleased fwcd main c8ac3d26 for `fun interface` support + new kotlin-fun-interface fixture in the corpus. Drift is both corpus-additive (the fixture) and grammar-driven (the new grammar parses `fun interface` as a class_declaration, not an ERROR node). Baselined to the NEW grammar's fingerprint, so this --check passes only once the regenerated prebuilds land \u2014 until then CI loads the committed 0.3.8 binary and the bench is red, same as the kotlin fun-interface integration tests. scaling ~0.83 (linear).",
|
||||
"_rebaselined_2522_review_fixes": "PR #2522 review fixes: fieldless assignment nodes decomposed positionally. Prior e856951c2a779163d555dadc8e1bf59304a86caed78ac1f450d9caa2b50f63d1 -> 4b31f46cfb004ba769a96feeb06ae4ef109c77410f54e7aaab4a688df599b112; scaling ratio re-verified within budget.",
|
||||
"_rebaselined_2550_instance_model": "PR #2549 (#2545): anonymous object expressions (object_literal) emit @scope.class, and the kotlin-object-literal-scope fixture joined the corpus. Prior 4b31f46cfb004ba769a96feeb06ae4ef109c77410f54e7aaab4a688df599b112 -> a6fce0dff00e88d41d85023eaf3f35016b5217c7e5225f24a598e4c70bb63091; scaling 0.951 < 1.5.",
|
||||
"_rebaselined_2563_instance_ownership": "#2563: kotlin-instance-ownership adds unrelated, inherited, outer-instance, and anonymous-object coverage. Prior a6fce0dff00e88d41d85023eaf3f35016b5217c7e5225f24a598e4c70bb63091 -> 9f159f8810d342ef1c821f466efd6920dad9a190f06000056e6cd2815861b195; scaling 1.257 < 1.5.",
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged — the tag is added to existing call matches, never a new match — so this is digest drift only. Prior 9f159f8810d342ef1c821f466efd6920dad9a190f06000056e6cd2815861b195 -> d3c4d2fa0d82d248a2299cfc888b067187ad1faf2c87a97f93c6ed835eefc3f1.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|…` instead of `1|…`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior d3c4d2fa0d82d248a2299cfc888b067187ad1faf2c87a97f93c6ed835eefc3f1 -> c1f0cc9058ab11b7cd6fc8b440deb6db2b2f530f2eb21178923e68a3d0796c4b.",
|
||||
"_rebaselined_2766_await_subscript_emission": "#2766: extractMixedChain now walks THROUGH await and subscript nodes and peels transparent wrappers at loop entry, so sites whose receiver is `repos[0]` or `(await f())` mint a receiver chain where they previously minted none. EMISSION CHANGE: more sites carry `@reference.receiver-chain`; no existing chain changed shape. Only go and kotlin drifted of 15 — the two whose fixture corpora contain such receivers. Prior c1f0cc9058ab11b7cd6fc8b440deb6db2b2f530f2eb21178923e68a3d0796c4b -> efd5dbf80ffcd3bab2834d1010f6fe2b239dcc5d58229938dea9cff8d0f380f2."
|
||||
"_rebaselined_receiver_chain_2747": "#2747 receiver-chain rollout: call matches whose receiver is itself an expression now carry `@reference.receiver-chain`, a compact encoding of the receiver's structure, so resolution types it by folding instead of re-parsing receiver source text. Capture GROUP counts are unchanged \u2014 the tag is added to existing call matches, never a new match \u2014 so this is digest drift only. Prior 9f159f8810d342ef1c821f466efd6920dad9a190f06000056e6cd2815861b195 -> d3c4d2fa0d82d248a2299cfc888b067187ad1faf2c87a97f93c6ed835eefc3f1.",
|
||||
"_rebaselined_2766_receiver_chain_wire_v2": "#2766: receiver-chain wire format v1 -> v2 (name-free `await` / `index` step kinds). The VERSION prefix is part of every emitted `@reference.receiver-chain` capture, so every chain-minting language's capture text changed. WIRE-FORMAT CHANGE, NOT A CAPTURE-SET CHANGE: the same chains are minted for the same sites, spelled `2|\u2026` instead of `1|\u2026`. Exactly the 12 chain-minting languages drifted; c, cobol and dart did not, which is the check that this is the prefix and not a capture regression. Accompanied by SCHEMA_BUMP 34 -> 37 and INCREMENTAL_SCHEMA_VERSION 28 -> 31 so a stale index is rejected rather than replaying chains a v2 decoder refuses. Prior d3c4d2fa0d82d248a2299cfc888b067187ad1faf2c87a97f93c6ed835eefc3f1 -> c1f0cc9058ab11b7cd6fc8b440deb6db2b2f530f2eb21178923e68a3d0796c4b.",
|
||||
"_rebaselined_2766_await_subscript_emission": "#2766: extractMixedChain now walks THROUGH await and subscript nodes and peels transparent wrappers at loop entry, so sites whose receiver is `repos[0]` or `(await f())` mint a receiver chain where they previously minted none. EMISSION CHANGE: more sites carry `@reference.receiver-chain`; no existing chain changed shape. Only go and kotlin drifted of 15 \u2014 the two whose fixture corpora contain such receivers. Prior c1f0cc9058ab11b7cd6fc8b440deb6db2b2f530f2eb21178923e68a3d0796c4b -> efd5dbf80ffcd3bab2834d1010f6fe2b239dcc5d58229938dea9cff8d0f380f2.",
|
||||
"capture_groups_small": 4753,
|
||||
"capture_groups_large": 15203,
|
||||
"capture_groups_fp": 2334,
|
||||
"fixture_count": 137
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -209,10 +209,22 @@ const LANGS = [
|
|||
// Heritage-bearing: `: public Base, public Mixin` (single + multiple
|
||||
// inheritance) drives emitCppInheritanceCaptures (#1951) at scale. Added
|
||||
// (was unbenched); adding it exposed + fixed the same O(n²) root-walk (#1956).
|
||||
//
|
||||
// Also GENERIC-MEMBER-bearing (#2833): `Repo<Entity_n> repo;` is a member
|
||||
// whose declared type is a bare `template_type`, and
|
||||
// `std::vector<Entity_n> items;` is the far commoner spelling where a
|
||||
// `qualified_identifier` WRAPS that template_type. Both were absent, and
|
||||
// their absence is why two successive rounds of `field_declaration`
|
||||
// type-binding rules landed with a byte-identical cpp fingerprint: the gate
|
||||
// could not see a member field it had no instance of. With them present,
|
||||
// reverting either round of rules drifts the fingerprint, which is the
|
||||
// property that makes the gate worth running.
|
||||
header:
|
||||
'#include <string>\n\nclass Base {\n public:\n long baseId() const { return 0; }\n};\n\nclass Mixin {\n public:\n void mix() {}\n};\n\n',
|
||||
'#include <string>\n#include <vector>\n\ntemplate <typename T>\nclass Repo {\n public:\n void save(T v) {}\n};\n\nclass Base {\n public:\n long baseId() const { return 0; }\n};\n\nclass Mixin {\n public:\n void mix() {}\n};\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} : public Base, public Mixin {\n public:\n long id;\n std::string name;\n` +
|
||||
` Repo<Entity${n}> repo;\n` +
|
||||
` std::vector<Entity${n}> items;\n` +
|
||||
` long getId() const { return id; }\n` +
|
||||
` void setName(std::string v) { name = v; }\n};\n\n`,
|
||||
},
|
||||
|
|
@ -255,14 +267,15 @@ const LANGS = [
|
|||
fixturePrefix: 'java',
|
||||
exts: ['.java'],
|
||||
file: 'bench.java',
|
||||
// Java was previously unbenched. Heritage-bearing: extends Base + implements
|
||||
// Marker (both forms) so the @reference.inherits synth (#1951) is driven at scale.
|
||||
// Java was previously unbenched. Class and record heritage both implement
|
||||
// Marker so the @reference.inherits synth (#1951, #2900) is driven at scale.
|
||||
header: 'package generated;\n\nclass Base {}\n\ninterface Marker {}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base implements Marker {\n` +
|
||||
` long id = 0L;\n String name = "";\n` +
|
||||
` public long getId() { return this.id; }\n` +
|
||||
` public void setName(String v) { this.name = v; }\n}\n\n`,
|
||||
` public void setName(String v) { this.name = v; }\n}\n\n` +
|
||||
`record RecordEntity${n}(long id) implements Marker {}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'java-local-types',
|
||||
|
|
|
|||
|
|
@ -503,7 +503,7 @@ function buildStaleIndexHint(gitNexusDir, cwd) {
|
|||
|
||||
if (currentHead === lastCommit) return '';
|
||||
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings, indexOnly: true });
|
||||
return (
|
||||
`[GitNexus] index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
`Run \`${analyzeCmd}\` to refresh the knowledge graph.`
|
||||
|
|
|
|||
|
|
@ -523,7 +523,7 @@ function handlePostToolUse(input) {
|
|||
// If HEAD matches last indexed commit, no reindex needed
|
||||
if (currentHead && currentHead === lastCommit) return;
|
||||
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings, indexOnly: true });
|
||||
sendHookResponse(
|
||||
'PostToolUse',
|
||||
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
|
|
|
|||
|
|
@ -276,7 +276,13 @@ function formatBunxCommand(gitnexusArgs) {
|
|||
}
|
||||
|
||||
function formatAnalyzeCommand(options = {}, deps = {}) {
|
||||
const suffix = options.embeddings ? ' --embeddings' : '';
|
||||
// `--index-only` is what a routine "your index is stale" nudge wants: it
|
||||
// reindexes without rewriting AGENTS.md / CLAUDE.md / skills, so an agent
|
||||
// following the nudge on every commit cannot churn the tracked agent guides
|
||||
// (#2907). Callers that actually want the docs refreshed omit it.
|
||||
const suffix = `${options.indexOnly ? ' --index-only' : ''}${
|
||||
options.embeddings ? ' --embeddings' : ''
|
||||
}`;
|
||||
// Keep the stale-index hook budget tight by querying each tool at most once.
|
||||
// The memoized `probe` is a spawn-free PATH scan (resolveOnPath) shared with
|
||||
// resolveInvocationMode, so `gitnexus` is scanned only once and no subprocess
|
||||
|
|
|
|||
68
gitnexus/package-lock.json
generated
68
gitnexus/package-lock.json
generated
|
|
@ -10,7 +10,7 @@
|
|||
"hasInstallScript": true,
|
||||
"license": "PolyForm-Noncommercial-1.0.0",
|
||||
"dependencies": {
|
||||
"@ladybugdb/core": "^0.18.3",
|
||||
"@ladybugdb/core": "^0.19.0",
|
||||
"@modelcontextprotocol/sdk": "^1.0.0",
|
||||
"@scarf/scarf": "^1.4.0",
|
||||
"busboy": "^1.6.0",
|
||||
|
|
@ -79,7 +79,7 @@
|
|||
"version": "1.0.0",
|
||||
"dev": true,
|
||||
"devDependencies": {
|
||||
"typescript": "^6.0.3"
|
||||
"typescript": "^7.0.2"
|
||||
}
|
||||
},
|
||||
"node_modules/@babel/code-frame": {
|
||||
|
|
@ -1254,9 +1254,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/@ladybugdb/core": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core/-/core-0.18.3.tgz",
|
||||
"integrity": "sha512-XjpPKW4MrL28D2gYGTZuIjiEcPx12L21lx58QggrdrItw8o/e9Lmg/Ejoo4Kz08lZj+rIcC1Fu9thzIYOTUlJw==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core/-/core-0.19.1.tgz",
|
||||
"integrity": "sha512-8W2g6xUi4jm96fs4EayyMcsvEEtIb8vboZhw9/YG98881cIcmZjmqAN91XGUp4vb8NqoFr3Wp7wcu3dqJk0b7w==",
|
||||
"hasInstallScript": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
|
|
@ -1265,17 +1265,17 @@
|
|||
"node-addon-api": "^6.0.0"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@ladybugdb/core-darwin-arm64": "0.18.3",
|
||||
"@ladybugdb/core-darwin-x64": "0.18.3",
|
||||
"@ladybugdb/core-linux-arm64": "0.18.3",
|
||||
"@ladybugdb/core-linux-x64": "0.18.3",
|
||||
"@ladybugdb/core-win32-x64": "0.18.3"
|
||||
"@ladybugdb/core-darwin-arm64": "0.19.1",
|
||||
"@ladybugdb/core-darwin-x64": "0.19.1",
|
||||
"@ladybugdb/core-linux-arm64": "0.19.1",
|
||||
"@ladybugdb/core-linux-x64": "0.19.1",
|
||||
"@ladybugdb/core-win32-x64": "0.19.1"
|
||||
}
|
||||
},
|
||||
"node_modules/@ladybugdb/core-darwin-arm64": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-darwin-arm64/-/core-darwin-arm64-0.18.3.tgz",
|
||||
"integrity": "sha512-DGZTOlvSS4esEb1vTekY5IDoAvZAeYzR5cXVkECtQj9BVkk05zsvCAdTPo1Rz1BuI0qvqUVF+2WlIerI67iA2g==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-darwin-arm64/-/core-darwin-arm64-0.19.1.tgz",
|
||||
"integrity": "sha512-VGQs1NThAygMsoOlxud05pqKA9xfUptl55iYkwvW45As5MSI7+M86WN0Pp0VdPEfw8vNQJehrlHR5LvVAuWc2Q==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
|
|
@ -1286,9 +1286,9 @@
|
|||
]
|
||||
},
|
||||
"node_modules/@ladybugdb/core-darwin-x64": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-darwin-x64/-/core-darwin-x64-0.18.3.tgz",
|
||||
"integrity": "sha512-Qp6j0CM/orBlK6KD0p/s4ofkIhNUwi1hdCgMw+fj81UHugWHkVLiYV4grRBdHhyplw+snchZpTxvfpxFbkG1Cw==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-darwin-x64/-/core-darwin-x64-0.19.1.tgz",
|
||||
"integrity": "sha512-CGfM6ostxDS5jztxwjkXXtxrjMDgsFMRoyr5HZDCFw1+iXC1rIzmK/Y7RIw+KbQ49aPzSmkhBC447mFviBJxoA==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
|
|
@ -1299,9 +1299,9 @@
|
|||
]
|
||||
},
|
||||
"node_modules/@ladybugdb/core-linux-arm64": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-linux-arm64/-/core-linux-arm64-0.18.3.tgz",
|
||||
"integrity": "sha512-F9miYjBuS43I7uNG199FNMqwdHJ98WA6dU3v2SZCeLXmXCdRzmYcuHQWlbNr2Tba9CX58w2XvBZoUaXZKJ/yKQ==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-linux-arm64/-/core-linux-arm64-0.19.1.tgz",
|
||||
"integrity": "sha512-BZUQwlkvNXENc5GVyXdfRF0Dv9JX8XMlcdMMiB5GKrEhTCpajQ3D58woHPVvn0JEjw7Ms3tHo6kXUAMZKYXIVg==",
|
||||
"cpu": [
|
||||
"arm64"
|
||||
],
|
||||
|
|
@ -1312,9 +1312,9 @@
|
|||
]
|
||||
},
|
||||
"node_modules/@ladybugdb/core-linux-x64": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-linux-x64/-/core-linux-x64-0.18.3.tgz",
|
||||
"integrity": "sha512-AfG5RDp/f/IDctDMpTAT5+2MYNtlWT191xiQNjSaWD4X85DhY3Dzps8Qu5VteIAPih5d6mmoaKGs8q0XIjfkFA==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-linux-x64/-/core-linux-x64-0.19.1.tgz",
|
||||
"integrity": "sha512-LDx+E1UHlmNXSb3F9QmvdBgZGfB3wI/DcrHzfOwXgT3BP8C4ScB2tZdpiYQiuPp8MiSZ9kuuGqos8A4tQKQu8Q==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
|
|
@ -1325,9 +1325,9 @@
|
|||
]
|
||||
},
|
||||
"node_modules/@ladybugdb/core-win32-x64": {
|
||||
"version": "0.18.3",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-win32-x64/-/core-win32-x64-0.18.3.tgz",
|
||||
"integrity": "sha512-bHuFk0m9cnq0WGd9I4D8or8g6cC/BS58iatMtilqM3JpDPIQIFk6MQl6exL7P4xyWbkLwQgsrv2ToDnyoQNKvg==",
|
||||
"version": "0.19.1",
|
||||
"resolved": "https://registry.npmjs.org/@ladybugdb/core-win32-x64/-/core-win32-x64-0.19.1.tgz",
|
||||
"integrity": "sha512-2spst1g+Z050Fz/5z7Pc6Fuc5dVXLzekOuWW4lP+mEGCL3tkv3QWYxk37DiFy+O9fDFxiWK2f3aauab58/f9kQ==",
|
||||
"cpu": [
|
||||
"x64"
|
||||
],
|
||||
|
|
@ -1939,9 +1939,9 @@
|
|||
"license": "MIT"
|
||||
},
|
||||
"node_modules/@types/node": {
|
||||
"version": "26.1.2",
|
||||
"resolved": "https://registry.npmjs.org/@types/node/-/node-26.1.2.tgz",
|
||||
"integrity": "sha512-Vu4a5UFA9rIIFJ7rB/Vaafh9lrCQszopTCx6KjFboXTGQbPNasehVR5TEiithSDGyd1DEiUByggTZsg8jukeIg==",
|
||||
"version": "26.2.0",
|
||||
"resolved": "https://registry.npmjs.org/@types/node/-/node-26.2.0.tgz",
|
||||
"integrity": "sha512-5IviulTZeRNp2vAJ514cc/HUlY5nZ9fCbq9DMyC52BrhFZACo3nI0R7qBxhQmo/d27NFe96ur/b7Wwxklda+kg==",
|
||||
"devOptional": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
|
|
@ -3018,9 +3018,9 @@
|
|||
}
|
||||
},
|
||||
"node_modules/express-rate-limit": {
|
||||
"version": "8.6.1",
|
||||
"resolved": "https://registry.npmjs.org/express-rate-limit/-/express-rate-limit-8.6.1.tgz",
|
||||
"integrity": "sha512-0D493aP61w0TJ2A0wy27riRsO7FMQ7FK+KUHOKCSfPvYo0R55aiC6emCVgFUeShH0fq0ICPVzNcgoS+BsbXQCA==",
|
||||
"version": "8.6.2",
|
||||
"resolved": "https://registry.npmjs.org/express-rate-limit/-/express-rate-limit-8.6.2.tgz",
|
||||
"integrity": "sha512-YH4ru+eOJxQABscKFfRCy9R7x9QFGdezclVMwwgFFndzS2Xnm0uo6B0ABZsLhcpeptGv2qvuJVWlQr9gQZoC3A==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"debug": "^4.4.3",
|
||||
|
|
@ -5430,9 +5430,9 @@
|
|||
"license": "0BSD"
|
||||
},
|
||||
"node_modules/tsx": {
|
||||
"version": "4.23.4",
|
||||
"resolved": "https://registry.npmjs.org/tsx/-/tsx-4.23.4.tgz",
|
||||
"integrity": "sha512-ZiUQ8oT/KzN51mJUWPqARYqwFLFJZtGZipRkw1ynHMr9vy3eU77m5yfF3Gzm6meEg/beW+lUu3fHYgskTN2oVQ==",
|
||||
"version": "4.23.12",
|
||||
"resolved": "https://registry.npmjs.org/tsx/-/tsx-4.23.12.tgz",
|
||||
"integrity": "sha512-FDf4L4sYzKtzWYhU/Xm0AQFdTjdIxNo9ElTf2mxXM6k8YMHXzYUe4yODVaXP4V9uMFbVg8c0qyBccK2OOxb45Q==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
|
|
|
|||
|
|
@ -56,7 +56,7 @@
|
|||
"version": "node scripts/sync-plugin-manifests.mjs"
|
||||
},
|
||||
"dependencies": {
|
||||
"@ladybugdb/core": "^0.18.3",
|
||||
"@ladybugdb/core": "^0.19.0",
|
||||
"@modelcontextprotocol/sdk": "^1.0.0",
|
||||
"@scarf/scarf": "^1.4.0",
|
||||
"busboy": "^1.6.0",
|
||||
|
|
|
|||
|
|
@ -47,6 +47,13 @@ export const WINDOWS_WEIGHTS_SEC: Readonly<Record<string, number>> = {
|
|||
'test/integration/cli-e2e.test.ts': 361,
|
||||
'test/integration/worker-pool.test.ts': 222,
|
||||
'test/unit/incremental-vector-extension-ordering.test.ts': 87,
|
||||
// ESTIMATE, not a measurement (#2841): this suite drives more full
|
||||
// `runFullAnalysis` cycles than the VECTOR sibling above, so the 8 s
|
||||
// PER_FILE_OVERHEAD floor would badly under-charge it and skew the Windows
|
||||
// split — the failure mode that produced the job timeouts this table exists
|
||||
// to prevent. Scaled from the sibling's measured 87 s by analyze-run count.
|
||||
// Replace with a real figure after the first green Windows matrix run.
|
||||
'test/unit/incremental-index-extension-dml-gate.test.ts': 180,
|
||||
'test/integration/cli-limit-e2e.test.ts': 75,
|
||||
'test/unit/hooks.test.ts': 26,
|
||||
'test/integration/analyze-heap-oom-e2e.test.ts': 23,
|
||||
|
|
|
|||
|
|
@ -36,6 +36,16 @@ const PLATFORM_LOGIC = [
|
|||
// must exercise the Windows backslash branch, so run it on the OS matrix (#2394).
|
||||
'test/unit/cli-entry.test.ts',
|
||||
'test/unit/platform-capabilities.test.ts',
|
||||
// The gitnexus-plan safe writer resolves every name through a per-platform
|
||||
// backend: Linux anchors through /proc/self/fd, macOS resolves lexically and
|
||||
// verifies each step against descriptors it holds open. Publication is link(2)
|
||||
// on both. #2905 shipped the Darwin backend after the suite had silently
|
||||
// skipped on every non-Linux runner, so this file must run on the OS matrix or
|
||||
// the macOS half is unverified by construction — and the flag, trailing-
|
||||
// separator and hard-link fixtures assert kernel behaviour that only a real
|
||||
// Darwin kernel can confirm. Windows is refused by the capability gate; the
|
||||
// suite asserts that refusal rather than skipping it.
|
||||
'test/unit/evidence-provenance-helper.test.ts',
|
||||
// Windows drive-letter case variance in the analyzer runner-identity path
|
||||
// fields (#2668): normalizeAnalyzerRootPath is a POSIX no-op, so the
|
||||
// "identity path fields are normalizer-stable" fixpoint guard only bites on
|
||||
|
|
@ -154,6 +164,18 @@ const LBUG_NATIVE = [
|
|||
// proven on the windows-latest native addon, not just Ubuntu. Budget: ~25s
|
||||
// on Linux → expect ~2min on the slowest Windows shard.
|
||||
'test/unit/incremental-vector-extension-ordering.test.ts',
|
||||
// #2841: the FTS half of that same gate, plus the both-extensions-blocked
|
||||
// case — and it needs this matrix for two reasons the VECTOR sibling above
|
||||
// does not cover. The reported failure environment is a machine where the
|
||||
// extension stopped LOADING, which is the #2374 class and Windows-reported
|
||||
// (the same reason fts-extension-e2e.test.ts is registered below), so the
|
||||
// FTS-unavailable branch has to run on a real Windows/macOS runner rather
|
||||
// than only on Ubuntu where FTS always loads. And its both-blocked case is
|
||||
// gated on GITNEXUS_REQUIRE_VECTOR=1, which ci-tests.yml sets ONLY on this
|
||||
// job — everywhere else an unavailable VECTOR extension skips instead of
|
||||
// failing. Budget: four real analyze runs, so expect it to sit alongside the
|
||||
// VECTOR sibling's ~87s Windows measurement.
|
||||
'test/unit/incremental-index-extension-dml-gate.test.ts',
|
||||
];
|
||||
|
||||
// Process spawning and CLI tests — exercise child_process with real
|
||||
|
|
|
|||
|
|
@ -60,7 +60,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
|||
| Flag | Effect |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| `--force` | Force full regeneration |
|
||||
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
|
|
|
|||
|
|
@ -53,6 +53,14 @@ description: "Use when the user wants to know what will break if they change som
|
|||
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||
| >15 symbols or many processes | HIGH |
|
||||
| Critical path (auth, payments) | CRITICAL |
|
||||
| **Zero callers found** | **UNKNOWN** |
|
||||
|
||||
`UNKNOWN` is not a low rung on this scale — it means the walk could not answer.
|
||||
An empty caller set is equally consistent with "genuinely unused" and "the
|
||||
callers are not resolvable by the index" (plain-object property access, dynamic
|
||||
dispatch, cross-language calls), so few-callers ⇒ LOW does **not** apply. The
|
||||
result carries a `riskNote` saying so. Confirm with a text search before
|
||||
treating the symbol as safe to change or delete.
|
||||
|
||||
## Tools
|
||||
|
||||
|
|
@ -84,6 +92,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth, and reads like
|
||||
`UNKNOWN` above: a zero there means unseen, not unaffected. Re-run it rather
|
||||
than tick the pre-commit check.
|
||||
|
||||
## Example: "What breaks if I change validateUser?"
|
||||
|
||||
```
|
||||
|
|
|
|||
|
|
@ -124,12 +124,17 @@ phase that needs them.
|
|||
statement-level claims (never reconstructs fake edges).
|
||||
- No GitNexus at all → fallback mode: targeted grep/read exploration, findings
|
||||
labelled **source-derived**, with a recommendation to index.
|
||||
- Reading or publishing a plan requires Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
and `O_NOFOLLOW`; publication also requires a validated absolute Python 3
|
||||
PATH candidate with libc `renameat2(RENAME_NOREPLACE)` support, a
|
||||
writable target repository, and a shared filesystem for the plan and
|
||||
Git-admin vault. The writer fails closed when those guarantees are
|
||||
unavailable; it never redirects the plan elsewhere.
|
||||
- Reading or publishing a plan requires `O_DIRECTORY` and `O_NOFOLLOW`, plus
|
||||
`/proc/self/fd` on Linux; every other platform is refused. No interpreter is
|
||||
spawned and no native code is loaded. Publication is `link(2)`, which fails
|
||||
rather than replaces when the destination name is taken. Linux resolves every
|
||||
name against a held descriptor, so a parent swapped mid-write cannot redirect
|
||||
the operation; macOS has no equivalent path and instead pins each directory
|
||||
with an open descriptor and re-proves the chain either side of every step,
|
||||
which detects such a swap and aborts. Publishing also needs a writable target
|
||||
repository and a shared filesystem for the plan and Git-admin vault. The
|
||||
writer fails closed when those guarantees are unavailable; it never redirects
|
||||
the plan elsewhere.
|
||||
|
||||
## Limitations
|
||||
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -87,6 +87,11 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth: a short or empty
|
||||
list is not proof that only the expected files changed. Re-run it rather than
|
||||
treat the refactor as verified.
|
||||
|
||||
**cypher** — custom reference queries:
|
||||
|
||||
```cypher
|
||||
|
|
|
|||
|
|
@ -216,7 +216,10 @@ Work through plan §7 step by step, in order. For each step:
|
|||
`detect_changes` → commit as one unbroken sequence from the repository
|
||||
root — interleaving other work between the gate and the commit is how
|
||||
the gate gets skipped. Unexpected
|
||||
affected flows → investigate before committing, not after.
|
||||
affected flows → investigate before committing, not after. A result
|
||||
flagged `partial` (a graph query failed) or `truncated` (the symbol
|
||||
listing was capped) blocks the commit the same way: the gate did not
|
||||
see every changed symbol, so re-run it rather than read it as clean.
|
||||
|
||||
A relationship-affecting implementation edit or commit invalidates the
|
||||
procedure's prior proof. The next step must perform the required inter-step
|
||||
|
|
|
|||
|
|
@ -98,8 +98,11 @@ excluded.
|
|||
|
||||
## Safe existing-plan read contract
|
||||
|
||||
`read-plan` fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`, and
|
||||
`O_NOFOLLOW` are available. It resolves the exact Git top-level, opens the
|
||||
`read-plan` fails closed unless the host platform can resolve names against a
|
||||
held directory descriptor: Linux `/proc/self/fd` with `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, or macOS `O_DIRECTORY`/`O_NOFOLLOW`. Every other platform is
|
||||
refused outright — an unverified read is not a degraded read, it is a different,
|
||||
racy operation. It resolves the exact Git top-level, opens the
|
||||
repository root and every plan parent as held no-follow directory descriptors,
|
||||
rejects missing, symlink, non-directory, and escaping parents, and opens the
|
||||
leaf with `O_NOFOLLOW`. It reads at most 16 MiB from that held file descriptor,
|
||||
|
|
@ -109,13 +112,17 @@ Neither Deepen nor work may parse bytes obtained before or outside this receipt.
|
|||
|
||||
## Safe generated-plan write contract
|
||||
|
||||
The writer fails closed unless Linux `/proc/self/fd`, `O_DIRECTORY`,
|
||||
`O_NOFOLLOW`, and Python 3 with libc `renameat2(RENAME_NOREPLACE)` support are
|
||||
available. Python may live in `/usr/local`, a Nix profile, or another absolute
|
||||
PATH directory, but the helper accepts only a resolved executable and
|
||||
containing directory owned by root or the current user and not writable by
|
||||
group/other. The resolved executable is opened without following links and
|
||||
invoked through that held descriptor. Relative PATH entries are ignored. The plan parent and the
|
||||
The writer fails closed unless the host platform offers `O_DIRECTORY` and
|
||||
`O_NOFOLLOW`, plus `/proc/self/fd` on Linux. It spawns no interpreter and loads
|
||||
no native code: publication is `link(2)`, which is atomic, fails `EEXIST` when
|
||||
the destination name is taken, and refuses a symlinked destination without
|
||||
following it — the same no-replace guarantee `renameat2(RENAME_NOREPLACE)` and
|
||||
`renameatx_np(RENAME_EXCL)` provide, available through `fs.linkSync` on every
|
||||
supported platform. The temporary name is unlinked once the link succeeds; the
|
||||
published file is the same inode the writer created and verified, so every
|
||||
identity check downstream holds by construction. A link that succeeds followed
|
||||
by an unlink that fails leaves the plan published and is reported as success,
|
||||
because it is one. The plan parent and the
|
||||
repository's Git-admin directory must also share a filesystem. It resolves
|
||||
the target repository's exact Git top-level, opens that root and every
|
||||
destination parent as held no-follow directory descriptors, creates missing
|
||||
|
|
@ -128,15 +135,45 @@ The writer creates a random exclusive temporary file relative to the held final
|
|||
parent descriptor and keeps its no-follow descriptor open. It writes and
|
||||
flushes the bytes, binds the temporary name to the opened inode, and hashes the
|
||||
open file before publication. Immediately before publication it revalidates
|
||||
the parent and the temporary path, inode, size, and digest. Publication uses an
|
||||
atomic no-replace move relative to the held directory descriptor. Initial mode
|
||||
therefore cannot overwrite a destination that appears after the absent check.
|
||||
the parent and the temporary path, inode, size, and digest. Publication links
|
||||
the temporary name to the destination relative to the held directory
|
||||
descriptor, which fails rather than replaces if the destination is taken.
|
||||
Initial mode therefore cannot overwrite a destination that appears after the
|
||||
absent check.
|
||||
The writer then flushes the directory and revalidates the committed path by
|
||||
opening it with `O_NOFOLLOW`, hashing both the original temporary fd and the
|
||||
path-bound fd, and performing a second descriptor-anchored path identity check
|
||||
after hashing. A detected mutation or replacement aborts instead of accepting
|
||||
mixed-era output.
|
||||
|
||||
### Linux anchors, macOS verifies
|
||||
|
||||
The two platforms reach the same destination by different proofs, and the
|
||||
difference is real enough to state rather than smooth over.
|
||||
|
||||
On Linux every name resolves through `/proc/self/fd/<fd>/<child>`, a magic link
|
||||
the kernel resolves against the inode the descriptor already holds. The names
|
||||
above it are never re-walked, so an attacker who renames a parent between the
|
||||
check and the use cannot redirect the operation. The race is impossible, not
|
||||
merely detected.
|
||||
|
||||
macOS has no such path. `/dev/fd/<fd>` is a devfs node, not a magic link: it can
|
||||
be opened, but nothing can be resolved through it. `open("/dev/fd/<fd>/child")`
|
||||
returns `ENOENT`, and `realpath` of it returns `/dev/fd/<fd>` rather than the
|
||||
directory's path — measured on macOS 26, not inferred. Node exposes no `openat`,
|
||||
no `dir_fd` parameter, and no FFI, so on macOS the writer resolves names
|
||||
lexically with `O_NOFOLLOW` at every component, holds an open descriptor on
|
||||
every directory in the chain for the whole operation, and proves before *and*
|
||||
after each step that the chain still names exactly the inodes it is holding.
|
||||
Holding the descriptors is what makes the recorded inode numbers trustworthy:
|
||||
an open descriptor pins its inode, so a freed number cannot be recycled beneath
|
||||
the walk.
|
||||
|
||||
What that buys is detection rather than prevention. A parent swapped inside the
|
||||
window between a check and its use is caught by the check that follows, and the
|
||||
operation aborts having written nothing — but on Linux it could not have
|
||||
happened at all. No published byte escapes verification on either platform.
|
||||
|
||||
`--replace` accepts only a pre-existing regular file and is reserved for
|
||||
Deepen; without it, accidental overwrite is rejected. It also requires the
|
||||
exact canonical `generated_plan_path` and `plan_digest` from the same session's
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -9,7 +9,7 @@
|
|||
import fs from 'fs/promises';
|
||||
import path from 'path';
|
||||
import { fileURLToPath } from 'url';
|
||||
import { type GeneratedSkillInfo } from './skill-gen.js';
|
||||
import { type GeneratedSkillInfo } from './generated-skill.js';
|
||||
import { STANDARD_SKILL_CATALOG } from './standard-skills.js';
|
||||
import { logger } from '../core/logger.js';
|
||||
|
||||
|
|
@ -157,7 +157,10 @@ export function generateGitNexusContent(
|
|||
? generatedSkills
|
||||
.map(
|
||||
(s) =>
|
||||
`| Work in the ${s.label} area (${s.symbolCount} symbols) | \`.claude/skills/${s.name}/SKILL.md\` |`,
|
||||
// The per-cluster count is as volatile as the header parenthetical,
|
||||
// so --no-stats drops it too (#2907) — otherwise the flag that
|
||||
// promises "omit volatile symbol counts" left a churning one behind.
|
||||
`| Work in the ${s.label} area${noStats ? '' : ` (${s.symbolCount} symbols)`} | \`.claude/skills/${s.name}/SKILL.md\` |`,
|
||||
)
|
||||
.join('\n')
|
||||
: '';
|
||||
|
|
@ -195,12 +198,23 @@ ${tableBody}`
|
|||
`No \`${runnerPath}\` yet? Bootstrap with \`npx\`, \`bunx\`, or \`pnpm dlx\` — ` +
|
||||
'e.g. `bunx gitnexus@latest analyze` (npm 11 npx crash; #1939).';
|
||||
|
||||
// This block is injected into every user's repo and its total size is capped
|
||||
// by test (ai-context.test.ts, #856) — a new bullet or clause has to be paid
|
||||
// for by trimming an existing one.
|
||||
//
|
||||
// The detect_changes bullet carries the degraded-result rule (#2915): a run
|
||||
// that sets `partial` (a graph query failed) or `truncated` (the changed-symbol
|
||||
// listing was capped) is not the pre-commit gate passing, and `partial` pairs
|
||||
// routinely with changed_count:0 — the exact shape that printed "No changes
|
||||
// detected." and exited 0 on a broken analysis. Same reasoning as the
|
||||
// `risk: UNKNOWN` bullet below: the tool could not answer, so its zero is not
|
||||
// an all-clear.
|
||||
return `${GITNEXUS_START_MARKER}
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}. Use GitNexus graph tools to understand code, assess impact, and navigate safely.
|
||||
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}.
|
||||
|
||||
> Index stale? Run \`${runner} analyze\` from the project root — it auto-selects an available runner. ${bootstrapNote}
|
||||
> Index stale? Run \`${runner} analyze --index-only\` from the project root — it auto-selects an available runner. ${bootstrapNote}
|
||||
|
||||
## Always Do
|
||||
|
||||
|
|
@ -209,8 +223,9 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
|
|||
? ` For unified PDG impact, add \`mode: "pdg"\` with optional \`line: <N>\` — it returns statement-level \`affectedStatements\` over CDG + REACHING_DEF and inter-procedural symbols in \`interproceduralByDepth\`/\`byDepth\`; no-layer/degraded PDG results are UNKNOWN-risk notes (\`--pdg\` layer). CLI equivalent: \`${runner} impact "symbolName" --direction upstream --mode pdg --line <N> --repo .\`.`
|
||||
: ''
|
||||
}
|
||||
- **MUST analyze graph changes before committing.** Use \`detect_changes({scope: "all"})\` (MCP) or \`${runner} detect-changes --scope all --repo .\` (CLI fallback). For regression review: \`detect_changes({scope: "compare", base_ref: ${JSON.stringify(markdownSafeBranch(defaultBranch))}})\` or \`${runner} detect-changes --scope compare --base-ref ${JSON.stringify(markdownSafeBranch(defaultBranch))} --repo .\`.
|
||||
- **MUST analyze graph changes before committing.** Use \`detect_changes({scope: "all"})\` (MCP) or \`${runner} detect-changes --scope all --repo .\` (CLI fallback). \`partial: true\` or \`truncated: true\` is not a clean check — a zero means unseen, not unaffected; re-run it. For regression review: \`detect_changes({scope: "compare", base_ref: ${JSON.stringify(markdownSafeBranch(defaultBranch))}})\` or \`${runner} detect-changes --scope compare --base-ref ${JSON.stringify(markdownSafeBranch(defaultBranch))} --repo .\`.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- **MUST treat \`risk: UNKNOWN\` as unresolved, not as low.** An empty caller set is not evidence the symbol is unused — it can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). \`impact\` pairs \`UNKNOWN\` with a \`riskNote\` saying so. Confirm with a text search before treating the symbol as safe to change or delete; do not proceed on the strength of a zero.
|
||||
- When exploring unfamiliar code, use \`query({search_query: "concept"})\` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use \`context({name: "symbolName"})\`.
|
||||
- For security review, \`explain({target: "fileOrSymbol"})\` lists taint findings (source→sink flows; needs \`analyze --pdg\`).${
|
||||
|
|
@ -222,7 +237,7 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
|
|||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis, and never read \`UNKNOWN\` as an all-clear — it means the walk could not answer, which is the one verdict that requires confirming by other means.
|
||||
- NEVER rename symbols with find-and-replace — use \`rename\` which understands the call graph.
|
||||
- NEVER commit before MCP/CLI graph change analysis.
|
||||
|
||||
|
|
@ -266,11 +281,33 @@ async function fileExists(filePath: string): Promise<boolean> {
|
|||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Replace the block's volatile counts — the header parenthetical and the
|
||||
* per-cluster symbol counts in the skills table — with fixed placeholders, so
|
||||
* two renderings that differ only in those numbers compare equal.
|
||||
*
|
||||
* Placeholders rather than deletions: `--no-stats` REMOVES the parenthetical,
|
||||
* which must still be written through. Deleting instead of substituting would
|
||||
* make a with-counts block and a without-counts block compare equal, and the
|
||||
* flag would silently stop taking effect on an already-injected file.
|
||||
*/
|
||||
function stripVolatileCounts(section: string): string {
|
||||
return section
|
||||
.replace(/ \(\d+ symbols, \d+ relationships, \d+ execution flows\)/g, ' (<counts>)')
|
||||
.replace(/ \(\d+ symbols\)/g, ' (<count>)');
|
||||
}
|
||||
|
||||
/**
|
||||
* Create or update GitNexus section in a file
|
||||
* - If file doesn't exist: create with GitNexus content
|
||||
* - If file exists without GitNexus section: append
|
||||
* - If file exists with GitNexus section: replace that section
|
||||
* - If file exists with GitNexus section: replace that section, UNLESS the only
|
||||
* delta is the volatile counts (#2907). AGENTS.md and CLAUDE.md are the agent
|
||||
* guides teams commit, and the counts move with any code change, so a
|
||||
* count-only rewrite dirties a tracked file on every reindex for no reader
|
||||
* benefit. Live counts stay available from `gitnexus status` and
|
||||
* `gitnexus://repo/{name}/context`; the committed block keeps whichever
|
||||
* numbers it was last materially updated with.
|
||||
*/
|
||||
async function upsertGitNexusSection(
|
||||
filePath: string,
|
||||
|
|
@ -282,7 +319,10 @@ async function upsertGitNexusSection(
|
|||
const exists = await fileExists(filePath);
|
||||
|
||||
if (!exists) {
|
||||
await fs.writeFile(filePath, content, 'utf-8');
|
||||
// Same `.trim() + '\n'` shape the update paths write. Creating without the
|
||||
// trailing newline made the NEXT analyze dirty a freshly committed file
|
||||
// even at unchanged counts, purely to append it (#2907).
|
||||
await fs.writeFile(filePath, content.trim() + '\n', 'utf-8');
|
||||
return 'created';
|
||||
}
|
||||
|
||||
|
|
@ -343,6 +383,11 @@ async function upsertGitNexusSection(
|
|||
|
||||
if (statsPattern.test(existingSection)) {
|
||||
const updatedSection = existingSection.replace(statsPattern, statsLine);
|
||||
// Count-only delta — leave the committed lean block alone (#2907). A
|
||||
// project rename, or --no-stats dropping the parenthetical, still writes.
|
||||
if (stripVolatileCounts(updatedSection) === stripVolatileCounts(existingSection)) {
|
||||
return 'preserved';
|
||||
}
|
||||
const before = existingContent.substring(0, startIdx);
|
||||
const after = existingContent.substring(endIdx + GITNEXUS_END_MARKER.length);
|
||||
await fs.writeFile(filePath, (before + updatedSection + after).trim() + '\n', 'utf-8');
|
||||
|
|
@ -354,7 +399,11 @@ async function upsertGitNexusSection(
|
|||
return 'preserved';
|
||||
}
|
||||
|
||||
// No keep marker — replace existing section with full verbose content
|
||||
// No keep marker — replace existing section with full verbose content,
|
||||
// unless the counts are the only thing that moved (#2907).
|
||||
if (stripVolatileCounts(existingSection) === stripVolatileCounts(content)) {
|
||||
return 'preserved';
|
||||
}
|
||||
const before = existingContent.substring(0, startIdx);
|
||||
const after = existingContent.substring(endIdx + GITNEXUS_END_MARKER.length);
|
||||
const newContent = before + content + after;
|
||||
|
|
|
|||
|
|
@ -30,7 +30,7 @@
|
|||
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import type { AnalyzeOptions } from './analyze.js';
|
||||
import type { AnalyzeOptions } from './analyze-options.js';
|
||||
|
||||
export const GITNEXUS_RC_FILENAME = '.gitnexusrc';
|
||||
|
||||
|
|
|
|||
131
gitnexus/src/cli/analyze-options.ts
Normal file
131
gitnexus/src/cli/analyze-options.ts
Normal file
|
|
@ -0,0 +1,131 @@
|
|||
/**
|
||||
* CLI-facing `analyze` option shape.
|
||||
*
|
||||
* This is the *flag* shape: it mirrors what Commander parses off the command
|
||||
* line and what `.gitnexusrc` may set, before `analyze` translates it into the
|
||||
* core orchestrator's own `AnalyzeOptions` (`core/run-analyze.ts`) — a
|
||||
* different, deliberately separate interface (`stats` here vs `noStats`
|
||||
* there, `embeddings?: boolean | string` here vs a resolved
|
||||
* `embeddingsNodeLimit` there).
|
||||
*
|
||||
* It lives in this leaf module because both `analyze.ts` (which consumes the
|
||||
* flags) and `analyze-config.ts` (which maps `.gitnexusrc` keys onto them)
|
||||
* need it, and `analyze.ts` already imports the config loader — a type import
|
||||
* back the other way put the two files, plus `core/run-analyze.ts`, in an
|
||||
* import cycle. `analyze.ts` re-exports the type for existing importers.
|
||||
*/
|
||||
export interface AnalyzeOptions {
|
||||
force?: boolean;
|
||||
repairFts?: boolean;
|
||||
/**
|
||||
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
|
||||
* - `undefined` when the flag is omitted
|
||||
* - `true` when passed without an argument (use default 50K node cap)
|
||||
* - a string when passed with an argument (`--embeddings 0` disables the
|
||||
* cap, `--embeddings <n>` uses `<n>` as the cap)
|
||||
*/
|
||||
embeddings?: boolean | string;
|
||||
/**
|
||||
* Explicitly drop existing embeddings on rebuild instead of preserving
|
||||
* them. Without this flag, a routine `analyze` keeps any embeddings
|
||||
* already present in the index even when `--embeddings` is omitted.
|
||||
*/
|
||||
dropEmbeddings?: boolean;
|
||||
skills?: boolean;
|
||||
verbose?: boolean;
|
||||
/** Skip AGENTS.md and CLAUDE.md gitnexus block updates. */
|
||||
skipAgentsMd?: boolean;
|
||||
/**
|
||||
* Build the control-flow-graph / PDG substrate (#2081 M1). Opt-in; off by
|
||||
* default. Threaded to both the worker (CFG build) and scope-resolution
|
||||
* (BasicBlock/CFG emit).
|
||||
*/
|
||||
pdg?: boolean;
|
||||
/**
|
||||
* Stats inclusion in AGENTS.md and CLAUDE.md.
|
||||
*
|
||||
* Commander.js represents `--no-stats` as `stats: boolean` (default
|
||||
* `true`; `false` when the user passes `--no-stats`), NOT as
|
||||
* `noStats: boolean`. Reading the negated form would always be
|
||||
* `undefined` and the flag would silently no-op (#1477). Consumers
|
||||
* that want "did the user request --no-stats?" should compare with
|
||||
* `=== false` to distinguish the explicit-off case from the
|
||||
* default-on case.
|
||||
*/
|
||||
stats?: boolean;
|
||||
/**
|
||||
* Opt-in auto-commit of any AGENTS.md/CLAUDE.md changes this `analyze` run
|
||||
* makes. Scoped to only those two files (never `git add -A`); no-ops
|
||||
* silently if neither exists, neither changed, or the commit step itself
|
||||
* fails (e.g. no git identity configured). See #2639.
|
||||
*/
|
||||
selfCommit?: boolean;
|
||||
/** Skip installing standard GitNexus skill files directly under .claude/skills/. */
|
||||
skipSkills?: boolean;
|
||||
/**
|
||||
* Default branch for the generated regression-compare example (#243). From
|
||||
* `--default-branch`; may also be supplied via `.gitnexusrc`. Resolved to a
|
||||
* concrete branch (CLI > `.gitnexusrc` > auto-detected origin/HEAD > "main")
|
||||
* before being threaded into the generated AGENTS.md / CLAUDE.md content.
|
||||
*/
|
||||
defaultBranch?: string;
|
||||
/**
|
||||
* Index-branch selector (#2106). From `--branch`. Distinct from
|
||||
* `defaultBranch` (cosmetic base_ref): this routes the index to a per-branch
|
||||
* slot. NOT sourced from `.gitnexusrc` — the `.gitnexusrc` `branch` key is an
|
||||
* alias for `defaultBranch` and must not change index placement. Defaults to
|
||||
* the checked-out branch inside `runFullAnalysis` when omitted.
|
||||
*/
|
||||
branch?: string;
|
||||
/** Pure index mode: skip all file injection (AGENTS.md, CLAUDE.md, skills). */
|
||||
indexOnly?: boolean;
|
||||
/** Index the folder even when no .git directory is present. */
|
||||
skipGit?: boolean;
|
||||
/**
|
||||
* Override the default basename-derived registry `name` with a
|
||||
* user-supplied alias (#829). Disambiguates repos whose paths share a
|
||||
* basename. Persisted — subsequent re-analyses of the same path without
|
||||
* `--name` preserve the alias.
|
||||
*/
|
||||
name?: string;
|
||||
/**
|
||||
* Allow registration even when another path already uses the same
|
||||
* `--name` alias (#829). Intentionally a distinct flag from `--force`
|
||||
* because the user may want to coexist under the same name WITHOUT
|
||||
* paying the cost of a pipeline re-index. Maps to registerRepo's
|
||||
* `allowDuplicateName` option end-to-end.
|
||||
*/
|
||||
allowDuplicateName?: boolean;
|
||||
/**
|
||||
* Override the walker's large-file skip threshold (#991). Value in KB;
|
||||
* clamped downstream to the tree-sitter 32 MB ceiling. Sets
|
||||
* `GITNEXUS_MAX_FILE_SIZE` for the rest of the pipeline.
|
||||
*/
|
||||
maxFileSize?: string;
|
||||
/** Override worker sub-batch idle timeout in seconds. */
|
||||
workerTimeout?: string;
|
||||
/** Control LadybugDB WAL auto-checkpoint threshold during analyze. */
|
||||
walCheckpointThreshold?: string;
|
||||
/** Parse worker pool size (>=1); 0 is rejected (no sequential mode). */
|
||||
workers?: string;
|
||||
embeddingThreads?: string;
|
||||
embeddingBatchSize?: string;
|
||||
embeddingSubBatchSize?: string;
|
||||
embeddingDevice?: string;
|
||||
/**
|
||||
* Extra fetch-wrapper function names to treat as HTTP consumers (#1589/#1852
|
||||
* residual). Supplied via `.gitnexusrc` `fetchWrappers: [...]`. Threaded into
|
||||
* the routes phase, where the cross-file consumer scan unions them with the
|
||||
* auto-detected `fetch()` wrappers so a custom/axios-based wrapper named
|
||||
* outside the built-in convention still produces `route_map` consumers.
|
||||
*/
|
||||
fetchWrappers?: string[];
|
||||
/** OpenAI-compatible embeddings base URL (incl. /v1). Overrides GITNEXUS_EMBEDDING_URL. */
|
||||
embeddingBaseUrl?: string;
|
||||
/** Embedding model name. Overrides GITNEXUS_EMBEDDING_MODEL. */
|
||||
embeddingModel?: string;
|
||||
/** Bearer token for the embeddings endpoint. Overrides GITNEXUS_EMBEDDING_API_KEY. Never logged. */
|
||||
embeddingAuthToken?: string;
|
||||
/** Embedding vector dimensions (positive integer string). Overrides GITNEXUS_EMBEDDING_DIMS. */
|
||||
embeddingDims?: string;
|
||||
}
|
||||
|
|
@ -50,6 +50,7 @@ import {
|
|||
validateBranchName,
|
||||
GitNexusRcError,
|
||||
} from './analyze-config.js';
|
||||
import type { AnalyzeOptions } from './analyze-options.js';
|
||||
import { runFullAnalysis } from '../core/run-analyze.js';
|
||||
import { getRuntimeFingerprint } from '../core/platform/capabilities.js';
|
||||
import { getMaxFileSizeBannerMessage } from '../core/ingestion/utils/max-file-size.js';
|
||||
|
|
@ -661,121 +662,14 @@ const restoreAnalyzeEnv = (snap: AnalyzeEnvSnapshot): void => {
|
|||
}
|
||||
};
|
||||
|
||||
export interface AnalyzeOptions {
|
||||
force?: boolean;
|
||||
repairFts?: boolean;
|
||||
/**
|
||||
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
|
||||
* - `undefined` when the flag is omitted
|
||||
* - `true` when passed without an argument (use default 50K node cap)
|
||||
* - a string when passed with an argument (`--embeddings 0` disables the
|
||||
* cap, `--embeddings <n>` uses `<n>` as the cap)
|
||||
*/
|
||||
embeddings?: boolean | string;
|
||||
/**
|
||||
* Explicitly drop existing embeddings on rebuild instead of preserving
|
||||
* them. Without this flag, a routine `analyze` keeps any embeddings
|
||||
* already present in the index even when `--embeddings` is omitted.
|
||||
*/
|
||||
dropEmbeddings?: boolean;
|
||||
skills?: boolean;
|
||||
verbose?: boolean;
|
||||
/** Skip AGENTS.md and CLAUDE.md gitnexus block updates. */
|
||||
skipAgentsMd?: boolean;
|
||||
/**
|
||||
* Build the control-flow-graph / PDG substrate (#2081 M1). Opt-in; off by
|
||||
* default. Threaded to both the worker (CFG build) and scope-resolution
|
||||
* (BasicBlock/CFG emit).
|
||||
*/
|
||||
pdg?: boolean;
|
||||
/**
|
||||
* Stats inclusion in AGENTS.md and CLAUDE.md.
|
||||
*
|
||||
* Commander.js represents `--no-stats` as `stats: boolean` (default
|
||||
* `true`; `false` when the user passes `--no-stats`), NOT as
|
||||
* `noStats: boolean`. Reading the negated form would always be
|
||||
* `undefined` and the flag would silently no-op (#1477). Consumers
|
||||
* that want "did the user request --no-stats?" should compare with
|
||||
* `=== false` to distinguish the explicit-off case from the
|
||||
* default-on case.
|
||||
*/
|
||||
stats?: boolean;
|
||||
/**
|
||||
* Opt-in auto-commit of any AGENTS.md/CLAUDE.md changes this `analyze` run
|
||||
* makes. Scoped to only those two files (never `git add -A`); no-ops
|
||||
* silently if neither exists, neither changed, or the commit step itself
|
||||
* fails (e.g. no git identity configured). See #2639.
|
||||
*/
|
||||
selfCommit?: boolean;
|
||||
/** Skip installing standard GitNexus skill files directly under .claude/skills/. */
|
||||
skipSkills?: boolean;
|
||||
/**
|
||||
* Default branch for the generated regression-compare example (#243). From
|
||||
* `--default-branch`; may also be supplied via `.gitnexusrc`. Resolved to a
|
||||
* concrete branch (CLI > `.gitnexusrc` > auto-detected origin/HEAD > "main")
|
||||
* before being threaded into the generated AGENTS.md / CLAUDE.md content.
|
||||
*/
|
||||
defaultBranch?: string;
|
||||
/**
|
||||
* Index-branch selector (#2106). From `--branch`. Distinct from
|
||||
* `defaultBranch` (cosmetic base_ref): this routes the index to a per-branch
|
||||
* slot. NOT sourced from `.gitnexusrc` — the `.gitnexusrc` `branch` key is an
|
||||
* alias for `defaultBranch` and must not change index placement. Defaults to
|
||||
* the checked-out branch inside `runFullAnalysis` when omitted.
|
||||
*/
|
||||
branch?: string;
|
||||
/** Pure index mode: skip all file injection (AGENTS.md, CLAUDE.md, skills). */
|
||||
indexOnly?: boolean;
|
||||
/** Index the folder even when no .git directory is present. */
|
||||
skipGit?: boolean;
|
||||
/**
|
||||
* Override the default basename-derived registry `name` with a
|
||||
* user-supplied alias (#829). Disambiguates repos whose paths share a
|
||||
* basename. Persisted — subsequent re-analyses of the same path without
|
||||
* `--name` preserve the alias.
|
||||
*/
|
||||
name?: string;
|
||||
/**
|
||||
* Allow registration even when another path already uses the same
|
||||
* `--name` alias (#829). Intentionally a distinct flag from `--force`
|
||||
* because the user may want to coexist under the same name WITHOUT
|
||||
* paying the cost of a pipeline re-index. Maps to registerRepo's
|
||||
* `allowDuplicateName` option end-to-end.
|
||||
*/
|
||||
allowDuplicateName?: boolean;
|
||||
/**
|
||||
* Override the walker's large-file skip threshold (#991). Value in KB;
|
||||
* clamped downstream to the tree-sitter 32 MB ceiling. Sets
|
||||
* `GITNEXUS_MAX_FILE_SIZE` for the rest of the pipeline.
|
||||
*/
|
||||
maxFileSize?: string;
|
||||
/** Override worker sub-batch idle timeout in seconds. */
|
||||
workerTimeout?: string;
|
||||
/** Control LadybugDB WAL auto-checkpoint threshold during analyze. */
|
||||
walCheckpointThreshold?: string;
|
||||
/** Parse worker pool size (>=1); 0 is rejected (no sequential mode). */
|
||||
workers?: string;
|
||||
embeddingThreads?: string;
|
||||
embeddingBatchSize?: string;
|
||||
embeddingSubBatchSize?: string;
|
||||
embeddingDevice?: string;
|
||||
/**
|
||||
* Extra fetch-wrapper function names to treat as HTTP consumers (#1589/#1852
|
||||
* residual). Supplied via `.gitnexusrc` `fetchWrappers: [...]`. Threaded into
|
||||
* the routes phase, where the cross-file consumer scan unions them with the
|
||||
* auto-detected `fetch()` wrappers so a custom/axios-based wrapper named
|
||||
* outside the built-in convention still produces `route_map` consumers.
|
||||
*/
|
||||
fetchWrappers?: string[];
|
||||
/** OpenAI-compatible embeddings base URL (incl. /v1). Overrides GITNEXUS_EMBEDDING_URL. */
|
||||
embeddingBaseUrl?: string;
|
||||
/** Embedding model name. Overrides GITNEXUS_EMBEDDING_MODEL. */
|
||||
embeddingModel?: string;
|
||||
/** Bearer token for the embeddings endpoint. Overrides GITNEXUS_EMBEDDING_API_KEY. Never logged. */
|
||||
embeddingAuthToken?: string;
|
||||
/** Embedding vector dimensions (positive integer string). Overrides GITNEXUS_EMBEDDING_DIMS. */
|
||||
embeddingDims?: string;
|
||||
}
|
||||
/**
|
||||
* CLI `analyze` flag shape. Defined in `./analyze-options.js` so
|
||||
* `analyze-config.ts` can reference it without importing this module back —
|
||||
* that type import closed a cycle over `analyze` → `analyze-config` and
|
||||
* `analyze` → `run-analyze` → `analyze-config`. Re-exported here because this
|
||||
* is where callers have always imported it from.
|
||||
*/
|
||||
export type { AnalyzeOptions };
|
||||
|
||||
/**
|
||||
* Whether the post-index skill step should run.
|
||||
|
|
@ -1624,6 +1518,27 @@ const analyzeCommandImpl = async (
|
|||
|
||||
// ── Summary ────────────────────────────────────────────────────
|
||||
const s = result.stats;
|
||||
// A collapsed graph write is NOT a successful index. The other incomplete
|
||||
// reasons (`incremental-in-progress`, `embedding-checkpoint-pending`)
|
||||
// describe a run that did what it said and left work for next time; this
|
||||
// one means most of your edges are gone, so every query answers a confident
|
||||
// empty and the exit code is the only thing automation reads. Printing
|
||||
// "indexed successfully" and exiting 0 here would be the same class of
|
||||
// false certainty the check itself was written to remove.
|
||||
if (result.graphWriteCollapsed) {
|
||||
const { expected, persisted } = result.graphWriteCollapsed;
|
||||
console.log(`\n Repository indexed INCOMPLETELY (${totalTime}s)\n`);
|
||||
console.log(
|
||||
` Graph write collapsed: the pipeline produced ${expected.toLocaleString()} relationships\n` +
|
||||
` but only ${persisted.toLocaleString()} are readable from the index. Queries will answer\n` +
|
||||
` with missing edges rather than an error.\n\n` +
|
||||
` The index is recorded INCOMPLETE (graph-write-collapsed). Re-run\n` +
|
||||
` \`gitnexus analyze --force\`; if it recurs, check disk space and run \`gitnexus doctor\`.`,
|
||||
);
|
||||
console.log(` ${repoPath}`);
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
console.log(`\n Repository indexed successfully (${totalTime}s)\n`);
|
||||
console.log(
|
||||
` ${(s.nodes ?? 0).toLocaleString()} nodes | ${(s.edges ?? 0).toLocaleString()} edges | ${s.communities ?? 0} clusters | ${s.processes ?? 0} flows`,
|
||||
|
|
@ -1644,9 +1559,14 @@ const analyzeCommandImpl = async (
|
|||
);
|
||||
} else {
|
||||
console.log(
|
||||
// NOT "then rerun" (#2841 §5.C): this run stamped `lastCommit`, so a
|
||||
// plain rerun on an unchanged tree takes the up-to-date fast path and
|
||||
// returns before Phase 3 could rebuild anything — the advice would be
|
||||
// ineffective exactly when the user follows it. `--repair-fts` is the
|
||||
// verb that rebuilds the search indexes without re-parsing the repo.
|
||||
`\n Warning: full-text/BM25 search is disabled — the LadybugDB FTS extension was unavailable.\n` +
|
||||
` Install it once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) then rerun, or\n` +
|
||||
` run \`gitnexus analyze --repair-fts\` when connected. Run \`gitnexus doctor\` for details.`,
|
||||
` Install it once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto), then run\n` +
|
||||
` \`gitnexus analyze --repair-fts\` to build the search indexes. Run \`gitnexus doctor\` for details.`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,4 +1,5 @@
|
|||
import { t } from './i18n/index.js';
|
||||
import { formatSymbolLine } from './format-symbol.js';
|
||||
|
||||
type DetectChangesSummary = {
|
||||
changed_files?: number;
|
||||
|
|
@ -25,6 +26,8 @@ type AffectedProcess = {
|
|||
|
||||
type DetectChangesResult = {
|
||||
error?: unknown;
|
||||
partial?: boolean;
|
||||
truncated?: boolean;
|
||||
summary?: DetectChangesSummary;
|
||||
changed_symbols?: ChangedSymbol[];
|
||||
affected_processes?: AffectedProcess[];
|
||||
|
|
@ -35,11 +38,28 @@ export function formatDetectChangesResult(result: unknown): string {
|
|||
if (payload.error) return t('common.error', { message: String(payload.error) });
|
||||
|
||||
const summary = payload.summary ?? {};
|
||||
// A swallowed query failure sets `partial` and leaves the counts at zero
|
||||
// (#2283). Printing only "No changes detected." turns a degraded run into a
|
||||
// clean bill of health for the pre-commit gate, so say so either way.
|
||||
// `truncated` is its sibling flag: the backend caps the changed_symbols
|
||||
// LISTING (never the counts), so a short list is not proof of a short diff.
|
||||
// Both lead the output — a caveat printed after the summary is read too late.
|
||||
const notes: string[] = [];
|
||||
if (payload.partial) notes.push(t('tool.detectChanges.partial'));
|
||||
// The plain truncation note reassures that the counts are whole. That is only
|
||||
// true when the run did NOT also degrade — `changed_count` sums the batches
|
||||
// that succeeded — so the two flags together get a different sentence.
|
||||
if (payload.truncated)
|
||||
notes.push(
|
||||
t(payload.partial ? 'tool.detectChanges.truncatedDegraded' : 'tool.detectChanges.truncated'),
|
||||
);
|
||||
|
||||
if ((summary.changed_count ?? 0) === 0) {
|
||||
return t('tool.detectChanges.noChanges');
|
||||
return [...notes, t('tool.detectChanges.noChanges')].join('\n');
|
||||
}
|
||||
|
||||
const lines: string[] = [];
|
||||
if (notes.length > 0) lines.push(...notes, '');
|
||||
lines.push(
|
||||
t('tool.detectChanges.changesSummary', {
|
||||
files: summary.changed_files ?? 0,
|
||||
|
|
@ -59,7 +79,7 @@ export function formatDetectChangesResult(result: unknown): string {
|
|||
lines.push(t('tool.detectChanges.changedSymbols'));
|
||||
const shown = changed.slice(0, 15);
|
||||
for (const symbol of shown) {
|
||||
lines.push(` ${symbol.type ?? 'Symbol'} ${symbol.name ?? '?'} → ${symbol.filePath ?? '?'}`);
|
||||
lines.push(formatSymbolLine(symbol.type, symbol.name, symbol.filePath));
|
||||
}
|
||||
// Overflow is measured against the TRUE total (summary.changed_count), not
|
||||
// the array length — the array may already be `--limit`-sliced, so using its
|
||||
|
|
|
|||
|
|
@ -45,6 +45,7 @@ import {
|
|||
import { logger } from '../core/logger.js';
|
||||
import { cliInfo, cliWarn, cliError } from './cli-message.js';
|
||||
import { formatDetectChangesResult } from './detect-changes-format.js';
|
||||
import { formatSymbolLine } from './format-symbol.js';
|
||||
|
||||
export { formatDetectChangesResult } from './detect-changes-format.js';
|
||||
|
||||
|
|
@ -209,7 +210,7 @@ export function formatQueryResult(result: any): string {
|
|||
if (defs.length > 0) {
|
||||
lines.push(`Standalone definitions:`);
|
||||
for (const d of defs.slice(0, 8)) {
|
||||
lines.push(` ${d.type || 'Symbol'} ${d.name} → ${d.filePath || '?'}`);
|
||||
lines.push(formatSymbolLine(d.type, d.name, d.filePath));
|
||||
}
|
||||
if (defs.length > 8) lines.push(` ... and ${defs.length - 8} more`);
|
||||
}
|
||||
|
|
|
|||
22
gitnexus/src/cli/format-symbol.ts
Normal file
22
gitnexus/src/cli/format-symbol.ts
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
/**
|
||||
* Symbol listing line — the one rendering of `Type name → path` shared by every
|
||||
* formatter that lists symbols. Kept in its own tool-neutral module so a new
|
||||
* consumer does not have to import it from another tool's formatter.
|
||||
*/
|
||||
|
||||
/**
|
||||
* One indented `Type name → path` listing line for a symbol. Shared by the
|
||||
* `detect_changes` CLI formatter and the eval-server `query` formatter so the
|
||||
* two renderings cannot drift apart.
|
||||
*
|
||||
* `||`, not `??`: a node whose label came back as an EMPTY STRING (several node
|
||||
* types do — see enrichCandidateLabels) still needs the placeholder, and `??`
|
||||
* would print the empty string instead.
|
||||
*/
|
||||
export function formatSymbolLine(
|
||||
type: string | undefined,
|
||||
name: string | undefined,
|
||||
filePath: string | undefined,
|
||||
): string {
|
||||
return ` ${type || 'Symbol'} ${name || '?'} → ${filePath || '?'}`;
|
||||
}
|
||||
17
gitnexus/src/cli/generated-skill.ts
Normal file
17
gitnexus/src/cli/generated-skill.ts
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
/**
|
||||
* Metadata for one repo-specific skill file generated from a detected
|
||||
* community.
|
||||
*
|
||||
* Produced by `skill-gen`'s `generateSkillFiles` and consumed by `ai-context`
|
||||
* when it lists the generated skills in AGENTS.md / CLAUDE.md. It lives in this
|
||||
* leaf module rather than in either of those so the consumer does not have to
|
||||
* import the producer for a type — `ai-context` already supplies the
|
||||
* `.agents/` mirror check that `skill-gen` calls, and the two directions
|
||||
* together made an import cycle.
|
||||
*/
|
||||
export interface GeneratedSkillInfo {
|
||||
name: string;
|
||||
label: string;
|
||||
symbolCount: number;
|
||||
fileCount: number;
|
||||
}
|
||||
|
|
@ -65,6 +65,14 @@ export const en = {
|
|||
'tool.warn.unknownKind':
|
||||
"--kind '{{kind}}' is not a known symbol kind (e.g. Function, Class, Method); it will not narrow the result.",
|
||||
'tool.detectChanges.noChanges': 'No changes detected.',
|
||||
'tool.detectChanges.partial':
|
||||
'PARTIAL RESULT: a graph query failed, so changed symbols may be missing. Do not read this as a clean pre-commit check.',
|
||||
'tool.detectChanges.truncated':
|
||||
'LISTING CAPPED: the changed-symbol list was capped, so it does not name every changed symbol. The counts and risk level still cover all of them.',
|
||||
// The reassurance above is only true on its own. When the run also degraded,
|
||||
// `changed_count` was summed from the batches that SUCCEEDED, so it is a floor.
|
||||
'tool.detectChanges.truncatedDegraded':
|
||||
'LISTING CAPPED: the changed-symbol list was capped. The run also degraded, so the counts are a lower bound, not a total.',
|
||||
'tool.detectChanges.changesSummary': 'Changes: {{files}} files, {{symbols}} symbols',
|
||||
'tool.detectChanges.affectedProcesses': 'Affected processes: {{count}}',
|
||||
'tool.detectChanges.riskLevel': 'Risk level: {{risk}}',
|
||||
|
|
@ -226,16 +234,15 @@ export const en = {
|
|||
'Clean parked LadybugDB recovery sidecars (missing-shadow WAL quarantines and dirty-recovery parks)',
|
||||
'help.option.wiki.force': 'Force full regeneration even if up to date',
|
||||
'help.option.wiki.provider':
|
||||
'LLM provider: openai, openrouter, atlascloud, azure, custom, cursor, claude, codex, or opencode (default: openai)',
|
||||
'help.option.wiki.model': 'LLM model or Azure deployment name (default: minimax/minimax-m2.5)',
|
||||
'LLM provider: minimax, openai, openrouter, atlascloud, azure, custom, cursor, claude, codex, or opencode (default: minimax)',
|
||||
'help.option.wiki.model': 'LLM model or deployment name (default: MiniMax-M3)',
|
||||
'help.option.wiki.baseUrl':
|
||||
'LLM API base URL. Azure v1: https://{resource}.openai.azure.com/openai/v1',
|
||||
'help.option.wiki.apiKey': 'LLM API key or Azure api-key (saved to ~/.gitnexus/config.json)',
|
||||
'help.option.wiki.apiVersion':
|
||||
'Azure api-version query param, e.g. 2024-10-21 (legacy Azure API only)',
|
||||
'help.option.wiki.reasoningModel':
|
||||
'Mark deployment as reasoning model (o1/o3/o4-mini) — strips temperature, uses max_completion_tokens',
|
||||
'help.option.wiki.noReasoningModel': 'Disable reasoning model mode (overrides saved config)',
|
||||
'help.option.wiki.reasoningModel': 'Enable reasoning mode; MiniMax-M3 uses adaptive thinking',
|
||||
'help.option.wiki.noReasoningModel': 'Disable reasoning mode; MiniMax-M3 disables thinking',
|
||||
'help.option.wiki.concurrency': 'Parallel LLM calls (default: 3)',
|
||||
'help.option.wiki.timeout': 'LLM request timeout in seconds (default: disabled)',
|
||||
'help.option.wiki.retries': 'Max LLM retry attempts per request (default: 3)',
|
||||
|
|
|
|||
|
|
@ -69,6 +69,12 @@ export const zhCN = {
|
|||
'tool.warn.unknownKind':
|
||||
"--kind '{{kind}}' 不是已知的符号类型(如 Function、Class、Method),不会用于缩小结果范围。",
|
||||
'tool.detectChanges.noChanges': '未检测到变更。',
|
||||
'tool.detectChanges.partial':
|
||||
'结果不完整:图查询失败,可能遗漏已变更符号。请勿将其视为通过的提交前检查。',
|
||||
'tool.detectChanges.truncated':
|
||||
'列表已截断:已变更符号列表被截断,未列出全部变更符号。计数与风险等级仍涵盖全部符号。',
|
||||
'tool.detectChanges.truncatedDegraded':
|
||||
'列表已截断:已变更符号列表被截断。本次运行同时不完整,因此计数为下限而非总数。',
|
||||
'tool.detectChanges.changesSummary': '变更:{{files}} 个文件,{{symbols}} 个符号',
|
||||
'tool.detectChanges.affectedProcesses': '受影响流程:{{count}}',
|
||||
'tool.detectChanges.riskLevel': '风险等级:{{risk}}',
|
||||
|
|
@ -214,15 +220,14 @@ export const zhCN = {
|
|||
'清理已暂存的 LadybugDB 恢复 sidecar(missing-shadow WAL 隔离文件与 dirty-recovery 暂存文件)',
|
||||
'help.option.wiki.force': '即使已是最新也强制完整重新生成',
|
||||
'help.option.wiki.provider':
|
||||
'LLM 提供商:openai、openrouter、atlascloud、azure、custom、cursor、claude、codex 或 opencode(默认:openai)',
|
||||
'help.option.wiki.model': 'LLM 模型或 Azure deployment 名称(默认:minimax/minimax-m2.5)',
|
||||
'LLM 提供商:minimax、openai、openrouter、atlascloud、azure、custom、cursor、claude、codex 或 opencode(默认:minimax)',
|
||||
'help.option.wiki.model': 'LLM 模型或 deployment 名称(默认:MiniMax-M3)',
|
||||
'help.option.wiki.baseUrl':
|
||||
'LLM API base URL。Azure v1:https://{resource}.openai.azure.com/openai/v1',
|
||||
'help.option.wiki.apiKey': 'LLM API key 或 Azure api-key(保存到 ~/.gitnexus/config.json)',
|
||||
'help.option.wiki.apiVersion': 'Azure api-version 查询参数,例如 2024-10-21(仅旧版 Azure API)',
|
||||
'help.option.wiki.reasoningModel':
|
||||
'标记 deployment 为 reasoning model(o1/o3/o4-mini)— 去除 temperature,使用 max_completion_tokens',
|
||||
'help.option.wiki.noReasoningModel': '禁用 reasoning model 模式(覆盖已保存配置)',
|
||||
'help.option.wiki.reasoningModel': '启用 reasoning 模式;MiniMax-M3 使用自适应 thinking',
|
||||
'help.option.wiki.noReasoningModel': '禁用 reasoning 模式;MiniMax-M3 关闭 thinking',
|
||||
'help.option.wiki.concurrency': '并行 LLM 调用数(默认:3)',
|
||||
'help.option.wiki.timeout': 'LLM 请求超时时间(秒,默认:禁用)',
|
||||
'help.option.wiki.retries': '每个请求的最大 LLM 重试次数(默认:3)',
|
||||
|
|
|
|||
|
|
@ -303,9 +303,9 @@ program
|
|||
.option('-f, --force', 'Force full regeneration even if up to date')
|
||||
.option(
|
||||
'--provider <provider>',
|
||||
'LLM provider: openai, openrouter, atlascloud, azure, custom, cursor, claude, codex, or opencode (default: openai)',
|
||||
'LLM provider: minimax, openai, openrouter, atlascloud, azure, custom, cursor, claude, codex, or opencode (default: minimax)',
|
||||
)
|
||||
.option('--model <model>', 'LLM model or Azure deployment name (default: minimax/minimax-m2.5)')
|
||||
.option('--model <model>', 'LLM model or deployment name (default: MiniMax-M3)')
|
||||
.option(
|
||||
'--base-url <url>',
|
||||
'LLM API base URL. Azure v1: https://{resource}.openai.azure.com/openai/v1',
|
||||
|
|
@ -315,11 +315,8 @@ program
|
|||
'--api-version <version>',
|
||||
'Azure api-version query param, e.g. 2024-10-21 (legacy Azure API only)',
|
||||
)
|
||||
.option(
|
||||
'--reasoning-model',
|
||||
'Mark deployment as reasoning model (o1/o3/o4-mini) — strips temperature, uses max_completion_tokens',
|
||||
)
|
||||
.option('--no-reasoning-model', 'Disable reasoning model mode (overrides saved config)')
|
||||
.option('--reasoning-model', 'Enable reasoning mode; MiniMax-M3 uses adaptive thinking')
|
||||
.option('--no-reasoning-model', 'Disable reasoning mode; MiniMax-M3 disables thinking')
|
||||
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
|
||||
.option('--timeout <seconds>', 'LLM request timeout in seconds (default: disabled)')
|
||||
.option('--retries <n>', 'Max LLM retry attempts per request (default: 3)')
|
||||
|
|
|
|||
|
|
@ -14,6 +14,7 @@ import { CommunityNode, CommunityMembership } from '../core/ingestion/community-
|
|||
import { ProcessNode } from '../core/ingestion/process-processor.js';
|
||||
import { KnowledgeGraph } from '../core/graph/types.js';
|
||||
import { shouldMirrorSkillsToAgents } from './ai-context.js';
|
||||
import type { GeneratedSkillInfo } from './generated-skill.js';
|
||||
|
||||
const GENERATED_SKILL_PREFIX = 'gitnexus-area-';
|
||||
const MAX_SKILL_NAME_LENGTH = 64;
|
||||
|
|
@ -23,13 +24,6 @@ const MAX_COMMUNITY_NAME_LENGTH = MAX_SKILL_NAME_LENGTH - GENERATED_SKILL_PREFIX
|
|||
// TYPES
|
||||
// ============================================================================
|
||||
|
||||
export interface GeneratedSkillInfo {
|
||||
name: string;
|
||||
label: string;
|
||||
symbolCount: number;
|
||||
fileCount: number;
|
||||
}
|
||||
|
||||
interface AggregatedCommunity {
|
||||
label: string;
|
||||
rawIds: string[];
|
||||
|
|
|
|||
|
|
@ -42,9 +42,18 @@ async function getBackend(): Promise<LocalBackend> {
|
|||
* and write directly to the real stdout fd (#324).
|
||||
*
|
||||
* Falls back to stderr if the fd write fails (e.g., broken pipe).
|
||||
*
|
||||
* `render` is for the commands that print prose instead of JSON: they hand over
|
||||
* the STRUCTURED result and a formatter, so the payload stays visible to the
|
||||
* exit-code test below — pre-formatting it into a string would hide the very
|
||||
* fields that test reads.
|
||||
*/
|
||||
function output(data: any): void {
|
||||
const text = typeof data === 'string' ? data : JSON.stringify(data, null, 2);
|
||||
function output<T>(data: T, render?: (data: T) => string): void {
|
||||
const text = render
|
||||
? render(data)
|
||||
: typeof data === 'string'
|
||||
? data
|
||||
: JSON.stringify(data, null, 2);
|
||||
try {
|
||||
writeSync(1, text + '\n');
|
||||
} catch (err: any) {
|
||||
|
|
@ -56,18 +65,34 @@ function output(data: any): void {
|
|||
// Fallback: stderr (previous behavior, works on all platforms)
|
||||
process.stderr.write(text + '\n');
|
||||
}
|
||||
// Backend failures come back as `{ error }` payloads rather than throws
|
||||
// (#2469). Every tool command routes its result through here, so this is
|
||||
// the one place that keeps scripted callers honest: print the payload,
|
||||
// then exit non-zero.
|
||||
if (
|
||||
data &&
|
||||
typeof data === 'object' &&
|
||||
'error' in data &&
|
||||
typeof data.error === 'string' &&
|
||||
data.error.trim().length > 0
|
||||
) {
|
||||
process.exitCode = 1;
|
||||
// Every tool command routes its result through here, so this is the one place
|
||||
// that keeps scripted callers honest — `gitnexus impact … && <edit>` and
|
||||
// `gitnexus detect-changes && git commit` must not proceed on a result that
|
||||
// did not complete. Two shapes say so, and both exit non-zero:
|
||||
//
|
||||
// • `error` — a backend failure, returned as a payload rather than thrown
|
||||
// (#2469).
|
||||
// • `partial` — a step failed and was SWALLOWED (#2915), so the counts and
|
||||
// risk level are lower bounds a caller would otherwise read as clean. It
|
||||
// is cross-tool vocabulary, not detect_changes' private flag: `query`
|
||||
// raises it for degraded enrichment or a partial FTS failure, and `impact`
|
||||
// for an interrupted traversal or capped per-symbol enrichment — a short
|
||||
// caller set and an under-ranked risk, on the tool AGENTS.md makes a MUST
|
||||
// gate before every edit.
|
||||
//
|
||||
// One code for both, because `&&` cannot tell two apart and a "softer" code
|
||||
// for `partial` would invite exempting it again.
|
||||
//
|
||||
// NOT here: `truncated`, where only the LISTING is capped while the counts and
|
||||
// risk are computed over the full set — the verdict is sound, so failing on it
|
||||
// would fire on every large-but-healthy diff. Nor `partialProbe`, a narrower
|
||||
// per-candidate flag on ambiguous impact targets.
|
||||
if (data && typeof data === 'object') {
|
||||
const payload = data as { error?: unknown; partial?: unknown };
|
||||
const failed =
|
||||
(typeof payload.error === 'string' && payload.error.trim().length > 0) ||
|
||||
payload.partial === true;
|
||||
if (failed) process.exitCode = 1;
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -337,7 +362,9 @@ export async function detectChangesCommand(options?: {
|
|||
if (Array.isArray(result.affected_processes))
|
||||
result.affected_processes = result.affected_processes.slice(0, limit);
|
||||
}
|
||||
output(formatDetectChangesResult(result));
|
||||
// Hand over the structured result plus its formatter, not the formatted text:
|
||||
// `output()` reads `error` / `partial` off the payload to set the exit code.
|
||||
output(result, formatDetectChangesResult);
|
||||
}
|
||||
|
||||
export async function checkCommand(options?: {
|
||||
|
|
@ -359,21 +386,44 @@ export async function checkCommand(options?: {
|
|||
repo: options.repo,
|
||||
branch: options.branch,
|
||||
});
|
||||
// A rendering guard, not an exit-code decision — `output()` owns that. An
|
||||
// error payload carries no `cycles` array, so the prose branch below would
|
||||
// throw on it; print the structured payload and stop.
|
||||
if (result?.error) {
|
||||
output(result);
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
if (options.json) {
|
||||
output(result);
|
||||
} else if (result.cycleCount === 0) {
|
||||
} else if (result.status === 'clean') {
|
||||
output('No circular imports found.');
|
||||
} else {
|
||||
output(
|
||||
result.cycles.map((cycle: { files: string[] }) => cycle.files.join(' -> ')).join('\n'),
|
||||
);
|
||||
// Past the enumeration cap the tool reports one representative cycle per
|
||||
// component instead of every elementary cycle. Say so, or the short list
|
||||
// reads as the whole truth on exactly the repositories where it is not.
|
||||
if (result.enumeration === 'component-representatives') {
|
||||
// Phrased to need no plural: `checkCommand` predates the `t()` i18n
|
||||
// layer and none of its output goes through it, so inventing a plural
|
||||
// here by hand would be the only one in the file.
|
||||
output(
|
||||
`\n(showing one representative cycle per circular component — ` +
|
||||
`${result.componentCount} in total; the full enumeration exceeded the safety limit.)`,
|
||||
);
|
||||
}
|
||||
}
|
||||
if (result.cycleCount > 0) process.exitCode = 1;
|
||||
// Policy, not degradation: a clean run that FOUND cycles is `check` failing
|
||||
// its own check, so `output()` — which fails closed on `error` and `partial`
|
||||
// — deliberately knows nothing about it.
|
||||
//
|
||||
// Keyed on `status`, NOT on `cycleCount`. Past the enumeration cap the
|
||||
// report carries `cycleCount: null` on purpose, because a partial count must
|
||||
// not read as a real one — and `null > 0` is false, so counting here would
|
||||
// exit 0 on precisely the repositories with the most cycles. `status`
|
||||
// answers "were any found" in both enumeration modes.
|
||||
if (result.status === 'cycles_found') process.exitCode = 1;
|
||||
} catch (error) {
|
||||
output({ error: error instanceof Error ? error.message : String(error) });
|
||||
process.exitCode = 1;
|
||||
|
|
|
|||
|
|
@ -21,6 +21,8 @@ import {
|
|||
ATLAS_CLOUD_BASE_URL,
|
||||
ATLAS_CLOUD_DEFAULT_MODEL,
|
||||
getProviderEnvApiKey,
|
||||
MINIMAX_MODEL_IDS,
|
||||
MINIMAX_OPENAI_BASE_URLS,
|
||||
parseLLMAllowedInsecureHttpHosts,
|
||||
resolveLLMConfig,
|
||||
type LLMProvider,
|
||||
|
|
@ -220,6 +222,14 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
) {
|
||||
const existing = await loadCLIConfig();
|
||||
const updates: Partial<typeof existing> = {};
|
||||
const providerChanged = !!options.provider && options.provider !== existing.provider;
|
||||
if (providerChanged) {
|
||||
updates.apiKey = undefined;
|
||||
updates.baseUrl = undefined;
|
||||
updates.model = undefined;
|
||||
updates.apiVersion = undefined;
|
||||
updates.isReasoningModel = undefined;
|
||||
}
|
||||
if (options.apiKey) updates.apiKey = options.apiKey;
|
||||
if (options.baseUrl) updates.baseUrl = options.baseUrl;
|
||||
if (options.provider) updates.provider = options.provider;
|
||||
|
|
@ -230,6 +240,17 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
}
|
||||
if (options.apiVersion) updates.apiVersion = options.apiVersion;
|
||||
if (options.reasoningModel !== undefined) updates.isReasoningModel = options.reasoningModel;
|
||||
if (options.provider === 'minimax') {
|
||||
if (providerChanged && options.reasoningModel === undefined) {
|
||||
updates.isReasoningModel = undefined;
|
||||
}
|
||||
if (!options.baseUrl && (providerChanged || !existing.baseUrl)) {
|
||||
updates.baseUrl = MINIMAX_OPENAI_BASE_URLS.global_en;
|
||||
}
|
||||
if (!options.model && (providerChanged || !existing.model)) {
|
||||
updates.model = MINIMAX_MODEL_IDS[0];
|
||||
}
|
||||
}
|
||||
// Save model to appropriate field based on provider.
|
||||
if (options.model) {
|
||||
const targetProvider = options.provider ?? existing.provider;
|
||||
|
|
@ -246,7 +267,7 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
const savedConfig = await loadCLIConfig();
|
||||
const hasSavedConfig = !!(
|
||||
isLocalProvider(savedConfig.provider) ||
|
||||
(savedConfig.apiKey && savedConfig.baseUrl)
|
||||
(savedConfig.apiKey && (savedConfig.baseUrl || savedConfig.provider === 'minimax'))
|
||||
);
|
||||
const hasCLIOverrides = !!(
|
||||
options?.apiKey ||
|
||||
|
|
@ -279,7 +300,7 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
if (!llmConfig.apiKey && !isLocalProvider(llmConfig.provider)) {
|
||||
console.log(' Error: No LLM API key found.');
|
||||
console.log(
|
||||
' Set ATLASCLOUD_API_KEY, OPENAI_API_KEY, or GITNEXUS_API_KEY environment variable,',
|
||||
' Set MINIMAX_API_KEY, ATLASCLOUD_API_KEY, GITNEXUS_API_KEY, or OPENAI_API_KEY,',
|
||||
);
|
||||
console.log(' or pass --api-key <key>, or use --provider cursor|claude|codex|opencode.\n');
|
||||
process.exitCode = 1;
|
||||
|
|
@ -289,7 +310,7 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
} else {
|
||||
console.log(" No LLM configured. Let's set it up.\n");
|
||||
console.log(
|
||||
' Supports OpenAI, OpenRouter, Atlas Cloud, Azure, any OpenAI-compatible API, Cursor CLI, Claude CLI, Codex CLI, or OpenCode CLI.\n',
|
||||
' Supports MiniMax, OpenAI, OpenRouter, Atlas Cloud, Azure, custom OpenAI-compatible APIs, and local agent CLIs.\n',
|
||||
);
|
||||
|
||||
// Check if local agent CLIs are available.
|
||||
|
|
@ -308,7 +329,9 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
console.log(' [3] Azure OpenAI');
|
||||
console.log(' [4] Atlas Cloud (api.atlascloud.ai)');
|
||||
console.log(' [5] Custom endpoint');
|
||||
let nextChoice = 6;
|
||||
console.log(' [6] MiniMax Global (api.minimax.io)');
|
||||
console.log(' [7] MiniMax China (api.minimaxi.com)');
|
||||
let nextChoice = 8;
|
||||
if (hasCursor) {
|
||||
const choice = String(nextChoice++);
|
||||
localChoices.push({
|
||||
|
|
@ -429,10 +452,10 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
provider: 'azure',
|
||||
};
|
||||
} else {
|
||||
// OpenAI-compatible provider (OpenAI, OpenRouter, Atlas Cloud, Custom)
|
||||
// OpenAI-compatible provider setup
|
||||
if (choice === '2') {
|
||||
baseUrl = 'https://openrouter.ai/api/v1';
|
||||
defaultModel = 'minimax/minimax-m2.5';
|
||||
defaultModel = '';
|
||||
provider = 'openrouter';
|
||||
} else if (choice === '4') {
|
||||
baseUrl = ATLAS_CLOUD_BASE_URL;
|
||||
|
|
@ -447,6 +470,11 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
}
|
||||
defaultModel = 'gpt-4o-mini';
|
||||
provider = 'custom';
|
||||
} else if (choice === '6' || choice === '7') {
|
||||
baseUrl =
|
||||
choice === '7' ? MINIMAX_OPENAI_BASE_URLS.cn_zh : MINIMAX_OPENAI_BASE_URLS.global_en;
|
||||
defaultModel = MINIMAX_MODEL_IDS[0];
|
||||
provider = 'minimax';
|
||||
} else {
|
||||
baseUrl = 'https://api.openai.com/v1';
|
||||
defaultModel = 'gpt-4o-mini';
|
||||
|
|
@ -454,8 +482,15 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
}
|
||||
|
||||
// Model
|
||||
const modelInput = await prompt(` Model (default: ${defaultModel}): `);
|
||||
const modelInput = await prompt(
|
||||
defaultModel ? ` Model (default: ${defaultModel}): ` : ' Model: ',
|
||||
);
|
||||
const model = modelInput || defaultModel;
|
||||
if (!model) {
|
||||
console.log('\n No model provided. Aborting.\n');
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
|
||||
// API key — pre-fill hint if env var exists
|
||||
const envKey = getProviderEnvApiKey(provider);
|
||||
|
|
@ -478,7 +513,15 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
}
|
||||
|
||||
// Save
|
||||
await saveCLIConfig({ apiKey: key, baseUrl, model, provider });
|
||||
await saveCLIConfig({
|
||||
...savedConfig,
|
||||
apiKey: key,
|
||||
baseUrl,
|
||||
model,
|
||||
provider,
|
||||
apiVersion: undefined,
|
||||
isReasoningModel: undefined,
|
||||
});
|
||||
console.log(' Config saved to ~/.gitnexus/config.json\n');
|
||||
|
||||
llmConfig = { ...llmConfig, apiKey: key, baseUrl, model, provider };
|
||||
|
|
|
|||
|
|
@ -11,6 +11,7 @@
|
|||
* via `AbortSignal.timeout` on the underlying fetch.
|
||||
*/
|
||||
|
||||
import { chunk } from '../../lib/utils.js';
|
||||
import {
|
||||
CircuitOpenError,
|
||||
ResilientFetchExhaustedError,
|
||||
|
|
@ -566,9 +567,7 @@ export const httpEmbed = async (
|
|||
const url = `${config.baseUrl}/embeddings`;
|
||||
const allVectors: Float32Array[] = [];
|
||||
|
||||
for (let i = 0; i < texts.length; i += HTTP_BATCH_SIZE) {
|
||||
const batch = texts.slice(i, i + HTTP_BATCH_SIZE);
|
||||
const batchIndex = Math.floor(i / HTTP_BATCH_SIZE);
|
||||
for (const [batchIndex, batch] of chunk(texts, HTTP_BATCH_SIZE).entries()) {
|
||||
const items = await httpEmbedBatch(
|
||||
url,
|
||||
batch,
|
||||
|
|
|
|||
|
|
@ -1,11 +1,568 @@
|
|||
/**
|
||||
* Elementary import-cycle enumeration.
|
||||
*
|
||||
* ## What is reported
|
||||
*
|
||||
* Every *elementary* cycle of the file-import graph — a closed walk that visits
|
||||
* no file twice — is reported exactly once. Self-imports (`a -> a`) and
|
||||
* two-file cycles count. Cycles that are nested inside, or that overlap with,
|
||||
* other cycles are each reported separately: a strongly connected component
|
||||
* with three mutually-importing files contributes five cycles, not one.
|
||||
*
|
||||
* This replaces an earlier implementation that returned ONE representative
|
||||
* cycle per cyclic strongly connected component. That made the reported count a
|
||||
* count of tangles, not of cycles, and it hid every cycle in a component but
|
||||
* the first — including cycles that a reader would have to break separately.
|
||||
* The tangle count is still available, as `componentCount`, under a name that
|
||||
* says what it is.
|
||||
*
|
||||
* The scale of what the old shape hid, measured on GitNexus itself (2,079
|
||||
* files, 5,320 initialization-forcing import edges): it reported 11 cycles.
|
||||
* There are 27,939, spread across those same 11 components. It showed 11 of
|
||||
* them and 27,928 were invisible.
|
||||
*
|
||||
* ## Algorithm
|
||||
*
|
||||
* Donald B. Johnson, "Finding all the elementary circuits of a directed graph",
|
||||
* SIAM J. Comput. 4(1), 1975 — SCC decomposition plus a backtracking search
|
||||
* guarded by the `blocked` flag and the `B` sets, which together guarantee that
|
||||
* no fruitless path is explored twice between two circuit outputs. That is what
|
||||
* buys the O((n + e)(c + 1)) bound for `c` circuits: the cost is proportional
|
||||
* to the answer, not to the size of the search space.
|
||||
*
|
||||
* SCCs come from an iterative Tarjan pass rather than the Kosaraju pass this
|
||||
* module used before. Johnson recomputes SCCs on each induced subgraph as the
|
||||
* root advances, and Tarjan needs only the forward adjacency, so nothing has to
|
||||
* rebuild a reverse graph once per root.
|
||||
*
|
||||
* ## How this differs from madge
|
||||
*
|
||||
* madge's `circular()` walks depth-first from every node carrying its ancestor
|
||||
* path and records `ancestors.slice(indexOf(dep))` whenever it reaches an
|
||||
* ancestor. It also marks nodes visited *globally* and skips them on later
|
||||
* walks, so once a node has been traversed, cycles reachable only by entering
|
||||
* it from a different predecessor are never seen. madge therefore reports many
|
||||
* cycles but not all of them, and which ones it misses depends on iteration
|
||||
* order. Johnson's is strictly stronger: it is complete.
|
||||
*
|
||||
* The practical consequence is that GitNexus reports MORE cycles than madge on
|
||||
* the same graph, and the two counts should not be expected to agree. Anyone
|
||||
* reconciling the two is not looking at a bug here.
|
||||
*
|
||||
* ## Determinism
|
||||
*
|
||||
* Adjacency lists and the node order are sorted (default string order, matching
|
||||
* `Array.prototype.sort`), the search visits neighbours in that order, and the
|
||||
* finished list is sorted element-wise. Same input, same output, byte for byte.
|
||||
*
|
||||
* ## Rotation normalization
|
||||
*
|
||||
* `[a, b, c, a]` and `[b, c, a, b]` are the same cycle and must be emitted
|
||||
* once. That is structural here rather than a post-hoc dedup pass: Johnson's
|
||||
* search for circuits rooted at `s` runs on the subgraph induced by the nodes
|
||||
* that sort at or after `s`, so every node of an emitted circuit sorts at or
|
||||
* after its root. Each cycle is therefore emitted exactly once, rooted at — and
|
||||
* closed back onto — its own lexicographically smallest node. No other rotation
|
||||
* of it can ever be produced.
|
||||
*
|
||||
* ## Bounds
|
||||
*
|
||||
* The number of elementary cycles is exponential in the worst case, so the
|
||||
* search is bounded twice: by the number of cycles (`IMPORT_CYCLE_LIMIT`) and
|
||||
* by the work spent finding them (`IMPORT_CYCLE_WORK_LIMIT`). The second is not
|
||||
* redundant — Johnson's is output-sensitive, so a graph that yields few cycles
|
||||
* per root can burn unbounded time while staying far under the cycle cap.
|
||||
*
|
||||
* Exceeding either bound abandons the enumeration. What a partial run had
|
||||
* accumulated is discarded rather than returned, because a partial list of
|
||||
* elementary cycles is indistinguishable from a complete one at the call site
|
||||
* and would be read as "these are all of them". What is returned instead is a
|
||||
* different KIND of list — one representative cycle per cyclic component, the
|
||||
* old pre-enumeration answer — under `enumeration: 'component-representatives'`
|
||||
* so the difference is machine-readable and not merely documented. Only a run
|
||||
* that dies inside the decomposition itself reports nothing at all.
|
||||
*/
|
||||
|
||||
import { compareCodeUnits } from '../../lib/utils.js';
|
||||
|
||||
interface ImportEdge {
|
||||
source: string;
|
||||
target: string;
|
||||
}
|
||||
|
||||
function findCyclePath(component: string[], adjacency: Map<string, string[]>): string[] {
|
||||
/**
|
||||
* Elementary cycles reported before the search fails closed.
|
||||
*
|
||||
* The binding constraint is response size, not time. THE MEASUREMENT THAT SETS
|
||||
* THIS NUMBER, on GitNexus itself — 2,079 files, 5,320 initialization-forcing
|
||||
* import edges: complete enumeration finds 27,939 elementary cycles across 11
|
||||
* components in 241ms. Fast. But those cycles average 13 files each, so
|
||||
* serializing them is 400,877 path entries — a 21.8 MB JSON response for a tool
|
||||
* whose result is read by an agent. Time was never going to stop that, and
|
||||
* neither was the work budget (the same run spends 6.3M of its 10M).
|
||||
*
|
||||
* Keep that measurement next to this constant. Without it the cap looks like an
|
||||
* arbitrary round number and gets raised or deleted by someone who has only
|
||||
* ever seen it not fire.
|
||||
*
|
||||
* So the cap is set where the answer stops being consumable rather than where
|
||||
* the machine stops coping. Past 10,000 cycles the ten-thousandth path tells a
|
||||
* reader nothing the first hundred did not, and what a reader acts on is
|
||||
* `componentCount` plus one cycle per component — which is exactly what a
|
||||
* report over the cap degrades to, rather than to nothing.
|
||||
*/
|
||||
export const IMPORT_CYCLE_LIMIT = 10_000;
|
||||
|
||||
/**
|
||||
* Units of search effort allowed before the search fails closed — edges
|
||||
* examined, nodes scanned per root, and emitted cycle nodes at
|
||||
* `EMITTED_NODE_COST` each — counted across the SCC passes and the circuit
|
||||
* search alike.
|
||||
*
|
||||
* The cycle cap alone does NOT bound this. Johnson's is output-sensitive at
|
||||
* O((n + e)(c + 1)), so producing `c` cycles still scales with the graph: a
|
||||
* component of mutually-importing neighbours yields one or two cycles per root,
|
||||
* so an SCC pass runs per node and the total is quadratic while the cycle count
|
||||
* stays low. `check` admits import graphs up to 100k edges, so that shape is
|
||||
* reachable, and there it is minutes of work under a cycle cap that never
|
||||
* trips. The reverse gap is just as real: a single 50k-file component produces
|
||||
* cycles 50k files long, and 10k of those exhaust the heap. One bound cannot
|
||||
* see both, which is why there are two.
|
||||
*
|
||||
* Measured on this implementation against the mutual-import chain — the shape
|
||||
* that spends the whole budget, where every unit buys a fresh SCC pass over a
|
||||
* barely-smaller component — the rate is 2.9-5.2M units/second (5.2M at 10k
|
||||
* nodes, 2.9M at 50k; it falls as the component grows). So 10M buys roughly
|
||||
* 1.9-3.5s of enumeration on this hardware. That is the one shape where a user
|
||||
* waits, and it is the number to re-measure if this constant is ever moved.
|
||||
*
|
||||
* It sits far above what real import graphs cost: a 100k-file acyclic graph
|
||||
* spends 220k units, and 20k independent three-file tangles spend 576k. Only a
|
||||
* component both large and densely tangled reaches the cap, and that
|
||||
* component's honest answer is "too tangled to enumerate", not a
|
||||
* silently-shortened list.
|
||||
*/
|
||||
export const IMPORT_CYCLE_WORK_LIMIT = 10_000_000;
|
||||
|
||||
/**
|
||||
* Work charged per node of an emitted cycle, relative to one edge examination.
|
||||
*
|
||||
* Emitted nodes are retained for the lifetime of the call and then sorted and
|
||||
* serialized, so they are the term that decides peak memory, while examined
|
||||
* edges cost nothing but time. Without a weight here, a graph whose cycles are
|
||||
* tens of thousands of files long exhausts the heap while both bounds still
|
||||
* read as comfortably unspent.
|
||||
*/
|
||||
const EMITTED_NODE_COST = 10;
|
||||
|
||||
/** Which bound stopped the search. */
|
||||
export type ImportCycleLimit = 'cycles' | 'work';
|
||||
|
||||
/**
|
||||
* The result of an enumeration.
|
||||
*
|
||||
* `enumeration` is the union's discriminant rather than a sibling flag,
|
||||
* deliberately: a caller cannot reach `cycles` without first narrowing on what
|
||||
* kind of list it is holding. A partial enumeration and a complete one are
|
||||
* indistinguishable by inspection — both are arrays of real cycles — so the
|
||||
* difference has to be carried in the type, not in a comment or a count that
|
||||
* happens to look small.
|
||||
*/
|
||||
export type ImportCycleReport =
|
||||
| {
|
||||
readonly enumeration: 'complete';
|
||||
/**
|
||||
* Every elementary cycle, each as `[n0, n1, ..., nk, n0]` — the first
|
||||
* node repeated at the end so the closing edge is explicit. Sorted.
|
||||
*/
|
||||
readonly cycles: readonly string[][];
|
||||
/**
|
||||
* Number of cyclic strongly connected components — the count of
|
||||
* independent tangles. This is what the previous implementation called
|
||||
* the cycle count; it is NOT the number of cycles.
|
||||
*/
|
||||
readonly componentCount: number;
|
||||
}
|
||||
| {
|
||||
/**
|
||||
* A bound was hit, so the enumeration is abandoned — but the SCC
|
||||
* decomposition had already finished, so every tangle is known and each
|
||||
* one gets a representative. This is strictly more useful than an error:
|
||||
* a CI job can act on "these 11 components are cyclic, here is one cycle
|
||||
* through each", and cannot act on nothing at all.
|
||||
*
|
||||
* What is NOT carried is any count of cycles. `componentCount` is exact;
|
||||
* the number of elementary cycles is unknown and stays unknown.
|
||||
*/
|
||||
readonly enumeration: 'component-representatives';
|
||||
/** One cycle per component, same shape and ordering as the complete list. */
|
||||
readonly cycles: readonly string[][];
|
||||
readonly componentCount: number;
|
||||
readonly reason: ImportCycleLimit;
|
||||
readonly limit: number;
|
||||
}
|
||||
| {
|
||||
/**
|
||||
* A bound was hit inside the decomposition itself, so not even the tangle
|
||||
* count is known. There is genuinely nothing to report.
|
||||
*/
|
||||
readonly enumeration: 'none';
|
||||
readonly reason: ImportCycleLimit;
|
||||
readonly limit: number;
|
||||
};
|
||||
|
||||
/** Sorted forward adjacency plus the set of nodes that import themselves. */
|
||||
interface ImportGraph {
|
||||
readonly adjacency: ReadonlyMap<string, readonly string[]>;
|
||||
readonly nodes: readonly string[];
|
||||
readonly selfLoops: ReadonlySet<string>;
|
||||
}
|
||||
|
||||
function buildGraph(edges: readonly ImportEdge[]): ImportGraph {
|
||||
const targetsBySource = new Map<string, Set<string>>();
|
||||
for (const { source, target } of edges) {
|
||||
if (!source || !target) continue;
|
||||
const targets = targetsBySource.get(source) ?? new Set<string>();
|
||||
targets.add(target);
|
||||
targetsBySource.set(source, targets);
|
||||
if (!targetsBySource.has(target)) targetsBySource.set(target, new Set());
|
||||
}
|
||||
|
||||
const adjacency = new Map<string, readonly string[]>();
|
||||
const selfLoops = new Set<string>();
|
||||
for (const [source, targets] of targetsBySource) {
|
||||
adjacency.set(source, [...targets].sort());
|
||||
if (targets.has(source)) selfLoops.add(source);
|
||||
}
|
||||
return { adjacency, nodes: [...adjacency.keys()].sort(), selfLoops };
|
||||
}
|
||||
|
||||
interface CircuitSearch {
|
||||
readonly cycles: string[][];
|
||||
readonly cycleLimit: number;
|
||||
readonly workLimit: number;
|
||||
/** Search effort so far, across the SCC passes and the circuit search alike. */
|
||||
work: number;
|
||||
/** Non-null once a bound is hit; every loop unwinds on it. */
|
||||
exceeded: ImportCycleLimit | null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Charge `amount` units of search effort. Returns true once the budget is
|
||||
* spent, which every caller must honour — a bulk charge that is not checked
|
||||
* would let the search run on past the bound it just crossed.
|
||||
*/
|
||||
function overBudget(search: CircuitSearch, amount = 1): boolean {
|
||||
search.work += amount;
|
||||
if (search.work <= search.workLimit) return false;
|
||||
search.exceeded = 'work';
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Strongly connected components of the subgraph induced by `allowed`, via an
|
||||
* iterative Tarjan. Iterative because import graphs reach 10^5 files and a
|
||||
* recursive walk would blow the stack long before that.
|
||||
*
|
||||
* `roots` fixes the order the outer loop starts from, which is what makes the
|
||||
* component set — and so Johnson's choice of root — deterministic.
|
||||
*
|
||||
* Abandons the pass and returns a partial list if the work budget runs out, so
|
||||
* every caller must check `search.exceeded` before using the result.
|
||||
*/
|
||||
function stronglyConnectedComponents(
|
||||
roots: readonly string[],
|
||||
adjacency: ReadonlyMap<string, readonly string[]>,
|
||||
allowed: ReadonlySet<string>,
|
||||
search: CircuitSearch,
|
||||
): string[][] {
|
||||
const index = new Map<string, number>();
|
||||
const lowLink = new Map<string, number>();
|
||||
const onStack = new Set<string>();
|
||||
const pending: string[] = [];
|
||||
const components: string[][] = [];
|
||||
let counter = 0;
|
||||
// One pass over the roots happens even for a component with no edges left.
|
||||
if (overBudget(search, roots.length)) return components;
|
||||
|
||||
for (const root of roots) {
|
||||
if (index.has(root)) continue;
|
||||
index.set(root, counter);
|
||||
lowLink.set(root, counter);
|
||||
counter += 1;
|
||||
pending.push(root);
|
||||
onStack.add(root);
|
||||
const frames = [{ node: root, nextIndex: 0 }];
|
||||
|
||||
while (frames.length > 0) {
|
||||
const frame = frames[frames.length - 1];
|
||||
const neighbors = adjacency.get(frame.node) ?? [];
|
||||
if (frame.nextIndex < neighbors.length) {
|
||||
const next = neighbors[frame.nextIndex];
|
||||
frame.nextIndex += 1;
|
||||
if (overBudget(search)) return components;
|
||||
if (!allowed.has(next)) continue;
|
||||
if (!index.has(next)) {
|
||||
index.set(next, counter);
|
||||
lowLink.set(next, counter);
|
||||
counter += 1;
|
||||
pending.push(next);
|
||||
onStack.add(next);
|
||||
frames.push({ node: next, nextIndex: 0 });
|
||||
} else if (onStack.has(next)) {
|
||||
lowLink.set(frame.node, Math.min(lowLink.get(frame.node)!, index.get(next)!));
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
frames.pop();
|
||||
const node = frame.node;
|
||||
if (lowLink.get(node)! === index.get(node)!) {
|
||||
const component: string[] = [];
|
||||
for (;;) {
|
||||
const member = pending.pop()!;
|
||||
onStack.delete(member);
|
||||
component.push(member);
|
||||
if (member === node) break;
|
||||
}
|
||||
components.push(component);
|
||||
}
|
||||
if (frames.length > 0) {
|
||||
const parent = frames[frames.length - 1].node;
|
||||
lowLink.set(parent, Math.min(lowLink.get(parent)!, lowLink.get(node)!));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return components;
|
||||
}
|
||||
|
||||
/** A component that contains at least one cycle: two-plus members, or a self-import. */
|
||||
function isCyclic(component: readonly string[], selfLoops: ReadonlySet<string>): boolean {
|
||||
return component.length > 1 || selfLoops.has(component[0]);
|
||||
}
|
||||
|
||||
function leastNode(nodes: readonly string[]): string {
|
||||
let least = nodes[0];
|
||||
for (const node of nodes) if (node < least) least = node;
|
||||
return least;
|
||||
}
|
||||
|
||||
/** Order components by their least node. Components are disjoint, so this is total. */
|
||||
function byLeastNode(left: readonly string[], right: readonly string[]): number {
|
||||
const leftLeast = leastNode(left);
|
||||
const rightLeast = leastNode(right);
|
||||
return compareCodeUnits(leftLeast, rightLeast);
|
||||
}
|
||||
|
||||
/**
|
||||
* Johnson's `UNBLOCK`, iterative. Lifts `node` and everything transitively
|
||||
* waiting on it out of `blocked`, so a path that was abandoned as fruitless
|
||||
* becomes explorable again once the reason it was fruitless is gone.
|
||||
*/
|
||||
function unblock(node: string, blocked: Set<string>, blockedBy: Map<string, Set<string>>): void {
|
||||
const stack = [node];
|
||||
while (stack.length > 0) {
|
||||
const current = stack.pop()!;
|
||||
blocked.delete(current);
|
||||
const waiting = blockedBy.get(current);
|
||||
if (waiting === undefined || waiting.size === 0) continue;
|
||||
for (const dependent of waiting) {
|
||||
if (blocked.has(dependent)) stack.push(dependent);
|
||||
}
|
||||
waiting.clear();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Johnson's `CIRCUIT`, iterative — enumerate the elementary circuits rooted at
|
||||
* `root` inside `allowed`.
|
||||
*
|
||||
* Every circuit found here starts and ends at `root`, and `root` is the least
|
||||
* node of `allowed` by construction, which is where the rotation guarantee in
|
||||
* the module docblock comes from.
|
||||
*/
|
||||
function enumerateCircuitsFrom(
|
||||
root: string,
|
||||
adjacency: ReadonlyMap<string, readonly string[]>,
|
||||
allowed: ReadonlySet<string>,
|
||||
search: CircuitSearch,
|
||||
): void {
|
||||
const blocked = new Set<string>([root]);
|
||||
const blockedBy = new Map<string, Set<string>>();
|
||||
const path = [root];
|
||||
// `neighbors` rides the frame: the list is fixed for a node, while this loop
|
||||
// re-enters per DFS STEP — ~2.6M iterations against 395k pushes on this
|
||||
// repository, so looking it up per iteration re-hashes the path each time.
|
||||
const frames = [
|
||||
{ node: root, nextIndex: 0, foundCircuit: false, neighbors: adjacency.get(root) ?? [] },
|
||||
];
|
||||
|
||||
while (frames.length > 0) {
|
||||
// Budget spent: return rather than unwind. Everything this function owns is
|
||||
// local and it returns void, so draining the stack would run the full
|
||||
// `blockedBy` bookkeeping (or an `unblock` walk) per frame, to no effect,
|
||||
// on exactly the graphs already judged too expensive.
|
||||
if (search.exceeded !== null) return;
|
||||
const frame = frames[frames.length - 1];
|
||||
const neighbors = frame.neighbors;
|
||||
|
||||
if (frame.nextIndex < neighbors.length) {
|
||||
const next = neighbors[frame.nextIndex];
|
||||
frame.nextIndex += 1;
|
||||
if (overBudget(search)) continue;
|
||||
if (!allowed.has(next)) continue;
|
||||
if (next === root) {
|
||||
// `path` is the elementary path root -> ... -> frame.node; closing it
|
||||
// back onto the root yields the cycle in the documented shape. A
|
||||
// self-import lands here on the first step with path === [root].
|
||||
search.cycles.push([...path, root]);
|
||||
frame.foundCircuit = true;
|
||||
// One-past, matching the edge-limit guard in `check`: `cycleLimit`
|
||||
// cycles is an acceptable answer, and it takes finding one MORE to
|
||||
// prove the graph overflowed. Stopping at `>= cycleLimit` would fail a
|
||||
// graph that has exactly that many cycles and could have been reported
|
||||
// in full.
|
||||
//
|
||||
// Tested before the emission charge so that a graph over both bounds
|
||||
// reports the cycle cap, which is the one a reader can act on, rather
|
||||
// than whichever happened to trip first.
|
||||
if (search.cycles.length > search.cycleLimit) {
|
||||
search.exceeded = 'cycles';
|
||||
continue;
|
||||
}
|
||||
// A found cycle is not merely traversed: it is copied, retained until
|
||||
// the call returns, sorted, and serialized into an MCP response. So it
|
||||
// is charged at EMITTED_NODE_COST per node, not 1. This is what bounds
|
||||
// MEMORY as well as time — 10,000 cycles is a modest cap when cycles
|
||||
// are four files long and a heap-exhausting one when a single strongly
|
||||
// connected component is 50,000 files around.
|
||||
overBudget(search, (path.length + 1) * EMITTED_NODE_COST);
|
||||
continue;
|
||||
}
|
||||
if (!blocked.has(next)) {
|
||||
blocked.add(next);
|
||||
path.push(next);
|
||||
frames.push({
|
||||
node: next,
|
||||
nextIndex: 0,
|
||||
foundCircuit: false,
|
||||
neighbors: adjacency.get(next) ?? [],
|
||||
});
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Leaving `frame.node`. If it reached the root, it may lie on further
|
||||
// circuits, so it and its waiters go back in play. If it did not, it is
|
||||
// recorded as a dead end on each of its successors: it stays blocked until
|
||||
// one of them is unblocked, which is the pruning that makes Johnson's
|
||||
// output-sensitive rather than exponential in the graph size.
|
||||
frames.pop();
|
||||
path.pop();
|
||||
if (frame.foundCircuit) {
|
||||
unblock(frame.node, blocked, blockedBy);
|
||||
} else {
|
||||
for (const next of neighbors) {
|
||||
if (!allowed.has(next)) continue;
|
||||
// `set` only when the entry is created: re-setting an existing key
|
||||
// re-hashes the path string for no effect, and this runs once per
|
||||
// out-edge of every unwound frame — measured at 1.06M redundant
|
||||
// `Map.set` calls on this repository's own import graph.
|
||||
let waiting = blockedBy.get(next);
|
||||
if (waiting === undefined) {
|
||||
waiting = new Set<string>();
|
||||
blockedBy.set(next, waiting);
|
||||
}
|
||||
waiting.add(frame.node);
|
||||
}
|
||||
}
|
||||
if (frames.length > 0 && frame.foundCircuit) {
|
||||
frames[frames.length - 1].foundCircuit = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Johnson's outer loop over one cyclic component: search the circuits rooted at
|
||||
* the component's least node, drop that node, and repeat on whatever cyclic
|
||||
* components the remainder falls into.
|
||||
*
|
||||
* Dropping the root is the whole rotation guarantee. Every cycle left after the
|
||||
* drop consists of nodes greater than every root taken so far, so when the
|
||||
* component holding it finally has that cycle's own minimum as its least node,
|
||||
* the cycle is emitted once, rooted there. No other rotation is reachable,
|
||||
* because the other rotations' starting nodes have already been excluded or are
|
||||
* not the component's least.
|
||||
*
|
||||
* Re-decomposing the REMAINDER rather than the original node range also keeps
|
||||
* each pass proportional to what is left: a tangle that falls apart when its
|
||||
* busiest file is removed stops costing anything immediately.
|
||||
*/
|
||||
function enumerateComponentCycles(
|
||||
component: readonly string[],
|
||||
graph: ImportGraph,
|
||||
search: CircuitSearch,
|
||||
): void {
|
||||
// Components still to search. Pushed so that they pop in increasing order of
|
||||
// least node — see the sort below.
|
||||
const stack: string[][] = [[...component]];
|
||||
|
||||
while (stack.length > 0 && search.exceeded === null) {
|
||||
const current = stack.pop()!;
|
||||
const root = leastNode(current);
|
||||
// Scanning for the root and materializing the allowed set both cost one
|
||||
// pass over the component, and both happen once per root, so they are the
|
||||
// O(n^2) term on a component that never splits. Charged, or the budget
|
||||
// would not see the work it exists to bound.
|
||||
if (overBudget(search, current.length)) return;
|
||||
enumerateCircuitsFrom(root, graph.adjacency, new Set(current), search);
|
||||
if (search.exceeded !== null) return;
|
||||
|
||||
const remaining = current.filter((node) => node !== root);
|
||||
if (remaining.length === 0) continue;
|
||||
// Deliberately NOT re-sorted: the SCC set is independent of the order its
|
||||
// roots are visited in, `leastNode` picks Johnson's root regardless, and
|
||||
// the finished cycle list is sorted at the end. Sorting here would add an
|
||||
// O(n log n) term to every root for no observable difference.
|
||||
const decomposed = stronglyConnectedComponents(
|
||||
remaining,
|
||||
graph.adjacency,
|
||||
new Set(remaining),
|
||||
search,
|
||||
);
|
||||
// Same rule as above the call: once the budget is spent the `while` will
|
||||
// refuse to pop whatever we push, so the filter/decorate/sort is waste.
|
||||
if (search.exceeded !== null) return;
|
||||
const subComponents = decomposed
|
||||
.filter((subComponent) => isCyclic(subComponent, graph.selfLoops))
|
||||
.map((subComponent) => ({ least: leastNode(subComponent), nodes: subComponent }))
|
||||
// Descending, so the stack pops them in increasing order of least node —
|
||||
// Johnson's root order, and what makes a budget-stopped run stop at a
|
||||
// deterministic point rather than wherever iteration happened to be.
|
||||
.sort((a, b) => -compareCodeUnits(a.least, b.least));
|
||||
for (const subComponent of subComponents) stack.push(subComponent.nodes);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The shortest cycle through a component's least node, by breadth-first search
|
||||
* across the component.
|
||||
*
|
||||
* This is the fallback when a bound stops the full enumeration: one concrete,
|
||||
* checkable cycle naming each tangle. It is also exactly what this module
|
||||
* returned for every component before elementary enumeration existed, so the
|
||||
* degraded answer is no worse than the old complete answer.
|
||||
*
|
||||
* Linear in the component, and it runs only after the decomposition has already
|
||||
* succeeded, so it cannot fail the way the enumeration did. The budget is
|
||||
*/
|
||||
function representativeCycle(
|
||||
component: readonly string[],
|
||||
adjacency: ReadonlyMap<string, readonly string[]>,
|
||||
): string[] {
|
||||
const allowed = new Set(component);
|
||||
const start = component[0];
|
||||
const start = leastNode(component);
|
||||
const parents = new Map<string, string | null>([[start, null]]);
|
||||
const queue = [start];
|
||||
|
||||
|
|
@ -29,82 +586,73 @@ function findCyclePath(component: string[], adjacency: Map<string, string[]>): s
|
|||
}
|
||||
}
|
||||
|
||||
throw new Error('Invariant violation: no cycle found through SCC root.');
|
||||
// Unreachable: every component reaching here is cyclic, and BFS from its
|
||||
// least node inside the component must close. Thrown rather than returned
|
||||
// empty so a future change that breaks the invariant is loud.
|
||||
throw new Error('Invariant violation: no cycle found through cyclic component root.');
|
||||
}
|
||||
|
||||
/** Element-wise lexicographic order, so the finished list is byte-stable. */
|
||||
function compareCycles(left: readonly string[], right: readonly string[]): number {
|
||||
const shared = Math.min(left.length, right.length);
|
||||
for (let index = 0; index < shared; index += 1) {
|
||||
const order = compareCodeUnits(left[index], right[index]);
|
||||
if (order !== 0) return order;
|
||||
}
|
||||
return left.length - right.length;
|
||||
}
|
||||
|
||||
/**
|
||||
* Return one deterministic concrete cycle for every cyclic strongly connected
|
||||
* component in the file import graph.
|
||||
* Enumerate every elementary cycle in the file import graph.
|
||||
*
|
||||
* The result is discriminated on `enumeration`; see `ImportCycleReport` for
|
||||
* what each variant carries. Past either bound the enumeration is discarded
|
||||
* rather than truncated — see the module docblock for the algorithm, the
|
||||
* rotation rule, and why a partial cycle list is not a safe thing to return.
|
||||
*/
|
||||
export function findImportCycles(edges: ImportEdge[]): string[][] {
|
||||
const adjacency = new Map<string, Set<string>>();
|
||||
for (const { source, target } of edges) {
|
||||
if (!source || !target) continue;
|
||||
const targets = adjacency.get(source) ?? new Set<string>();
|
||||
targets.add(target);
|
||||
adjacency.set(source, targets);
|
||||
if (!adjacency.has(target)) adjacency.set(target, new Set());
|
||||
export function findImportCycles(
|
||||
edges: readonly ImportEdge[],
|
||||
cycleLimit: number = IMPORT_CYCLE_LIMIT,
|
||||
workLimit: number = IMPORT_CYCLE_WORK_LIMIT,
|
||||
): ImportCycleReport {
|
||||
const graph = buildGraph(edges);
|
||||
const allNodes = new Set(graph.nodes);
|
||||
const search: CircuitSearch = { cycles: [], cycleLimit, workLimit, work: 0, exceeded: null };
|
||||
|
||||
const decomposition = stronglyConnectedComponents(graph.nodes, graph.adjacency, allNodes, search);
|
||||
// Only a decomposition that ran to completion has a trustworthy count; one
|
||||
// abandoned mid-pass would undercount silently.
|
||||
const decompositionComplete = search.exceeded === null;
|
||||
const cyclicComponents = decomposition
|
||||
.filter((component) => isCyclic(component, graph.selfLoops))
|
||||
.sort(byLeastNode);
|
||||
|
||||
for (const component of cyclicComponents) {
|
||||
if (search.exceeded !== null) break;
|
||||
enumerateComponentCycles(component, graph, search);
|
||||
}
|
||||
|
||||
const sortedAdjacency = new Map(
|
||||
[...adjacency].map(([node, targets]) => [node, [...targets].sort()] as const),
|
||||
);
|
||||
const reverseAdjacency = new Map<string, string[]>();
|
||||
for (const node of sortedAdjacency.keys()) reverseAdjacency.set(node, []);
|
||||
for (const [source, targets] of sortedAdjacency) {
|
||||
for (const target of targets) reverseAdjacency.get(target)!.push(source);
|
||||
if (search.exceeded !== null) {
|
||||
const reason = search.exceeded;
|
||||
const limit = reason === 'cycles' ? cycleLimit : workLimit;
|
||||
if (!decompositionComplete) return { enumeration: 'none', reason, limit };
|
||||
// Whatever the abandoned enumeration accumulated is discarded — it is a
|
||||
// partial list of elementary cycles and would read as a complete one.
|
||||
// Representatives are a different KIND of list, one per component, and the
|
||||
// report says so in the type.
|
||||
return {
|
||||
enumeration: 'component-representatives',
|
||||
cycles: cyclicComponents
|
||||
.map((component) => representativeCycle(component, graph.adjacency))
|
||||
.sort(compareCycles),
|
||||
componentCount: cyclicComponents.length,
|
||||
reason,
|
||||
limit,
|
||||
};
|
||||
}
|
||||
for (const sources of reverseAdjacency.values()) sources.sort();
|
||||
|
||||
const visited = new Set<string>();
|
||||
const finishOrder: string[] = [];
|
||||
const components: string[][] = [];
|
||||
|
||||
for (const start of [...sortedAdjacency.keys()].sort()) {
|
||||
if (visited.has(start)) continue;
|
||||
visited.add(start);
|
||||
const stack = [{ node: start, nextIndex: 0 }];
|
||||
while (stack.length > 0) {
|
||||
const frame = stack[stack.length - 1];
|
||||
const neighbors = sortedAdjacency.get(frame.node) ?? [];
|
||||
if (frame.nextIndex < neighbors.length) {
|
||||
const next = neighbors[frame.nextIndex++];
|
||||
if (!visited.has(next)) {
|
||||
visited.add(next);
|
||||
stack.push({ node: next, nextIndex: 0 });
|
||||
}
|
||||
} else {
|
||||
finishOrder.push(frame.node);
|
||||
stack.pop();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
visited.clear();
|
||||
for (let index = finishOrder.length - 1; index >= 0; index -= 1) {
|
||||
const start = finishOrder[index];
|
||||
if (visited.has(start)) continue;
|
||||
const component: string[] = [];
|
||||
const stack = [start];
|
||||
visited.add(start);
|
||||
while (stack.length > 0) {
|
||||
const node = stack.pop()!;
|
||||
component.push(node);
|
||||
for (const next of reverseAdjacency.get(node) ?? []) {
|
||||
if (visited.has(next)) continue;
|
||||
visited.add(next);
|
||||
stack.push(next);
|
||||
}
|
||||
}
|
||||
component.sort();
|
||||
components.push(component);
|
||||
}
|
||||
|
||||
return components
|
||||
.filter(
|
||||
(component) =>
|
||||
component.length > 1 || (sortedAdjacency.get(component[0]) ?? []).includes(component[0]),
|
||||
)
|
||||
.sort((a, b) => (a[0] < b[0] ? -1 : a[0] > b[0] ? 1 : 0))
|
||||
.map((component) => findCyclePath(component, sortedAdjacency));
|
||||
return {
|
||||
enumeration: 'complete',
|
||||
cycles: search.cycles.sort(compareCycles),
|
||||
componentCount: cyclicComponents.length,
|
||||
};
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
import fsp from 'node:fs/promises';
|
||||
import path from 'node:path';
|
||||
import { createHash, randomBytes } from 'node:crypto';
|
||||
import { createHash } from 'node:crypto';
|
||||
import lbug from '@ladybugdb/core';
|
||||
import type { LbugValue } from '@ladybugdb/core';
|
||||
import type { BridgeHandle, BridgeMeta, StoredContract, CrossLink, RepoSnapshot } from './types.js';
|
||||
|
|
@ -12,7 +12,7 @@ import {
|
|||
} from '../lbug/lbug-config.js';
|
||||
import { dedupeContracts, dedupeCrossLinks } from './normalization.js';
|
||||
import { createLogger } from '../logger.js';
|
||||
import { retryRename } from '../../storage/fs-atomic.js';
|
||||
import { retryRename, writeFileAtomic } from '../../storage/fs-atomic.js';
|
||||
|
||||
const bridgeLogger = createLogger('bridge-db', {
|
||||
debugEnvVar: 'GITNEXUS_DEBUG_BRIDGE',
|
||||
|
|
@ -647,30 +647,7 @@ export async function closeBridgeDb(handle: BridgeHandle): Promise<void> {
|
|||
/* ------------------------------------------------------------------ */
|
||||
|
||||
export async function writeBridgeMeta(groupDir: string, meta: BridgeMeta): Promise<void> {
|
||||
const target = path.join(groupDir, 'meta.json');
|
||||
// Unpredictable suffix + O_EXCL via `'wx'` flag closes the symlink/
|
||||
// pre-create attack window. The third argument `0o600` is the
|
||||
// user-only mode mask — CodeQL's `js/insecure-temporary-file` query
|
||||
// sources its verdict from the `mode` argument, NOT from `flags`:
|
||||
// its `isSecureMode(mode)` predicate requires the low 6 bits to be
|
||||
// zero (no group/world bits). Without an explicit mode the file is
|
||||
// created with the process umask (typically 0o644 = group/world
|
||||
// readable), which the query treats as the actual vulnerability.
|
||||
// Both `'wx'` (runtime O_EXCL) AND `0o600` (CodeQL-credited mode)
|
||||
// are needed: one closes the symlink race, the other closes the
|
||||
// permissions exposure.
|
||||
const tmp = `${target}.tmp.${randomBytes(8).toString('hex')}`;
|
||||
const handle = await fsp.open(tmp, 'wx', 0o600);
|
||||
try {
|
||||
await handle.writeFile(JSON.stringify(meta, null, 2), 'utf-8');
|
||||
} finally {
|
||||
await handle.close();
|
||||
}
|
||||
// Use retryRename for consistency with writeBridge's atomic swap — on
|
||||
// Windows a concurrent reader can cause EBUSY/EPERM even on a tiny
|
||||
// meta.json, and we don't want meta write to be less robust than the
|
||||
// bridge.lbug swap it accompanies.
|
||||
await retryRename(tmp, target);
|
||||
await writeFileAtomic(path.join(groupDir, 'meta.json'), JSON.stringify(meta, null, 2));
|
||||
}
|
||||
|
||||
export async function readBridgeMeta(groupDir: string): Promise<BridgeMeta> {
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Reference in a new issue