mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-01 02:01:24 +00:00
527 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d2e51c144e
|
Merge branch 'main' into feat/3371-ipynb-python-indexing | ||
|
|
51fb64c976
|
fix(swift): model Swift modules like the compiler (nested packages, Xcode targets, linear visibility) (#3387)
* fix(swift): discover nested Package.swift manifests for module grouping A Swift monorepo laid out as Core/<pkg>/Package.swift has no root manifest and no root Sources/, so every Swift file fell into one __default__ module. The Swift resolver now walks the repo (bounded, skipping dot, ignored, and Xcode bundle directories) and adds each nested package's targets, keyed by repo-relative directory and ordered deepest-first so first-match grouping picks the most specific target. Import resolution keeps the root view. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(swift): bound implicit IMPORTS edges by a total budget Implicit same-module IMPORTS are n*(n-1) edges per module, and the graph keeps every relationship in one Map capped by V8 at 2^24 entries. Emission now fills a 4M-edge budget smallest module first and skips, with a warning, any module that does not fit, so no module layout can crash analyze. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(swift): skip pairwise sibling passes for oversized modules The target-siblings and sibling-type-bindings passes copy every file's declarations into every other file of a module, so heap grows as n^2: about 3.8 GB at 1,000 files with 15 defs each, about 15 GB at 2,000. Modules over 1,000 files now skip both passes with a warning and resolve through the global name fallback. GITNEXUS_SWIFT_MAX_MODULE_FILES changes the ceiling. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(swift): cover nested SwiftPM packages end to end A fixture with three nested Package.swift manifests (two declaring a target named Net) and no root manifest. Each target gets its own implicit IMPORTS, none cross packages, and Config() resolves to the caller's own package. The Swift capture golden and scope-capture fingerprint grow with the new fixture corpus only. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(review): apply review findings - Rebase nested target paths with the existing Zig path helpers, which also reject Windows drive paths and paths that resolve to the repo root (whose empty prefix would group every file into one target). - Document GITNEXUS_SWIFT_MAX_MODULE_FILES in the README env table. - State that skipped modules resolve through lower-confidence fallback edges, why nested-type fragments still run for oversized modules, and fix a stale loader name in the target-grouping header. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(swift): cover every skipped Xcode bundle suffix in the package walk Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(swift): model modules the way the compiler does Replace the caps from the first round with representations that stay linear, and derive module identity from the same sources the compiler uses. - Same-module visibility is one File -> Module IMPORTS edge per file (reason `module-membership`) instead of an edge per ordered file pair. Incremental importer expansion treats files sharing a Module hub as importers of each other. The 4M edge budget is gone. - Sibling declarations and type bindings live in one shared table per module in the namespace channel C# uses since #1871, instead of being copied into every file. The 1,000-file ceiling and GITNEXUS_SWIFT_MAX_MODULE_FILES are gone. - Modules come from root and nested SwiftPM manifests (Sources, Source, src, srcs; plugins under Plugins; the newest Package@swift-X.Y.swift) and from Xcode native targets in project.pbxproj, including Xcode 16 synchronized folders. Target paths are matched from the repo root, so a vendored copy of the same layout is no longer grouped into a root target (this reverses #2931's floating match). - `import X` resolves to the modules named X instead of any folder named X. A name no module carries is external when every manifest and project was read completely. - The global-name-fallback veto uses the same membership, so test targets, custom-path targets and Xcode targets have module identity. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(bench): move the Swift package bench to module grouping The bench still imported the removed groupSwiftFilesBySpmTarget. Its first-wins probe keeps its meaning under root-anchored grouping: the clash file sits under Sources/Mod0, and the Sources/Mod1 further down its path is a vendored copy. Plugins are now non-importable modules under Plugins/, so parse_targets counts importable source targets (still 3) and parse_binary_skipped also checks that the plugin is recorded that way. Baseline values unchanged; the notes say why the definitions moved. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(swift): match compiler module names, manifests, and target filters - Module names follow the compiler's c99 mangling (`my-lib` imports as `my_lib`); Xcode targets use a literal PRODUCT_MODULE_NAME or PRODUCT_NAME when the project sets one. - Package manifests are one-file modules, as SwiftPM compiles them. Files outside every target are one-file modules once every manifest and project was read; otherwise they keep the shared __default__ module. - Xcode 16 synchronized-folder exceptions add a file to a target that does not list the folder, and remove it from one that does. - SwiftPM `sources:` / `exclude:` narrow a target; a computed list marks the manifest unreadable. - The default target folder is chosen once per package, as SwiftPM does, instead of per target. Refs #3355 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address PR review feedback (#3387) - Inferred Swift folders keep SwiftPM's first predefined parent when a target name repeats (Sources before srcs). - The pbxproj parser rejects a \U escape without four hex digits instead of decoding garbage, so the project reads as incomplete. - Extension owners are stamped in every Xcode membership of a shared file. - A plugin name never reaches a plugin: neither the fallback veto nor the folder-index fallback treats a known non-importable module as imported. - A file an Xcode target compiles keeps that membership even when it also lies under a SwiftPM target directory. - The workspace scan bounds the queue, not only the directories read. - Integration tests require each module hub to exist before comparing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
e8b5feb36b
|
Merge branch 'main' into feat/3371-ipynb-python-indexing | ||
|
|
0261982d9a
|
fix(analyze): make --memory-budget set the real heap and report rebuild reasons once (#3386)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* feat(analyze): --memory-budget flag with heap-limit override and worker-pool degradation (#3137) Adds an explicit `--memory-budget <mb>` CLI flag that overrides the RAM/cgroup auto-sized main-thread heap ceiling for the parse phase: - CLI validation (integer >= 200 MB) before bar.start(), matching the --workers pattern - Threaded CLI → runFullAnalysis → PipelineOptions → parse-impl as memoryBudgetBytes - parse-impl resolves the heap limit as budget ?? v8.heap_size_limit, so both the preflight projection warning and the #2649 mid-loop abort probe honor the budget - Graceful degradation: when the projected heap need exceeds the budget at the computed pool size, the pool shrinks (never below 1, never above the operator's --workers) before sub-batch math and pool construction, so all downstream consumers see the degraded size Omitting the flag keeps the auto-sizer path byte-identical. Refs #3137 * feat(analyze): collapse rebuild-gate log into one summary + persist needsFullRebuild verdict (#3137) The nine meta-mismatch rebuild gates (pdg mode, content retention, schema fingerprint, graph-write collapse, analysis features, Spring vendor prefixes, runner identity, FTS CJK mode, embedding dims) each logged individually and set force:true independently. An upgrade that trips several at once printed a scattered wall of near-identical warnings. - Gates now collect into rebuildReasons[]; a single summary block prints them (inline for one, numbered for many) and sets force once. Per-gate Tip text is preserved verbatim inside the entries. - The verdict persists to meta (needsFullRebuild: {reasons, recordedAt}) BEFORE the rebuild starts. If the rebuild is interrupted, the next run announces the recorded reasons up front instead of quietly attempting an incremental write on a half-rebuilt index — the gates may not all re-fire against a wiped DB. - The verdict is cleared on the next successful completion (the final meta does not carry the field forward). Semantics unchanged: every gate was already evaluated (none early-returns), force is idempotent, and a rebuild happens iff at least one reason fired. Refs #3137 * refactor(cli): share one integer flag parser across analyze, watch, and wiki Replace the duplicated Number.isInteger checks for --workers, --embeddings, the positive env-backed analyze flags, the watch interval flags, and wiki's --timeout/--retries with parseIntegerOption (per-flag minimum, optional scale for the safe-integer bound). User-facing messages are unchanged. * fix(analyze): make --memory-budget set the real V8 heap through the respawn The budget now drives ensureHeap's existing respawn instead of a parse-phase override, so the #2649 preflight, mid-loop abort, remedy text, and GC pacing all see one heap limit. The respawn sizes old space plus three semi-spaces to equal the budget, the child resolves as already at the budget (no second respawn), and GITNEXUS_HEAP_LIMIT_SOURCE drives budget-aware OOM advice. Budget validation moves to the preAction hook so analyze and watch reject a bad value before any respawn. Removes the pool-shrink block and the memoryBudgetBytes plumbing through PipelineOptions and run-analyze. * docs(analyze): describe --memory-budget accurately and translate its help The help text claimed graceful worker-pool degradation, which no longer exists; it now says the flag sets the main-thread V8 heap and that parse workers keep their own caps. Wires the option through the help i18n map with en and zh-CN strings, and documents it in both READMEs and the out-of-memory troubleshooting section. * feat(analyze): add a pure rebuild-reason collector One collector per run holds keyed rebuild reasons, merges by key, flattens reasons stored by an interrupted rebuild into one recovery entry, validates stored reasons on read, and formats the single up-front summary plus one follow-up line for reasons added after the pipeline. * fix(analyze): route every forced rebuild through one reason collector Every path that forces a full rebuild (the nine meta gates, --force, --skills, --no-parse-cache, --drop-embeddings, --repair-fts retention, Spring Actuator, AsyncAPI, shared-store graph gaps, dirty-flag recovery, the post-pipeline capability gate, and the #2409 escalation) now adds a keyed reason to one collector. The rebuild decision is applied from the collector at fixed checkpoints, one summary prints right before the pipeline, and late reasons print one follow-up line. The escalation stays non-forcing. runFullAnalysis returns the collected keys, which replaces the runner-identity source-regex test with a behavior test. Removes the separate needsFullRebuild field and its announcement, and stops folding --skills and --no-parse-cache into --force. * fix(analyze): persist rebuild reasons on the existing crash marker Every incrementalInProgress writer (the full-rebuild stamp before the wipe, the incremental pre-write, saveIncrementalDirtyState including the #2409 escalation, and buildFtsDirtyStamp) now carries the collected reasons into the active slot's metaDir, so an interrupted rebuild explains itself on the next run through one merged recovery entry. A successful run still clears the marker and its reasons; the FTS-park recovery clears them without forcing. * test(analyze): cover every rebuild-reason key through runFullAnalysis Add a coverage table that the typechecker keeps complete: every RebuildReasonKey maps to a test file that drives it through runFullAnalysis and asserts the returned key. Adds the missing graph-write-collapse and drop-embeddings drivers, asserts the key in the existing pdg-mode, spring-vendor-prefixes, cjk-segmentation, and embedding-dims tests, and removes plan-local IDs from test names and comments. * fix(review): apply review findings - A --max-old-space-size pin equal to --memory-budget no longer counts as the exact budget heap (V8 adds the young generation on top); only the budget-respawned child skips the respawn, so the limit really equals the budget. - Snapshot the analyze env before ensureHeap and restore GITNEXUS_HEAP_LIMIT_SOURCE, so a kept process does not leak its heap source into a later programmatic analyzeCommand call. - --skills and --no-parse-cache keep the forced storage requirements they had before force stopped being folded from them. - Merge the duplicated follow-up announcement into one helper and fix a stale --drop-embeddings comment. - The rebuild-reason coverage table no longer greps driver files for the key string; add tests for a programmatic invalid budget and the multi-cause interrupted-rebuild text. * fix(review): don't announce the escalated write as a full rebuild The #2409 escalation is a non-forcing reason, but its follow-up line used the 'Full rebuild also required' lead. A follow-up that carries only non-forcing reasons now leads with 'Write plan changed'. * docs(analyze): document GITNEXUS_HEAP_LIMIT_SOURCE in the env table CONTRIBUTING requires every new GITNEXUS_* variable to have a row; this one is internal (set by analyze itself) and exists so OOM advice points at --memory-budget. * fix(review): address GitNexus review threads on #3386 - heapPressureRemedy measures pressure against the real auto-sized cap (heapCapMbFor) instead of a flat 0.75 x RAM, and no longer tells a GITNEXUS_MEMORY=off run with no pin to drop a pin that does not exist. - toStored() persists the interrupted rebuild's reasons first, as documented. - ensureHeap's doc names which paths leave GITNEXUS_HEAP_LIMIT_SOURCE unset. - The heap-respawn suite restores the caller's GITNEXUS_MEMORY. - The non-forcing follow-up test rejects any 'full rebuild' wording. * fix(review): require both budget flags and check key coverage at runtime - A budget-respawned child is recognized only when the inherited heap-source marker comes with both the budget's old-space and semi-space flags; the marker alone is an inherited env var, not proof. The old-space parser is generalized to any V8 size flag instead of copying its regex. - REBUILD_REASON_KEYS is exported and RebuildReasonKey derives from it, so the coverage table is checked at runtime (CI does not type-check test files). * fix(review): don't claim a full rebuild in a non-forcing summary formatSummary and formatFollowUp now share one leadFor helper, so a block of only non-forcing reasons reads 'Write plan changed' in both. --------- Co-authored-by: ChunxueLi <mecoloud@users.noreply.gitee.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> |
||
|
|
bb5c9c84cf
|
Merge branch 'main' into feat/3371-ipynb-python-indexing | ||
|
|
5fc518d2cb
|
fix(dart): resolve package imports by pubspec identity (#3369)
* fix(dart): resolve package imports by pubspec identity * fix(dart): keep package-identity edges out of the cycle check Pubspec identity edges invalidate importers when a manifest changes. They cannot form an init cycle, so the cycle query excludes them before the row cap. Discovery reads each manifest once, with a size bound, and resolution shares one package-URI parser with those edges. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3369) - Reject package URIs with an empty library path so they do not emit identity edges - Skip the pubspec permission test where chmod cannot deny reads - Document that the Dart heap probe is not a uniqueTarget spelling Note: pre-existing failure in test/unit/incremental-index-extension-dml-gate.test.ts not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): list package directories through a no-follow descriptor A directory replaced by a symlink between the parent listing and the next visit must not be traversed. The walk opens it with O_DIRECTORY|O_NOFOLLOW and lists that inode. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * fix(mcp): keep the Dart identity reason out of MCP startup The cycle query still excludes the same reason string. The constant now lives with the other non-initializing import reasons, so MCP startup does not load a language provider. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3369) Open discovered pubspecs and child directories through the parent directory inode on Linux, so replacing that directory with a symlink cannot redirect the walk. Note: pre-existing failure in test/unit/incremental-index-extension-dml-gate.test.ts (worker pool startup timeout) not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): reject Windows junctions during pubspec walk * Address PR review feedback (#3369) Refuse pubspec discovery that cannot set O_NOFOLLOW, and verify macOS child opens against the pinned directory chain instead of reopening a mutable path. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): bound live pubspec descriptors and close the macOS check-then-open A deep directory chain held one descriptor per level until open failed with EMFILE, and macOS child opens statted the path before using it. Refuse the next directory at 64 live handles, and stat only the descriptor opened with O_NOFOLLOW. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): open pubspecs non-blocking so a FIFO cannot hang discovery A listed pubspec can be replaced by a FIFO before open. O_RDONLY alone waits inside open for a writer, so the file-type check never runs. O_NONBLOCK returns immediately and the walk rejects the non-regular file. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): cap names read from each pubspec directory readdir kept every entry before the visit budget could run, so one huge directory could allocate without bound. Read the listing one name at a time and fail closed past 100,000 entries. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dart): reject a pubspec that grows while its descriptor is read The size cap was taken from the stat before the read, so a file that grew in that window could be parsed from a short prefix. Re-stat the same descriptor afterward and fail closed when the size no longer matches the bytes captured. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): rebaseline the Dart scope-capture fingerprint for package-import fixtures The benchmark hashes every dart-* fixture. The new package-import corpus adds six Dart files and 33 capture groups. Parking that directory restores the previous fingerprint, so this is corpus growth, not a capture change. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
6f1a9eda39 |
Merge remote-tracking branch 'origin/main' into feat/3371-ipynb-python-indexing
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # gitnexus/src/storage/parse-cache.ts # gitnexus/test/unit/incremental-parse-cache.test.ts |
||
|
|
e06c2dd13e
|
fix(impact): fail closed on id-less targets and follow ??/||/?: callable values (#3354) (#3373)
* fix(impact): fail closed on id-less targets and follow ??/||/?: callable values (#3354) #3354 reports `impact` returning a byte-identical 1037/CRITICAL/`exact` result for three unrelated targets, with only `target.name` differing. The reporter's `target` had no `id` and no `filePath`, which the traversal path always emits, and a synthetic reproduction of their monorepo (pnpm, Cloudflare worker-configuration.d.ts in five packages, Hono, a Durable Object) resolves every target correctly on main. So the identical result was not reproduced. The repro did surface two real gaps and one hardening point: - `_runImpactBFS` now throws when the target has no node id. Every caller already catches, so impact reports `impactedCount: null, risk: UNKNOWN` instead of a normal-looking blast radius that cannot be about this symbol. - Callable-value flow followed only a single designator on the RHS, so `const sweep = env.__sweep ?? runSweep; await sweep(env)` produced no flow and `scheduled` was missing as a caller of `runSweep`, while the result still claimed `epistemic: exact`. Each branch of `??`, `||`, `or`, and `?:` now flows into the binding (language-neutral: operator and field names, no language checks). Parse cache bumped 104 -> 112 (105-111 are claimed by open PR #3326). - A whitespace-only `target_uid` (strict adapters materialize omitted optional strings) is treated as omitted and falls back to the name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(mcp): treat a non-string uid as omitted instead of throwing (#3354) Review follow-up on #3373. `query.uid?.trim()` called `.trim()` on a client-supplied value, and the MCP envelope is not type-validated, so `context({uid: 42})` threw a TypeError that `context()` does not catch. Before #3373 the same input ended as a structured not_found. A non-string uid now counts as omitted, which matches how normalizeToolParams already treats a non-string `target_uid`. The impact and trace not-found messages print a trimmed uid only when it is a non-blank string, so a strict adapter sending " " sees the name it searched for instead of `' '`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ingestion): expand ??/?:/ternary branches through a provider hook (#3354) Review follow-up on #3373. The shared value-alternatives rule keys on tree-sitter field names, and several grammars spell the same construct differently, so the expansion never fired for them: - Kotlin `elvis_expression` has no fields. - Swift uses `value`/`if_nil` and `if_true`/`if_false`. - Dart uses `first`/`second`, and its conditional has no `condition` field. - Python `a if c else b` has no fields. The `'?:'` operator entry was dead, since no bundled grammar emits it. Add an optional `valueAlternatives` hook to CallableFlowCaptureOptions, consulted before the shared rule, and implement it in the Kotlin, Swift, Dart, Python and Ruby providers. Shared code still names no language. Ruby's statement-bodied `if`/`unless`/`elsif` also carries `condition`/`consequence`/`alternative`, so the shared ternary rule dug an identifier out of an arbitrary statement (`g = h; 0` flowed `h`) and produced a wrong CALLS edge. The Ruby hook now keeps a multi-statement branch as one opaque source, as before #3373. Tests: new provider fixtures for Python, Kotlin, Swift, Dart and Ruby (including a Ruby negative), and TS chain, parenthesized, callable-left and `&&` negative cases. Captures goldens gain one entry each for the new fixtures. SCHEMA_BUMP stays 112 (same unreleased PR). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ingestion): keep long ||/??/?: chains linear in callable-flow capture (#3354) Expanding value-selecting sources into per-branch flows made long chains super-linear and, past a few thousand operands, a stack overflow. A generated `w === "k0" || w === "k1" || ...` keyword table took 6.7 s at 1000 operands and 48 s at 2000, and threw RangeError at 8000. Two causes, both fixed without changing any emitted capture: - valueAlternatives recursed once per operator level and spread the partial results at every level. It is now an explicit stack that writes to one output array. The left-to-right order, the paren unwrapping, and the provider-hook contract (`[node]` means opaque) are unchanged. - Every alternative ran its visibility walks from its own leaf, which can be as deep as the chain is long, all the way to the root. Also, tree-sitter's `parent` re-descends from the root, so each step costs the node's depth. The walks now jump between "anchor" nodes, the only nodes any check can match: region ids and formal owners. The nearest anchor is memoized per node across the file, and parents come from a map recorded by the one DFS the synthesizer already does. After the fix: 0.34 s / 0.23 s / 1.6 s at 1000 / 2000 / 8000 operands. That is within about 1.2x of main without the expansion; what remains is tree-sitter query time. Capture fingerprints are byte-identical before and after the fix for all 16 scope-capture bench languages and for the Python harness. A `typescript-deep-chain` case added to the scope-capture bench guards the scaling: 1.11-1.22 now, 7.5-8.0 before. A unit test pins a 10000-operand chain: it must still yield the seed for a callable operand, the copy for a formal at the deepest leaf, and the invoke. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(kotlin): see through braced if-branches in callable alternatives (#3354) tree-sitter-kotlin wraps a braced branch as `control_structure_body > statements > <expr>`, so for any non-empty block kotlinValueAlternatives saw exactly one named child (`statements`) and pushed the wrapper itself as the branch value. operandSyntax emits nothing for a `statements` node, so `val run = if (c) { ::f } else { ::g }; run()` produced no flow edges, and the "multi-statement block stays opaque" guard could never fire. Descend one level through `statements` and require exactly one non-comment expression there. Empty blocks (`{}` has no named children), multi-statement blocks, and `if` without `else` still return the whole `if` as one opaque source. The doc comment now describes that. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ruby): skip only the multi-statement branch in callable alternatives (#3354) rubyValueAlternatives returned `[node]` (opaque) as soon as one branch held more than one statement. Two things went wrong because of that: - At the top level, `run = if c then g = h; 0 else method(:f) end` dropped the single-statement `else`, so `f` got no flow edge. - In an elsif chain, the outer `if` pushed the `elsif` node as a branch. The shared loop called the hook on it again, got `[elsif]` back, and used the whole elsif subtree as one source. So in `if a then method(:run_a) elsif b then g = h; 0 else method(:run_b) end` the `run_b` edge was lost. The hook now walks the elsif chain itself and skips each multi-statement branch, while every single-statement branch still becomes an alternative. This cannot add a wrong edge: each emitted alternative is a value the conditional really evaluates to, and nothing is taken from the skipped branch. Before, that branch did not contribute a resolvable value either. A whole conditional used as a source becomes a qualified-name seed, which resolves to nothing when it has more than one identifier leaf. With a single identifier leaf, it could even seed the condition variable. When no branch is a single statement, the conditional still stays one opaque source, as before. Ruby captures golden: ruby-callable-alternatives/app.rb goes from 33 to 58 capture groups. 23 of them come from the new fixture functions (checked against the old source). The other 2 are the new branch alternatives, `statement_if -> run_sweep` and `elsif_chain -> run_b`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ingestion): flow the right operand of && / and into callable bindings (#3354) `a && b` / `a and b` yields `a` when it is falsy and `b` otherwise. A falsy value is never a callable, so the right operand is the only one that can be invoked later. The capture left `&&` unexpanded, and the compound source became a qualified seed that resolves to nothing: `const run = x && f; run()` gained no edge to `f` while impact claimed `exact`. Python `x and f or g` reached only `g`. The shared expansion now maps `&&` / `and` to the right branch only. It recurses, so `x and f or g` reaches both `f` and `g`. Where `&&` yields a boolean (Java, C#, Go, Rust, C, C++, PHP, Zig), the destination cannot be invoked, so the flow never meets a call. A before/after CALLS diff over all 83 lang-resolution fixtures that contain `&&` or `and` shows exactly one new edge, logicalAnd -> runAndRight. PHP and Ruby bind low-precedence `and` looser than `=`, so `$g = $x and $y` never reaches this rule. The Python golden digest changes only for the extended python-callable-alternatives fixture. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ingestion): keep operator branches of a value-selecting source opaque (#3354) The fan-out sent every branch of `??` / `||` / `or` / `?:` to emitAssignmentFact on its own, including branches that compute a value. `Handlers.fallback === run || fb` then emitted the comparison as a seed whose qualified text sliced to receiver `Handlers`, member `run`, and resolveBoundMemberCandidates minted a CALLS edge to `Handlers.run`. The same happened for `this.state !== run ?? this.fallback`, `this.state + run || this.fallback`, and Python `self.state != run or self.fallback`. Before the fan-out, the whole compound source was one opaque seed. A branch that is a binary operator expression now contributes nothing. The check uses the field vocabulary valueBranches already reads (a `left`/`right` pair, or an `operator`/`operators`/`op` token after the expression start), not grammar type names. Member accesses that field their `.`/`->` as `operator` (Ruby `call`, C/C++ `field_expression`) also field a member name through the list memberParts uses, now shared as memberNameNode, so they stay designators. Unary `&f`/`*fp` lead with their operator and stay designators. Call results were already dropped by emitAssignmentFact, and lambdas and callable references are unchanged. CALLS-edge diff over 87 lang-resolution fixtures (505 -> 501 edges): only the four false edges above were removed, and none were added. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(ingestion): pin the LEFT operand of ?? / or / elvis per provider (#3354) The Python, Kotlin, Swift and Dart "override ?? fn" cases put an unresolvable parameter on the left, so a fan-out that kept only the last operand still passed them. Each fixture now adds a callableLeft case with a real function on the left and a parameter on the right (`run_left or fallback`, `::runLeft ?: fallback`, `runLeft ?? fallback`), and each language asserts `callableLeft → runLeft`. Mutation check: dropping the left branch from the shared `??`/`||`/`or` rule and making the four provider hooks return only their last operand fails all four new tests, while the existing right-operand tests still pass. Capture goldens were regenerated with UPDATE_GOLDEN=1: - python app.py: 35 -> 85 groups. 35 -> 75 is stale drift from |
||
|
|
353c30d33f |
fix(review): keep notebook snippets parseable and File nodes honest
Parse Python snippets whose path is still .ipynb, treat the extract cache as an LRU, and stop expecting skipped notebooks to lack File nodes. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
5535df854d |
fix(review): keep notebook JSON coordinates and extract once
Review findings: last-wins cells arrays, source-string line anchors, skip cells without source, reuse worker extraction for scope captures, and fail closed in embeddings. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
020bf6e4dc |
fix(review): keep notebook graph lines aligned with JSON
Join cells so tree-sitter rows match segment maps, remap CFG and scope captures from 1-based extract lines, and highlight .ipynb as JSON. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
d2acefe016 |
feat: index Jupyter notebooks as Python via code-cell extraction
Extract .ipynb code cells in-process (Vue-style) and parse them with the existing Python provider so notebook functions and imports join the graph without executing kernels. Fixes #3371. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
233ca28492
|
fix(ruby): model block-taking class factories (#3376)
* fix(ruby): model block-taking class factories * fix(ruby): cover braced factory blocks * handle brace factory blocks * fix nested Ruby factory ownership * test(ruby): pin factory ownership by node id * test(ruby): cover brace factory variants --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> |
||
|
|
ad5c7364e1
|
feat(storage): share one index store across linked worktrees and sibling clones (#3374)
* feat(storage): resolve the shared sibling-store identity and layout (#3352) Linked worktrees of one repository resolve to one store under GITNEXUS_HOME/stores/<key>, keyed by the canonical git common dir. The resolver reads the .git entry directly, so hot paths spawn no git. Slot naming moves to a leaf module so storage-resolver and shared-store do not import each other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): resolve a shared checkout's graph and existing store slot (#3352) getStoragePaths reads a flat slot's recorded graphPath only for checkout slots under the stores directory; other paths keep <storagePath>/lbug with no I/O. A recorded path outside the store's commit graphs is ignored. resolveStoragePath falls back to an existing store slot for an unregistered checkout, so reads never move to an empty slot. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): share one immutable commit graph across clean worktrees (#3352) A linked worktree writes its own slot in the shared store. A new slot is seeded with a pointer to the nearest commit graph, so a clean checkout at an indexed commit takes the up-to-date path and writes no graph. A checkout with local changes gets a copy-on-write private graph before its first write. After a successful run, a clean checkout at HEAD publishes its graph into commits/ under a store lock, or drops it when that commit graph already exists. Commit graphs are never written after publish. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): seed a worktree's graph from the nearest index (#3352) A new shared slot is seeded from the store's commit graph nearest to HEAD, else from this checkout's or the main checkout's repository-local index (copied under that index's lock, source left in place). The run that follows is up to date or incremental instead of a full build. If a pointed-at shared graph has been removed, analyze falls back to a full build instead of an incremental update over a missing baseline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): keep one parse cache per shared store (#3352) Linked worktrees read and write the parse cache and durable ParsedFile store under the store's caches/ directory. Before pruning, a run folds in the chunk keys recorded by every member slot and commit graph, and the fold, prune and save run under a store-wide cache lock so one member never evicts another's live chunks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(clean): remove only what no shared-store member references (#3352) Deleting a shared checkout slot (clean, clean --all, remove, and the server delete route) recounts references under the store's publish lock and deletes commit graphs no member points at, then the store itself once empty. A graph that cannot be deleted (open on Windows) is reported and kept for the next pass. clean --gc also drops member slots whose worktree is gone or no longer resolves to the store. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): let a clone opt in to a shared store with --share-with (#3352) An independent clone joins a linked worktree's shared store only with analyze --share-with <repo>, and only when its normalized origin URL matches that member's (credentials stripped, as #2054 compares). The registry remembers the choice. --no-share moves an opted-in clone back to its own .gitnexus and reclaims its old slot; linked worktrees always share and are pointed at GITNEXUS_SHARED_STORE=off instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): count a shared checkout's commit graph as its code index (#3352) A clean shared checkout reads a commit graph and owns no graph file, so the code-index presence check made status report it unindexed and registry validation skip it. The check now follows the slot's validated graphPath. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): adopt existing worktree indexes and report shared-store state (#3352) A shared checkout gets <repo>/.gitnexus/store.json pointing at its store slot; resolution follows it only when it names that checkout's own slot. An existing local index seeds the slot and is left in place. status (text and --json) and doctor report the store, whether the graph is shared or private, and any leftover local index, which clean --local-index removes while keeping the pointer. With GITNEXUS_SHARED_STORE=off a previously shared checkout indexes into its own .gitnexus again and never writes a commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(mcp): read the graph a shared checkout points at (#3352) MCP, the HTTP API, group sync, augmentation, and the Claude hook (all three byte-identical copies) resolve a flat slot's graph through resolveGraphPath instead of joining 'lbug' onto the storage path, so checkouts at one commit share one open database. The embeddings writers (embeddings sync and the server embed job) take a private copy first and never write an immutable commit graph. The post-analyze settle probe accepts fresh metadata that points at an existing commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: describe the shared worktree index store (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(storage): reuse helpers across the shared-store code (#3352) One graph-clone helper replaces two copy-then-rename blocks; clean and status reuse formatSlotSize; leaving a store reuses removeSharedStorePointer; withStoreLock is imported from its own module instead of a re-export. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review findings in the shared store (#3352) - Never publish a graph whose build saw dirty files, and never trust a local-index seed on the up-to-date path; it may hold reverted edits. - Record slot pointers and reclaim under the publish lock, and reclaim right after each publish, so a commit graph is never deleted between publish and pointer save and superseded graphs don't pile up. - clean --gc decides membership from the registry, so opted-in clones and a main checkout without worktrees are not dropped. - Leaving a store re-registers first, so an up-to-date run cannot leave the registry pointing at a deleted slot. - Drop pinned branch summaries when an entry moves into a store slot. - Keep run.cjs and the AGENTS.md runner path inside the checkout. - MCP handles follow the slot's current graph, and branch scoping reads the slot's own metadata. - Cache resolveGraphPath by metadata file identity; look up opted-in entries with canonical registry paths. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): localize help for the shared-store flags (#3352) --help renders option text from the i18n catalog, so --share-with, --no-share, clean --gc and clean --local-index need keys in both locales, not just index.ts literals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): lock the embed-job graph copy and skip it on forced rebuilds (#3352) The server embed job copies a shared checkout's graph under the slot's index lock, so a CLI analyze in another process cannot interleave. A forced rebuild with no embeddings to carry over drops the pointer without copying a graph it would discard unread. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(agents): bump AGENTS.md version for the shared store notes (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep shared-store file reads inside checked paths (#3374) Address CodeQL js/path-injection and js/file-system-race on shared-store.ts: every filesystem read rebuilds its path under a fixed parent with an inline path.relative barrier, the .git probe is one read (EISDIR marks a directory) instead of stat-then-read, and the graph pointer cache stats and reads through one file descriptor. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review feedback on the shared store (#3374) - Hold the slot index lock for the whole server embedding job, not just the graph copy. - --no-share re-registers and deletes the old slot under that slot's index lock, so a running shared analyze cannot re-register it. - clean --gc previews without --force, like every other destructive arm. - status reports a pinned branch index as private. - Slot names prefix Windows device names that carry an extension. - A slot or store named ..<name> is a legal direct child. - Shared-store suites run in the serialized lbug-db vitest project and clear an inherited GITNEXUS_SHARED_STORE. - Doc and test accuracy fixes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(storage): pin shared-store behavior for worktrees of a bare repository (#3374) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address second review round on the shared store (#3374) - Reject --no-share in a linked worktree before taking any lock or indexing anything. - clean --all and remove also delete each shared checkout's pointer. - Delete an empty store only while also holding its cache lock, and re-check emptiness under it. - A store pointer is trusted only when the slot's own metadata names this checkout; the editable pointer file just says where to look. - Device names with an extension are prefixed on Windows only, so POSIX slot names stay stable. - Test fixtures use the gitnexus-test- prefix the stale-sidecar sweep recognizes; help text and the private-graph label are accurate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address third review round on the shared store (#3374) - clean --gc never collects a slot whose index lock is held: an analyze holds it until it registers the checkout, so a seeded but not yet registered slot is busy, not orphaned. - Reclaim compares absolute paths, so a relative GITNEXUS_HOME does not make live commit graphs look unreferenced. - graphPath is followed only into a published <commit>-<featureKey> dir, never .publish-* staging (TS and all three hook copies). - clean --gc fails on an unreadable stores root instead of reporting nothing to collect. - --no-share help states it is for opted-in clones. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): land the English --no-share help and private-graph label (#3374) These two strings were meant for |
||
|
|
b92c14cdd0
|
feat: add Factory AI (Droid) integration (#2543)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* feat(setup): add Factory Droid (MCP + skills) to gitnexus setup Register 'droid' in the editor-targets abstraction so `gitnexus setup -c droid` writes the MCP server to ~/.factory/mcp.json and installs skills to ~/.factory/skills/ from the single canonical skills/ source (no per-editor copies). uninstall.ts is target-driven, so removal is covered automatically. Adds unit + round-trip coverage. * feat(plugin): add gitnexus-factory-plugin for droid plugin install * docs: add Factory Droid to editor support table and setup docs * fix(factory-plugin): guard augment hook against fan-out and DB contention Reuse the Claude adapter's acquireHookSlot and LadybugDB owner probe (bundled byte-identical, kept in lockstep by a drift test) instead of running an unguarded augment. Add direct tests for the hook and manifests. * docs: align Factory row in editor support table * fix(factory-plugin): honor GITNEXUS_HOOK_CLI_PATH so augment runs on Windows * docs(hooks): point bundled guard copies at their drift tests * docs(factory-plugin): note the Execute tokenizer's quoting limit * docs(readme): clarify the Full tier and group the Factory row * docs(hooks): trim drift note to a single line * test(ci): run factory-plugin tests on the windows cross-platform lane * refactor(hooks): drop the drift-note comments, the tests already enforce it * fix(factory-plugin): pin CLI version and parse quoted shell patterns - Pin mcp.json and the hook's npx fallback to gitnexus@<version> from the plugin manifest, registered with the release sync script so a mutable @latest can never execute on MCP connect or augment fallback - Port the #2938 shell tokenizer (tokenizeShellWords + parseRgGrepPattern) so quoted, backslash-escaped, --regexp=, -eVALUE, and -- patterns survive - Add the #2938 regression matrix and pin assertions to factory-plugin.test.ts * docs: add Factory Droid to published npm README * fix(factory-plugin): wire marketplace so droid installs the Factory plugin Add .factory-plugin/marketplace.json sourcing ./gitnexus-factory-plugin. Droid reads it before .claude-plugin/marketplace.json, so `droid plugin install` now delivers the Factory plugin (Execute matcher, pinned mcp.json) instead of the translated Claude plugin (Bash matcher, gitnexus@latest). Register the surface in the version-sync script and cover the wiring in the factory and sync test suites. * fix(factory-plugin): use registry lookup for index resolution Bundle registry-query.cjs so external indexes resolve (#3060); re-pin to 1.6.12. * fix(factory-plugin): sync Execute parser with Cursor hook Fixes echo-rg and -f false positives; tighten test env isolation. * fix(factory-plugin): stop no-match augment from re-running via npx A PATH `gitnexus` that finds no match exits 0 with empty stderr, which fell through to a second `npx -y gitnexus@<pin> augment` with its own 8s timeout (16s worst case vs the 10s hook budget). Fall through to npx only when the PATH launcher is missing (ENOENT); any launched PATH binary, including a timeout or non-zero exit, now ends the augment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(factory-plugin): filter augment stderr to the [GitNexus] block runAugment returned raw child stderr, so npm/Node/LadybugDB warnings leaked into additionalContext and noise-only stderr counted as success. Port the Claude adapter's extractAugmentContext (verbatim, with isDebugEnabled) and apply it on every launch tier before the success decision. Adds a drift test against the Claude copy and PATH-tier noise/noise-only behavior tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(factory-plugin): quote DROID_PLUGIN_ROOT in hook command An unquoted plugin root containing spaces (e.g. a Windows user profile path) split into multiple argv words, so the PostToolUse hook silently never ran. Quote it like the Claude plugin does, and pin the exact quoted command in the hooks.json wiring test. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(release): stage Factory plugin manifests in the release commit The rc release job stages only the original four manifest surfaces in the detached release commit, so the v<version> tag tree carried the Factory plugin.json, mcp.json and marketplace.json at the previous version while --check (working tree) passed. Stage them too, and guard the git add block against the synced surfaces in sync-plugin-manifests.test.ts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(factory-plugin): simplify hook gates, spawn tiers and tests - main(): resolve the repo only after the tool-name and pattern gates, matching the Claude/Cursor hook order (skips fs/git work on no-op calls). - runAugment(): share one spawnAugment helper between the GITNEXUS_HOOK_CLI_PATH and npx tiers; PATH tier ENOENT logic unchanged. - factory-plugin test: pre-filter comment lines instead of `continue`. - sync-plugin-manifests test: hoist EXECUTABLE_MCP_FILES and derive TOTAL_SURFACES from its length. - fnSource(): throw when the function or its closing brace is not found. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): list Factory Droid in localized setup help `localizeCliHelp` overwrites the `setup` command description with the `help.command.setup.description` i18n key, so the literal edited in index.ts never reached `gitnexus setup --help`. Add Factory Droid to the en and zh-CN keys. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address PR review feedback (#2543) - factory hook: run every augment tier under the bundled Unix timeout guard (npx tier group-kills), keeping exactly-one-tier fall-through - hook-db-lock-probe: trim GITNEXUS_HOOK_{LSOF,PS}_PATH once so a padded override is used, not silently replaced (all 3 copies) - hook-lock: evict a stale slot via rename-to-tombstone + identity check, so a concurrently recreated fresh lock is never deleted (all 4 copies) - registry-query: a set-but-invalid storage override resolves no repo instead of falling back to the registry storagePath (all 4 copies) - publish.yml: stage the ten skill mcp.json manifests in the rc release commit; the staging test now requires every synced surface Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address PR review feedback round 2 (#2543) - hook-lock: replace rename-to-tombstone eviction with an O_EXCL per-slot `.evicting` marker plus an identity re-check before unlink, so a live lock is never moved, and a crashed evictor leaves only a self-expiring marker (all 4 copies) - hook-db-lock-probe: clamp GITNEXUS_HOOK_PROC_CMDLINE_MAX to a named 256 KiB ceiling and require an integer, so an oversized override can no longer fail the buffer allocation and miss a live owner (all 3 copies) - registry-query: treat an empty GITNEXUS_STORAGE_PATH/ROOT as set but invalid, matching the CLI's `!== undefined` rule (all 4 copies); the factory test env now deletes those keys instead of blanking them Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address PR review feedback round 3 (#2543) - hook-db-lock-probe: a capped /proc cmdline read stops early only once both the GitNexus token and the mcp/serve mode are present (or at EOF, the ceiling, or the budget), so a mode word such as `--require mcp` before the GitNexus path no longer hides a live owner (all 3 copies) - registry-query: correct the override comment; a filesystem root is invalid only for GITNEXUS_STORAGE_PATH, not GITNEXUS_STORAGE_ROOT (all 4 copies, comment only) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address PR review feedback round 4 (#2543) - hook-db-lock-probe: an fd-directory read error other than ENOENT or ENOTDIR on an identified server candidate now fails closed ('timeout') instead of reporting not-owned (EMFILE/ENFILE/ENOMEM/EINTR) - hook-db-lock-probe: resolve GITNEXUS_HOOK_TIMEOUT_PATH to an absolute path before validating and caching it, so callers that spawn with a request cwd can still execute the guard - hook-db-lock-probe: document the chunked cmdline read's actual stop conditions (all 3 copies) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Harden hook-lock eviction marker lifecycle (#2543) Per the chosen option (B) for the stale-slot eviction race: - `.evicting` markers carry a per-call owner token (pid + random hex) - an evictor re-reads its token immediately before the slot identity check and unlink; a stalled evictor whose marker was broken backs off - `finally` removes the marker only while it still holds our token - an orphaned marker is broken only if, re-checked just before unlink, its bigint identity and token are unchanged from when judged stale - doc comment states the two remaining two-syscall windows (slot lstat->unlink, marker token->unlink); POSIX has no conditional unlink, and the worst case is one extra concurrent augment All four byte-identical hook-lock copies updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix CodeQL file-system race in hook-lock orphan-marker check (#2543) breakOrphanedMarker stat'd the marker by path and then read it by path, which CodeQL flags (js/file-system-race): the file could be replaced between the two calls. Take the stat and the token from one open descriptor (readMarkerSnapshot, O_NOFOLLOW where available) for both the "judged stale" snapshot and the pre-unlink re-check. All four hook-lock copies updated. The replaced-marker test injected its swap via a readFileSync(path) spy, which no longer fires; it now swaps the marker just before its second open, counting opens of the marker path only (a per-path counter fired early on slot-0 and let a mutant pass). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Unregister hook-lock exit listener on release (#2543) Each acquireHookSlot registered `release` as a process 'exit' listener that was never removed, so a long-lived process acquiring and releasing slots repeatedly would accumulate listeners (MaxListenersExceededWarning) and retain every closure. release() now removes itself. All four hook-lock copies updated; a test asserts 12 acquire/release cycles leave the 'exit' listener count unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
084ac4514d
|
fix(query): send hub content once across process_symbols rows (#3356)
* fix(query): send hub content once across process_symbols rows
With include_content, a symbol in several execution flows carried its
full source text on every (id, process_id) row. Keep content on the
first row for each symbol id and omit it from later rows. Membership
fields, is_entry_point, and symbol_count are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(query): point agents to the content row and lock the dedup in tests
The query tool text now says content is kept once per symbol id across
the whole process_symbols array, possibly under a different process_id,
and names context({uid, include_content: true}) as the fallback.
Tests: the integration suite asserts func:validate has two rows with
content on exactly one. Unit tests cover a later entry-point row that
is flagged and stripped, a hub in three processes, and that the shaper
does not mutate its input.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(query): state the content-once rule on include_content and in the query hint
The query tool's include_content property now says content is sent once
per symbol id, on its first process_symbols row, and that
context({uid: "<id>", include_content: true}) returns it for any row.
When a query asked for content, the Next hint adds that same fallback;
without include_content the hint is unchanged. The example call now
uses the "<id>" placeholder style used elsewhere in the tool text.
Tests: pin the property sentence, check the hint with and without
include_content, and cover a max_symbols slice that moves the content
row to a later process and a first row that is also the entry point.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
aa2d6aaf7d
|
fix(query): keep process_symbols attaches per (id, process_id) (#3353)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / ci (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(query): keep process_symbols attaches per (id, process_id) Hubs that belong to several processes were collapsed to one row by unique-id first-wins, so later process cards advertised symbol_count with nothing to expand. Slice, pair-key, and recount now live in one shaper so each listed process can open its own attaches. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(query): drop leftover unique-id attach bind name The shaper already returns process_symbols; keep that name through the query response so unique-id vocabulary does not linger at the call site. Co-authored-by: Cursor <cursoragent@cursor.com> * style(query): wrap process attach call for prettier CI quality/format rejects the previous call shape. Co-authored-by: Cursor <cursoragent@cursor.com> * test(query): assert hub lines under each process card A global duplicate count still passes when both lines render under one process. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(query): recount group symbol_count after the service prefix The single-repo (id, process_id) join was readable as if @group results included process_symbols. Scope that sentence, and set symbol_count from the filtered attaches when a service prefix is set. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(query): document the attach contract and count prefix rows once The Unreleased notes and query tool text now say how to join process_symbols, including the @group follow-up. Service-prefix filtering counts those rows in one pass. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
c2ca132620
|
fix: web citation/code panel bugs and serve analyze/route hardening (#3348)
* chore: ignore local Vercel link artifacts
Keep .vercel and env files out of the repo after a local preview link.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): repair citation chips, code panel line math, stale agent state
Audit findings in the web client, each verified against the source:
- RightPanel: citation chips ([[path:10-20]], [[Class:Foo]]) were wired to a
stub resolver that always returned null, so clicking any citation did
nothing. Expose resolveFilePath from useAppState and use it.
- CodeReferencesPanel: graph startLine/endLine are 1-based but were treated
as 0-based, so the highlighted range and scroll target were off by one
line; AI citation cards always rendered "code not available" because the
snippet loader was a stub. Fetch per-citation snippets via /api/file.
- useAppState: sendChatMessage read llmSettings.activeProvider outside its
deps (stale provider capabilities after switching provider);
initializeAgent trapped projectName at '' for callers without an override
(system prompt labelled the codebase "project"); the embeddings 409 dedup
matched a message the server never sends for same-repo jobs.
- tools.ts impact: for path targets every symbol defined in the file shares
the filePath, so the disambiguation always picked the first row and could
analyze an arbitrary symbol while reporting a file impact. Prefer the File
node.
- useSigma: the layout timeout called stop() but never kill(), leaking one
ForceAtlas2 Web Worker plus four graph listeners per completed layout.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): release repo lock on cancel, IPC-first worker cancel, route hardening
Audit findings in gitnexus/src, each verified against the source:
- analyze-launch: cancelJob marks the job failed before the worker exits,
so the exit handler's terminal early-return skipped releaseLockOnce and
the repo stayed locked ("Another job is already active") until restart.
Release on terminal exit and when forkWorker bails on a terminal job.
- analyze-job: cancellation now sends { type: 'cancel' } over IPC first and
signals only after a 15s grace. On Windows child.kill('SIGTERM') is a
forceful termination, so leading with it could kill the worker inside a
LadybugDB write. Mirrors core/auto-sync/analysis-worker-launch.
- api resolveRepo: a job that FAILED during the hold-queue wait fell through
to the { __timedOut } sentinel ("taking longer than expected") instead of
404; only /api/repo checked the sentinel, so graph/query/search/file/grep/
embed/delete crashed on entry.storagePath with a 500 after a 5 minute hang.
Return null on failed jobs and check the sentinel in every consumer.
- api processes/process/clusters/cluster: resolve ?repo= through the HTTP
resolver (documented policy on resolveRegisteredRepoEntry) and pass the
registered absolute path to the backend; add the standard rate limiter.
The raw param previously reached the MCP resolver, which runs a
cwd-relative realpathSync probe + registry refresh on a bare-name miss
and accepts unambiguous partial names.
- api body handling: Express 5 leaves req.body undefined without a JSON
content type, turning "Missing X" 400s into TypeError 500s; body-parser
4xx errors (malformed JSON, over-limit) were also reported as 500.
- /api/file: the lexical path.relative check cannot see symlinks; re-check
containment on realpath so a cloned repo containing evil -> /etc/passwd
cannot read outside the root.
- repo-manager unregisterRepo: used the lenient reader, so a transient read
error (EBUSY/EPERM racing another process's atomic rename) turned into
writing [] and deregistering every repo. Use the strict-if-present reader.
- clean --branch: compared registry paths with raw path.resolve instead of
the canonical registryPathEquals used everywhere else (macOS /private/var,
Windows short names / drive-letter case) and reported indexed branches as
not indexed.
Tests: cancelJob IPC-before-signal contract; /api/file symlink escape (403)
and in-repo symlink (200), skipped where the host cannot create symlinks.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore: keep example env files visible and ignore local Cursor config
.env* also hid gitnexus/.env.example and eval/.env.example. .vercel was already ignored. The web app's .cursor/ stays local, including its MCP file.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Add realtime execution ops dashboard for Vercel monitoring.
Expose /api/ops snapshots over the local serve process and a ?view=ops SPA panel so analyze/embed jobs can be watched live from the hosted web UI.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: address gitnexus-check review on cite/embed/resolve paths
Pass req into resolveRepo for process/cluster routes, tighten same-repo embed 409 handling, guard empty citation paths, fix snippet retry races, and drop the lone-File ambiguity fallback.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops): harden CORS, redact paths, and bound ops streams
Restrict Vercel CORS to exact production hosts, omit raw repo paths/URLs from the unauthenticated ops feed, rate-limit and cap SSE connections, and fix dashboard SSE/poll edge cases.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): stabilize citation fetches and file impact matching
Retry cancelled snippet loads without duplicate in-flight reads, cap range-less citation downloads, and make impact file matching unique-suffix-aware with a synthetic File target when LIMIT drops the File node.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web/ops): close gitnexus-check review threads on SSE and redaction
Cap citation reads, reconnect ops on applied server URL, skip SSE onError after abort, and strip URL query/fragment from public repoName.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops): abort SSE on poll fallback and harden repoName parsing
Prevent dual SSE+poll after a failed safety snapshot, skip overlapping poll ticks, ignore aborted streamSSE onError, and basename Windows drive-letter URLs so ops never leaks path segments.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops/web): close gitnexus-check threads on SSE budget and credential leak
Keep finite SSE retries across short 200s, strip backend URL userinfo before ?server=, and redact progress messages on the public ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Restore 0-based GraphNode line math, redact public job poll/error fields, fix omit-?repo= 400, hold the analyze lock across cancel-during-settle, and restore the Vercel shared compile.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(cli): contain leftover-slot reclaim to the slot and keep --stale status honest
Preview and force now share one branches/ containment rule, nested junctions cannot walk a sibling index, and a mid-loop git failure no longer claims leftovers were not deleted after a successful rm.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(cli): share leftover-slot helpers without changing reclaim behavior
Pull the rolling I/O pool and owned-cwd storage lookup into one place so clean --stale/--branch and leftover listing stop restating the same ownership and concurrency paths.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* Address PR review feedback (#3348)
Keep the omitted-repo snapshot instead of re-listing, stop citation and ops races, redact full public repo URLs, and restore fake timers in teardown.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Store graph node citation lines as 0-based offsets and correct the default-port origin comment.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Redact public SSE progress text, stop citation append retries from
cancelling in-flight reads, and keep the ops dashboard from showing a
stale snapshot or clearing a failed Connect.
Note: pre-existing failure in incremental-index-extension-dml-gate and other lbug/env unit tests not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Prune citation snippets when AI refs are cleared, and poll /api/ops at 2s so the fallback stays under the 60/min limit.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ci): apply prettier class order for format check (#3348)
CI quality/format runs root-only npm ci, so prettier-plugin-tailwindcss
sorts scrollbar-thin without the web Tailwind catalog.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
- Require a unique suffix match for graph-backed citation paths so
ambiguous names like index.ts no longer open the first graph file.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
- Require a path-component boundary so unique citation suffixes cannot match filename substrings like myindex.ts
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): redact filesystem paths in ops text
Unauthenticated /api/ops and poll replay worker errors. URLs were
scrubbed but home-directory and Windows paths still leaked.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): omit repoPath from public SSE frames
/api/ops lists job ids, so the unauthenticated progress stream
must not replay the analyzed filesystem path.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): treat embed lock 409 as a busy error
Analyze and embed share the same lock string. Mapping that 409 to
embedding hid an in-flight analyze as a successful embed start.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): bound impact File path suffix matches
Unbounded endsWith let lib/foo.ts select src/mylib/foo.ts. Require
an exact path or a unique /suffix, matching resolveUniqueIndexedPath.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): keep canceled analyze slot until exit
Marking failed before the worker exited let a second POST start
cloneOrPull against a LadybugDB file still being written.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): skip Windows SIGTERM on job dispose
child.kill('SIGTERM') is TerminateProcess there. Ask over IPC first
and leave the 15s grace timer to SIGKILL.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): let process routes skip the analyze hold
GET /api/processes and /api/clusters always waited up to 300s.
?awaitAnalysis=false fails fast; default still waits like /api/repo.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(web): cover public ops and analyze SSE user flows
Lock the unauthenticated dashboard and analyze complete/fail/cancel paths so a leaked repoPath, token, or home path cannot ship unnoticed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Keep caller cancel reasons over the worker's generic IPC, skip publish while cancel is pending, and redact scp-style remotes on the public ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Release the analyze slot when a worker fails to spawn, and omit branch refs from the unauthenticated ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): do not reuse an analyze job that is pending cancel
A dying same-repo job still occupies the single slot; 202-reuse would
attach a new client to a cancel in flight.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): hold the analyze lock until the worker exits after cancel
Cancel error IPC used to drop the repo lock while the child was still
checkpointing. Abort settle immediately on pending cancel so the slot
is not held for a 60s disk poll.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): redact known repo paths with spaces in public ops text
Known repoPath/repoUrl literals are replaced first so a clone dir with
spaces cannot leak past the whitespace-bounded path regex.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): return the live job on analyze and embed DELETE
Hard-coding failed made clients retry immediately and 409 while the
child still occupied the slot. resolveRepo now returns not-found as
soon as that job fails instead of waiting out the hold timeout.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(web): drive public analyze and ops through a live gitnexus serve
Spawn the real backend and observe requests instead of intercepting
them, so slot occupancy, redaction, and reconnect stay honest.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(web): reuse code-panel helpers and drop dead UI aliases
Citation fetches already had selectedNodeFileRange and snippetRepoKey;
the impact File suffix filter already handled exact paths.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* Address PR review feedback (#3348)
- Skip createJob reuse after cancel IPC is consumed while the child remains
- Hold the repo lock until exit when complete IPC races a pending cancel
- Scrub full remote URLs before known repoUrl prefixes in public ops text
- Reject unique impact File suffix matches from a truncated LIMIT 10 page
- Make live e2e helpers bound probes, clean up failed startups, and wait out the cancel slot
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): open live e2e pages on the Vite host CI actually bound
Absolute 127.0.0.1:5173 navigation refused on Actions because wait-on
and Vite use localhost (often ::1). Honor FRONTEND_URL when set, else
pick the first of localhost / 127.0.0.1 that answers.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Hold the analyze lock until worker exit when cancel aborts settle, and assert GitLab failure chrome does not leak host or path.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): keep live analyze e2e under the analyze rate limit
POST /api/analyze allows 10 requests per minute per IP. The slot-free
helper re-POSTed every 400ms while a cancelled worker was exiting, spent
that budget, and the lock test's hold request got 429 instead of 202.
- postAnalyze waits out a 429 using the RateLimit reset and retries
- slot polling backs off to 2s and leaves a small POST budget for callers
- the slot probe is a clone that fails before any worker fork, so the
probe itself no longer holds the slot after reporting failed
- the lock test holds the slot with a real local analyze
- token and GitLab tests wait for a free slot before posting from the UI
- request fetches carry a timeout; teardown signals the serve process group
- an empty FRONTEND_URL falls back to the default base URL
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- ops view: a `?server=` link no longer auto-connects to another origin
while a deploy token is held; it prefills and waits for Connect
- analyze completion: a local-path run reconnects by the path this client
submitted, so duplicate basenames stay collision-safe without repoPath
on the public SSE frame
- stale slot cleanup: revalidate each nested directory (lstat + realpath)
right before readdir, so a mid-cleanup junction swap aborts instead of
walking an outside tree; list phases run sequentially so the slot-I/O
cap is global
- e2e: 429 backoff honours the caller deadline; the unreachable-backend
ops test navigates through the resolved frontend URL
- drop the unused `repo` member from the clean integration helper
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(server): reconnect after analyze by an opaque repo id
Public job views and the SSE terminal frame no longer carry repoPath, and
repoName is not unique, so a post-analyze reconnect by name could load a
same-named sibling. The server now issues `repoId`: an HMAC of the
canonical registry path under a per-process random key. It is set on
complete jobs (ops view, analyze poll, SSE terminal frame) and matches
the new `id` on `GET /api/repos` entries. The web client resolves it to
the exact entry path on completion; unknown ids fall back to the name.
This covers URL clones and folder uploads, and replaces the local-path
only fallback.
With reconnect off the label, public `repoName` for a branch-pinned URL
clone is the repository name, not the `<repo>__<branch slug>` registry
name, so the requested branch stays off the unauthenticated feed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- stale slot cleanup: a descendant that vanishes before its unlink is
treated as removed instead of aborting the reclaim
- e2e teardown: escalate to SIGKILL on the process group when the live
backend ignores SIGTERM for 5s
- ops view: drop the dead initial value CodeQL flagged
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- public redaction: a known repoPath now also consumes its descendant
tail, so `<repoPath>/src/secret.ts` becomes `[path]` instead of
`[path]/src/secret.ts`; a same-prefix sibling is left to the path scrub
- e2e: the cancel test waits for the analyze slot the previous failed
local-path job still holds; `fetchOps` carries the request timeout
- docs: SSE terminal payload and RepoAnalyzer `onComplete` describe
`repoId` resolution
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- RepoAnalyzer: drop a completion that resolves after unmount, so a slow
/api/repos lookup cannot switch repos after the sheet was dismissed
- e2e: select the local-path input by test id (the placeholder differs on
Windows); the slot probe's job wait honours the caller's deadline
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(server): assert the public SSE terminal frame in analyze-api
The #2790 terminality tests still expected `repoPath` on the terminal
frame. This PR replaced it with the opaque `repoId`, so assert that shape
and that the analyzed path never appears in the stream.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(web): wait for the analyze slot between duplicate-repo setup runs
The server now keeps the single analyze slot until the worker exits, even
after its job reports complete. repo-path-identity posted the second
duplicate's analyze immediately and got 409 in CI. Use the shared
slot-aware POST, which also waits out a 429.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
5be80b1fc3
|
fix(lbug): self-heal read-only opens refused by an interrupted checkpoint (#3340)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(lbug): self-heal read-only opens refused by an interrupted checkpoint Homelab repro 2026-09-19 (image 1.6.10-20260917, @ladybugdb/core 0.19.x): a wiki pod killed mid-CHECKPOINT left the engine's checkpoint artifacts on disk (lbug.wal.checkpoint / lbug.shadow / checkpoint intent+apply locks), and every later READ-ONLY open refused with 'Cannot open database in read-only mode while checkpoint is in progress' — permanently, until a writable open (any gitnexus analyze) happened to run. Two defects, both fixed: 1. The refusal was unclassified. ensureReadOnlyConnectionUsable (direct adapter) and openReadOnlyDatabase (pool) recovered missing-shadow and shadow-replay errors but rethrew this one raw, so 'LadybugDB unavailable for __wiki__' repeated forever. Add isReadOnlyCheckpointInProgressError (LADYBUGDB-CONTRACT, live-verified against 0.19.1) and route it through the same writable-open recovery — including at OPEN time, where the refusal fires before any probe can run (pool: init() moved inside the try; direct: doInitLbug catch). 2. The existing shadow-replay recovery was not durable. Reproduction matrix on 0.19.1: the writable probe replays the WAL in MEMORY only — without an explicit CHECKPOINT the engine drops the pages at close and the follow-up read-only open silently serves the pre-checkpoint state. Both recovery paths now CHECKPOINT after the probe, which applies the replay, consumes the sidecars, and clears the checkpoint locks. Verified end-to-end: real engine 0.19.1, killed-mid-CHECKPOINT state → exact refusal → pool adapter self-heals → 80,800 rows intact, sidecars consumed. The planted-signature integration test runs on any engine version (refusal asserted only where the engine emits it, 0.19+). * docs(architecture): list checkpoint-in-flight artifacts and the read-path self-heal * fix(lbug): address review findings — shared cursor closer, tighten structural guard Both findings from the gitnexus-check review on this PR: 1. pool-adapter.ts: the recovery CHECKPOINT closed its cursor with a bare unawaited result.close?.(). Use the shared best-effort closer (closeQueryResults) the repo already funnels both adapters through, so a cursor-close failure stays cleanup and cannot escape as an unhandled rejection. 2. sidecar-recovery.test.ts: the structural regex omitted the leading negation, so as a substring match it also accepted the inverted predicate (recover ONLY shadow-replay, exclude checkpoint) — the exact regression the guard exists to prevent. Pin the full '!isReadOnlyShadowReplayError(err) && !isReadOnlyCheckpointInProgressError(err)' throw-through shape. * test(lbug): wire the recovery plant into lbug-db/LBUG_NATIVE; force the refusal on any engine pin Address the tri-review findings (all four): P1 — the planted-signature integration test was collected by the parallel 'default' project and never by the serialized 'lbug-db' project, and Windows/macOS CI never ran it: register it next to its sibling in vitest.config.ts (lbug-db include + default exclude) and in cross-platform-tests.ts LBUG_NATIVE, per TESTING.md's rule for native @ladybugdb/core suites. Verified via 'vitest list --project lbug-db'. P2 — on the committed 0.18.3 pin the plant passes as 'pool opens and count(n)=300' without ever exercising the new classifier or recovery CHECKPOINT. Add forced-refusal behavioral tests for BOTH adapters: a mocked native layer whose first read-only Database refuses with the canonical 0.19 message, asserting the exact self-heal shape (ro-refused -> writable open -> CHECKPOINT -> ro retry) and that a healthy db never triggers a writable open. The direct adapter is lazy, so its refusal is scripted at the first probe query rather than init(). P2 — the LADYBUGDB-CONTRACT header on isReadOnlyCheckpointInProgressError claimed '^0.18.0' like its siblings; only 0.19.x emits this string (0.18.3 tolerates the state). State the first-observed version so a bump reviewer validates the right binary. P2 — the pool's writable replay recovery quarantined on missing-shadow even when the error came from the post-probe CHECKPOINT, where the main file has already changed: mirror the direct adapter's probeSucceeded guard (replaySucceeded) and fail closed instead of parking a live sidecar. Also assert walBuffer.byteLength > 0 in the plant so the fixture cannot silently degrade into an empty shell on tolerant engines. * test(lbug): fix hosted-CI failures — Windows handle release, version-gated plant, prettier Address the CHANGES_REQUESTED review of the hosted run: Windows blocker (Win32 Error 33, locked file region at the pooled reopen): the fixture now makes handle release explicit — waitForFixtureRelease probe-reads the db and its residual WAL with bounded retries after every native close (plant, raw refusal probe, pool close), mirroring the adapter's own Windows handle-release probing. The WAL is re-planted from the captured bytes instead of renamed: a close-time auto-checkpoint can consume the live .wal out from under the rename (ENOENT, second flake). Version-gate the plant itself: on < 0.19 engines that tolerate the planted signature, its synthetic sidecars are not a consistent staging state for the old engine (double-apply replays surfaced as 'Person already exists in catalog' — the third flake), and no refusal can be forced there anyway. The plant now skips below 0.19 with that rationale in-file; the behavioral contract on every pin stays with the forced-refusal units, and this native suite remains registered in lbug-db + LBUG_NATIVE so Win/macOS exercise it as soon as the pin moves off 0.18.3. Also: prettier on the two flagged files (format gate). * fix(lbug): keep failed checkpoint heals from re-entering CHECKPOINT Wrap recovery failures without repeating native refusal text, skip already-wrapped errors in doInitLbug, and pin constructor-time heal plus the Windows reopen skip. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3340) - Clean each forced-heal tmpDir in afterEach so earlier cases do not leak - Compare major.minor when version-gating the interrupted-checkpoint plant Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3340) Reset the pool forced-refusal mock Database sequence in native.reset() so later tests can still script the first construction as the checkpoint victim. * Address PR review feedback (#3340) Count every MATCH on the pool forced-refusal mock and assert the writable replay probe so CHECKPOINT-without-probe cannot stay green. --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
620fd18a5c
|
feat(cli): reclaim leftover per-branch indexes after branch delete (#3338)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* feat(storage): classify leftover per-branch index slots Operators need a shared enumerator for deleted-branch leftovers before clean --stale or doctor can reclaim or report them. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(storage): reclaim a per-branch slot and empty branches/ Named clean --branch now shares one rm-then-registry helper so the last leftover slot can drop the empty branches directory, and a failed rm still keeps the retryable summary. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(cli): add clean --stale to reclaim leftover branch indexes Operators can drop per-branch slots whose recorded branch is gone without remembering each name, while a git-list failure stays a no-op. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(cli): report leftover branch indexes in doctor Operators can see cwd orphaned per-branch slots and their size, then reclaim them with clean --stale, without doctor deleting anything. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(review): keep unreadable branch slots out of stale reclaim A stat error other than ENOENT/ENOTDIR must not look like a missing directory, or clean --stale --force drops the registry row and leaves the slot on disk. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(cli): keep leftover-slot display helpers in the CLI layer Preview and doctor share one size formatter and an i18n path for registry-only rows, so storage no longer owns display copy. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(review): keep leftover reclaim moving after a registry drop failure Catch removeBranchIndex rejections so --stale continues, match doctor registry rows through canonicalizePath, and size leftover slots sequentially. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(review): match leftover-slot registry rows with canonicalizePath Use the repo-manager path contract so clean --stale and --branch still see registry-only leftover rows when cwd and the stored path differ. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(review): contain leftover-slot deletes and re-check live heads Refuse symlink and junction escapes under branches/, unlink slot links instead of removing through them, and skip --stale --force when a name is a local head again. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3338) - Bound listLocalHeads spawnSync with GIT_PATH_LIST_MAX_BUFFER. - Clarify that doctor leftover reporting is cwd-only, not registry-wide. - Drop the MCP/serve assumption from clean --stale delete failures. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3338) - Describe disk-only leftover slots in --stale help, not only recorded branches. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(cli): stop doctor reclaim copy when heads cannot be listed Doctor was naming clean --stale for leftover rows even when git cannot list local heads, which is a no-op. Print the retry-git message instead (#3337). Co-authored-by: Cursor <cursoragent@cursor.com> * revert: drop Unreleased changelog notes from this branch Co-authored-by: Cursor <cursoragent@cursor.com> * fix(cli): keep live branch pins when a tag shares the name %(refname:short) disambiguates against tags, so clean --stale treated still-local heads as leftover. Fail closed on obstructed slots and unlistable branches/ directories. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3338) Bound leftover-slot listing, revalidate paths immediately before delete, and keep registry rows when a stray disk-only directory claims a recorded branch. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
a578747455
|
fix(swift): resolve nested constructors in extensions (#3308)
* fix(swift): resolve nested constructors in extensions * fix(swift): preserve qualified extension owners * Address PR review feedback (#3308) Recover qualified Swift extension owners through public / attribute prefixes, and stop last-dot-guessing when source text is present. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(swift): keep attribute text from stealing extension owners Bound header recovery so @available messages cannot rekey a fragment, and still inject nested types when the Class scope has no bindings. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(swift): nest comments and keep the public extension fixture valid Review follow-up: skip nested /* */ in the header scan, put the Inner.Entry decoy in a parsed file, and mark Outer/Container/Entry public so the live fixture is valid Swift. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(swift): recover Unicode identifiers as extension owners The header regex was ASCII-only, so extension Café.Container keyed as Caf and dropped nested-type siblings. Match ID_Start/ID_Continue segments instead. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(swift): skip raw strings and decode UTF-8 scope columns Header recovery treated #"..."# as an ordinary quote and sliced Tree-sitter byte columns as JS offsets, so a same-line Café prefix or a raw attribute message could steal or drop the extension owner. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
795cf0e151
|
fix(swift): resolve inherited protocol extension calls (#3309)
* fix(swift): resolve inherited protocol extension calls * fix(scope): gate inherited implicit receiver lookup * fix(swift): resolve call result types by exact callee * test(swift): align cache and local call expectations * fix(scope): reconcile replay diagnostics * fix(swift): preserve exact callable return types * fix(swift): capture throwing async call results * test(swift): refresh capture golden * test(swift): refresh scope capture baseline * fix(swift): require explicit callable returns * fix(scope): preserve duplicate return metadata * chore(scope): align index documentation * Address PR review feedback (#3309) Stamp only protocol/class-extension members (nested QN + SPM buckets), arity-narrow implicit-this across MRO, and keep Swift type peeling out of shared workspace-index via stripTypePreservingDecoration. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3309) Stamp extension members even when the extension declares a nested type, keep inherited class members ahead of protocol-extension defaults, and report replay-only interface-dispatch fan-out drops. Note: pre-existing failure in gitnexus tsc against an older gitnexus-shared dist not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3309) Walk inherited implicit-this owners nearest-first so a nearer override wins, and tighten Swift owner-stamp tests. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
ba39d5c009
|
feat(analyze): expose process-detection budget overrides (#3324)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
* feat(analyze): expose process-detection budget overrides (#3313) Operators can raise or lower process count, branching, trace depth, and the entry-point candidate pool via CLI, .gitnexusrc, or GITNEXUS_* without changing shipped defaults. A budget-only change re-detects flows on the next analyze without --force. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(review): say invalid budget flags still honor env A rejected --max-processes value was described as falling back to the built-in default even when GITNEXUS_MAX_* still won the next precedence tier. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(analyze): share process-detection defaults and skip unused walks Keep DEFAULT_CONFIG aligned with the budget resolver and count symbols only when maxProcesses is still dynamic. Co-authored-by: Cursor <cursoragent@cursor.com> * style(analyze): wrap process-detection budget files for prettier Co-authored-by: Cursor <cursoragent@cursor.com> * docs(analyze): name the real process-detection default formula Co-authored-by: Cursor <cursoragent@cursor.com> * docs(analyze): stop calling maxProcesses*2 a hard trace quota Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): say invalid env budget tokens fall back to defaults Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): recertify process-detection after in-place FTS abort (#3324) Persist processDetection.uncertified on the in-place FTS dirty stamp when the budget mismatched so a flagless retry cannot keep rewritten flows. Qualify .gitnexusrc fail-fast copy and tighten related tests. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): skip live dirty stamp on atomic incremental (#3324) POSIX atomic incremental mutates a staging copy, so stamping live incrementalInProgress before swap made a crash force-rebuild a healthy index. Align analyze --help with CLI > .gitnexusrc > env > default. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(changelog): drop the atomic-incremental dirty-stamp note The code fix stays; Unreleased no longer lists that recovery change. Co-authored-by: Cursor <cursoragent@cursor.com> * test(cli): survive FTS SIGSEGV in --limit e2e CREATE_FTS_INDEX can kill the setup analyze on some WSL hosts (status null). Rebuild with --skip-fts and skip BM25-only query --limit cases unless GITNEXUS_REQUIRE_FTS=1. Refs #3324 Co-authored-by: Cursor <cursoragent@cursor.com> * test(cli): mark update-check child at import Writing refresh-started from fetch() raced a 30s poll against cold tsx boot on a loaded default-project worker. Refs #3324 Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3324) Isolate default-budget FTS crash-marker tests from GITNEXUS_MAX_* env, assert uncertify-before-FTS order and deferred flow detection on park recovery, drop the dangling "then" from entry-point help, and correct stale streamGraphEmit docs without skipping the process-detection stamp. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(changelog): drop Unreleased process-detection notes Keep the #3313 / #3322 code; Unreleased changelog matches main until release. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
9d95af9fc3
|
refactor(analyze): move detected-branch sanitization off CLI config (#3325)
* refactor(analyze): load detected-branch sanitization from core git-ref Keep the never-throw helper next to validateBranchName so run-analyze no longer imports CLI config parsing. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): warn once when a checkout name cannot label the index After the write lock settles, emit a single onLog warning and keep writing the workspace slot. Pin that run-analyze does not import CLI analyze-config. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): escape hidden checkout names in the detect-reject warning Keep the rejected ref visible in onLog without replaying bidi or quote characters, and document that sanitizeDetectedBranch rethrows unexpected errors. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): keep detect-reject warnings on one line (#3325) Git-legal U+2028/U+2029 checkout names were rejected as whitespace but left raw in the new onLog warning, so the message split across two lines. Escape those code points in the formatter without changing validateBranchName. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * fix(analyze): keep C1 and Unicode spaces in detect-reject warnings Escape NEL and remaining whitespace as \uXXXX so stripControlCharacters cannot drop or disguise the rejected checkout name. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
a2e1710ac0
|
fix(embeddings): stop Caching embeddings OOM on large incremental analyze (#3310)
* fix(embeddings): spill cached vectors to a Float32 temp file Keep restore metadata in RAM and write embeddings once the in-memory row limit is exceeded so incremental analyze can survive large caches without a full-table number[] heap (#3306). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): stream CodeEmbedding cache under the connection lock Spill vectors once the in-memory limit is crossed and fail the load instead of adopting an empty snapshot, so incremental analyze cannot OOM or quietly drop the restore cache (#3306). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(analyze): restore cached embeddings from a streamed spill snapshot Hold row metadata across wipe, materialize 200-row batches, and treat cache-load failures as warn-and-continue so incremental analyze can preserve vectors without a full-table heap (#3306). Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3310) Loop spill writes until the full vector lands, keep materialize failures out of the insert catch and the Phase 4 hash skip-set, and assert spilled restore subsets by node id instead of scan order. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3310) Discard only this analyze run's embedding spills so a concurrent analyze on another index keeps its restore file, and isolate the default in-memory limit test from inherited env. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3310) Mark a node stale when any restore batch fails so leftover chunks are deleted and rembedded, and exercise a full-length bad-magic spill header. --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
8b21ce3b95
|
feat(auto-sync): preserve PDG indexes across updates (#3290)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
Skill copy sync / shipped skills drift guard (push) Has been cancelled
* feat(auto-sync): preserve PDG indexes across updates * docs(auto-sync): document durable PDG synchronization * Address PR review feedback (#3290) - Correct requestedPdg state docs for threshold-skipped syncs - Defer coalesced follow-up and skip failure-threshold counts for leftover-worker / retryable lock waits - Document the pdg tri-state and caveat the 30m/5m example Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): raise Windows Ladybug #605 hang budget off the CI tail Windows 3/3 typically finishes this native race in ~15s but has a 56s tail; 60s false-positives as deadlock. Keep the POSIX 60s detector and the completion/.shadow/row-count contract. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
0ede6ae501
|
fix(rust): respect Cargo target boundaries in name fallback (#3294)
* fix(rust): respect Cargo target boundaries in name fallback * fix(rust): read Cargo sources through validated file descriptors * fix(rust): require matching imports for crate-root guesses * fix(rust): require Cargo root identity for cross-target imports * fix(rust): enforce Cargo identity across fallback imports * fix(rust): keep Cargo membership on typical derive/macros (#3294) Abort only genuine unknown expansion, ignore Cargo artifact layouts instead of every `target` path segment, and require a covering import for unanswered crate-root name fallback so #3253 still holds on ordinary Rust sources. Co-authored-by: Cursor <cursoragent@cursor.com> * test(rust): gate Cargo target membership in CI (#3253) Add a parse-dispatch-rounds-style bench so derive/macro abort, target-path globs, and superlinear membership walks fail in CI instead of staying graph-invisible. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(rust): verify Cargo macro and re-export evidence * style(bench): format Cargo membership benchmark * fix(rust): require imports for cross-file root guesses --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
ebbcd5b0b4
|
fix(mcp): name the indexed ref on hot read tool staleness (#3291) (#3293) | ||
|
|
21a52af1d4
|
fix(lbug): ship FTS per-platform and recover in-place native aborts (#3274)
* fix(lbug): pin Ladybug core so Dependabot cannot ship a skewed FTS artifact The extension version is a separate upstream constant. Ignore daily core bumps and fail the pairing gate when the committed manifest does not name the installed core. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): make doctor and CI FTS gates resolve the packaged artifact Doctor and the REQUIRE_FTS file gates still treated an empty ~/.lbdb as unavailable, which would turn three CI jobs red once analyze stops installing into that tree. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): name native-abort and tuple-missing so analyze cannot mis-advise The CLI summary's trailing else treated every unknown skip reason as a missing extension. New crash and platform causes must get their own remedies, not a network-install hint. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): delete the dead read-path FTS index create ensureFTSIndex had no production callers and swallowed read-only CREATE_FTS_INDEX failures, which hid the only signal that a reader tried to write. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): vendor per-platform FTS artifacts so analyze needs no host install Keyword search depended on a CDN fetch into ~/.lbdb. Shipping the five published tuples inside the package makes air-gapped and ignore-scripts installs load the same artifact the publish gate checksums. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): load the packaged FTS artifact before any network install Analyze still required a CDN fetch into ~/.lbdb even when the package already shipped the file. FTS now path-loads the vendored tuple first and records source labels so a later truncated home copy cannot steal the diagnosis. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): diagnose a core/extension version skew instead of a missing runtime A structurally valid FTS artifact whose path version disagrees with the packaged pin must name both versions, not prescribe VC++ or OpenSSL. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): stamp an FTS phase so repair stays usable after an in-place abort A native CREATE_FTS_INDEX abort leaves no skip reason; the next run infers it from the dirty flag, and --repair-fts must not treat that phase as a half-written graph. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): park an in-place FTS crash WAL without wiping the graph An FTS abort after a successful checkpoint must reopen the live index on macOS, Windows, and Linux. Staging never parks the live WAL; readers keep today's large-WAL refusal. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): refuse read-only opens of an FTS-poisoned WAL MCP and serve cannot repair a leftover in-place abort. Fail before the native open and name --repair-fts, on macOS, Windows, and Linux. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): name a vendor-neutral Windows OpenSSL prerequisite OQ1 is unanswered here so GitNexus does not ship OpenSSL DLLs. Windows FTS now asks for a system OpenSSL 3 runtime instead of Git Bash PATH. Co-authored-by: Cursor <cursoragent@cursor.com> * test(lbug): inject the FTS vendor root and redact it on HTTP and MCP Path-loaded artifacts no longer vary with HOME. Tests pass an injected vendor tree and assert search warnings never leak a filesystem path. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(lbug): document load-only as the global FTS install default Analyze still overrides to auto. Packaged per-platform artifacts load before any network install on macOS, Windows, and Linux. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(lbug): format the FTS install-policy README table Prettier does not run on Markdown in pre-commit, so the U10 table wrap needs its own formatting commit. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): skip FTS CREATE after a persisted native abort A recovered analyze run was retrying CREATE_FTS_INDEX from skipReason alone. Keep that skip until --repair-fts, fail closed on unsupported tuples, and honor the checkpoint warrant for park/repair. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): honor checkpoint flushed warrant and align FTS tests with packaged vendor A no-op CHECKPOINT must not satisfy the FTS park warrant, and CI still asserted HOME-only FTS isolation after analyze started path-LOADing the packaged artifact. Co-authored-by: Cursor <cursoragent@cursor.com> * test(lbug): accept a nonempty incremental write set in the #2790 recovery check FTS-phase recovery can incremental-add files (changed=0, added=1). That is not the #2790 empty-diff wipe skip. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): compare FTS home versions to the core pin and tighten the publish filename gate Ladybug's ~/.lbdb/extension directory is the runtime/core version; treating it as the artifact version false-diagnosed skew. The publish guard now rejects a path-escaping filename the same way the fetch script does. Co-authored-by: Cursor <cursoragent@cursor.com> * test(lbug): seed FTS e2e fixtures from the packaged vendor artifact A machine with no ~/.lbdb copy should still run the vendor-survivorship cases; the seed no longer depends on HOME or a network install. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3274) Keep in-place FTS abort evidence after persist so a second CREATE abort cannot fail-open readers, and close the CLI, loader, embed, and e2e gaps the review called out. Note: full npm test hit Ladybug worker-pool startup failures under memory pressure; tsc and 180 targeted unit tests passed. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3274) Run the vendored-path symlink guard on the OS matrix, put e2e HOME fixtures on Ladybug's real extension layout, pin the embed crash-WAL gate before the writable open, and let analyze writers park through missing-shadow recovery. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(lbug): keep --repair-fts CI green after vendored-first FTS Never-installed warning fixtures must not inspect a packaged vendor binary, and a failed dirty restamp must not abort an otherwise successful --repair-fts run. Co-authored-by: Cursor <cursoragent@cursor.com> * test(cli): give the #1169 analyze e2e the same 90s Windows budget as its sibling The first #1169 persist-meta case was still on a 60s spawn/it budget and was killed banner-only on windows-latest after the FTS warning fixture no longer failed the shard first. Co-authored-by: Cursor <cursoragent@cursor.com> * test(ci): reweight Windows shards after the FTS e2e grew Vendored-first HOME fixtures pushed fts-extension-e2e to ~6 minutes on windows-latest, so the old 146s weight packed it with skills-e2e and blew the 20-minute watchdog. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
79543c8f83
|
feat(storage): add configurable index storage and content retention tiers (#3060)
* feat(storage): add configurable index storage and content retention tiers Rebase #3060 onto current origin/main. Keep GITNEXUS_STORAGE_PATH, GITNEXUS_STORAGE_ROOT, and GITNEXUS_CONTENT_RETENTION, and fold in main's FTS skip, embed-session, and help-text updates. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3060) Keep legacy registry rows on the local storage fallback, resolve symlinks before the destructive-path guard, and align hook lookup with CLI branch slugs, branch-slot metadata, and longest-path match. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3060) Only list swept upload directories after a successful removal so callers cannot treat a permission or transient rm failure as gone. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3060) Document that getStoragePath may consult registered storage while this module still does not mutate the global registry. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(storage): close review findings for external indexes and retention Re-inspect ownership under the analyze lock, fail-closed when the registry file is missing, and keep skip-git hook discovery plus retention fields on HTTP/MCP list surfaces. /api/file stays 410 unless contentRetention is full. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * Address PR review feedback (#3060) Treat lock-only index dirs as empty, honor HTTP --force storage policy, and prefer registered plus branch-aware slots in hooks and augment. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3060) Keep hook fallbacks inside the current worktree, compare foreign-local slots canonically, and make storage fixtures survive ownership validation. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix macOS hook test expecting realpath'd registry paths. resolveHookRepo returns the written registry path, not a filesystem realpath, so the assertion must match that. * Address gitnexus-check warnings on hook install docs and slot tests. The Cursor troubleshooting list omitted registry-query.cjs, and the writable-slot test only checked that isDirectory exists instead of that the path is a directory. * Align the HTTP catalog source-scan with skippable resolveRepo validation. resolveRepo lists fresh repos with validate: options.validateStorage !== false so DELETE can skip prune; the test still required a literal validate: true. * Harden storage path sinks so CodeQL path-injection and ReDoS alerts clear. Contain every filesystem probe inside the resolved storage slot with the inline path.relative idiom, reject filesystem-root slots, and trim slot basenames in linear time. * Settle bridge stamps before writing so CI size/mtime matches stay stable. LadybugDB can still flush into bridge.lbug after close+rename; persist whole-millisecond mtimes and wait for consecutive stats to agree so a freshly written pair matches. * Type the settled bridge stat as fs.Stats so tsc does not see bigint. Awaited<ReturnType<typeof fsp.stat>> collapsed the bigint overload and broke prepare/typecheck on CI. * Keep the bridge mtime stamp exact so same-size swaps still fail the pair check. Co-authored-by: Cursor <cursoragent@cursor.com> * Wrap the bridge stamp predicate so prettier --check stays green. Co-authored-by: Cursor <cursoragent@cursor.com> * Require a quiet interval before stamping a settled bridge file. Co-authored-by: Cursor <cursoragent@cursor.com> * Reuse shared storage and settle helpers instead of local copies. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
ceaff27c1e
|
fix(parse-cache): retire a chunk whose durable generation could not be reset (#3271)
* fix(parse-cache): retire a chunk whose durable generation could not be reset #3200 skipped the parse-cache write when `prepareDurableParsedFileChunk` failed, but the chunk hash was already in `usedKeys` from the lookup. When a previous generation existed on disk — reachable because the coherence gate re-dispatches a chunk whose `.v8` shard is live but whose durable shards are unreadable — `saveParseCache` copied that old shard forward, and the durable prune, which keeps exactly the saved keys, retained the mixed directory. The next run then served a warm hit out of a directory the previous run had already decided it could not account for. Retire the hash instead of only skipping the write: - `ParseCache.staleKeys` is a transient set that `saveParseCache` filters out of its key list. Filtering at save is what makes it survive the post-parse key merges in run-analyze (#2106 sibling fold, unreadable-meta retention), and it reaches both stores at once because the durable prune keeps exactly the keys `saveParseCache` returns. - The hash is retired at the reset-failure site, which runs unconditionally. The parse-cache write branch sits behind `rawResults.length > 0`, so a chunk whose worker round returns nothing would never have been retired there. - Worker-quarantined chunks get the same treatment for the same reason: they also reach the save with no in-memory entry, which is what triggers the copy-forward. That branch was previously unreachable when the worker died on the chunk and returned no results. - Guard the durable prune's non-survivor `fs.rm`. The causes that break the reset break that delete too, and it sat outside the validation try — one undeletable directory aborted the loop and cost every remaining chunk its index entry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(review): apply review findings Retire a chunk only when a generation nobody cleared is still on disk. `prepareDurableParsedFileChunk` is rm-then-mkdir, and the catch could not tell the two apart: an rm that succeeded before a failing mkdir leaves NO directory, so the workers recreate it and write a clean generation. Retiring there discarded a good `.v8` for no safety gain — and under a correlated failure (an empty durable index turns every chunk into a re-dispatched miss, then a descriptor burst rejects the resets en masse) it would have wiped both shared stores for every branch, where the pre-#3204 posture cost only the writes. `durableChunkHasStaleShards` is the discriminator. Finish the delete guard on the path that runs before it. The staged→live overlay in `mergeStagedDurableParsedFileStore` awaited `replaceDurableChunkDir` unguarded, so on the cold-rebuild path one undeletable directory threw out of the merge before the prune ever ran — the durable index was never rewritten and a retired chunk kept its directory. Same log-and-continue treatment, plus best-effort handling of the two `.replacing` backup removals. Aggregate the prune's delete-failure warning: a store-wide cause hits every non-survivor, and one line per directory buries the message that matters. Tests: - Guard the chmod-based prune test with the repo's `skipIf` for root/Windows and assert the directory survived, so it cannot pass vacuously where the delete succeeds. - Add a two-chunk control: one chunk's reset fails, and the sibling must stay warm through the next run. One chunk plus a global spawn marker could not tell "retires the failing chunk" from "retires everything". - Add the rm-succeeded/mkdir-failed case, which must NOT retire. - Model both post-parse merges in the R4 test (the sibling fold re-adds the key, the unreadable-meta fallback unions `entries`), and move the in-memory `entries` assertion to a direct helper test — the sharded path never populates `entries`, so the old assertion proved nothing. - Type the cache factory as `ParseCache`; the `staleKeys` assertions were TS2339 and `?? false` read as a pass regardless. - Register the store test in the cross-platform filesystem list. Correct two comments that still described a quarantined chunk by the premise this fix disproves. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(review): stop swallowing the backup removal in replaceDurableChunkDir Swallowing that `fs.rm` manufactured the very hazard this PR removes. When a non-empty `${to}.replacing` survives, the following `fs.rename(to, backup)` cannot overwrite it and is suppressed as "dest was missing", so `backedUp` stays false and the `fs.cp(from, to)` fallback merges the staged generation INTO the live directory — old shards alongside new, which the prune then indexes as one valid survivor. Let it throw; the per-entry guard added to `mergeStagedDurableParsedFileStore` already stops one such chunk from costing the others their prune. The post-publish backup cleanup stays best-effort, where an undeletable leftover really is litter. Also: - Make the sibling-isolation test perform the run its title claims. It asserted index membership and stopped; an index entry does not exercise the warm-hit path, so it would have passed even if the sibling re-dispatched. Each chunk now runs alone so the single spawn marker names which one re-parsed. - Correct two comments that outran the implementation: retirement is gated on shards actually surviving, and an undeletable directory is dropped from the index rather than removed from disk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f5ca8cdb7
|
fix(storage): guard stale file-lock reclamation (#3234)
* fix(storage): guard stale file-lock reclamation * fix(storage): close lock recovery failure paths * Address PR review feedback (#3234) - Flush the lock-child stderr diagnostic before process.exit - Treat explicit NaN timeouts as the default ceiling - Document non-retryable guard timeouts on the worker IPC contract Co-authored-by: Cursor <cursoragent@cursor.com> * fix(storage): stop mislabeling live lock waits as orphan recovery A brief peer inspect must not attach guardPath or send operators to RUNBOOK delete steps. Refuse lock-free embeddings sync, stop --watch only on a true guard timeout, and drop an unreadable self-created guard before failing closed. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * Address PR review feedback (#3234) - Verify lock/guard absence before degrading a denied main-lock create - Launch the third contender from unlinkSync, not the dead rename path - Document group-lock timeouts for unrecoverable guard leftovers Co-authored-by: Cursor <cursoragent@cursor.com> * style: prettier index-lock reclaim guard tests Co-authored-by: Cursor <cursoragent@cursor.com> * fix(storage): treat O_EXCL as the lock-file presence check CodeQL flagged existsSync-then-wx on analyze.lock. Create with wx first and only reclaim unreadable leftovers after grace, so a successor is never unlinked from a lost race. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
f8036ac349
|
feat(analyze): add explicit FTS opt-out (#3205)
* feat(analyze): add explicit FTS opt-out * chore(autofix): apply prettier + eslint fixes via /autofix command * fix(analyze): treat flag and env FTS opt-out as one mode Avoid a same-commit rebuild when only the skipReason discriminator changes, and advertise disablement on dirty meta before leftover indexes are wiped. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
d1a3edd333
|
perf(mcp): avoid O(n) git spawns on tools/list with many repos (#3259)
* perf(mcp): avoid O(n) git spawns on tools/list with many repos toolSchemaRepoRequirements called listAllowedRepos -> listRepos -> checkStalenessAsync for every registered repo. With ~200 repos, this spawned 200 parallel git rev-list processes on every tools/list discovery call, causing a ~30s delay. Replaced with a lightweight countRepos() method that reads the registry file once without spawning git processes, preserving full staleness checks for list_repos. * Address PR review feedback (#3259) Count the validated registry in countRepos so tools/list cannot advertise a multi-repo schema for ENOENT ghosts, and update the unrestricted listTools mocks to that contract. Note: pre-existing failure in update-notice.test.ts (missing dist/cli/mcp.js) not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * Add a 200-repo tools/list bench for the countRepos path (#3259) Pin the #1363 comparison (listRepos git fan-out vs validated countRepos / listTools) in-tree so the latency claim can be re-run. Also drop the change-history comments on the unrestricted schema arm. Co-authored-by: Cursor <cursoragent@cursor.com> * Gate the tools/list bench with baselines and CI --check (#3259) Exact registry/schema floors plus ratio timing, no millisecond ceiling, so restoring listRepos() on tools/list fails CI. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3259) Align unrestricted tools/list schema flags with the refreshed registry snapshot, and make the bench reject a non-positive BENCH_REPS and isolate fixtures by N. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * Address PR review feedback (#3259) Isolate the tools/list bench from GITNEXUS_MCP_READ_ONLY and create the default fixture under mkdtempSync so CodeQL is not looking at a predictable /tmp path. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
8bd71c8335
|
feat(staleness): report diverged and unknown index state instead of fresh (#3257)
* feat(staleness): report diverged and unknown index state instead of fresh
A staleness check collapsed every git failure into { isStale: false,
commitsBehind: 0 }. On a branch-pinned serve clone, a failed re-index
leaves the recorded commit orphaned by the --depth 1 fetch; once the
reflog expires and gc prunes it, rev-list fails and the index silently
reads as fresh while still behind.
checkStaleness / checkStalenessAsync now return an additive status:
current, behind, diverged (rev-list failed but HEAD resolved and differs
from lastCommit) or unknown (HEAD unresolvable, no lastCommit, timeout).
isStale and commitsBehind keep their values in every case.
One payload builder (core/staleness-status.ts) feeds MCP list_repos, the
hot read tools and /api/repos + /api/repo. diverged carries a hint and no
invented count; unknown appears on listings only. The helpers live in a
pure module so existing vi.mock stubs of git-staleness stay valid.
gitnexus group status no longer prints "-1 commits behind".
Fixes #3256
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Address PR review feedback (#3257)
- Document the `current` arm `fromHead` can return after a failed `rev-list`.
The `staleness-status.ts` status list, the `fromHead` docstring and the
`stalenessForTool` comment each named only `diverged`/`unknown`, so all three
described a contract the helpers do not have.
- Document both `unknown` rows on the group status type: no recorded commit
(`indexStale: true`, `commitsBehind: -1`) and a git probe that could not
answer (`indexStale: false`, `commitsBehind: 0`), which do not agree on
either field.
- Update both shipped `gitnexus-guide` copies to the new wire shape
(`{ status, commitsBehind?, hint? }`), add a `diverged` example carrying no
count, and state `unknown` is reported only by the `list_repos` listing.
Presence now means "not `current`", not "behind N". Pinned in
shipped-skills-sync.
- Lock the timed-out `rev-list` short-circuit with staleness-timeout.test.ts:
it asserts `unknown` from exactly one git spawn, so removing the `killed`
guard (which would probe HEAD again and double the #3232 bound) now fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Claim only the uncountable gap in the diverged hint (#3257)
`fromHead` reaches the `diverged` branch on any non-timeout `rev-list` failure
where `rev-parse HEAD` resolves to a different SHA. It compares the two SHAs and
runs no reachability check, so a transient object-read failure lands there while
`lastCommit` is still an ancestor of HEAD — and the hint told the operator the
commit was gone from history. `staleness-status.ts` already documents `diverged`
as only "provably not at HEAD; the count is unknown", so the string was also the
one place contradicting its own contract.
The hint now states what the check established and hedges the usual cause:
"Index is not at HEAD and the commit gap could not be counted — the recorded
commit may no longer be in this clone's history."
Updated with it: the `staleness.test.ts` diverged assertion, and the `diverged`
example plus its lead-in sentence in both shipped `gitnexus-guide` copies (still
byte-identical). `mcp/resources.ts` needs no change — it interpolates
`staleness.hint`, holding no copy of the text.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(staleness): cover the remaining #3256 review test gaps
The tri-review's lower-priority gaps, each now locked:
- staleness-fallback.test.ts: after a failed (not timed-out) rev-list,
both helpers report `current` when HEAD alone still resolves to the
indexed commit, `diverged` when it resolves elsewhere, and `unknown`
when it cannot be read, each from exactly rev-list + rev-parse. This
complements staleness-timeout.test.ts: a timeout spawns once, any other
failure probes HEAD.
- list_repos carries the listing-only `unknown` and a count-free
`diverged`; `current` carries nothing.
- group status: a resolvable repo with no recorded commit composes to
indexStale: true, commitsBehind: -1, status: unknown and renders as
"STALE (? commits behind)", through the real groupStatus path rather
than a hand-built row.
Each test was checked against a mutation of the code it guards (the
fromHead current arm, the group no-commit literal, list_repos
includeUnknown); every mutation fails its test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: tech-admin3 <tech-admin@kodnest.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
|
||
|
|
0edf9ce0ff
|
fix(ruby): guard gem requires with dependency metadata (#3096)
* docs(plans): add ruby gem require boundary plan * fix(ruby): guard gem requires with dependency metadata * fix(ruby): scope gem sources by manifest * test(ruby): model resolved lockfile specs * fix(ruby): stop local gem suffix fallthrough * docs: remove Ruby resolution plan * test(ruby): gate gem resolution correctness and scaling * test(ruby): align gem benchmark baseline with ratio gates * Address PR review feedback (#3096) - Strip a trailing .rb so local and external gem prefix matching follows Ruby's optional-suffix require. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
e7141ab0c1
|
fix(schema): persist Spring constructor-to-bean injection edges (#3239)
Co-authored-by: Jayson Albert <momeijw@gamil.com> Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> |
||
|
|
3236e2fbcd
|
fix(cobol): prefer copybook dirs so COPY EXTERNAL does not hit vendor decoys (#3240)
* fix(cobol): prefer copybook dirs so COPY EXTERNAL does not hit vendor decoys
COPY of an out-of-repo member first-won any same-named .cpy, so
vendor/EXTERNAL.cpy became a live cobol-copy IMPORTS edge. Share one
resolver between census and the regex processor: prefer copybooks/cpy/copy
plus the importer dir when present, else fail-open. Drop COBOL KNOWN_GAPS.
Fixes #2967
* bench(cobol): update depth budget for copybook-dir preference (#2967)
COBOL resolver now prefers well-known copybook directories (copybooks/,
cpy/, copy/, plus importer dir) over vendor paths when resolving COPY
statements. This intentional behavior change moves the resolver from
depth-free (prior measured ~0.885) to depth-sensitive (measured 1.751
on CI run 34394116972), because the new preferredCopybookDirs check
walks path components.
- Raise depth_budget from 1.6 to 2.4 (~1.37x the measured ratio)
- Update _measured.depth_ratio from 0.885 to 1.751
- Add _cobol_copybook_dir_preference_2967 note documenting the change
The COBOL fingerprints already reflect the new target set behavior
(vendor/EXTERNAL.cpy correctly returns null when a copybook dir is
present) per commit
|
||
|
|
79f210c5b5
|
fix(go, workspace): resolve test siblings and tighten package discovery (#3191)
* fix(go): resolve test helpers through package sibling tables * fix(workspace): discover source entries from scoped static configuration * Address PR review feedback (#3191) - Align sibling comments with the no-bare-name partition and drop the stale same-dir fallback claim. - Pin `_test.go` dot-import wildcard augmentation so a revert to nonTestFiles cannot stay green. Note: pre-existing failure in worker-pool startup crashes in the full vitest suite not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: tighten workspace discovery and index Go sibling bindings Skip leftover test/ workspace roots and extra Vite configs. Publish same-package Go names from per-package indexes instead of pairing every file. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
2220f4d851
|
fix(resolution): label fallback guesses and preserve export visibility (#3190)
* fix(resolution): distinguish name guesses and preserve export visibility * test(go): keep method enrichment fixture in one package * fix(resolution): address split review edge cases and evidence reporting * fix(exports): recognize imported and expression-local receivers * fix(resolution): align export and target evidence with language scope * fix(ingestion): preserve lexical import provenance through resolution * test(ci): rebalance Windows shards from measured slow suites * Address PR review feedback (#3190) - Label constructor unique-name guesses as global-name-fallback and run language vetoes - Tighten Go qualified, Rust crate::, and Ruby class-reopen fallback guards - Ignore for-loop shadowed CommonJS receivers and exclude guesses from the resolved-call census - Refresh FinalizeOutput and hook docs for lexical binding scopes Co-authored-by: Cursor <cursoragent@cursor.com> * Tighten review-feedback leftovers on fallback visibility. Qualified Go calls still respect export and test-package rules, nested Rust src/ stays a module segment, and top-level conditional this.x is treated as CommonJS. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command * Address PR review feedback (#3190) - Distinguish Swift package prefixes when comparing target modules - Document lexical import binding and handledSites refusal marking - Drop the stale ci-scope-parity workflow claim and prototype-safe export verdicts Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3190) Supply caller source on Ruby visibility cases so they exercise the named allow branches instead of the missing-text bypass. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(resolution): label unique constructor types as name guesses A workspace-unique class hit in findClassBindingInScope was treated as an in-scope bind, so Go/JS constructor-form sites skipped the guess label and the Go unexported veto. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(resolution): keep qualified constructors precise after unique-name split Bare constructor unique-name hits stay guesses so Go can veto an unexported type. A written qualifier is now carried as rawQualifiedName so `new pkg.Foo()` and `models.Box[T]{}` can still recover the unique class without that veto. Co-authored-by: Cursor <cursoragent@cursor.com> * test(bench): rebaseline Go/Java scope-capture fingerprints for constructor qualifiers Generic Go composite literals and qualified Java `new pkg.Foo()` now carry @reference.qualified-name on existing constructor matches. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(resolution): treat import-reached unique constructors as precise C++ #include and Rust re-exports do not mint a lexical class binding. A unique type in an imported file (or imported directory) is therefore a real bind, not a name guess, so those CALLS edges stay import-resolved. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(resolution): require named or resolved imports for constructor precision Bare Go Box[T]{} is not package-qualified, and a sibling-file import of a different name is not constructor visibility. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
506432017f
|
fix(zig): model callable-value references, and stop reporting their absence as exact (#3219)
Some checks are pending
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
Scorecard / Scorecard analysis (push) Waiting to run
* fix(impact): stop reporting 'exact' over unmodelled callable-value references
A function named in VALUE position — `bridge.accessor(Element.getNamespaceUri,
null, .{})`, `{ onClick: handler }`, a comparator handed to a sort — is
registered somewhere rather than called. The registration is modelled (a
`value-ref` site becomes a USES edge; Kythe `ref` vs `ref/call`, Joern
METHOD_REF), but the invocation THROUGH the stored value is not: it happens
later via a struct field, a registry lookup or comptime reflection.
impact()/context() nonetheless reported such a target as `epistemic: 'exact'`.
For lightpanda-io/browser that meant the DOM `Element.namespaceURI` accessor
came back with two internal callers, LOW risk and a claim of completeness —
worse than no answer, because 'exact' tells the reader not to look further.
tools.ts defines 'lower-bound' as "the walk provably missed callers", which is
precisely this case.
computeEpistemicBoundary now probes for inbound USES edges stamped with the
value-ref reason and hedges when it finds any, contributing a boundary note and
a new `causes.callableValueReferences` (unit: distinct referrer symbols).
Read from the graph, not from index metadata: unlike a dropped receiver — which
leaves no edge to find and therefore needs a persisted summary — a value
reference IS in the graph. So the signal needs no re-index and no analyzer
change, and it works on indexes written before this commit.
The cause gets its own slot rather than joining `dispatchBoundary`: a value
referrer is neither an implementation nor an interface-level consumer, and the
two differ in what the reader should do about them — a dispatch boundary is
irreducible, a callable value usually becomes traceable once the provider
models the store/load that carries it.
Language-neutral: every provider that emits a `value-ref` capture participates.
The writer and the reader now share VALUE_REF_EDGE_REASON, because drift
between them would fail silently — the probe would match nothing and every
answer would go back to claiming certainty.
* fix(zig): emit `value-ref` for a callable named in value position
`mapReferenceKindToEdgeType` has handled the registration-vs-invocation case
since #2437, and TypeScript, JavaScript and C++ all emit `value-ref`. Zig
emitted zero — so a Zig function handed somewhere as a VALUE was absent from
the graph entirely.
That is not a corner: Zig's JS bridge is built out of this one shape,
pub const namespaceURI = bridge.accessor(Element.getNamespaceUri, null, .{});
and `bridge.{accessor,function,indexed,…}` appears 2,047 times across 257 files
in lightpanda-io/browser — the project's whole JS<->Zig surface, none of it
reaching the graph. `impact` on `Element.getNamespaceUri`, which IS the DOM
`Element.namespaceURI` accessor, answered with its two in-file callers.
Three query rules, tagging `@reference.value-ref` on: a bare identifier in
argument position (`bridge.accessor(_tagName, …)`), a qualified one
(`bridge.accessor(Element.getNamespaceUri, …)`), and a const binding initialiser
(`pub const defaultHandler = onReset;`). Everything downstream already existed —
no new edge type, no schema change, no capture-machinery change, no baseline
edited.
One grammar detail is load-bearing: in tree-sitter-zig, call arguments are
DIRECT children of `call_expression` — there is no `arguments` node, only
builtins have one — so the callee has to be consumed explicitly by `function:`.
Without that binding the same rule also matches the callee of `foo(bar)` and
mints a USES edge shadowing the call's own CALLS edge. There is a test for it.
The rules are deliberately broad — `js.Bridge(Element)` and `register(count)`
match too — because the callable gate in the property-dispatch pass
(Function/Method/Constructor only) is the filter, the same design that keeps
TypeScript's `{ port: DEFAULT_PORT }` from registering anything. Measured on the
real corpus: all 3,169 emitted edges land on a Method (3,146) or Function (23).
Deliberately NOT modelled: the terminal invoke. The value reaches
`Accessor.init` -> a struct field -> `Factory.zig` reflection (`inline for`,
`@typeInfo`) -> `Caller.zig`'s `@call(.auto, func, args)` over `func: anytype`.
Resolving that needs comptime evaluation. The point is to stop dropping the
reference; the shortfall is now reported as `epistemic: "lower-bound"` by the
preceding commit instead of being papered over.
Measured on lightpanda-io/browser (698 .zig files), wiped index both times:
30,222 nodes / 71,070 edges -> 30,222 nodes / 74,229 edges. The delta is
entirely `USES` (3,169 after vs 10 before, and every one is a value-ref); CALLS,
ACCESSES, IMPORTS, HAS_METHOD, DEFINES and the rest are unchanged — purely
additive, no node invented.
The fixture joins the existing `zig-idioms` corpus rather than adding a new
one, because `bench/receiver-resolution` uses `test/fixtures/lang-resolution` as
its `--check` corpus. That gate, `scope-capture` (15 languages),
`zig-cross-file-resolution` and every other bench `--check` pass unchanged.
* fix(review): resolve qualified value references through their owner, and stop
hedging on registrations the analyzer already followed
Five findings from the review bot on #3219; four valid, all addressed.
1. QUALIFIED VALUE REFERENCES RESOLVED BY TAIL NAME (the serious one).
`bridge.accessor(Element.getNamespaceUri, …)` was resolved with
`findCallableBindingInScope(site.inScope, site.name, …)`, which never sees
the receiver and gives LOCAL bindings precedence. Reproduced:
const Element = @This();
pub fn getThing(...) // main.getThing
pub const JsApi = struct {
fn getThing(...) // JsApi.getThing
pub const thing = bridge.accessor(Element.getThing, null, .{});
};
emitted `USES JsApi → JsApi.getThing` — a WRONG edge, which is worse than
the missing edge this PR set out to fix.
`resolveValueRefTarget` now resolves a site carrying an explicit receiver
through `findClassBindingInScope` → `findOwnedMember` (the machinery
`receiver-bound-calls` already uses), gated on `CALL_TARGET_TYPES` because
`findOwnedMember` also answers with fields. A bare site keeps the lexical
walk, which is what an unqualified name means. When the owner cannot be
resolved the site is DECLINED rather than falling back: declining costs a
reference, falling back mints a confident edge to the wrong function, and
the missing reference is now reported as `lower-bound` anyway.
Cost on lightpanda-io/browser: value-ref edges 3,169 → 2,799 (−12%). Those
370 were tail-name coincidences, not registrations — all 94 `Element.zig`
`JsApi` entries survive, the cross-container case resolves and is correctly
attributed (`IntersectionObserverEntry.JsApi` → `IntersectionObserverEntry.
getTarget#0`, not the enclosing observer), and both acceptance probes are
unchanged: `getNamespaceUri` 7/3/LOW/lower-bound, `getTagNameLower`
31/10/HIGH/exact.
2. A FAILED PROBE READ AS "NO BOUNDARY". `.catch(() => [])` turned an
unanswerable query into count 0 and no note, so `exact` could be published
on the strength of a question that was never asked. It now returns `null`
and emits a boundary note. This file's own `loadMeta` comment states the
rule: a probe failing must never read as certainty.
3. NOT EVERY VALUE REFERENCE IS AN UNMODELLED INVOCATION. Where
`emitPropertyDispatchCalls` sweep 2 synthesized the dispatch (reason
`property-dispatch`), the walk did NOT provably miss the caller, so hedging
was noise over an answer that was computed — and a signal that fires on
every JS/TS hook table stops carrying information. A second probe excludes
those targets. Zig never sets a property key, so the motivating case is
untouched. The exclusion is symbol-level, not per-edge, because the graph
does not record which registration produced which synthesized call; that
residual is documented at the call site.
4. `LIMIT 50` SILENTLY UNDERSTATED THE COUNT. `rows.length` over a capped row
set published a ceiling as the documented symbol count. Replaced with
`COUNT(DISTINCT other.id)` — bounded work without a bounded answer, the
shape `countByType` in the same file already uses.
5. THE TEST DID NOT PIN THE CALLER COUNT its own comment promised. Added
`expect(result.impactedCount).toBe(2)`.
New regression tests: qualified references bind the written owner and not the
nearer lexical match; an unresolvable receiver emits nothing rather than a
wrong edge; a dispatch-modelled registration stays `exact`; a probe that cannot
run hedges instead of claiming certainty.
Known limitation, pinned by a test rather than left implicit: when a `@This()`
alias's NAME differs from its container's (`const Element = @This();` inside
`main.zig`), the receiver does not resolve and the reference is declined. The
provider has `rewriteZigThisAlias` for this, but it is applied to type nodes
and extending it to reference receivers would change existing CALL-site
behaviour. Lightpanda's `const Foo = @This();` in `Foo.zig` convention makes
the names coincide, which is why the corpus is unaffected.
* fix(review): bump the parse-cache schema and resolve module-owned value references
Review round 2 on #3219. Four inline items, all reproduced against the
worktree before deciding.
P1 — the new captures could stay INERT on a warm parse cache. Adding
`@reference.value-ref` rules to `ZIG_SCOPE_QUERY` changes
`ParsedFile.referenceSites`, which is a PARSE-TIME fact, but `SCHEMA_BUMP`
stayed at 93. A repo indexed before this branch and re-analyzed after it
replays the old, empty site list for every unchanged `.zig` file — `--force`
included, since shards are content-addressed — so no USES edge is emitted, the
boundary probe measures a real zero, and `impact` on a registered accessor goes
back to `epistemic: "exact"`. #3399 un-fixed on the incremental path most users
are on, with every cold-run test still green. DECISIONS D1-1's "zero changes to
… the schema" conflated the graph schema with the cache schema.
Bumped 93 -> 98, not 94: #3190 claims 94 and #3179 claims 94 through 97 in one
PR. `incremental-parse-cache.test.ts` re-pinned, with 93-97 added to the taken
list. Re-check against origin/main and open PRs immediately before merging.
P2 — a qualified value reference through a MODULE was declined with no hedge.
R1-2 resolves a written receiver with `findClassBindingInScope`, which requires
`isClassLike`; a namespace-only `@import` handle is not class-like, so
const dom_utils = @import("dom_utils.zig"); // no `@This()` in that file
pub const comparator = bridge.accessor(dom_utils.compare, null, .{});
resolved to nothing. That is not the conservative half of R1-2's trade-off: a
declined site emits NO edge, so there is nothing for the probe to read and
`impact` on `compare` reports `exact`. Silence, not a hedge — and the pass
comment claiming otherwise was wrong on this path.
`resolveValueRefTarget` now tries the second kind of owner a qualified name can
have. `findNamespaceValueRefTarget` reads the file's `namespace` import edges
for the handle and the target module's own `origin: 'local'` module-scope
bindings for the member — the same channel `receiver-bound-calls` Case 1
already trusts for `dom_utils.compare()`, with Case 1's three guards for Case
1's reasons: `isNamespaceNameShadowed`, local-origin bindings only, and two
distinct defs under one name resolve nothing. `CALL_TARGET_TYPES` gates module
owners exactly as it gates container owners. Language-neutral: it reads generic
namespace import edges, names no language.
Still declined, deliberately: a receiver this index knows under no name at all
(the `@This()`-alias case, owners outside the workspace). There the alternative
is a confident edge to a lexically-nearer function the source did not name.
P2 — `impact-callable-value-references.test.ts` was in neither vitest list.
It opens a real engine via `withTestLbugDB(poolAdapter: true)`, so TESTING.md
puts it in the `lbug-db` include list and the `default` exclude list; it was in
neither, so `default` also collected it into the parallel pool. Added next to
its `impact-epistemic-lower-bound` sibling in both arrays.
P3 — DECISIONS.md was stale against HEAD and embedded host paths. D2-3 still
described the `LIMIT 50` that R1-5 removed and D1-3 still described the
receiver-blind resolution that R1-2 replaced; both now carry explicit
"superseded by" pointers. The `~/code/...` and mise-node paths are replaced
with placeholders.
Two smaller corrections the new cause made necessary:
- `formatImpactResult`'s `lower-bound` header hard-coded "callers binding via
DI / dynamic dispatch", which now contradicts the value-reference bullet
printed directly under it. The bullets carry the cause; the header only
states that the count is a floor.
- `tools.ts` said a `causes.callableValueReferences` of 0 means "nothing was
missed". The dispatch exclusion is symbol-level, not edge-level, so a target
with both a followed registration and an unfollowed escape also reads 0. The
docs now say a 0 means "no unfollowed registration was proven".
Fixture: `src/webapi/dom_utils.zig` (namespace-only) plus three cases in
`Element.zig` — the module-qualified registration, a non-callable module member,
and a `u8` parameter shadowing the handle. The shadow case was verified to FAIL
with the guard disabled, so it is not passing for an unrelated reason.
Gates: `tsc --noEmit` clean, `npm run build` clean, prettier clean.
`resolvers/` 3,630 passed / 3 skipped; `unit/scope-resolution` 2,010 passed;
`impact-callable-value-references` 7 passed under `lbug-db`;
`incremental-parse-cache` 40 passed; eval formatters 104 passed. Bench --check
all PASS with no baseline edited: receiver-resolution, zig-cross-file-resolution,
scope-emission, callable-value-flow, scope-capture (15 languages), python-scope.
* fix(review): guard the class receiver against a shadowing binding, and pin the dispatchability partition
Review round 3 (`gitnexus-check` bot on `cf53bbaa`). Three findings, each
reproduced against the code before deciding.
R3-1 (Error, valid, REPRODUCED) — a CLASS receiver could resolve through a
shadowing value binding. `findClassBindingInScope` is a class-only walk: it
filters the scope chain by `isClassLike`, so it steps over a nearer binding that
is a value and keeps climbing — and past the chain entirely, into a
qualified-name fallback that answers with the unique workspace definition of the
name. A `u8` parameter named `Ticker`, in a file that neither declares nor
imports the `Ticker` container another file defines, therefore emitted
`register(Ticker.fire)` as a confident USES edge to that container's method.
That is exactly the wrong-edge failure R1-2 exists to prevent, arriving through
the class channel instead of the lexical one.
Fixed with `isOwnerNameShadowedBySomethingElse` — a sibling of
`isNamespaceNameShadowed` with one extra clause. The plain namespace guard could
NOT be reused: a container is often its own local declaration
(`fn make() { const Local = struct {…}; register(Local.go); }`), and reading that
binding as its own shadow suppresses precisely the resolutions this path exists
to make — the #2723 mistake, one channel over. So a scope that binds the name
answers immediately, and the answer is "not shadowed" only when one of that
scope's own bindings IS the def just resolved.
Both halves are pinned and both were verified to fail when the guard is
weakened: the parameter case fails with no guard at all, the local-container
case fails with the plain `isNamespaceNameShadowed`.
R3-2 (Warning; mechanism correct, unreachable today; fragility fixed instead) —
the dispatch exclusion could suppress an unfollowed registration. The bot is
right about the code: sweep 2 synthesizes CALLS only for a registration whose
site carried a `propertyKey`, while the exclusion zeroes the note on ANY inbound
`property-dispatch` CALLS edge. It is not reachable in the current rule set, and
the reason is measured rather than assumed: `@reference.value-ref` is emitted by
exactly three languages — JavaScript (2 rules), TypeScript (2), Zig (3) — every
JS/TS rule also captures `@reference.property-key` (both are object-literal
shapes) and no Zig rule does. A dispatchable registration is therefore always a
JS/TS one, an undispatchable one always a Zig one, and they cannot meet on one
symbol.
Rejected: splitting the edge `reason` into dispatchable / undispatchable. It is
the precise fix, but it is a graph-content change that churns whichever side
keeps the old literal — the Zig, TypeScript and probe suites all pin
'scope-resolution: value-ref' by hand as a drift canary — and it buys nothing
against a case no rule can produce.
What was actually wrong is that the exclusion's soundness rested on a
coincidence recorded nowhere, in files nobody reading `local-backend.ts` would
open. Fixed at both ends: the exclusion site now states the invariant, the three
facts it rests on and the two options for when it breaks; and
`value-ref-dispatchability.test.ts` fails the day it does — a JS/TS rule for a
bare callback argument, a Zig rule that grows a key, or a fourth language
emitting `value-ref` at all. Verified to fire (adding a property key to a Zig
value-ref rule fails the Zig case), and its rule splitter has its own guard test
so the suite cannot pass vacuously. `ZIG_SCOPE_QUERY` is exported for that test
only.
R3-3 (Nit, valid) — a test comment claimed the wrong epistemic result. The
`declines a qualified reference whose receiver cannot be resolved` case said the
shortfall shows up as `lower-bound`. It does not: with no edge there is no
evidence and the target stays `exact`. The pass docstring was corrected in round
2 and this comment was missed. It now says the decline costs the reference AND
the hedge, and why that is still the right trade.
Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,632 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission PASS with no baseline edited; callable-value-flow failed once on
its TIMING budget (2.006 > 1.9) with a byte-identical fingerprint, then passed
twice at 1.788 / 1.813 — machine load, not a regression.
* fix(review): let a file's own `@import` outrank the workspace-wide class fallback
Found by running the local preflight harness before pushing rather than after.
One of its five findings is a real hole in R2-2; the rest are documentation that
overclaimed.
R4-1 (valid, REPRODUCED) — the container channel preempted the file's own
`@import`. R2-2 added the module channel as a FALLBACK after
`findClassBindingInScope`, and that order is wrong: `findClassBindingInScope`
does not stop at the scope chain. When its `isClassLike` walk misses — and a
namespace handle binds a Module, so it always misses — it falls back to
`scopes.qualifiedNames`, a workspace-wide index, and answers with the unique def
of that name anywhere in the repo. A container named `dom_utils` in a file
`Element.zig` never imports therefore captured
`bridge.accessor(dom_utils.compare, …)`, binding
`Method:src/webapi/decoy.zig:dom_utils.compare#2` while the module channel that
would have answered correctly was never reached. R3-1's shadow guard cannot
catch it: the import binds at MODULE scope, which that guard treats as the floor.
Fixed by trying the module channel FIRST. An `@import` written in this file is
the strongest available statement about what the name means here and outranks a
global uniqueness guess; when the handle is not an import of this file the
channel answers nothing and the container path runs exactly as before. Pinned by
`decoy.zig` and a strengthened assertion on the existing module-owner test,
which fails on the old order.
R4-2 — the cause documentation named shapes nothing captures. `tools.ts`
illustrated `callableValueReferences` with "a callback argument", "a stored
function pointer" and `qsort(xs, n, sz, compareItems)`. Only Zig captures a call
argument or a const initialiser; JS/TS capture only object-literal property
values, and C has no value-ref rule, so the `qsort` example is counted in no
language. Both cause blocks now name the captured shapes and say that a bare
JS/TS callback argument is not among them, so a 0 does not rule it out.
R4-3 — the same block exempted itself from the re-index caveat this PR proves it
needs. "Read from the graph, so it needs no index-time metadata" is true of the
probe and false of the edges: an index built before a language emitted these
captures has none and reports 0 — the warm-cache failure R2-1 bumped
SCHEMA_BUMP for. It now says to re-analyze before reading a 0 as measured.
R4-4 — the dispatchability canary was narrower than its own promise. Its header
claimed it fails on "a fourth language emitting value-ref at all"; it reads
`languages/<dir>/query.ts`, so Vue — which owns no query and borrows
`emitTsScopeCaptures` / `emitJsScopeCaptures` — emits value-refs while the
assertion lists three languages, and a capture synthesized in code is invisible
to it. The case now asserts on query OWNERS, which is what it checks and is
sound because a delegating language inherits the rules it borrows; the header
states the synthesized-capture gap instead of letting a green tick imply it away.
R4-5 — the BARE docstring described a lexical walk that is not one.
`findCallableBindingInScope` applies the callable predicate WHILE walking, so a
nearer parameter or local is stepped over: the defect R3-1 fixed on the container
channel, unguarded here, pre-existing since #2437 and reachable in JS/TS. Out of
scope for #3399, so behaviour is unchanged and the sentence now says what the
walk does rather than implying a guarantee it does not give.
Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,632 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.
* fix(review): honour the hub opt-in for value references, so `hub.fn` means one thing
Review round 5 (`gitnexus-check` bot on `1c7c05ff`). One finding, valid and
reproduced before fixing.
`findNamespaceValueRefTarget` accepted only `ref.origin === 'local'`. R2-2
recorded that as deliberate — "the `namespaceExportsIncludeImportedNames` hub
opt-in is a provider decision this language-neutral pass does not make" — and
that reasoning was wrong twice over. Zig sets the flag
(`languages/zig/scope-resolver.ts:41`, measured on ghostty and tigerbeetle before
it landed), and the pass does have the provider in scope: `runScopeResolution`
takes one and already forwards several of its hooks to other passes.
The result was the exact asymmetry R2-2 argued against in its own first
paragraph. A Zig hub declares nothing — every name it publishes it imported — so
requiring a local declaration declines every member reached through one:
// hub.zig
pub const scale = @import("dom_utils.zig").scale;
// Element.zig
const hub = @import("hub.zig");
pub fn callsThroughTheHub(v: u8) u8 { return hub.scale(v); } // resolved
pub const scaled = bridge.accessor(hub.scale, null, .{}); // declined
One name meaning two different things depending on whether a `(` follows it.
Fixed by forwarding `provider.namespaceExportsIncludeImportedNames` into the pass
and consulting the published channel when it is set — the same question
`receiver-bound-calls` Case 1 asks, through the same `lookupBindingsAt` read
`findExportedDefIncludingImportedNames` performs for the CALL form. The pass
still names no language; it asks the provider, which is the sanctioned hook.
Precedence is unchanged where it mattered: a locally declared member still wins
over a republished one, two distinct defs under one name still resolve nothing,
and `CALL_TARGET_TYPES` still gates the answer — pinned by `hub.DEFAULT_NS`, a
re-exported CONSTANT, which stays unregistered. Languages that do not opt in are
unaffected: the parameter defaults to `false`, and passing `false` was verified
to fail the new hub test, so it is load-bearing. The fixture republishes a member
(`scale`) that nothing else uses, so the hub assertion cannot be satisfied by an
edge another case emitted.
Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,634 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.
Also recorded in DECISIONS.md: `/autofix` returned "No successful autofix run"
because the `PR Autofix` run on `1c7c05ff` failed at `actions/upload-artifact`
with a 403 from GitHub's artifact storage, not because of anything in this
branch. This push triggers a fresh run.
* fix(review): inspect the module scope in the owner-shadow guard, and check the TSX suffix
The `gitnexus-check` pass on `bd6e577e` carried three findings of its own, and I
answered the later `1c7c05ff` pass without noticing them. Recording that as a
process failure too: bot reviews are per-head, and a later pass does not
necessarily repeat an earlier one's findings.
R6-1 (Error, valid, REPRODUCED) — the shadow guard stopped one rung short.
`isOwnerNameShadowedBySomethingElse` returned `false` on reaching the module
scope, justified as "a container declared there IS the binding, and the caller
already resolved it". That holds when the owner came from the scope chain and
fails when it came from the workspace-wide qualified-name fallback:
// Gauge.zig — never imported by Element.zig
const Gauge = @This();
pub fn read(self: *Gauge) u8 { … }
// Element.zig
const Gauge = @import("dom_utils.zig").DEFAULT_NS; // NOT a container
pub const level = bridge.accessor(Gauge.read, null, .{}); // → Gauge.zig's read
`findClassBindingInScope` steps over the module-scope binding because it is not
class-like, answers from `scopes.qualifiedNames`, and the guard waved it through.
Worth recording: the first fixture attempt did NOT reproduce. A local
`const Gauge: u8 = 3;` also claims the workspace qualified name `Gauge`, leaving
two candidates, and the fallback refuses to guess between two — so the shape
defeated itself. Binding the name by IMPORT claims no qualified name, the
fallback stays unique, and it fires. A negative result on the first shape was not
evidence the finding was wrong.
Fixed by inspecting the module scope as the last rung instead of skipping it.
The identity exemption is what makes that safe where `isNamespaceNameShadowed`
cannot do it (#2723: a namespace import writes its own name into the module scope
and would read as its own shadow) — the binding that IS the owner exempts itself,
and only a binding to something else answers `true`. `lookupBindingsAt` is
consulted at that scope and only there, because an imported alias lives in the
finalized channel rather than in `scope.bindings`.
R6-2 — the hub re-export finding, already fixed in `5d8fe9d8`; the same defect
restated on the later head.
R6-3 (valid, fixed) — the dispatchability canary omitted the TSX suffix.
`getTsScopeQuery` analyzes a `.tsx` file with `TYPESCRIPT_SCOPE_QUERY +
TSX_JSX_QUERY_SUFFIX`, and the test read only the base, so a `value-ref` rule
added to the suffix would be emitted in TSX analysis with the canary green. The
suffix is now exported and concatenated into the check; verified load-bearing by
adding an unkeyed `jsx_expression` value-ref rule to it, which fails the
TypeScript case.
Gates: tsc --noEmit clean, npm run build clean, prettier clean.
`test/integration/resolvers` 3,635 passed / 3 skipped (70 files);
`test/unit/scope-resolution` 2,015 passed (120 files);
`impact-callable-value-references` 7 passed under `lbug-db`. Bench --check:
receiver-resolution, zig-cross-file-resolution, scope-capture (15 languages),
scope-emission, callable-value-flow all PASS with no baseline edited.
* perf(bench): gate callable-value reference resolution on linear scaling
`resolveValueRefTarget` (#3399) replaced one lexical walk with four
channels, and the last of them — a qualified receiver resolved through
`findClassBindingInScope` — falls back to `scopes.qualifiedNames`, a
WORKSPACE-WIDE index consulted once per site. Keyed that is O(1); scanned
it is O(files) per site, and a registration table that costs O(sites)
today costs O(sites x files) tomorrow. Nothing in the suite can see that:
the fixtures are single-file, and the corpus it actually matters on is
lightpanda-io/browser, where `bridge.{accessor,function,…}` appears 2,047
times across 257 files.
`bench/value-ref-resolution/measure.mjs` builds two synthetic Zig corpora
of identical shape 4x apart in file count, and times ONLY the per-site
resolution loop — extraction, `reconcileOwnership` and finalize are setup.
`linear_factor` is `(t_large/t_small)/(N_large/N_small)`: measured
1.00-1.09 across runs, with `us_per_site` flat at 1.4-1.5 between the arms.
Correctness comes first, because a timing gate alone is satisfied by a fast
wrong answer: exact site/resolved/declined counts per arm plus an
order-independent sha256 over every (site -> resolved target) pair. The
corpus exercises all four channels — CONTAINER, NAMESPACE, HUB and BARE —
and carries two DECLINE controls per module (a non-callable namespace
member, a non-callable bare argument), so a widened callable gate moves
`declined` instead of hiding inside the timing.
Verified load-bearing rather than assumed: making the qualified lookup scan
a workspace-sized collection leaves the fingerprint IDENTICAL and takes
`linear_factor` to 3.752 against a slack of 1.375 — the regression class
this exists for is exactly the one no correctness gate can see.
Zig is the corpus because it is the only language whose provider sets
`namespaceExportsIncludeImportedNames`, so it is the only one that can
exercise the hub channel at all; the pass itself names no language.
`resolveValueRefTarget` is exported for the bench. Timing
`emitPropertyDispatchCalls` instead would fold the signal into edge
emission, and re-implementing the channel order in the bench would pin the
bench's idea of the function rather than the function.
* feat(zig): index every build package in the repo, not only the root one
Zig already had workspace setup — `loadZigBuildConfig` has parsed
`build.zig.zon` `.path` deps, the root `build.zig`'s named modules and each
build module's own `addImport` table since #1432, threaded through
`ScopeResolver.loadResolutionConfig` exactly as tsconfig is for TypeScript.
What it did not have is the part `tsconfigFor` supplies: per-package scope.
`loadZigBuildConfig` reads `<repoRoot>/build.zig{,.zon}` and nothing else, so
a repo laying its packages out as `packages/<name>/build.zig` — no root build
files at all — got `null`, and EVERY bare `@import("<module>")` in it went
unresolved. Cross-file resolution silently degraded to relative imports.
Measured on a two-package probe before this change: `config = null`,
`@import("core")` from `packages/app/src/main.zig` → `null`.
`loadZigWorkspaceIndex` discovers packages the way `findTsconfigFiles`
discovers configs — one bounded breadth-first walk that skips the hardcoded
ignore set — and `zigPackageFor` is the `tsconfigFor` analogue: the nearest
enclosing package governs a file, deepest-first, with no fall-through to an
outer package. Fall-through is what makes a vendored dependency's
`@import("config")` resolve to the outer repo's `config` module, the same
failure `loadTsconfigIndex` documents for a package declaring no `baseUrl`.
`resolveZigImportInternal` is UNCHANGED — impact analysis puts it at HIGH risk
with 6 dependents, and it does not need to move: it is handed one package's
config, and which config it receives is the only difference. Its 48 existing
tests pass untouched. A nested package's paths are rebased to repo-relative at
load time (`.path = "../core"` → `packages/core`, which the root-relative
reading rejects outright as an escape); the ROOT package keeps its raw
spelling, which is what `parseZigBuildZon` promises and its tests pin, so a
single-package repo is byte-identical.
`loadImportConfigs` — which runs unconditionally for every repo, Zig or not —
keeps calling the root-only loader, so no non-Zig repo pays for the walk. That
is the split TypeScript already has between the cheap `loadTsconfigPaths` and
the repo-walking `loadTsconfigIndex`, which `loadResolutionConfig` reaches only
during a language pass.
The `zig-monorepo` fixture carries the discriminating case rather than only the
happy path: `tool` binds the alias `core` to its OWN `src/core.zig`, so a
repo-wide flattened module map — the shape a workspace index invites — would
point `measure` at `packages/core`, a confident edge into a package `tool` does
not depend on. That is worse than the unresolved import this fixes, and only
per-package scoping keeps them apart.
Tests: 7 unit cases pinning the index (including `loadZigBuildConfig` answering
`null` on the same fixture, side by side) and 3 integration cases pinning the
EDGES — cross-package CALLS from the module root AND from a non-root file, and
`tool` not crossing over. Verified load-bearing: restricting the index to the
root package fails all three.
Second consumer caught by the integration test and fixed with it:
`populateZigWorkspaceStaticGating` reads the same `resolutionConfig`, per file
because two files of that pass can belong to different packages.
Gates: build clean, `tsc --noEmit` clean, prettier clean, eslint 0 errors.
`test/integration/resolvers` 3,768 passed / 4 skipped, `test/unit/scope-resolution`
2,015 passed. Bench `--check` with NO baseline edited: `receiver-resolution`
(whose corpus is `test/fixtures/lang-resolution`, where the fixture lands),
`zig-cross-file-resolution`, `value-ref-resolution`, `scope-capture` (15
languages), `import-target`, `callable-value-flow`, `scope-emission`.
* perf(zig): dequeue the package walk by head index, not `shift()`
`findZigPackageDirs` pushes children while it drains the queue, which keeps
the array in a mode where `Array.prototype.shift()` memmoves the whole
remainder rather than taking V8's left-trimming fast path — so the walk is
quadratic in the frontier, bounded only by `ZIG_SCAN_MAX_DIRS` (20,000).
Measured at that bound rather than estimated from the move count: 53 ms at
fan-out 4 and 81 ms at fan-out 20, against 0.8 ms with a head index — 66-106x,
and paid before a single config is read.
FIFO order is unchanged, and the package ordering never depended on the walk
anyway: `loadZigWorkspaceIndex` sorts by (dir length desc, localeCompare).
Memory is unchanged too — the entries were already retained by the pushes;
`shift()` only dropped the head.
`findTsconfigFiles` and `loadNodeWorkspacePackages` carry the same walk with
the same bound and are equally affected. Deliberately NOT fixed here: they are
pre-existing, outside this PR's diff, and reach TypeScript/Node import
resolution, which nothing in this change set covers. Filed separately.
Reported by gitnexus-check on
|
||
|
|
b60c21d05d
|
fix(group): extract NestJS GraphQL contracts on real indexes (#3201) (#3227)
* fix(group): extract NestJS GraphQL contracts against real 0-based indexes (#3201) Provider lookup used 1-based startLine while the graph stores tree-sitter rows, so every resolver missed. Also try PascalCased Document names and inline sibling FragmentDoc interpolations from graphql-codegen output. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(group): bind arrow-field providers and fail closed on interpolations Match Method startLine to the public_field_definition wrapper, decode template escape sequences, and reject FragmentDoc names that mix static and dynamic declarators. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(group): only inline interpolated templates under a gql tag Cooked reconstruction is not the runtime value for String.raw or unknown tags, so those interpolations stay fail-closed. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(group): tighten gql-tag trust and PascalCase Document lookup Only the identifier `gql` is a trusted interpolating tag. Underscored operation names now try the full pascal-case Document candidate graphql-codegen emits. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(group): memoize GraphQL interpolation source resolution Avoid exponential re-walks when the same fragment name is declared twice at each layer of a ${FragmentDoc} chain. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(group): fail closed on invalid tagged-template escapes Treat line continuations as empty cooked text and reject \8/\9 plus legacy octals so reconstructed gql source matches runtime. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
a839d029ac
|
feat(api): report branch and index freshness on the serve repo routes (#3232)
* feat(api): report branch and index freshness on the serve repo routes `GET /api/repos` and `GET /api/repo` now return the branch an index was built from, its `lastCommit`, and — where it can be computed — how far behind the working tree it is. All of it already existed: the fields are on `RegistryEntry`, and `checkStalenessAsync` is the helper MCP `list_repos` and `gitnexus status` already use. Only HTTP never asked. #3199 made this pointed. A branch-pinned analyze registers its own entry, so one repository yields two rows in /api/repos, and telling them apart over HTTP meant pattern-matching the clone-directory suffix — a layout detail that is trimmed for long refs and absent for path-registered repos. Staleness is reported in the shape `list_repos` already returns: present only when the index is behind, carrying commitsBehind and hint. It is meaningful for path-registered repos; a url-registered repo is cloned --depth 1, so its recorded commit is HEAD and a diverged history cannot be walked by rev-list anyway. Documented in repo-projection.ts rather than left to be discovered. The projections live in their own module so the field list is assertable. Route-inline, they were reachable only by booting a server, which is how `branch` stayed unexposed while `gitnexus list` printed it. /api/repos also gains the rate limiter it lacked: this change makes one unauthenticated GET cost a `git rev-list` per registered repo, the shape this codebase already limits elsewhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api): bound the staleness probe and keep liveness off the fan-out route Addresses the tri-review on #3232. P1. `checkStalenessAsync` caught every git ERROR but not a HANG — a working tree on a disconnected mount or behind a stuck lock never settles, and /api/repos fans that out once per registered repo. Worse, the web liveness probe used /api/repos on a 2s budget and re-polls on failure, so each failed probe stacked another N children on a server that was actually healthy. Two independent fixes, because either alone still leaves a sharp edge: - `execFileAsync` now carries a 5s timeout, so a hang is killed and routed into the same fail-closed "not stale" answer the existing catch already gives a bad SHA. This also protects MCP `list_repos`, which shares the helper. - `probeBackendStatus` (and `isDatabaseReady` through it) asks /api/health instead. Liveness should not cost one subprocess per indexed repo. Health sits behind the same /api/* edge gate, so the 401 "gated vs absent" distinction is preserved. P2. rate-limit.test.ts pins which routes carry a limiter so a dropped one fails a test before CodeQL has to catch it; /api/repos was newly limited but not pinned. Added, and verified it fails when the limiter is removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> |
||
|
|
4154b63131
|
feat(indexing): add Objective-C semantic indexing support (#3179)
* docs: add Objective-C fork provider notes * feat(objective-c): add deterministic provider and grammar * feat(objective-c): finalize provider MVP * fix(objective-c): harden provider integration * fix(objective-c): normalize bare macro markers * docs(objective-c): integrate provider documentation * fix(objective-c): harden resolution and header classification * fix(objective-c): complete provider follow-ups * fix: address Objective-C review follow-ups * chore: format Objective-C grammar sources * fix(objective-c): harden review follow-ups * Address PR review feedback (#3179) Keep Objective-C chunking and macro recovery aligned with the grammar, and stop Community MEMBER_OF edges from leaking into symbol context. Co-authored-by: Cursor <cursoragent@cursor.com> * Address follow-up review on ObjC chunking and language fallback. Keep preprocessor directive text from changing file-scope brace depth, group real ivar nodes, skip header modifiers, and restore Rakefile/Gemfile detection through getLanguageFromFilename. Co-authored-by: Cursor <cursoragent@cursor.com> * Parse Objective-C headers with the objc grammar in embeddings. ensureAndParse and structural extraction now use the same content classifier as ingest, including method snippets from .h files, so Protocol/Category/Class chunks are not re-parsed as C++. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3179) Keep file-scope macro elision off C line splices and @interface/@protocol/@implementation bodies, and attach ivar attributes to the following instance variable when chunking. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(bench): rebaseline Objective-C CSV emit * feat(objective-c): add workspace resolution and linear emit benches Plain .h files are classified as C++, so the ObjC pass could not resolve #import of those headers. Load a C/C#-style workspace once per pass, and keep protocol-candidate USES linear. Refs #3179 Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3179) - Compare LadybugDB labels() as a scalar when excluding Community MEMBER_OF edges. - Walk superclass members, skip file-static C sibling defs, and ignore comments in ObjC header/macro scans. Note: pre-existing failure in objective-c-provider integration (worker-pool ready timeout) not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3179) Emit Objective-C declaration captures so compilation-unit siblings can share header/implementation bindings, and keep class vs protocol visibility groups distinct. Note: pre-existing failure in worker-pool startup (GITNEXUS_WORKER_READY_TIMEOUT_MS) not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review feedback (#3179) Emit every comma-separated property/ivar declarator, and count @interface after a multiline block comment closes so in-declaration macros stay intact. Note: pre-existing failure in worker-pool startup (GITNEXUS_WORKER_READY_TIMEOUT_MS) not addressed by this PR. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: ximengkai <ximengkai@soyoung.com> Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
8ddab9aed5
|
feat(api): honor branch on POST /api/analyze (#3199)
* feat(api): honor branch on POST /api/analyze The serve route accepted a `branch` field in the request body, returned 202 and reported the job `complete` — while indexing the remote's default branch. Express drops unknown body fields, so the caller got no error and no warning; the only way to notice was to inspect the checked-out clone. Both ends of the plumbing already existed: CloneOrPullOptions.branch is honored by cloneOrPull, and AnalyzeOptions.branch already drives resolveBranchPlacement. Only the HTTP layer was missing, so this wires `branch` from the route through cloneOrPull and LaunchOptions into the worker's AnalyzeOptions. StartMessage.options is already typed as AnalyzeOptions, so the IPC protocol is unchanged. Validation reuses validateBranchName — the same function backing the CLI's `--branch` — so both entry points accept exactly the same refs and a malformed value is rejected with 400 before it can reach git. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api): make branch part of job identity and complete ref validation Addresses the review on #3199. 1. Job dedup ignored `branch`, so a request for branch B while branch A was in flight was answered with A's job and a 202. The caller would read that as "B is indexed" — the same silent wrong-branch outcome honoring `branch` was meant to remove. Branch is now part of the dedup identity; a different-branch request falls through to the single-slot guard and gets a truthful 409 instead. 2. validateBranchName implemented only a subset of git's ref rules, so `feature.lock`, `/feature`, `feature/`, `feature//next`, `@`, `@{` and dot-prefixed components passed validation and failed later in the git subprocess — a 202 plus a background failure rather than the advertised 400. The remaining `git check-ref-format` rules are now enforced at the same chokepoint, which fixes the CLI and `.gitnexusrc` paths too. No branch git can create is affected. 3. The LaunchOptions doc claimed an explicit branch always pins `branches/<slug>/`. resolveBranchPlacement keeps the run on the flat slot when that slot has no owner, or when its owner is already this label. Comment and CHANGELOG corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(server): describe cloneOrPull's branch path in its contract comment The header comment predated `options.branch` and still said an existing clone is only ever `git pull --ff-only`. The implementation has a second path: with a branch it fetches that ref and runs `checkout -B <branch> origin/<branch>`, so the requested branch — not the one already checked out — ends up in the working tree. The stale comment is actively misleading: a reviewer reading it concludes that requesting a branch on an existing clone silently analyzes the default branch, which is not what happens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(server): scope the origin check to the existing-clone path My previous comment said remote.origin is verified "in both cases", which is wrong: assertRemoteMatchesRequestedUrl runs inside `if (exists)`, so a fresh clone has no origin to check. Restructured around whether targetDir exists, which is what actually selects the behavior. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(server): cover branch job identity in JobManager's own suite The branch dedup tests were sitting in analyze-api.test.ts, but they exercise JobManager directly, so they belong beside the existing "returns existing job for same repoUrl when active" case in analyze-job.test.ts. Moved, and extended to cover the callers that omit branch entirely — the upload route, the embed manager and the existing tests — which compare undefined === undefined and are unaffected. Also pins that branch survives the whole clone -> analyze -> terminal update sequence, since it is now part of dedup identity and must not drift mid-flight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api): give a pinned branch its own clone and settle the slot it wrote Addresses the review on #3199. Per-branch clone directories (review option 2). One checkout per repo made `branch` one-shot: after any analyze the tree is dirty with generated AGENTS.md / CLAUDE.md / .claude/, so a pinned request 202'd and then died on cloneOrPull's porcelain refusal. Worse, a later request that OMITTED `branch` pulled whatever branch the last pin left checked out and indexed it as the default — silent wrong content, the same class as #3198 one request later. A pinned run now clones into `<repo>__<branchSlug>`, so the two requests no longer share a tree. The pinned clone registers under its directory name, because both dirs share an origin and the inferred name would otherwise collide; that name re-derives through getCloneDir, so DELETE still finds it. Finalization gate. registerRepo always records the flat `.gitnexus`, but a pinned run whose label differs from the flat slot's owner writes `branches/<slug>/`. The gate probed the flat path regardless, so it never settled, and the worker's normal exit 0 — sent ~500ms after `complete`, while the job is deliberately still non-terminal — was classified as a crash and a successful analysis was retried three times and failed. The gate now follows the placement the worker reports (isPrimaryBranch, added to the IPC allowlist under the rule that module already documents), and an exit after a terminal IPC counts as winding down, not dying. Reported by the maintainer and reproduced independently by @azizur100389. analyzeCloneOptions extracted so the token/branch combination is asserted. Inline, the branch-only case — a public URL with no token — was untested, and dropping it there would silently reindex the default branch while every other test stayed green. CHANGELOG: the previous entry claimed the newly-400'd payloads "would have failed at git", which is true of CLI --branch but wrong for HTTP, where they succeeded on the default branch. Documented as an explicit behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(server): bound the branch clone-dir name to one path component `validateBranchName` allows a 255-character ref and `branchSlug` appends a dash plus 8 hash characters, so `<repo>__<slug>` reached 267 — past the 255-byte component limit on ext4/APFS/NTFS. The clone would then fail to create its target directory, which the character-only regex could not catch. Only the readable half is trimmed. The hash is a digest of the full ref and is always kept, so two long branches sharing a prefix still resolve to different directories rather than silently sharing an index. `branchSlug` itself is untouched: the per-branch index slots already use those names on disk, and shortening them there would orphan existing indexes. Also corrects two comments: the forwarding test claimed a branch selector always pins `branches/<slug>/` (it keeps the flat slot when that slot has no owner or already owns the label), and the web client's `branch` doc said omitting it means the remote default — true for a `url` request, but a `path` request is never cloned and indexes whatever that tree has checked out. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(server): bound the branch clone-dir name and unblock same-branch re-index Two defects found by testing this branch end to end, plus the review nits. 1. Path length. `validateBranchName` allows a 255-character ref and `branchSlug` appends a dash plus 8 hash characters, so `<repo>__<slug>` reached 267 — past the 255-byte component limit on ext4/APFS/NTFS, and the clone could not create its directory. Only the readable half is trimmed; the hash is a digest of the full ref and is always kept, so two long branches sharing a prefix still get separate directories. `branchSlug` itself is untouched — the per-branch index slots already use those names on disk and shortening them there would orphan existing indexes. 2. Same-branch re-index. Per-branch clone dirs stopped branches from contaminating each other, but a REPEAT pin still failed: analyze writes AGENTS.md / CLAUDE.md / .claude/ into the clone, so the second pinned run met its own dirt at the porcelain check and asked for `overwrite_local_changes`. When HEAD already matches the requested branch there is nothing to switch, so the run now takes the same `pull --ff-only` path an unpinned request takes — review option (1), alongside (2). The refusal is untouched where it matters: a real switch, or a detached HEAD, still goes through the checkout path and can still refuse. Also: repositions getCloneDir's JSDoc, which an inserted constant had orphaned; gates the pinned `registryName` on the same condition as the clone, so supplying both `url` and `path` no longer renames the operator's local repo; corrects a test comment that claimed a branch selector always pins `branches/<slug>/`; and corrects the web client's `branch` doc, which said omitting it means the remote default — true for `url`, but a `path` request is never cloned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: drop the CHANGELOG edits from this PR Requested in review. Checking the history, no feat/fix PR here touches gitnexus/CHANGELOG.md — the only recent commit on it is `chore: release v1.6.11`, and CONTRIBUTING says release notes are generated from the merged PR title via .github/release.yml. Hand-editing an [Unreleased] section from a feature branch was my mistake, not the project's convention. The behavior change it documented (branch: null / "" / non-string now 400 where they were previously dropped and the default branch indexed) is stated in the PR description instead, which is what feeds the release notes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(server): align the clone/pull contract with the same-branch fast path Two comments I wrote went stale against my own later change. `cloneOrPull`'s header still said an existing clone with `options.branch` always fetches and runs `checkout -B`. Since the same-branch fast path landed that is only true when the branch actually differs; when it is already checked out the run takes `pull --ff-only` and no dirty-tree check applies. The header now splits on whether the branch differs, which is what the code branches on. The web client's `branch` doc said omitting it on a `url` request clones the remote default. That holds only when there is no clone yet — an existing unpinned clone is pulled on whatever branch it already has checked out. Comments only; no behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address PR review feedback (#3199) Pin a same-branch re-index to `git pull --ff-only origin <branch>` so the job cannot follow an unverified `branch.<name>.merge` while still skipping the dirty-tree refuse. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(server): close the #3199 holes on pinned analyze re-index Same-ref updates were a raw pull dest (force-fetch via +) or a shallow ff-merge that could not move, and a tag pin compared the tag-object SHA so re-index refused a dirty tree. Fetch the mapped remote-tracking ref, stay put on a peeled SHA match, restore only GitNexus overlays, skip the 60s settle on alreadyUpToDate, and keep branch validation in core so the HTTP route does not import the CLI. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(autofix): apply prettier + eslint fixes via /autofix command --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
bddbb0ff9f
|
test(cli): prove detached refresh by ordering, not by wall clock (#3221)
`lets the parent exit without waiting for a detached refresh child` bet twice on absolute wall-clock budgets and lost both bets on loaded CI runners: - `expect(elapsed).toBeLessThan(1_800)` bounded the *parent's* cold `node --import tsx` boot, inferring "did not wait" from a 1050ms margin over the mock fetch's 750ms sleep. Reproduced failing on Linux under 40-way load at 1904ms. - The 15s poll for the cache file had to cover the whole detached child: node boot, a tsx transpile of 22 source files, an `acquireFileLock` that shells out to `ps` (POSIX) or `powershell.exe -Command Get-CimInstance Win32_Process` (Windows), the mocked fetch, and the atomic write. That chain measures ~1.0s locally but has no bounded upper limit on a contended runner, and it is what timed out in CI. Replace both budgets with an ordering proof. The preloaded mock fetch now parks the refresh child until the test releases it, so the sequence asserted is: parent exited -> child provably still mid-refresh (started marker present, cache absent) -> release -> cache written. That is strictly stronger than the old elapsed-time inference, and it holds at any machine speed. The remaining `expect.poll` timeouts no longer carry the assertion's meaning; they are only "is the child dead" safety nets. The mock's wait is bounded at 60s so an abandoned child (test failed before releasing, temp home already removed) still exits instead of spinning forever. Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1c1cbf111e
|
fix(zig): resolve cross-file static gates (#3185)
Some checks failed
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
Gitleaks / gitleaks (push) Has been cancelled
Publish / Classify release event (push) Has been cancelled
Scorecard / Scorecard analysis (push) Has been cancelled
Trivy Image Scan / Trivy (gitnexus-cli) (push) Has been cancelled
Trivy Image Scan / Trivy (gitnexus-web) (push) Has been cancelled
Publish / Build & Push RC Docker images (push) Has been cancelled
Publish / RC guard (marker + release-PR skip) (push) Has been cancelled
Publish / ci (push) Has been cancelled
Publish / Publish to npm (push) Has been cancelled
* fix(zig): resolve cross-file static gates * Fix Zig workspace import alias enrichment * Handle extensionless Zig workspace imports * Reuse Zig import resolution for static gates * Document Zig workspace static gating * Harden Zig workspace reference enrichment * Benchmark Zig cross-file static gating * Enforce linear Zig benchmark scaling --------- Co-authored-by: Gergő Magyar <gergomagyar@icloud.com> |
||
|
|
8f006bd759
|
perf(parse): batch cache packs into one dispatch round (#3196)
* perf(parse): batch cache packs into one dispatch round `WorkerPool.dispatch` is a barrier, so dispatching one parse-cache pack at a time leaves most slots idle for every round-trip. Packs are keyed by `(language, hash(path) % 128)`, so the byte budget rarely binds: this repo produces 1285 packs where the budget alone needs 16, and 549 of those hold a single file. In a real analyze, 76 of 221 dispatched chunks carried one file and cost 15.3s — 20% of the parse phase for 3.4% of the files. Chunks now accumulate into a round bounded by `GITNEXUS_PARSE_ROUND_BYTES` of cache-missing source (default: the chunk byte budget) and go out through a new `WorkerPool.dispatchGroups`. Jobs are still cut at pack boundaries, so each job carries exactly one `chunkHash` and every result stays attributable to the pack whose cache key owns it. Cache hits ride along as round entries, and rounds drain in `chunkIdx` order, so deferred aggregation stays deterministic. Cold `analyze --index-only` on this repo (2234 parseable files, 16 workers): 110.3s -> 70.5s total, parse phase 74.0s -> 40.5s, 221 dispatches -> 15. Graph output is unchanged: 51,286 nodes / 163,092 edges / 2106 clusters / 759 flows in both arms. Peak main-thread RSS 3372MB -> 3487MB (+3.4%). `dispatchGroups` also claims the pool synchronously and rejects a concurrent call. Two overlapping dispatches hand the same slots out twice and both stall; the first version of this change did exactly that, and the only symptom was every worker idle-timing out ~10s later with no indication of the cause. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(parse): bound what an open round holds, not just what it dispatches Follow-up to the review of #3196. Three reviewers independently found the same defect: `roundMissBytes` was the only in-loop close condition, but cache HITS were queued into the same round without contributing to it. A warm re-analyze misses nothing, so no round ever closed and every chunk's source plus its cached worker output stayed resident until the tail drain — the #2649 heap failure shape on a large repo. - Hit entries now carry a file COUNT, not the file array, so a replayed chunk never pins its source text. `applyChunkResults` only ever read `.length`. - Track `roundBufferedBytes` across hits and misses and close on either cap. Verified on a warm run: with the cap, draining starts as soon as 2MB is buffered; without it all 221 merges land in the final 10% of the phase. - Warm progress no longer freezes at the phase floor. `filesParsedSoFar` only advances at drain, so a new `queuedFilesSoFar` feeds the progress events while `filesParsedSoFar` stays the merge-accurate throughput number. - A throw from `drainRound` used to unwind straight to `terminate()` while the next round's workers were still busy — the #2432 mid-N-API abort hazard. Settle the in-flight round first, then propagate. - `dispatchGroups` returns one array per group; assert that length instead of `?? []`, which turned a contract break into a silently empty chunk. - Collapse `PendingWorkerChunk` into the `miss` RoundEntry it duplicated. - Repair two stale doc comments: `dispatch`'s JSDoc had been orphaned onto `dispatchGroups`, and `dispatchChunkParse` still described chunk overlap that now lives in parse-impl's round machinery. - New test: a round mixing a cache hit and a cache miss. `drainRound` walks entries in chunkIdx order but pulls results on a separate cursor, and no existing test put both kinds in one round with content assertions. Cold analyze unchanged: 71.3s, 15 rounds, 51,286 nodes / 163,092 edges. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
780cac7885
|
fix(parse): keep stable cache packs parallel (#3194)
Stable cache packs introduced by
|
||
|
|
c5c4fbe43c
|
fix(zig): vendor tree-sitter-zig so npm i -g no longer warns on peers (#3180)
* fix(zig): vendor tree-sitter-zig so npm i -g no longer warns on peers Published overrides do not apply to dependents, so the Zig optionalDependency kept warning that tree-sitter@0.21.1 does not satisfy peerOptional ^0.22.1. Load it from vendor/ like Dart/Kotlin/Swift instead. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): add zig parse snippet to prebuild validate The six zig prebuild jobs failed at "Validate the .node loads and parses" because snippets[GRAMMAR] was undefined and tree-sitter threw "Input must be a function". Co-authored-by: Cursor <cursoragent@cursor.com> * chore: drop Unreleased changelog note from the Zig vendor PR CHANGELOG.md is owned by the release process, not individual PRs. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33949409205 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33949616521 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33949829377 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33950025077 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33950220655 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33950400275 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33950607933 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33950882912 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33951170483 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33951386477 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33951624305 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33951813452 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33951998309 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33952225172 * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33952524393 * fix(ci): stop native prebuild rebuild loops PR path filters and source checks see the cumulative diff, so generated binaries kept rebuilding the original source change. Skip output-only synchronize events using their exact before/head range, failing closed when Git cannot compare it. Exercise the workflow against real commit histories, including multi-commit source pushes and merge-ref drift. * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33953226606 * fix: address Zig PR review feedback (#3180) Check for the vendored package without loading its native binding so a broken installed Zig grammar fails the parsing test instead of skipping. Match the optional child descriptor in the Zig metadata declaration and include Zig in the two optional/vendored grammar comments. Validation: 159 targeted tests, TypeScript, metadata type fixture, and formatting passed. Injected native-load failure now fails instead of skipping; explicit Zig opt-out still skips. * chore(vendor): rebuild native prebuilds (tree-sitter-zig) Built by https://github.com/abhigyanpatwari/GitNexus/actions/runs/33956342308 --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: gitnexus-release-bot[bot] <gitnexus-release-bot[bot]@users.noreply.github.com> |