mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-03 02:21:44 +00:00
2 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0261982d9a
|
fix(analyze): make --memory-budget set the real heap and report rebuild reasons once (#3386)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* feat(analyze): --memory-budget flag with heap-limit override and worker-pool degradation (#3137) Adds an explicit `--memory-budget <mb>` CLI flag that overrides the RAM/cgroup auto-sized main-thread heap ceiling for the parse phase: - CLI validation (integer >= 200 MB) before bar.start(), matching the --workers pattern - Threaded CLI → runFullAnalysis → PipelineOptions → parse-impl as memoryBudgetBytes - parse-impl resolves the heap limit as budget ?? v8.heap_size_limit, so both the preflight projection warning and the #2649 mid-loop abort probe honor the budget - Graceful degradation: when the projected heap need exceeds the budget at the computed pool size, the pool shrinks (never below 1, never above the operator's --workers) before sub-batch math and pool construction, so all downstream consumers see the degraded size Omitting the flag keeps the auto-sizer path byte-identical. Refs #3137 * feat(analyze): collapse rebuild-gate log into one summary + persist needsFullRebuild verdict (#3137) The nine meta-mismatch rebuild gates (pdg mode, content retention, schema fingerprint, graph-write collapse, analysis features, Spring vendor prefixes, runner identity, FTS CJK mode, embedding dims) each logged individually and set force:true independently. An upgrade that trips several at once printed a scattered wall of near-identical warnings. - Gates now collect into rebuildReasons[]; a single summary block prints them (inline for one, numbered for many) and sets force once. Per-gate Tip text is preserved verbatim inside the entries. - The verdict persists to meta (needsFullRebuild: {reasons, recordedAt}) BEFORE the rebuild starts. If the rebuild is interrupted, the next run announces the recorded reasons up front instead of quietly attempting an incremental write on a half-rebuilt index — the gates may not all re-fire against a wiped DB. - The verdict is cleared on the next successful completion (the final meta does not carry the field forward). Semantics unchanged: every gate was already evaluated (none early-returns), force is idempotent, and a rebuild happens iff at least one reason fired. Refs #3137 * refactor(cli): share one integer flag parser across analyze, watch, and wiki Replace the duplicated Number.isInteger checks for --workers, --embeddings, the positive env-backed analyze flags, the watch interval flags, and wiki's --timeout/--retries with parseIntegerOption (per-flag minimum, optional scale for the safe-integer bound). User-facing messages are unchanged. * fix(analyze): make --memory-budget set the real V8 heap through the respawn The budget now drives ensureHeap's existing respawn instead of a parse-phase override, so the #2649 preflight, mid-loop abort, remedy text, and GC pacing all see one heap limit. The respawn sizes old space plus three semi-spaces to equal the budget, the child resolves as already at the budget (no second respawn), and GITNEXUS_HEAP_LIMIT_SOURCE drives budget-aware OOM advice. Budget validation moves to the preAction hook so analyze and watch reject a bad value before any respawn. Removes the pool-shrink block and the memoryBudgetBytes plumbing through PipelineOptions and run-analyze. * docs(analyze): describe --memory-budget accurately and translate its help The help text claimed graceful worker-pool degradation, which no longer exists; it now says the flag sets the main-thread V8 heap and that parse workers keep their own caps. Wires the option through the help i18n map with en and zh-CN strings, and documents it in both READMEs and the out-of-memory troubleshooting section. * feat(analyze): add a pure rebuild-reason collector One collector per run holds keyed rebuild reasons, merges by key, flattens reasons stored by an interrupted rebuild into one recovery entry, validates stored reasons on read, and formats the single up-front summary plus one follow-up line for reasons added after the pipeline. * fix(analyze): route every forced rebuild through one reason collector Every path that forces a full rebuild (the nine meta gates, --force, --skills, --no-parse-cache, --drop-embeddings, --repair-fts retention, Spring Actuator, AsyncAPI, shared-store graph gaps, dirty-flag recovery, the post-pipeline capability gate, and the #2409 escalation) now adds a keyed reason to one collector. The rebuild decision is applied from the collector at fixed checkpoints, one summary prints right before the pipeline, and late reasons print one follow-up line. The escalation stays non-forcing. runFullAnalysis returns the collected keys, which replaces the runner-identity source-regex test with a behavior test. Removes the separate needsFullRebuild field and its announcement, and stops folding --skills and --no-parse-cache into --force. * fix(analyze): persist rebuild reasons on the existing crash marker Every incrementalInProgress writer (the full-rebuild stamp before the wipe, the incremental pre-write, saveIncrementalDirtyState including the #2409 escalation, and buildFtsDirtyStamp) now carries the collected reasons into the active slot's metaDir, so an interrupted rebuild explains itself on the next run through one merged recovery entry. A successful run still clears the marker and its reasons; the FTS-park recovery clears them without forcing. * test(analyze): cover every rebuild-reason key through runFullAnalysis Add a coverage table that the typechecker keeps complete: every RebuildReasonKey maps to a test file that drives it through runFullAnalysis and asserts the returned key. Adds the missing graph-write-collapse and drop-embeddings drivers, asserts the key in the existing pdg-mode, spring-vendor-prefixes, cjk-segmentation, and embedding-dims tests, and removes plan-local IDs from test names and comments. * fix(review): apply review findings - A --max-old-space-size pin equal to --memory-budget no longer counts as the exact budget heap (V8 adds the young generation on top); only the budget-respawned child skips the respawn, so the limit really equals the budget. - Snapshot the analyze env before ensureHeap and restore GITNEXUS_HEAP_LIMIT_SOURCE, so a kept process does not leak its heap source into a later programmatic analyzeCommand call. - --skills and --no-parse-cache keep the forced storage requirements they had before force stopped being folded from them. - Merge the duplicated follow-up announcement into one helper and fix a stale --drop-embeddings comment. - The rebuild-reason coverage table no longer greps driver files for the key string; add tests for a programmatic invalid budget and the multi-cause interrupted-rebuild text. * fix(review): don't announce the escalated write as a full rebuild The #2409 escalation is a non-forcing reason, but its follow-up line used the 'Full rebuild also required' lead. A follow-up that carries only non-forcing reasons now leads with 'Write plan changed'. * docs(analyze): document GITNEXUS_HEAP_LIMIT_SOURCE in the env table CONTRIBUTING requires every new GITNEXUS_* variable to have a row; this one is internal (set by analyze itself) and exists so OOM advice points at --memory-budget. * fix(review): address GitNexus review threads on #3386 - heapPressureRemedy measures pressure against the real auto-sized cap (heapCapMbFor) instead of a flat 0.75 x RAM, and no longer tells a GITNEXUS_MEMORY=off run with no pin to drop a pin that does not exist. - toStored() persists the interrupted rebuild's reasons first, as documented. - ensureHeap's doc names which paths leave GITNEXUS_HEAP_LIMIT_SOURCE unset. - The heap-respawn suite restores the caller's GITNEXUS_MEMORY. - The non-forcing follow-up test rejects any 'full rebuild' wording. * fix(review): require both budget flags and check key coverage at runtime - A budget-respawned child is recognized only when the inherited heap-source marker comes with both the budget's old-space and semi-space flags; the marker alone is an inherited env var, not proof. The old-space parser is generalized to any V8 size flag instead of copying its regex. - REBUILD_REASON_KEYS is exported and RebuildReasonKey derives from it, so the coverage table is checked at runtime (CI does not type-check test files). * fix(review): don't claim a full rebuild in a non-forcing summary formatSummary and formatFollowUp now share one leadFor helper, so a block of only non-forcing reasons reads 'Write plan changed' in both. --------- Co-authored-by: ChunxueLi <mecoloud@users.noreply.gitee.com> Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> |
||
|
|
ad5c7364e1
|
feat(storage): share one index store across linked worktrees and sibling clones (#3374)
* feat(storage): resolve the shared sibling-store identity and layout (#3352) Linked worktrees of one repository resolve to one store under GITNEXUS_HOME/stores/<key>, keyed by the canonical git common dir. The resolver reads the .git entry directly, so hot paths spawn no git. Slot naming moves to a leaf module so storage-resolver and shared-store do not import each other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): resolve a shared checkout's graph and existing store slot (#3352) getStoragePaths reads a flat slot's recorded graphPath only for checkout slots under the stores directory; other paths keep <storagePath>/lbug with no I/O. A recorded path outside the store's commit graphs is ignored. resolveStoragePath falls back to an existing store slot for an unregistered checkout, so reads never move to an empty slot. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): share one immutable commit graph across clean worktrees (#3352) A linked worktree writes its own slot in the shared store. A new slot is seeded with a pointer to the nearest commit graph, so a clean checkout at an indexed commit takes the up-to-date path and writes no graph. A checkout with local changes gets a copy-on-write private graph before its first write. After a successful run, a clean checkout at HEAD publishes its graph into commits/ under a store lock, or drops it when that commit graph already exists. Commit graphs are never written after publish. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): seed a worktree's graph from the nearest index (#3352) A new shared slot is seeded from the store's commit graph nearest to HEAD, else from this checkout's or the main checkout's repository-local index (copied under that index's lock, source left in place). The run that follows is up to date or incremental instead of a full build. If a pointed-at shared graph has been removed, analyze falls back to a full build instead of an incremental update over a missing baseline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): keep one parse cache per shared store (#3352) Linked worktrees read and write the parse cache and durable ParsedFile store under the store's caches/ directory. Before pruning, a run folds in the chunk keys recorded by every member slot and commit graph, and the fold, prune and save run under a store-wide cache lock so one member never evicts another's live chunks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(clean): remove only what no shared-store member references (#3352) Deleting a shared checkout slot (clean, clean --all, remove, and the server delete route) recounts references under the store's publish lock and deletes commit graphs no member points at, then the store itself once empty. A graph that cannot be deleted (open on Windows) is reported and kept for the next pass. clean --gc also drops member slots whose worktree is gone or no longer resolves to the store. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(analyze): let a clone opt in to a shared store with --share-with (#3352) An independent clone joins a linked worktree's shared store only with analyze --share-with <repo>, and only when its normalized origin URL matches that member's (credentials stripped, as #2054 compares). The registry remembers the choice. --no-share moves an opted-in clone back to its own .gitnexus and reclaims its old slot; linked worktrees always share and are pointed at GITNEXUS_SHARED_STORE=off instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): count a shared checkout's commit graph as its code index (#3352) A clean shared checkout reads a commit graph and owns no graph file, so the code-index presence check made status report it unindexed and registry validation skip it. The check now follows the slot's validated graphPath. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(storage): adopt existing worktree indexes and report shared-store state (#3352) A shared checkout gets <repo>/.gitnexus/store.json pointing at its store slot; resolution follows it only when it names that checkout's own slot. An existing local index seeds the slot and is left in place. status (text and --json) and doctor report the store, whether the graph is shared or private, and any leftover local index, which clean --local-index removes while keeping the pointer. With GITNEXUS_SHARED_STORE=off a previously shared checkout indexes into its own .gitnexus again and never writes a commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(mcp): read the graph a shared checkout points at (#3352) MCP, the HTTP API, group sync, augmentation, and the Claude hook (all three byte-identical copies) resolve a flat slot's graph through resolveGraphPath instead of joining 'lbug' onto the storage path, so checkouts at one commit share one open database. The embeddings writers (embeddings sync and the server embed job) take a private copy first and never write an immutable commit graph. The post-analyze settle probe accepts fresh metadata that points at an existing commit graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: describe the shared worktree index store (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(storage): reuse helpers across the shared-store code (#3352) One graph-clone helper replaces two copy-then-rename blocks; clean and status reuse formatSlotSize; leaving a store reuses removeSharedStorePointer; withStoreLock is imported from its own module instead of a re-export. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review findings in the shared store (#3352) - Never publish a graph whose build saw dirty files, and never trust a local-index seed on the up-to-date path; it may hold reverted edits. - Record slot pointers and reclaim under the publish lock, and reclaim right after each publish, so a commit graph is never deleted between publish and pointer save and superseded graphs don't pile up. - clean --gc decides membership from the registry, so opted-in clones and a main checkout without worktrees are not dropped. - Leaving a store re-registers first, so an up-to-date run cannot leave the registry pointing at a deleted slot. - Drop pinned branch summaries when an entry moves into a store slot. - Keep run.cjs and the AGENTS.md runner path inside the checkout. - MCP handles follow the slot's current graph, and branch scoping reads the slot's own metadata. - Cache resolveGraphPath by metadata file identity; look up opted-in entries with canonical registry paths. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): localize help for the shared-store flags (#3352) --help renders option text from the i18n catalog, so --share-with, --no-share, clean --gc and clean --local-index need keys in both locales, not just index.ts literals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): lock the embed-job graph copy and skip it on forced rebuilds (#3352) The server embed job copies a shared checkout's graph under the slot's index lock, so a CLI analyze in another process cannot interleave. A forced rebuild with no embeddings to carry over drops the pointer without copying a graph it would discard unread. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(agents): bump AGENTS.md version for the shared store notes (#3352) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep shared-store file reads inside checked paths (#3374) Address CodeQL js/path-injection and js/file-system-race on shared-store.ts: every filesystem read rebuilds its path under a fixed parent with an inline path.relative barrier, the .git probe is one read (EISDIR marks a directory) instead of stat-then-read, and the graph pointer cache stats and reads through one file descriptor. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review feedback on the shared store (#3374) - Hold the slot index lock for the whole server embedding job, not just the graph copy. - --no-share re-registers and deletes the old slot under that slot's index lock, so a running shared analyze cannot re-register it. - clean --gc previews without --force, like every other destructive arm. - status reports a pinned branch index as private. - Slot names prefix Windows device names that carry an extension. - A slot or store named ..<name> is a legal direct child. - Shared-store suites run in the serialized lbug-db vitest project and clear an inherited GITNEXUS_SHARED_STORE. - Doc and test accuracy fixes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(storage): pin shared-store behavior for worktrees of a bare repository (#3374) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address second review round on the shared store (#3374) - Reject --no-share in a linked worktree before taking any lock or indexing anything. - clean --all and remove also delete each shared checkout's pointer. - Delete an empty store only while also holding its cache lock, and re-check emptiness under it. - A store pointer is trusted only when the slot's own metadata names this checkout; the editable pointer file just says where to look. - Device names with an extension are prefixed on Windows only, so POSIX slot names stay stable. - Test fixtures use the gitnexus-test- prefix the stale-sidecar sweep recognizes; help text and the private-graph label are accurate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address third review round on the shared store (#3374) - clean --gc never collects a slot whose index lock is held: an analyze holds it until it registers the checkout, so a seeded but not yet registered slot is busy, not orphaned. - Reclaim compares absolute paths, so a relative GITNEXUS_HOME does not make live commit graphs look unreferenced. - graphPath is followed only into a published <commit>-<featureKey> dir, never .publish-* staging (TS and all three hook copies). - clean --gc fails on an unreadable stores root instead of reporting nothing to collect. - --no-share help states it is for opted-in clones. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): land the English --no-share help and private-graph label (#3374) These two strings were meant for |