Commit graph

3 commits

Author SHA1 Message Date
Felipe
c2ca132620
fix: web citation/code panel bugs and serve analyze/route hardening (#3348)
* chore: ignore local Vercel link artifacts

Keep .vercel and env files out of the repo after a local preview link.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): repair citation chips, code panel line math, stale agent state

Audit findings in the web client, each verified against the source:

- RightPanel: citation chips ([[path:10-20]], [[Class:Foo]]) were wired to a
  stub resolver that always returned null, so clicking any citation did
  nothing. Expose resolveFilePath from useAppState and use it.
- CodeReferencesPanel: graph startLine/endLine are 1-based but were treated
  as 0-based, so the highlighted range and scroll target were off by one
  line; AI citation cards always rendered "code not available" because the
  snippet loader was a stub. Fetch per-citation snippets via /api/file.
- useAppState: sendChatMessage read llmSettings.activeProvider outside its
  deps (stale provider capabilities after switching provider);
  initializeAgent trapped projectName at '' for callers without an override
  (system prompt labelled the codebase "project"); the embeddings 409 dedup
  matched a message the server never sends for same-repo jobs.
- tools.ts impact: for path targets every symbol defined in the file shares
  the filePath, so the disambiguation always picked the first row and could
  analyze an arbitrary symbol while reporting a file impact. Prefer the File
  node.
- useSigma: the layout timeout called stop() but never kill(), leaking one
  ForceAtlas2 Web Worker plus four graph listeners per completed layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): release repo lock on cancel, IPC-first worker cancel, route hardening

Audit findings in gitnexus/src, each verified against the source:

- analyze-launch: cancelJob marks the job failed before the worker exits,
  so the exit handler's terminal early-return skipped releaseLockOnce and
  the repo stayed locked ("Another job is already active") until restart.
  Release on terminal exit and when forkWorker bails on a terminal job.
- analyze-job: cancellation now sends { type: 'cancel' } over IPC first and
  signals only after a 15s grace. On Windows child.kill('SIGTERM') is a
  forceful termination, so leading with it could kill the worker inside a
  LadybugDB write. Mirrors core/auto-sync/analysis-worker-launch.
- api resolveRepo: a job that FAILED during the hold-queue wait fell through
  to the { __timedOut } sentinel ("taking longer than expected") instead of
  404; only /api/repo checked the sentinel, so graph/query/search/file/grep/
  embed/delete crashed on entry.storagePath with a 500 after a 5 minute hang.
  Return null on failed jobs and check the sentinel in every consumer.
- api processes/process/clusters/cluster: resolve ?repo= through the HTTP
  resolver (documented policy on resolveRegisteredRepoEntry) and pass the
  registered absolute path to the backend; add the standard rate limiter.
  The raw param previously reached the MCP resolver, which runs a
  cwd-relative realpathSync probe + registry refresh on a bare-name miss
  and accepts unambiguous partial names.
- api body handling: Express 5 leaves req.body undefined without a JSON
  content type, turning "Missing X" 400s into TypeError 500s; body-parser
  4xx errors (malformed JSON, over-limit) were also reported as 500.
- /api/file: the lexical path.relative check cannot see symlinks; re-check
  containment on realpath so a cloned repo containing evil -> /etc/passwd
  cannot read outside the root.
- repo-manager unregisterRepo: used the lenient reader, so a transient read
  error (EBUSY/EPERM racing another process's atomic rename) turned into
  writing [] and deregistering every repo. Use the strict-if-present reader.
- clean --branch: compared registry paths with raw path.resolve instead of
  the canonical registryPathEquals used everywhere else (macOS /private/var,
  Windows short names / drive-letter case) and reported indexed branches as
  not indexed.

Tests: cancelJob IPC-before-signal contract; /api/file symlink escape (403)
and in-repo symlink (200), skipped where the host cannot create symlinks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: keep example env files visible and ignore local Cursor config

.env* also hid gitnexus/.env.example and eval/.env.example. .vercel was already ignored. The web app's .cursor/ stays local, including its MCP file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add realtime execution ops dashboard for Vercel monitoring.

Expose /api/ops snapshots over the local serve process and a ?view=ops SPA panel so analyze/embed jobs can be watched live from the hosted web UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: address gitnexus-check review on cite/embed/resolve paths

Pass req into resolveRepo for process/cluster routes, tighten same-repo embed 409 handling, guard empty citation paths, fix snippet retry races, and drop the lone-File ambiguity fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): harden CORS, redact paths, and bound ops streams

Restrict Vercel CORS to exact production hosts, omit raw repo paths/URLs from the unauthenticated ops feed, rate-limit and cap SSE connections, and fix dashboard SSE/poll edge cases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): stabilize citation fetches and file impact matching

Retry cancelled snippet loads without duplicate in-flight reads, cap range-less citation downloads, and make impact file matching unique-suffix-aware with a synthetic File target when LIMIT drops the File node.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web/ops): close gitnexus-check review threads on SSE and redaction

Cap citation reads, reconnect ops on applied server URL, skip SSE onError after abort, and strip URL query/fragment from public repoName.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): abort SSE on poll fallback and harden repoName parsing

Prevent dual SSE+poll after a failed safety snapshot, skip overlapping poll ticks, ignore aborted streamSSE onError, and basename Windows drive-letter URLs so ops never leaks path segments.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops/web): close gitnexus-check threads on SSE budget and credential leak

Keep finite SSE retries across short 200s, strip backend URL userinfo before ?server=, and redact progress messages on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Restore 0-based GraphNode line math, redact public job poll/error fields, fix omit-?repo= 400, hold the analyze lock across cancel-during-settle, and restore the Vercel shared compile.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): contain leftover-slot reclaim to the slot and keep --stale status honest

Preview and force now share one branches/ containment rule, nested junctions cannot walk a sibling index, and a mid-loop git failure no longer claims leftovers were not deleted after a successful rm.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(cli): share leftover-slot helpers without changing reclaim behavior

Pull the rolling I/O pool and owned-cwd storage lookup into one place so clean --stale/--branch and leftover listing stop restating the same ownership and concurrency paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

Keep the omitted-repo snapshot instead of re-listing, stop citation and ops races, redact full public repo URLs, and restore fake timers in teardown.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Store graph node citation lines as 0-based offsets and correct the default-port origin comment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Redact public SSE progress text, stop citation append retries from
cancelling in-flight reads, and keep the ops dashboard from showing a
stale snapshot or clearing a failed Connect.

Note: pre-existing failure in incremental-index-extension-dml-gate and other lbug/env unit tests not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Prune citation snippets when AI refs are cleared, and poll /api/ops at 2s so the fallback stays under the 60/min limit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): apply prettier class order for format check (#3348)

CI quality/format runs root-only npm ci, so prettier-plugin-tailwindcss
sorts scrollbar-thin without the web Tailwind catalog.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a unique suffix match for graph-backed citation paths so
  ambiguous names like index.ts no longer open the first graph file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a path-component boundary so unique citation suffixes cannot match filename substrings like myindex.ts

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact filesystem paths in ops text

Unauthenticated /api/ops and poll replay worker errors. URLs were
scrubbed but home-directory and Windows paths still leaked.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): omit repoPath from public SSE frames

/api/ops lists job ids, so the unauthenticated progress stream
must not replay the analyzed filesystem path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): treat embed lock 409 as a busy error

Analyze and embed share the same lock string. Mapping that 409 to
embedding hid an in-flight analyze as a successful embed start.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): bound impact File path suffix matches

Unbounded endsWith let lib/foo.ts select src/mylib/foo.ts. Require
an exact path or a unique /suffix, matching resolveUniqueIndexedPath.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): keep canceled analyze slot until exit

Marking failed before the worker exited let a second POST start
cloneOrPull against a LadybugDB file still being written.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): skip Windows SIGTERM on job dispose

child.kill('SIGTERM') is TerminateProcess there. Ask over IPC first
and leave the 15s grace timer to SIGKILL.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): let process routes skip the analyze hold

GET /api/processes and /api/clusters always waited up to 300s.
?awaitAnalysis=false fails fast; default still waits like /api/repo.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): cover public ops and analyze SSE user flows

Lock the unauthenticated dashboard and analyze complete/fail/cancel paths so a leaked repoPath, token, or home path cannot ship unnoticed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Keep caller cancel reasons over the worker's generic IPC, skip publish while cancel is pending, and redact scp-style remotes on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Release the analyze slot when a worker fails to spawn, and omit branch refs from the unauthenticated ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): do not reuse an analyze job that is pending cancel

A dying same-repo job still occupies the single slot; 202-reuse would
attach a new client to a cancel in flight.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): hold the analyze lock until the worker exits after cancel

Cancel error IPC used to drop the repo lock while the child was still
checkpointing. Abort settle immediately on pending cancel so the slot
is not held for a 60s disk poll.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact known repo paths with spaces in public ops text

Known repoPath/repoUrl literals are replaced first so a clone dir with
spaces cannot leak past the whitespace-bounded path regex.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): return the live job on analyze and embed DELETE

Hard-coding failed made clients retry immediately and 409 while the
child still occupied the slot. resolveRepo now returns not-found as
soon as that job fails instead of waiting out the hold timeout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): drive public analyze and ops through a live gitnexus serve

Spawn the real backend and observe requests instead of intercepting
them, so slot occupancy, redaction, and reconnect stay honest.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(web): reuse code-panel helpers and drop dead UI aliases

Citation fetches already had selectedNodeFileRange and snippetRepoKey;
the impact File suffix filter already handled exact paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

- Skip createJob reuse after cancel IPC is consumed while the child remains
- Hold the repo lock until exit when complete IPC races a pending cancel
- Scrub full remote URLs before known repoUrl prefixes in public ops text
- Reject unique impact File suffix matches from a truncated LIMIT 10 page
- Make live e2e helpers bound probes, clean up failed startups, and wait out the cancel slot

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): open live e2e pages on the Vite host CI actually bound

Absolute 127.0.0.1:5173 navigation refused on Actions because wait-on
and Vite use localhost (often ::1). Honor FRONTEND_URL when set, else
pick the first of localhost / 127.0.0.1 that answers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Hold the analyze lock until worker exit when cancel aborts settle, and assert GitLab failure chrome does not leak host or path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): keep live analyze e2e under the analyze rate limit

POST /api/analyze allows 10 requests per minute per IP. The slot-free
helper re-POSTed every 400ms while a cancelled worker was exiting, spent
that budget, and the lock test's hold request got 429 instead of 202.

- postAnalyze waits out a 429 using the RateLimit reset and retries
- slot polling backs off to 2s and leaves a small POST budget for callers
- the slot probe is a clone that fails before any worker fork, so the
  probe itself no longer holds the slot after reporting failed
- the lock test holds the slot with a real local analyze
- token and GitLab tests wait for a free slot before posting from the UI
- request fetches carry a timeout; teardown signals the serve process group
- an empty FRONTEND_URL falls back to the default base URL

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- ops view: a `?server=` link no longer auto-connects to another origin
  while a deploy token is held; it prefills and waits for Connect
- analyze completion: a local-path run reconnects by the path this client
  submitted, so duplicate basenames stay collision-safe without repoPath
  on the public SSE frame
- stale slot cleanup: revalidate each nested directory (lstat + realpath)
  right before readdir, so a mid-cleanup junction swap aborts instead of
  walking an outside tree; list phases run sequentially so the slot-I/O
  cap is global
- e2e: 429 backoff honours the caller deadline; the unreachable-backend
  ops test navigates through the resolved frontend URL
- drop the unused `repo` member from the clean integration helper

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(server): reconnect after analyze by an opaque repo id

Public job views and the SSE terminal frame no longer carry repoPath, and
repoName is not unique, so a post-analyze reconnect by name could load a
same-named sibling. The server now issues `repoId`: an HMAC of the
canonical registry path under a per-process random key. It is set on
complete jobs (ops view, analyze poll, SSE terminal frame) and matches
the new `id` on `GET /api/repos` entries. The web client resolves it to
the exact entry path on completion; unknown ids fall back to the name.
This covers URL clones and folder uploads, and replaces the local-path
only fallback.

With reconnect off the label, public `repoName` for a branch-pinned URL
clone is the repository name, not the `<repo>__<branch slug>` registry
name, so the requested branch stays off the unauthenticated feed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- stale slot cleanup: a descendant that vanishes before its unlink is
  treated as removed instead of aborting the reclaim
- e2e teardown: escalate to SIGKILL on the process group when the live
  backend ignores SIGTERM for 5s
- ops view: drop the dead initial value CodeQL flagged

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- public redaction: a known repoPath now also consumes its descendant
  tail, so `<repoPath>/src/secret.ts` becomes `[path]` instead of
  `[path]/src/secret.ts`; a same-prefix sibling is left to the path scrub
- e2e: the cancel test waits for the analyze slot the previous failed
  local-path job still holds; `fetchOps` carries the request timeout
- docs: SSE terminal payload and RepoAnalyzer `onComplete` describe
  `repoId` resolution

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- RepoAnalyzer: drop a completion that resolves after unmount, so a slow
  /api/repos lookup cannot switch repos after the sheet was dismissed
- e2e: select the local-path input by test id (the placeholder differs on
  Windows); the slot probe's job wait honours the caller's deadline

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(server): assert the public SSE terminal frame in analyze-api

The #2790 terminality tests still expected `repoPath` on the terminal
frame. This PR replaced it with the opaque `repoId`, so assert that shape
and that the analyzed path never appears in the stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(web): wait for the analyze slot between duplicate-repo setup runs

The server now keeps the single analyze slot until the worker exits, even
after its job reports complete. repo-path-identity posted the second
duplicate's analyze immediately and got 409 in CI. Use the shared
slot-aware POST, which also waits out a 429.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 13:27:22 +01:00
Gergő Magyar
561f913a32
fix(embeddings): retry unparseable 200 responses and survive partial embedding failures (#2790) (#2795)
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Skill copy sync / shipped skills drift guard (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
* fix(embeddings): retry unparseable 200 responses and survive partial embedding failures (#2790)

A long-running embedding job against an OpenAI-compatible endpoint could lose
hours of work to a single transient glitch, then refuse to recover on the next
run. Four defects compounded:

1. An HTTP 200 carrying a truncated or non-JSON body was never retried.
   `classifyOutcome` treats any 2xx as success, and the `resp.json()` parse ran
   after `resilientFetch` had already returned, so the parse failure surfaced as
   a terminal error. Measured: a 503 got 3 attempts, a garbage 200 got 1.

   The parse and the response-shape check now run inside the `fetchImpl`
   callback, so a bad body is classified as a retryable failure and gets the
   same backoff as a 5xx. This also stops a garbage 200 from calling the circuit
   breaker's `recordSuccess()`, which previously erased accumulated failures and
   meant an endpoint alternating 5xx and garbage-200 could never trip it.

2. One failed `embedBatch` sub-batch aborted the entire pipeline. Failures are
   now tolerated: the sub-batch's node ids are collected and all of their
   embedding rows are deleted, so those nodes hold zero rows and are re-embedded
   later. Deleting rather than keeping partial rows is deliberate — chunk arrays
   are flat over a 16-node batch and sliced by 8, so a node's chunks can straddle
   a sub-batch boundary, and surviving rows carry the current content hash. The
   hash maps collapse per-chunk rows last-row-wins, so a partially embedded node
   would read as fresh forever and never regenerate its missing chunks.

   A run that fails 5 sub-batches in a row still aborts, and rethrows the first
   error of the streak rather than the last: after 3 failures the circuit breaker
   opens, so later errors degrade into "circuit open, retry in 30s" while the
   first still names the real defect.

3. The Phase 5 `embeddingCount === 0` fail-fast could not tell "wrote nothing"
   from "could not ask" — the count query's catch was silent. The count is now
   tri-state and only a known zero after real work is fatal. A non-numeric count
   previously bypassed the gate entirely, because `Number()` returns NaN and
   `NaN === 0` is false, and then serialized as `embeddings: null`. An unverified
   count no longer certifies `capabilities.vectorSearch.status`.

4. `saveEmbeddingCheckpoint` wrote a completion-shaped meta: it advanced
   `lastCommit`, wrote the new `fileHashes` and cleared `incrementalInProgress`.
   The first checkpoint window fires before a single embedding exists, and on a
   full rebuild the graph is still in a staging database that a crash discards.
   The next run then diffed against the advanced hashes, saw no changes and
   preserved the old graph — the "skipping wipe" symptom in the report. It now
   re-reads meta and replaces only the checkpoint, matching what the server
   endpoint already did.

A partially failed run keeps its checkpoint with the failed ids in
`pendingNodeIds`, so the next plain `analyze` regenerates them through the
existing resume path. Clearing it would have been silent data loss: a plain run
derives `shouldGenerateEmbeddings: false` once embeddings exist, so the pipeline
would never have run again. The old crash-and-abort self-healed only by accident,
via the checkpoint its crash left behind. `gitnexus status` reports the index
incomplete until the nodes recover, and `--drop-embeddings` still abandons them.

`POST /api/embed` is the pipeline's other caller and was discarding the result,
reporting "Embeddings complete" for a partial run. It now persists the pending
ids and reports the run as failed with the underlying endpoint error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): abort a run whose sub-batch failure ratio is too high (#2790)

The consecutive-failure ceiling only catches a total outage, because any
successful sub-batch resets it. An endpoint under load shedding that alternates
success and failure never trips it, so the run walks the whole corpus, deletes
every failed node's rows and exits 0 having dropped a large fraction of the
index. The retained checkpoint made that visible in `gitnexus status`, but a run
that drops a quarter of the corpus should tell the operator to fix their
endpoint, not leave them to notice a status flag.

Adds a cumulative guard: abort once more than 25% of attempted sub-batches have
failed, evaluated as the run progresses and gated behind a floor of 20 attempted
sub-batches. The shape follows Resilience4j's circuit breaker (failure rate plus
a minimum-sample floor) because it is the only one of the surveyed designs that
answers the small-repo case — a three node repo can fail one sub-batch and never
accumulate enough sample for a ratio to mean anything. The rate sits below a live
traffic breaker's 50% because a batch indexer's job is to index the whole corpus
rather than serve degraded traffic, and above Hadoop's single-digit
`failures.maxpercent` because tolerating transient hiccups is the point of the
change this follows.

The guard reuses the existing break-then-cleanup path, so the failed batch's
DELETE still runs before the rethrow, and it wraps the retained first-error-of-
streak rather than inventing a new one, so the message names both the ratio and
the underlying endpoint failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(server): record the embedding count after /api/embed so the next analyze cannot wipe it

`POST /api/embed` generated embeddings and wrote them to the database but never
wrote `stats.embeddings` into meta.json. Its checkpoint writer replaced only
`embeddingCheckpoint`, and the finalize write folded in nothing else.

So a repo embedded purely through the server kept whatever count the last CLI
`analyze` stamped, which is 0 for a repo analyzed without embeddings. The next
CLI run read `existingEmbeddingCount = 0`, `deriveEmbeddingMode` returned
`shouldLoadCache: false`, and `gitnexus analyze --force` wiped the database with
no cache load. Every server generated embedding was silently destroyed, with no
warning — the user just lost semantic search.

The route now measures the live count with the same query the CLI uses and folds
it into both meta writes. The measurement is tri-state and deliberately never
falls back to 0: an unverified count is written as absent rather than as zero,
because a wrong-low value is exactly what arms the wipe. It is taken after
`flushWAL()` and inside `withLbugDb`, so it describes durable rows and the
connection is still open. A partial run records its honest count too, alongside
the retained checkpoint, so the next CLI run preserves the partial index instead
of discarding it.

Found while working #2790; not part of that issue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): retry short 200 bodies and stop laundering body-phase timeouts

Two gaps in the #2790 retry fix, both found by review.

A 200 carrying `{"data": []}` or fewer vectors than inputs passed the
in-`fetchImpl` shape check, because `every(isEmbeddingItem)` is vacuously true
for an empty array. `resilientFetch` then classified it `success` and called
`recordSuccess()`, erasing the outage signal, and the cardinality check in
`httpEmbed` threw terminally one attempt later. That is exactly the pair of
properties #2790 was filed about, still broken for this body shape — and worse
than before the fix, since the pipeline now tolerates the error by deleting
those nodes' rows instead of aborting loudly. The count check moves inside the
retried callback; the outer one stays as a backstop.

The `.json()` catch also swallowed every rejection, not just parse errors.
`AbortSignal.any([caller, timeout])` is wired to the body stream, so a stalled
body rejects with a DOMException — which, wrapped in a plain Error, defeated
`classifyOutcome`'s terminal-network test. Measured: the same TimeoutError got
3 attempts and "unparseable response" when raised during the body read, but 1
attempt and "timed out after 180000ms" when raised by fetch itself, and three
such sub-batches opened the process-global breaker that `recordNeutral()`
exists to protect. Abort-like DOMExceptions are now re-raised unchanged.

The dimension check stays outside the loop deliberately: it validates against
`config.dimensions ?? DEFAULT_DIMS`, not the request-dimensions argument, and a
width mismatch is a configuration error where retrying only triples latency and
books failures against a healthy endpoint.

Adds the negative assertion the review found missing: response body text must
never reach the user-facing error string.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): scale the sub-batch failure-ratio floor to the run

The cumulative guard needed 20 attempted sub-batches before a failure rate could
abort anything — roughly 160 chunks, or ~80 embeddable nodes at the default
subBatchSize of 8. A 50-node repo whose endpoint sheds every other sub-batch
fails half of them and still exits 0: the ratio guard is below its floor, and
every intervening success resets the consecutive ceiling.

The floor was a good choice for a first run over a small repo, where one failure
out of one sub-batch is 100% and means nothing. The defect is that every resume
run has that shape by construction — its node set is only the pending ids — so
the guard was structurally off in the one run whose entire purpose is retrying
against the endpoint that already failed.

The floor is now sized to the run: clamp(ceil(totalNodes / 16), 5, 20). The
lower bound keeps the case the flat floor protected; the upper bound preserves
today's behavior above 320 nodes and avoids a proportional-only floor perversely
weakening the guard at scale, where a sixteenth of a 20k-node repo would be 1250
sub-batches of damage before a rate could fire. Resilience4j can use a constant
minimumNumberOfCalls because a breaker sits on an unbounded call stream; a batch
indexer has a finite budget, so a constant can exceed the whole run.

The ratio is still evaluated only inside the catch. That is already its local
maximum — both counters have just incremented — so sampling more often would
only ever observe lower ratios.

Also: a failing cleanup DELETE no longer swallows the abort, which was
discarding the retained first-error-of-the-streak that names the real endpoint
fault; `ceilingError` is renamed `abortError` since it carries the ratio abort
too; and three `{ error }` log keys become `{ err }` (#2114 — an arbitrary key
serializes to `{}`, losing message and stack).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): one tri-state embedding counter, and stop partial runs wedging later runs

The tri-state count doctrine this branch introduced was applied at two of its
three CLI sites, and the two implementations that were meant to mirror each
other had already drifted.

`measurePersistedEmbeddingCount` moves to `core/embedding-count.ts` — beside
`embedding-mode.ts`, with the same no-native-imports property, and outside
`core/embeddings/` so the lazy-embeddings convention (#2370) still holds. All
three call sites now share it.

  - The mid-run `onCheckpoint` counter ran the query bare. A throw there — DB
    busy, connection closed, read-only, the VECTOR DML lock (#2623) — rejected
    the callback out of `runEmbeddingPipeline` and killed the analyze before
    Phase 5 could apply the tri-state that exists for exactly this case. A
    non-numeric cell wrote `stats.embeddings: null` to disk mid-run.
  - Phase 5 used `?? 0` while the server used `?? Number.NaN`, under a comment
    asserting both measured the field the same way. `Number.isFinite(0)` is
    true, so a no-row answer became a *measured* zero and hard-failed a run
    whose embeddings had all persisted.
  - The unknown-count fallback read `existingMeta`, assigned once at run start,
    so it republished the pre-run figure over the fresher count the terminal
    checkpoint had already written. With a prior count of 0 that armed the wipe
    chain: hasExisting false, shouldLoadCache false, and the next --force
    discards live embeddings. It now re-reads the latest on-disk meta, and an
    unverifiable count retains a recovery marker instead of clearing it.

A completed-but-partial run also planted a landmine. Its checkpoint is stamped
with the run's embedding identity, so a later plain `gitnexus analyze` from a
hook, a CI job, or a shell without GITNEXUS_EMBEDDING_URL resolved provider
'local' and threw before any phase ran — after an exit-0 run, where previously
only a visible crash left that state. `--force` did not help: the resume gate
inspected only `--drop-embeddings`.

`RepoMeta.embeddingCheckpoint` gains `kind` to tell the two situations apart.
An 'interrupted' marker (or one with no kind, so markers already on disk keep
the stricter path) still fails closed — its nodes may be half-written, and
resuming under a foreign model would mix vector spaces. A 'partial' marker
names nodes the pipeline already deleted to zero rows, so nothing is at risk: an
identity mismatch drops the pending set with a warning and continues. `--force`
now discards a checkpoint, and `attempts` bounds the retry at
EMBEDDING_RESUME_MAX_ATTEMPTS (3, matching the HTTP embedder's and the WAL
driver's existing per-operation budgets) so a node the endpoint deterministically
rejects converges instead of keeping the repo incomplete forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(server): close the SSE stream on terminal job status, not a progress phase

A tolerated partial run reached SSE clients as a clean success — a regression in
this branch's own claim that /api/embed reports a partial run as failed.

The pipeline emits `phase:'ready'` unconditionally before returning, including
when it dropped nodes. The route mapped that to `'complete'`, and
`mountSSEProgress` treated a terminal-looking *progress phase* as terminal:
write the event, `res.end()`, `unsubscribe()`. The route's own
`updateJob({status:'failed'})` then fired into a stream with no listener, and
the web app had already shown "ready". Before this branch the pipeline threw,
which produced `phase:'error'` and did reach the client. Pollers on
GET /api/embed/:jobId were unaffected, so the two consumers disagreed.

Terminality is a property of the job, so the relay now asks the job. Remapping
`ready` alone would have left the trap armed: the `error -> 'failed'` mapping
has the identical shape and would emit `event: failed` with `error: undefined`
before the catch block fills the message in. `ready` is additionally remapped to
`finalizing` so a poller no longer sees `status:'analyzing'` next to
`progress.phase:'complete'`. The single-terminal-event property (#2264) is
preserved on both the clean and partial paths, and /api/analyze is unaffected —
its terminal progress phase is 'done', never 'complete'.

`AnalyzeJob` gains an optional `partial` payload so a client can tell a partial
run from a total failure without a new status member; it is absent on every
other job, so existing payloads stay byte-identical. Consuming it in
gitnexus-web is left to that app's owner — today it renders both as the same
red retry chip.

`resolveEmbedRunOutcome` moves to `embed-run-outcome.ts` and `mountSSEProgress`
to `sse-progress.ts`, both free of Express/LadybugDB/MCP imports, and the local
count copy is replaced by the shared `core/embedding-count.ts`. Reaching three
pure functions previously meant importing the whole server: measured at ~20s
against a 30s test timeout, with one observed timeout failure. That file is now
1.6s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document the partial embedding index and its recovery

A run can now finish exit 0 with a partial embedding index, which neither
operator doc described.

GUARDRAILS' "Embeddings vanished after analyze" Sign keys its trigger on
`stats.embeddings` being 0 and lists "the only ways to end up at zero". A
partial run stamps an honest non-zero count and sets `embeddingCheckpoint`, so
the operator's actual symptom is `incompleteReasons:
["embedding-checkpoint-pending"]` — a state that Sign cannot match. Adds a Sign
for it and drops the exhaustive framing from the existing one.

RUNBOOK gains the recovery path: a plain `gitnexus analyze` is correct and needs
no flag, because a retained checkpoint forces generation for the pending nodes
regardless of flags. Also corrects two stale claims — that `stats.embeddings` is
always freshly measured (it can carry forward when the count query cannot
answer, which is why `capabilities.vectorSearch.status` is the certified read),
and that later analyzes must always pass `--embeddings` or lose their vectors,
which contradicts Non-negotiable 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(embeddings): one owner for the checkpoint record and the abort predicate

Cleanup pass over the #2790 review fixes. No behavior change except where
noted; the two exceptions are both cases where the code was lying to the
operator or to the other half of itself.

The previous pass extracted `core/embedding-count.ts` because two hand-copied
bodies of "measure the embedding count" had drifted inside a single change. It
then created a second pair of hand-copied publishers — of
`RepoMeta.embeddingCheckpoint` — and those had drifted too: the CLI armed the
attempt counter only after clearing its identity gate, the server derived it
from the resumed marker alone. Only one of the two READERS implemented `kind`
at all, so a 'partial' marker written by `gitnexus analyze` and resumed through
POST /api/embed still hit the permanent wedge `kind` exists to remove.

`core/embedding-checkpoint.ts` now owns the record: `checkpointKind` (the one
home for absent-means-interrupted), the three minters, `nextAttemptCount`, and
`decideEmbeddingResume`, which both gates route through. Five mint sites and
two resume gates become one implementation each.

`resilient-fetch.ts` exports `isTerminalNetworkError` and `classifyOutcome`
calls it, replacing a caller-side copy of the same DOMException test whose
docstring promised it "mirrors classifyOutcome exactly" — an invariant enforced
by prose, where a divergence silently reverts body-phase timeouts to being
retried three times and charged to the shared breaker.

The ratio-guard floor now divides by the run's actual `subBatchSize` instead of
a constant 16 that assumed the default of 8. At `subBatchSize: 32` the old
formula demanded more sub-batches than the run contains, leaving the guard
structurally off — the exact failure the scaled floor was introduced to fix,
and sub-batch size is tuned mainly for the flaky endpoints it protects.

Two operator-facing corrections:

  - The count-recovery marker was stamped `kind: 'partial'` with an empty
    pending set, so `gitnexus status` reported "N node(s) lost their embeddings"
    where N is zero. It gets its own kind and its own incomplete reason.
  - `decideEmbeddingResume` initially keyed its skip-the-identity-gate branch on
    an empty pending set, assuming that meant the count-recovery marker. It does
    not: `onCheckpoint` mints an 'interrupted' marker with no pending nodes
    after every post-window save. That silently cleared an interrupted marker
    under a foreign provider instead of failing closed. Keyed on `kind` now,
    with a regression test.

Also: `isTerminalJobStatus` adopted at the seven sites that still hand-copied
it, including the one gating the single-terminal-event emit; `mountSSEProgress`
re-export dropped and `server-sse-payload.test.ts` repointed at the extracted
module, which takes it from 24.60s to 0.408s — the test that motivated the
extraction was still paying the cost it was meant to remove; the count-mismatch
message and the SSE test harness deduplicated; per-batch error strings made
lazy (~75k needless `new URL()` per large run); `retryable: true` dropped as a
field that can never be false; ~110 lines of restated rationale reduced to
pointers at their canonical home.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:43:59 +00:00
Parafee41
737a8cdb18
fix(web): use repo path identity in switcher (#2420)
* fix(web): use repo path identity in switcher

* keep repo URL project names stable

* fix server repo path resolution

* fix repo path miss resolution

* fix(server): guard clone-dir deletion with path ownership check

Deleting a registry entry derived its clone dir from the entry NAME with
no ownership check, so deleting a local repo that shares a display name
with a server-cloned sibling wiped the sibling's checkout. Gate the
removal on cloneDirBelongsToEntry (canonicalized path equality), the
same entry.path-driven rule the handler's step 2b already mandates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): fail closed on relative repo params and rate-limit GET /api/repo

Relative separator-containing ?repo= values (org/name, ./repo) were
canonicalized against the server CWD — an attacker-influenced
realpathSync probe on an un-rate-limited GET — before failing anyway.
Reject them immediately without touching the filesystem, drop the
redundant path.sep clause, document the resolver's two-tier contract,
and wire createRouteLimiter on GET /api/repo like its DELETE sibling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(server): lock repo resolver branches and register for Windows CI

Lock in the resolver's remaining branches: first-wins for ambiguous
bare names, Windows-shaped input as a fail-closed path claim, the
repos[0] default, and the case-insensitive name fallback. Register the
suite in cross-platform-tests.ts so windows-latest actually runs the
path-shape logic it exists to protect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): single repoIdentity helper with repoPath normalized end-to-end

The identity fallback chain was copy-pasted in Header and RepoLanding
while backend-client already owns BackendRepo and the repoPath
normalization. Export one repoIdentity helper, normalize fetchRepos
like fetchRepoInfo, and emit repoPath from GET /api/repos so the
scheme no longer silently relies on /api/repo.repoPath equalling
/api/repos.path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): persist and restore repo path identity in the URL

The URL persisted only ?project=<display name>, so refreshing after
switching to a duplicate-name repo silently restored the first
same-named sibling. Persist ?repo=<server-resolved path> alongside the
readable ?project= at both write sites, prefer it on restore (legacy
project-only URLs still work), keep failed path restores fail-visible
(no name fallback), and strip stale identity params when deleting the
active or last repo.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): analyze completion connects by path identity

RepoAnalyzer's completion callback passed the display name, so
analyzing a repo whose basename collides with an existing one
reconnected the first same-named sibling. The SSE terminal payload now
carries the job's repoPath (both emit sites), the analyzer passes that
identity to onComplete while the done screen keeps showing the display
name, and old servers without repoPath degrade to today's behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): scope code-reference file reads to the active repo identity

The code viewer passed the display name as the repo scope, so with
duplicate-name repos it rendered the wrong repo's file contents under
the right filename. Pass the active path identity (currentRepo) with
the display name as fallback, and collapse the two dead repo fields
that were already shadowed by the readFile spread.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): show display names instead of absolute paths in labels

The path-identity switch leaked raw filesystem paths into three
user-facing surfaces: the re-analyze progress label, the repo-switch
overlay, and the agent prompt's project name via loadGraphAnyway.
Resolve display names at render time (registry lookup, then basename
fallback) while state keeps holding the identity; loadGraphAnyway
passes the name explicitly because initializeAgent's empty-deps
closure would otherwise fall through to the literal 'project'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): stop initializeAgent from clobbering repo identity with display names

initializeAgent fell back to writing overrideProjectName (a display
name) into the repo identity, so any future name-only caller — the
pre-PR idiom — would silently kill the Active badge and re-admit the
duplicate-name ambiguity through the agent path. Only opts.repo may
write the identity now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(web): drop dead initializers flagged by CodeQL

pNameStr's and repoIdentity's initial values were never read: both are
assigned on the success path before any use and the catch returns
early. Bare declarations resolve CodeQL alerts 825/826.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(web): fix tailwind class order per root prettier plugin

The worktree pre-commit hook resolved prettier-plugin-tailwindcss
through symlinked node_modules and sorted scrollbar-thin differently
than CI's clean-room install. Re-formatted with the root lockfile
environment; no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(web): e2e coverage for every #2419 duplicate-name ambiguity

Provision two live repos with the same basename under different parents
via POST /api/analyze, then drive a real browser through each item of
the issue's "Actual behavior" list:

- duplicate rows render and the ACTIVE one is identifiable before and
  after switching (active-state must not compare repo.name)
- switching between duplicates swaps the loaded graph, verified by
  per-repo marker files (onSwitchRepo must not receive repo.name)
- re-analyze targets the clicked duplicate's exact path (POST body),
  tracks progress on that row only, and the completion reconnect
  requests that same path — never the same-named sibling
- delete requests target exactly the chosen duplicate's path; the
  sibling stays registered and loaded
- backend ?repo= resolution is path-first: landing selection loads the
  exact repo, ?repo= survives F5, and a stale path fails closed to the
  repo picker instead of retargeting the sibling

Adds four data-testids to Header (switcher trigger/row/reanalyze/
delete, rows expose data-active) so the spec has stable selectors, and
broadens the post-analyze reconnect retry in App to any BackendError:
the server may still be reinitializing when the SSE complete event
fires, and that surfaces as transient 5xx/binder errors, not only 404.

The re-analyze and delete tests deliberately assert identity at the
request level and tolerate two pre-existing server races that are
unrelated to the #2419 identity contract (freshly-analyzed DB briefly
unreadable after SSE complete; registry validate-prune clobbering a
concurrent unregister) — see the in-test comments.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(web): isolate repo-path-identity e2e onto a spec-owned backend

The spec is the only e2e file doing write operations (analyze,
re-analyze, delete). Running its force re-analysis against the shared
CI backend while parallel workers held connections took the whole
server down (run 29145679019: the jobId poll died with ECONNRESET and
every later test in every file failed to connect).

Spawn a dedicated `gitnexus serve` on port 4799 with an isolated
GITNEXUS_HOME in beforeAll instead: writes can no longer perturb the
other suites, a crash is contained to this spec (its output is captured
and printed, which CI otherwise loses), and the registry is hermetic by
construction — the previous leftover-purge and shared-registry cleanup
are gone. Every page is pointed at the spec backend through
useBackend's supported localStorage override, which covers both the
probe-driven landing flow and the ?server= auto-connect. Verified
self-sufficient (6/6 with no shared server running) and non-interfering
(full suite 39/39 with the shared server up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* stabilize repo path identity e2e

* fix(server): don't report analyze complete before the index is settled

The analyze worker reports `complete` over IPC before its on-disk
finalization (LadybugDB checkpoint, native handle release, metadata
write) is visible at the storage path — observed up to ~6.5s behind the
IPC message. The launcher's "reinitialize backend BEFORE marking
complete" ordering was meant to make the repo queryable by the time the
client sees the SSE complete event, but it never verified that: clients
reconnecting on that event read a database still being written. Locally
that surfaces as "Binder exception: Table CodeRelation does not exist"
or a silently empty graph, and the open can quarantine the in-flight
WAL; on slow CI runners the native layer racing the rewrite has killed
the whole server (signal exit, no output — run 29146867959).

Gate the complete transition on the index actually settling: LadybugDB
file and metadata both rewritten by THIS job (mtime >= job start — bare
existence is not enough, a re-analysis leaves the previous index in
place while it works) and no transient WAL/shadow/checkpoint sidecars
remaining. Bounded (60s) and proceed-on-timeout, so a job whose
analysis legitimately rewrites nothing cannot wedge. Also evict the
server's cached DB handle before reinitializing — same invalidation
DELETE /api/repo performs — so post-completion reads cannot be served
from a pre-rewrite handle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(web): assert re-analyze completion identity at the request level

The strict form (Ready + marker on the re-analyzed duplicate) still
trips a deeper pre-existing storage race that makes a freshly
re-analyzed database transiently unreadable to the reconnect even with
the settle gate in place — unrelated to the #2419 identity contract
this test covers. Keep the identity assertions (the reconnect targets
the exact duplicate's path and never the same-named sibling) and leave
a pointer to tighten once the storage race is fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): resolve the settle-gate path from the registry, not the request

CodeQL flagged the settle gate's stat/exists probes as js/path-injection:
the probed path derived from the user-provided analyze `path`. Resolve
it from the repo's registry entry instead — the user value is now only a
comparison key, and the probes run against the server-owned storagePath
record, which is also the authoritative path readers resolve through.
Re-resolved each poll round because the worker registers the repo as
part of the same finalization the gate is waiting out.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-11 11:43:47 +01:00