Commit graph

12 commits

Author SHA1 Message Date
Felipe
c2ca132620
fix: web citation/code panel bugs and serve analyze/route hardening (#3348)
* chore: ignore local Vercel link artifacts

Keep .vercel and env files out of the repo after a local preview link.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): repair citation chips, code panel line math, stale agent state

Audit findings in the web client, each verified against the source:

- RightPanel: citation chips ([[path:10-20]], [[Class:Foo]]) were wired to a
  stub resolver that always returned null, so clicking any citation did
  nothing. Expose resolveFilePath from useAppState and use it.
- CodeReferencesPanel: graph startLine/endLine are 1-based but were treated
  as 0-based, so the highlighted range and scroll target were off by one
  line; AI citation cards always rendered "code not available" because the
  snippet loader was a stub. Fetch per-citation snippets via /api/file.
- useAppState: sendChatMessage read llmSettings.activeProvider outside its
  deps (stale provider capabilities after switching provider);
  initializeAgent trapped projectName at '' for callers without an override
  (system prompt labelled the codebase "project"); the embeddings 409 dedup
  matched a message the server never sends for same-repo jobs.
- tools.ts impact: for path targets every symbol defined in the file shares
  the filePath, so the disambiguation always picked the first row and could
  analyze an arbitrary symbol while reporting a file impact. Prefer the File
  node.
- useSigma: the layout timeout called stop() but never kill(), leaking one
  ForceAtlas2 Web Worker plus four graph listeners per completed layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): release repo lock on cancel, IPC-first worker cancel, route hardening

Audit findings in gitnexus/src, each verified against the source:

- analyze-launch: cancelJob marks the job failed before the worker exits,
  so the exit handler's terminal early-return skipped releaseLockOnce and
  the repo stayed locked ("Another job is already active") until restart.
  Release on terminal exit and when forkWorker bails on a terminal job.
- analyze-job: cancellation now sends { type: 'cancel' } over IPC first and
  signals only after a 15s grace. On Windows child.kill('SIGTERM') is a
  forceful termination, so leading with it could kill the worker inside a
  LadybugDB write. Mirrors core/auto-sync/analysis-worker-launch.
- api resolveRepo: a job that FAILED during the hold-queue wait fell through
  to the { __timedOut } sentinel ("taking longer than expected") instead of
  404; only /api/repo checked the sentinel, so graph/query/search/file/grep/
  embed/delete crashed on entry.storagePath with a 500 after a 5 minute hang.
  Return null on failed jobs and check the sentinel in every consumer.
- api processes/process/clusters/cluster: resolve ?repo= through the HTTP
  resolver (documented policy on resolveRegisteredRepoEntry) and pass the
  registered absolute path to the backend; add the standard rate limiter.
  The raw param previously reached the MCP resolver, which runs a
  cwd-relative realpathSync probe + registry refresh on a bare-name miss
  and accepts unambiguous partial names.
- api body handling: Express 5 leaves req.body undefined without a JSON
  content type, turning "Missing X" 400s into TypeError 500s; body-parser
  4xx errors (malformed JSON, over-limit) were also reported as 500.
- /api/file: the lexical path.relative check cannot see symlinks; re-check
  containment on realpath so a cloned repo containing evil -> /etc/passwd
  cannot read outside the root.
- repo-manager unregisterRepo: used the lenient reader, so a transient read
  error (EBUSY/EPERM racing another process's atomic rename) turned into
  writing [] and deregistering every repo. Use the strict-if-present reader.
- clean --branch: compared registry paths with raw path.resolve instead of
  the canonical registryPathEquals used everywhere else (macOS /private/var,
  Windows short names / drive-letter case) and reported indexed branches as
  not indexed.

Tests: cancelJob IPC-before-signal contract; /api/file symlink escape (403)
and in-repo symlink (200), skipped where the host cannot create symlinks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: keep example env files visible and ignore local Cursor config

.env* also hid gitnexus/.env.example and eval/.env.example. .vercel was already ignored. The web app's .cursor/ stays local, including its MCP file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add realtime execution ops dashboard for Vercel monitoring.

Expose /api/ops snapshots over the local serve process and a ?view=ops SPA panel so analyze/embed jobs can be watched live from the hosted web UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: address gitnexus-check review on cite/embed/resolve paths

Pass req into resolveRepo for process/cluster routes, tighten same-repo embed 409 handling, guard empty citation paths, fix snippet retry races, and drop the lone-File ambiguity fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): harden CORS, redact paths, and bound ops streams

Restrict Vercel CORS to exact production hosts, omit raw repo paths/URLs from the unauthenticated ops feed, rate-limit and cap SSE connections, and fix dashboard SSE/poll edge cases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): stabilize citation fetches and file impact matching

Retry cancelled snippet loads without duplicate in-flight reads, cap range-less citation downloads, and make impact file matching unique-suffix-aware with a synthetic File target when LIMIT drops the File node.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web/ops): close gitnexus-check review threads on SSE and redaction

Cap citation reads, reconnect ops on applied server URL, skip SSE onError after abort, and strip URL query/fragment from public repoName.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): abort SSE on poll fallback and harden repoName parsing

Prevent dual SSE+poll after a failed safety snapshot, skip overlapping poll ticks, ignore aborted streamSSE onError, and basename Windows drive-letter URLs so ops never leaks path segments.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops/web): close gitnexus-check threads on SSE budget and credential leak

Keep finite SSE retries across short 200s, strip backend URL userinfo before ?server=, and redact progress messages on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Restore 0-based GraphNode line math, redact public job poll/error fields, fix omit-?repo= 400, hold the analyze lock across cancel-during-settle, and restore the Vercel shared compile.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): contain leftover-slot reclaim to the slot and keep --stale status honest

Preview and force now share one branches/ containment rule, nested junctions cannot walk a sibling index, and a mid-loop git failure no longer claims leftovers were not deleted after a successful rm.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(cli): share leftover-slot helpers without changing reclaim behavior

Pull the rolling I/O pool and owned-cwd storage lookup into one place so clean --stale/--branch and leftover listing stop restating the same ownership and concurrency paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

Keep the omitted-repo snapshot instead of re-listing, stop citation and ops races, redact full public repo URLs, and restore fake timers in teardown.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Store graph node citation lines as 0-based offsets and correct the default-port origin comment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Redact public SSE progress text, stop citation append retries from
cancelling in-flight reads, and keep the ops dashboard from showing a
stale snapshot or clearing a failed Connect.

Note: pre-existing failure in incremental-index-extension-dml-gate and other lbug/env unit tests not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Prune citation snippets when AI refs are cleared, and poll /api/ops at 2s so the fallback stays under the 60/min limit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): apply prettier class order for format check (#3348)

CI quality/format runs root-only npm ci, so prettier-plugin-tailwindcss
sorts scrollbar-thin without the web Tailwind catalog.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a unique suffix match for graph-backed citation paths so
  ambiguous names like index.ts no longer open the first graph file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a path-component boundary so unique citation suffixes cannot match filename substrings like myindex.ts

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact filesystem paths in ops text

Unauthenticated /api/ops and poll replay worker errors. URLs were
scrubbed but home-directory and Windows paths still leaked.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): omit repoPath from public SSE frames

/api/ops lists job ids, so the unauthenticated progress stream
must not replay the analyzed filesystem path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): treat embed lock 409 as a busy error

Analyze and embed share the same lock string. Mapping that 409 to
embedding hid an in-flight analyze as a successful embed start.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): bound impact File path suffix matches

Unbounded endsWith let lib/foo.ts select src/mylib/foo.ts. Require
an exact path or a unique /suffix, matching resolveUniqueIndexedPath.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): keep canceled analyze slot until exit

Marking failed before the worker exited let a second POST start
cloneOrPull against a LadybugDB file still being written.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): skip Windows SIGTERM on job dispose

child.kill('SIGTERM') is TerminateProcess there. Ask over IPC first
and leave the 15s grace timer to SIGKILL.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): let process routes skip the analyze hold

GET /api/processes and /api/clusters always waited up to 300s.
?awaitAnalysis=false fails fast; default still waits like /api/repo.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): cover public ops and analyze SSE user flows

Lock the unauthenticated dashboard and analyze complete/fail/cancel paths so a leaked repoPath, token, or home path cannot ship unnoticed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Keep caller cancel reasons over the worker's generic IPC, skip publish while cancel is pending, and redact scp-style remotes on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Release the analyze slot when a worker fails to spawn, and omit branch refs from the unauthenticated ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): do not reuse an analyze job that is pending cancel

A dying same-repo job still occupies the single slot; 202-reuse would
attach a new client to a cancel in flight.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): hold the analyze lock until the worker exits after cancel

Cancel error IPC used to drop the repo lock while the child was still
checkpointing. Abort settle immediately on pending cancel so the slot
is not held for a 60s disk poll.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact known repo paths with spaces in public ops text

Known repoPath/repoUrl literals are replaced first so a clone dir with
spaces cannot leak past the whitespace-bounded path regex.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): return the live job on analyze and embed DELETE

Hard-coding failed made clients retry immediately and 409 while the
child still occupied the slot. resolveRepo now returns not-found as
soon as that job fails instead of waiting out the hold timeout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): drive public analyze and ops through a live gitnexus serve

Spawn the real backend and observe requests instead of intercepting
them, so slot occupancy, redaction, and reconnect stay honest.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(web): reuse code-panel helpers and drop dead UI aliases

Citation fetches already had selectedNodeFileRange and snippetRepoKey;
the impact File suffix filter already handled exact paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

- Skip createJob reuse after cancel IPC is consumed while the child remains
- Hold the repo lock until exit when complete IPC races a pending cancel
- Scrub full remote URLs before known repoUrl prefixes in public ops text
- Reject unique impact File suffix matches from a truncated LIMIT 10 page
- Make live e2e helpers bound probes, clean up failed startups, and wait out the cancel slot

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): open live e2e pages on the Vite host CI actually bound

Absolute 127.0.0.1:5173 navigation refused on Actions because wait-on
and Vite use localhost (often ::1). Honor FRONTEND_URL when set, else
pick the first of localhost / 127.0.0.1 that answers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Hold the analyze lock until worker exit when cancel aborts settle, and assert GitLab failure chrome does not leak host or path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): keep live analyze e2e under the analyze rate limit

POST /api/analyze allows 10 requests per minute per IP. The slot-free
helper re-POSTed every 400ms while a cancelled worker was exiting, spent
that budget, and the lock test's hold request got 429 instead of 202.

- postAnalyze waits out a 429 using the RateLimit reset and retries
- slot polling backs off to 2s and leaves a small POST budget for callers
- the slot probe is a clone that fails before any worker fork, so the
  probe itself no longer holds the slot after reporting failed
- the lock test holds the slot with a real local analyze
- token and GitLab tests wait for a free slot before posting from the UI
- request fetches carry a timeout; teardown signals the serve process group
- an empty FRONTEND_URL falls back to the default base URL

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- ops view: a `?server=` link no longer auto-connects to another origin
  while a deploy token is held; it prefills and waits for Connect
- analyze completion: a local-path run reconnects by the path this client
  submitted, so duplicate basenames stay collision-safe without repoPath
  on the public SSE frame
- stale slot cleanup: revalidate each nested directory (lstat + realpath)
  right before readdir, so a mid-cleanup junction swap aborts instead of
  walking an outside tree; list phases run sequentially so the slot-I/O
  cap is global
- e2e: 429 backoff honours the caller deadline; the unreachable-backend
  ops test navigates through the resolved frontend URL
- drop the unused `repo` member from the clean integration helper

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(server): reconnect after analyze by an opaque repo id

Public job views and the SSE terminal frame no longer carry repoPath, and
repoName is not unique, so a post-analyze reconnect by name could load a
same-named sibling. The server now issues `repoId`: an HMAC of the
canonical registry path under a per-process random key. It is set on
complete jobs (ops view, analyze poll, SSE terminal frame) and matches
the new `id` on `GET /api/repos` entries. The web client resolves it to
the exact entry path on completion; unknown ids fall back to the name.
This covers URL clones and folder uploads, and replaces the local-path
only fallback.

With reconnect off the label, public `repoName` for a branch-pinned URL
clone is the repository name, not the `<repo>__<branch slug>` registry
name, so the requested branch stays off the unauthenticated feed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- stale slot cleanup: a descendant that vanishes before its unlink is
  treated as removed instead of aborting the reclaim
- e2e teardown: escalate to SIGKILL on the process group when the live
  backend ignores SIGTERM for 5s
- ops view: drop the dead initial value CodeQL flagged

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- public redaction: a known repoPath now also consumes its descendant
  tail, so `<repoPath>/src/secret.ts` becomes `[path]` instead of
  `[path]/src/secret.ts`; a same-prefix sibling is left to the path scrub
- e2e: the cancel test waits for the analyze slot the previous failed
  local-path job still holds; `fetchOps` carries the request timeout
- docs: SSE terminal payload and RepoAnalyzer `onComplete` describe
  `repoId` resolution

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- RepoAnalyzer: drop a completion that resolves after unmount, so a slow
  /api/repos lookup cannot switch repos after the sheet was dismissed
- e2e: select the local-path input by test id (the placeholder differs on
  Windows); the slot probe's job wait honours the caller's deadline

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(server): assert the public SSE terminal frame in analyze-api

The #2790 terminality tests still expected `repoPath` on the terminal
frame. This PR replaced it with the opaque `repoId`, so assert that shape
and that the analyzed path never appears in the stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(web): wait for the analyze slot between duplicate-repo setup runs

The server now keeps the single analyze slot until the worker exits, even
after its job reports complete. repo-path-identity posted the second
duplicate's analyze immediately and got 409 in CI. Use the shared
slot-aware POST, which also waits out a 429.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 13:27:22 +01:00
mengkaka
79543c8f83
feat(storage): add configurable index storage and content retention tiers (#3060)
* feat(storage): add configurable index storage and content retention tiers

Rebase #3060 onto current origin/main. Keep GITNEXUS_STORAGE_PATH,
GITNEXUS_STORAGE_ROOT, and GITNEXUS_CONTENT_RETENTION, and fold in
main's FTS skip, embed-session, and help-text updates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Keep legacy registry rows on the local storage fallback, resolve
symlinks before the destructive-path guard, and align hook lookup
with CLI branch slugs, branch-slot metadata, and longest-path match.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Only list swept upload directories after a successful removal so
callers cannot treat a permission or transient rm failure as gone.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Document that getStoragePath may consult registered storage while
this module still does not mutate the global registry.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(storage): close review findings for external indexes and retention

Re-inspect ownership under the analyze lock, fail-closed when the
registry file is missing, and keep skip-git hook discovery plus
retention fields on HTTP/MCP list surfaces. /api/file stays 410
unless contentRetention is full.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3060)

Treat lock-only index dirs as empty, honor HTTP --force storage policy, and prefer registered plus branch-aware slots in hooks and augment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3060)

Keep hook fallbacks inside the current worktree, compare foreign-local slots canonically, and make storage fixtures survive ownership validation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix macOS hook test expecting realpath'd registry paths.

resolveHookRepo returns the written registry path, not a filesystem realpath, so the assertion must match that.

* Address gitnexus-check warnings on hook install docs and slot tests.

The Cursor troubleshooting list omitted registry-query.cjs, and the writable-slot test only checked that isDirectory exists instead of that the path is a directory.

* Align the HTTP catalog source-scan with skippable resolveRepo validation.

resolveRepo lists fresh repos with validate: options.validateStorage !== false so DELETE can skip prune; the test still required a literal validate: true.

* Harden storage path sinks so CodeQL path-injection and ReDoS alerts clear.

Contain every filesystem probe inside the resolved storage slot with the inline path.relative idiom, reject filesystem-root slots, and trim slot basenames in linear time.

* Settle bridge stamps before writing so CI size/mtime matches stay stable.

LadybugDB can still flush into bridge.lbug after close+rename; persist whole-millisecond mtimes and wait for consecutive stats to agree so a freshly written pair matches.

* Type the settled bridge stat as fs.Stats so tsc does not see bigint.

Awaited<ReturnType<typeof fsp.stat>> collapsed the bigint overload and broke prepare/typecheck on CI.

* Keep the bridge mtime stamp exact so same-size swaps still fail the pair check.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Wrap the bridge stamp predicate so prettier --check stays green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Require a quiet interval before stamping a settled bridge file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Reuse shared storage and settle helpers instead of local copies.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-12 20:31:55 +00:00
Parafee41
737a8cdb18
fix(web): use repo path identity in switcher (#2420)
* fix(web): use repo path identity in switcher

* keep repo URL project names stable

* fix server repo path resolution

* fix repo path miss resolution

* fix(server): guard clone-dir deletion with path ownership check

Deleting a registry entry derived its clone dir from the entry NAME with
no ownership check, so deleting a local repo that shares a display name
with a server-cloned sibling wiped the sibling's checkout. Gate the
removal on cloneDirBelongsToEntry (canonicalized path equality), the
same entry.path-driven rule the handler's step 2b already mandates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): fail closed on relative repo params and rate-limit GET /api/repo

Relative separator-containing ?repo= values (org/name, ./repo) were
canonicalized against the server CWD — an attacker-influenced
realpathSync probe on an un-rate-limited GET — before failing anyway.
Reject them immediately without touching the filesystem, drop the
redundant path.sep clause, document the resolver's two-tier contract,
and wire createRouteLimiter on GET /api/repo like its DELETE sibling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(server): lock repo resolver branches and register for Windows CI

Lock in the resolver's remaining branches: first-wins for ambiguous
bare names, Windows-shaped input as a fail-closed path claim, the
repos[0] default, and the case-insensitive name fallback. Register the
suite in cross-platform-tests.ts so windows-latest actually runs the
path-shape logic it exists to protect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): single repoIdentity helper with repoPath normalized end-to-end

The identity fallback chain was copy-pasted in Header and RepoLanding
while backend-client already owns BackendRepo and the repoPath
normalization. Export one repoIdentity helper, normalize fetchRepos
like fetchRepoInfo, and emit repoPath from GET /api/repos so the
scheme no longer silently relies on /api/repo.repoPath equalling
/api/repos.path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): persist and restore repo path identity in the URL

The URL persisted only ?project=<display name>, so refreshing after
switching to a duplicate-name repo silently restored the first
same-named sibling. Persist ?repo=<server-resolved path> alongside the
readable ?project= at both write sites, prefer it on restore (legacy
project-only URLs still work), keep failed path restores fail-visible
(no name fallback), and strip stale identity params when deleting the
active or last repo.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): analyze completion connects by path identity

RepoAnalyzer's completion callback passed the display name, so
analyzing a repo whose basename collides with an existing one
reconnected the first same-named sibling. The SSE terminal payload now
carries the job's repoPath (both emit sites), the analyzer passes that
identity to onComplete while the done screen keeps showing the display
name, and old servers without repoPath degrade to today's behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): scope code-reference file reads to the active repo identity

The code viewer passed the display name as the repo scope, so with
duplicate-name repos it rendered the wrong repo's file contents under
the right filename. Pass the active path identity (currentRepo) with
the display name as fallback, and collapse the two dead repo fields
that were already shadowed by the readFile spread.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): show display names instead of absolute paths in labels

The path-identity switch leaked raw filesystem paths into three
user-facing surfaces: the re-analyze progress label, the repo-switch
overlay, and the agent prompt's project name via loadGraphAnyway.
Resolve display names at render time (registry lookup, then basename
fallback) while state keeps holding the identity; loadGraphAnyway
passes the name explicitly because initializeAgent's empty-deps
closure would otherwise fall through to the literal 'project'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): stop initializeAgent from clobbering repo identity with display names

initializeAgent fell back to writing overrideProjectName (a display
name) into the repo identity, so any future name-only caller — the
pre-PR idiom — would silently kill the Active badge and re-admit the
duplicate-name ambiguity through the agent path. Only opts.repo may
write the identity now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(web): drop dead initializers flagged by CodeQL

pNameStr's and repoIdentity's initial values were never read: both are
assigned on the success path before any use and the catch returns
early. Bare declarations resolve CodeQL alerts 825/826.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(web): fix tailwind class order per root prettier plugin

The worktree pre-commit hook resolved prettier-plugin-tailwindcss
through symlinked node_modules and sorted scrollbar-thin differently
than CI's clean-room install. Re-formatted with the root lockfile
environment; no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(web): e2e coverage for every #2419 duplicate-name ambiguity

Provision two live repos with the same basename under different parents
via POST /api/analyze, then drive a real browser through each item of
the issue's "Actual behavior" list:

- duplicate rows render and the ACTIVE one is identifiable before and
  after switching (active-state must not compare repo.name)
- switching between duplicates swaps the loaded graph, verified by
  per-repo marker files (onSwitchRepo must not receive repo.name)
- re-analyze targets the clicked duplicate's exact path (POST body),
  tracks progress on that row only, and the completion reconnect
  requests that same path — never the same-named sibling
- delete requests target exactly the chosen duplicate's path; the
  sibling stays registered and loaded
- backend ?repo= resolution is path-first: landing selection loads the
  exact repo, ?repo= survives F5, and a stale path fails closed to the
  repo picker instead of retargeting the sibling

Adds four data-testids to Header (switcher trigger/row/reanalyze/
delete, rows expose data-active) so the spec has stable selectors, and
broadens the post-analyze reconnect retry in App to any BackendError:
the server may still be reinitializing when the SSE complete event
fires, and that surfaces as transient 5xx/binder errors, not only 404.

The re-analyze and delete tests deliberately assert identity at the
request level and tolerate two pre-existing server races that are
unrelated to the #2419 identity contract (freshly-analyzed DB briefly
unreadable after SSE complete; registry validate-prune clobbering a
concurrent unregister) — see the in-test comments.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(web): isolate repo-path-identity e2e onto a spec-owned backend

The spec is the only e2e file doing write operations (analyze,
re-analyze, delete). Running its force re-analysis against the shared
CI backend while parallel workers held connections took the whole
server down (run 29145679019: the jobId poll died with ECONNRESET and
every later test in every file failed to connect).

Spawn a dedicated `gitnexus serve` on port 4799 with an isolated
GITNEXUS_HOME in beforeAll instead: writes can no longer perturb the
other suites, a crash is contained to this spec (its output is captured
and printed, which CI otherwise loses), and the registry is hermetic by
construction — the previous leftover-purge and shared-registry cleanup
are gone. Every page is pointed at the spec backend through
useBackend's supported localStorage override, which covers both the
probe-driven landing flow and the ?server= auto-connect. Verified
self-sufficient (6/6 with no shared server running) and non-interfering
(full suite 39/39 with the shared server up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* stabilize repo path identity e2e

* fix(server): don't report analyze complete before the index is settled

The analyze worker reports `complete` over IPC before its on-disk
finalization (LadybugDB checkpoint, native handle release, metadata
write) is visible at the storage path — observed up to ~6.5s behind the
IPC message. The launcher's "reinitialize backend BEFORE marking
complete" ordering was meant to make the repo queryable by the time the
client sees the SSE complete event, but it never verified that: clients
reconnecting on that event read a database still being written. Locally
that surfaces as "Binder exception: Table CodeRelation does not exist"
or a silently empty graph, and the open can quarantine the in-flight
WAL; on slow CI runners the native layer racing the rewrite has killed
the whole server (signal exit, no output — run 29146867959).

Gate the complete transition on the index actually settling: LadybugDB
file and metadata both rewritten by THIS job (mtime >= job start — bare
existence is not enough, a re-analysis leaves the previous index in
place while it works) and no transient WAL/shadow/checkpoint sidecars
remaining. Bounded (60s) and proceed-on-timeout, so a job whose
analysis legitimately rewrites nothing cannot wedge. Also evict the
server's cached DB handle before reinitializing — same invalidation
DELETE /api/repo performs — so post-completion reads cannot be served
from a pre-rewrite handle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(web): assert re-analyze completion identity at the request level

The strict form (Ready + marker on the re-analyzed duplicate) still
trips a deeper pre-existing storage race that makes a freshly
re-analyzed database transiently unreadable to the reconnect even with
the settle gate in place — unrelated to the #2419 identity contract
this test covers. Keep the identity assertions (the reconnect targets
the exact duplicate's path and never the same-named sibling) and leave
a pointer to tighten once the storage race is fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): resolve the settle-gate path from the registry, not the request

CodeQL flagged the settle gate's stat/exists probes as js/path-injection:
the probed path derived from the user-provided analyze `path`. Resolve
it from the repo's registry entry instead — the user value is now only a
comparison key, and the probes run against the server-owned storagePath
record, which is also the authoritative path readers resolve through.
Re-resolved each poll round because the worker registers the repo as
part of the same finalization the gate is waiting out.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-11 11:43:47 +01:00
ChamHerry
fc6007e70b
feat(i18n): make web and CLI language-aware (#1748) 2026-05-23 06:14:24 +01:00
Yogesh Singh
d87744fffc
fix: resolve false 404 errors and stale repo context during multi-repo switching on Windows (#633)
Some checks are pending
CI / quality (push) Waiting to run
CI / tests (push) Waiting to run
CI / e2e (push) Waiting to run
CI / Save PR Metadata (push) Blocked by required conditions
CI / CI Gate (push) Blocked by required conditions
* fix: resolve false 404s and stale repo context during multi-repo switching on Windows

* test(e2e): add repo-switching tests — hold-queue 503, ?project= URL, Windows path normalization

* test(e2e): fix repo-switching specs — use live backend with ?server= param
2026-04-10 19:59:16 +01:00
Abhigyan Patwari
57951a197b
fix(web): scope all backend calls to the active repo, not always the first (#644) 2026-04-04 11:55:34 +01:00
Gergő Magyar
bf09eab95b
feat: configure prettier with pre-commit hook (#563)
* feat: configure prettier with pre-commit hook integration

Add prettier, lint-staged, and prettier-plugin-tailwindcss at the repo
root with husky pre-commit hook integration. Moves husky from
gitnexus/ to root package.json for reliable hook installation.

- Root package.json with prepare/format/format:check scripts
- .prettierrc with endOfLine:lf and tailwindStylesheet for TW v4
- .prettierignore excluding fixtures, vendor, generated, *.d.ts, *.md
- .gitattributes enforcing LF line endings for Windows consistency
- Pre-commit hook uses direct node_modules/.bin/ paths (no npx)

* style: apply prettier formatting to entire codebase

One-time bulk format. No logic changes.
Use .git-blame-ignore-revs to skip this commit in git blame.

* chore: add .git-blame-ignore-revs for prettier format commit

* perf: pre-commit hook runs only tests related to staged files

Use vitest --related to scope test execution to tests that import
the changed files, instead of running the full suite on every commit.

* perf: remove vitest from pre-commit hook, keep in CI only

Pre-commit now runs lint-staged + tsc only. Tests run in CI
(ci-tests.yml) where they belong — keeps commits fast.

* ci: add prettier format check to quality workflow

PRs will now fail if code isn't formatted with prettier.
2026-03-28 14:58:04 +00:00
Gergő Magyar
fd7fb5bf1f
feat: unify web and cli ingestion pipeline (#536)
* feat: add server-side ingestion API (POST /api/analyze, SSE progress)

Extract core analysis orchestration from CLI into shared run-analyze.ts
module. Add server-side analyze endpoints so the web app can trigger
ingestion via HTTP instead of running the full pipeline in-browser.

New files:
- src/core/run-analyze.ts — shared runFullAnalysis() orchestrator
- src/server/analyze-job.ts — job manager (single-slot, dedup, SSE events)
- src/server/analyze-worker.ts — forked child process (8GB heap, IPC)
- src/server/git-clone.ts — shallow clone/pull with SSRF protection

API endpoints:
- POST /api/analyze — start analysis (returns 202 + jobId)
- GET /api/analyze/:jobId — poll job status
- GET /api/analyze/:jobId/progress — SSE progress stream

Security: URL validation blocks private IPs and non-HTTP schemes.
Path validation requires absolute paths. Git stderr not leaked to API.

* feat(web): add server-side analyze UI (Phase 2)

Add "Analyze on Server" flow to the web app's Server tab so users
can trigger server-side ingestion from the browser. On completion,
the graph is automatically loaded via the existing connectToServer flow.

New files:
- AnalyzeProgress.tsx — progress bar with phase label, elapsed time, cancel

Modified files:
- backend.ts — startAnalyze(), streamAnalyzeProgress() SSE client
- DropZone.tsx — analyze URL input + button below Connect section
- App.tsx — onServerAnalyze handler wires analyze -> connect flow

* feat: add job cancellation, timeout, and child process tracking (Phase 3)

- DELETE /api/analyze/:jobId — cancel running analysis (SIGTERM to worker)
- 30-minute timeout kills long-running workers automatically
- Child process refs tracked in JobManager for cleanup on shutdown
- dispose() kills all active children on SIGINT/SIGTERM
- Web cancel button now calls server DELETE endpoint
- cancelAnalyze() added to web backend client

* refactor(web): remove browser ingestion pipeline (Phase 4)

Delete 16 duplicated ingestion files, 2 unused service files
(git-clone, zip), and tree-sitter parser-loader from gitnexus-web.
All ingestion now runs server-side via POST /api/analyze.

Deleted (18 files, ~5,000 lines):
- core/ingestion/*.ts (16 pipeline processors)
- core/tree-sitter/parser-loader.ts (WASM tree-sitter loader)
- services/git-clone.ts (isomorphic-git client-side clone)
- services/zip.ts (JSZip extraction)

Simplified:
- DropZone.tsx — server-only (removed ZIP/GitHub tabs)
- ingestion.worker.ts — removed runPipeline/runPipelineFromFiles
- useAppState.tsx — removed pipeline callbacks
- App.tsx — removed handleFileSelect/handleGitClone
- main.tsx — removed Buffer polyfill for isomorphic-git
- types/pipeline.ts — removed PipelineResult/serialize helpers

Kept: cluster-enricher.ts (LLM enrichment, still used by worker)

Dependencies now removable: web-tree-sitter, isomorphic-git,
@isomorphic-git/lightning-fs, jszip (estimated 3-4MB bundle savings)

* refactor(web): sync graph schema from CLI + delete WASM grammars

Sync graph/types.ts and lbug/schema.ts from the CLI (source of truth)
to the web module so the browser LadybugDB can handle all node and
relationship types the server pipeline produces.

Synced types: Route, Tool, Section node labels; HANDLES_ROUTE, FETCHES,
HANDLES_TOOL, ENTRY_POINT_OF, WRAPS, QUERIES relationship types;
description fields on Function/Class/Interface/Method/CodeElement.

Deleted: public/wasm/ directory (14 tree-sitter WASM grammars + core).
Removed deps: web-tree-sitter, isomorphic-git, @isomorphic-git/lightning-fs,
jszip, buffer, @types/jszip (~3-4MB bundle savings).

* feat: create gitnexus-shared package for unified type definitions

Create a new gitnexus-shared package that is the single source of truth
for types shared between the CLI and web modules:

- SupportedLanguages enum (15 languages)
- Graph types: NodeLabel, NodeProperties, RelationshipType, GraphNode, GraphRelationship
- Schema constants: NODE_TABLES, REL_TYPES, REL_TABLE_NAME, EMBEDDING_TABLE_NAME
- Pipeline types: PipelinePhase, PipelineProgress

Both gitnexus (CLI) and gitnexus-web import from gitnexus-shared via
file: dependency. Each package re-exports and extends with platform-specific
additions (CLI: KnowledgeGraph with mutation methods; Web: simpler KnowledgeGraph).

This ensures types can never drift between packages — adding a new
language, node type, or relationship type in gitnexus-shared automatically
propagates to both consumers.

* refactor: import shared types directly from gitnexus-shared at call sites

Replace all re-export patterns with direct imports from gitnexus-shared.
72 files updated across CLI and web:

- SupportedLanguages: 49 CLI files now import from 'gitnexus-shared'
  instead of '../config/supported-languages.js'
- GraphNode, GraphRelationship, NodeLabel: 22 CLI + 10 web files now
  import from 'gitnexus-shared' instead of local re-export wrappers
- NODE_TABLES: api.ts imports from 'gitnexus-shared'
- PipelineProgress: useAppState.tsx imports from 'gitnexus-shared'

Local types.ts files now only define platform-specific KnowledgeGraph
(CLI has mutation methods, web has add-only). No more re-exports.

* fix: update lock files for gitnexus-shared, remove stale vite polyfills

Add gitnexus-shared@1.0.0 to lock files so npm ci succeeds in CI.
Remove buffer polyfill and global define from vite.config.ts (isomorphic-git was removed).

* fix(security): add write guard to HTTP /api/query, fix CORS proxy bypass

- Add isWriteQuery() check to POST /api/query handler — blocks CREATE,
  DELETE, SET, MERGE, DROP, etc. via HTTP API (guard was only in MCP
  pool adapter and browser-side, not the HTTP server path)
- Extend CYPHER_WRITE_RE with CALL, INSTALL, LOAD keywords
- Fix CORS proxy subdomain bypass: endsWith('github.com') allowed
  'evil-github.com'. Now requires exact match or '.github.com' suffix

* feat(server): enhance /api/search with enrichment, add /api/grep, strip graph content

- POST /api/search: add mode param (hybrid|semantic|bm25), server-side
  enrichment returns connections/cluster/processes per result in one call
  (collapses 31 sequential HTTP calls to 1 for the agent search tool)
- GET /api/grep: regex search across indexed file contents, eliminates
  need to transfer all file contents to browser
- GET /api/graph: strip content field by default (80-95% payload
  reduction). Use ?includeContent=true for backward compat
- Add LRU cache invalidation hook point for future caching

* feat(server): add /api/embed endpoint for server-side embedding generation

- POST /api/embed: triggers embedding pipeline via onnxruntime-node
  with JobManager for single-slot concurrency, timeout, and dedup
- GET /api/embed/:jobId: poll job status
- GET /api/embed/:jobId/progress: SSE stream with heartbeat, event IDs,
  and X-Accel-Buffering:no header for proxy compatibility
- DELETE /api/embed/:jobId: cancel running embedding job
- Maps embedding pipeline phases (ready→complete, error→failed) to
  JobManager status conventions

* feat(web): create consolidated BackendClient module

Single HTTP client replacing backend.ts, server-connection.ts, and
worker HTTP helpers. Includes:
- Typed methods: runQuery, search (enriched), grep, readFile, connect
- Generic streamSSE<T> utility extracted from analyze progress pattern
- BackendError with discriminated code field (network/server/client/timeout)
- Embed API: startEmbeddings, streamEmbeddingProgress, cancelEmbeddings
- Search with mode param (hybrid|semantic|bm25) and enrichment

* refactor(web): rewrite Graph RAG tools for backend-only HTTP queries

- Search tool: uses enriched /api/search (1 call replaces 31 sequential queries)
- Cypher tool: removes browser-side embedding; {{QUERY_VECTOR}} routes to
  /api/search with mode:'semantic' instead of local transformers.js
- Grep tool: uses /api/grep instead of in-memory fileContents map
- Read tool: uses /api/file instead of fileContents map lookup
- Impact tool: getCallSiteSnippet now async via /api/file
- createGraphRAGTools now accepts GraphRAGBackend interface instead of
  7 separate function params + fileContents map
- createGraphRAGAgent simplified to (config, backend, context?)
- Removed imports: embedder, lbug/schema (replaced with gitnexus-shared)
- Net: -205 lines

* refactor(web): delete WASM infrastructure, remove 7 packages (-5242 lines)

Delete browser-side LadybugDB, embeddings, search, and worker:
- gitnexus-web/src/core/lbug/ (adapter, csv-generator, schema, query-result)
- gitnexus-web/src/core/embeddings/ (embedder, pipeline, text-gen, types)
- gitnexus-web/src/core/search/ (bm25-index, hybrid-search)
- gitnexus-web/src/workers/ingestion.worker.ts (828 lines)
- gitnexus-web/src/services/server-connection.ts (merged into backend-client)
- gitnexus-web/src/types/lbug-wasm.d.ts

Remove packages: @ladybugdb/wasm-core, @huggingface/transformers,
comlink, minisearch, vite-plugin-wasm, vite-plugin-top-level-await,
vite-plugin-static-copy

Update vite.config.ts: remove WASM plugins, COOP/COEP headers,
worker config, optimizeDeps exclude

Update imports: App.tsx, DropZone, Header, AnalyzeProgress,
BackendRepoSelector, useBackend → backend-client

* refactor(web): replace Worker/Comlink with direct BackendClient calls

- useAppState: remove Worker instantiation, Comlink.wrap, apiRef.
  All queries now go through BackendClient HTTP functions directly.
- Agent runs on main thread (I/O-bound LLM streaming, not CPU-bound)
- initializeAgent: creates GraphRAGAgent with GraphRAGBackend interface
  bound to BackendClient methods (runQuery, search, grep, readFile)
- startEmbeddings: calls POST /api/embed + SSE progress instead of
  running browser-side transformers.js pipeline
- switchRepo: no longer loads graph into WASM DB or extracts fileContents
- App.tsx: handleServerConnect simplified (no fileContents, no loadServerGraph)
- Delete old backend.ts (replaced by backend-client.ts)
- Net: -396 lines

* fix(web): fix await-in-map build error in agent streaming

Move dynamic import of AIMessage outside .map() callback to avoid
"await can only be used inside an async function" build error.

* fix(web): remove stale apiRef references that broke chat functionality

sendChatMessage referenced apiRef.current (deleted Worker ref) which
would throw TypeError. Replaced with agentRef.current guard since agent
now runs on main thread.

* fix(server): dispose embedJobManager on shutdown, fix job mutation

- Add embedJobManager.dispose() to shutdown handler (was missing,
  causing cleanup timer to keep Node process alive)
- Replace direct job.repoName/status mutation with updateJob() to
  ensure SSE event emission for initial status change

* fix(server): parameterize Cypher, harden grep, unify SSE endpoints

- Search enrichment: replace string interpolation with executePrepared()
  using $nid parameter binding to prevent Cypher injection
- Add executePrepared() to core lbug-adapter (prepare/execute pattern)
- /api/grep: add 200-char pattern length limit (ReDoS protection),
  search files on disk instead of loading entire corpus into memory
  (constant memory usage regardless of repo size)
- Extract mountSSEProgress() shared helper for SSE streaming — both
  analyze and embed endpoints now have consistent heartbeat (30s),
  event IDs (reconnection support), and X-Accel-Buffering header

* refactor(web): remove dead code from Worker-era architecture

- Remove loadServerGraph no-op function, interface member, and all consumers
- Remove testArrayParams stub and interface member
- Remove fileContents state from GraphStateProvider (never populated in
  server-side architecture)
- Remove forceDevice parameter from startEmbeddings (server-side, no device choice)
- Replace phantom EmbeddingProgress type with inline { phase, percent }
- Replace resolvePathFromContents (needed fileContents Map) with graph-based
  file path resolution using filePathIndex built from graph nodes
- Fix: AI citation grounding ([[file.ts:10]]) now works via graph node lookup
  instead of broken fileContents-based resolution

* fix(web): use streamAgentResponse for full tool_call/reasoning streaming

Replace naive agent.stream() loop that only handled content chunks with
streamAgentResponse() generator from agent.ts. This properly routes:
- reasoning tokens (before/between tool calls)
- tool_call events (name, args, status)
- tool_result events (completed tool output)
- content tokens (final answer after all tools done)

Previously the onChunk handler for tool_call/tool_result/reasoning was
dead code since the streaming loop only emitted content events.

* fix(web): resolve CI type errors from dead code removal

- Import GraphNode/GraphRelationship from gitnexus-shared in graph.ts
  (not re-exported from local types.ts)
- Add Route, Tool entries to NODE_COLORS and NODE_SIZES constants
- Add PipelineResult type to web types/pipeline.ts
- Remove fileContents from CodeReferencesPanel and RightPanel
- Remove testArrayParams and forceDevice from EmbeddingStatus
- Remove forceDevice args from startEmbeddings() calls in App.tsx
- Fix embeddingProgress property accesses for simplified type

* fix(ci): add setup-gitnexus-web action, build shared once per job

- Remove prepare script from gitnexus-shared (tsc not available during
  npm ci of consuming packages)
- Create .github/actions/setup-gitnexus-web composite action: builds
  gitnexus-shared then runs npm ci for gitnexus-web
- setup-gitnexus action: already builds gitnexus-shared for CLI jobs
- ci-quality typecheck-web: uses setup-gitnexus-web (DRY)
- ci-e2e: uses setup-gitnexus-web (DRY)
- ci-tests: gitnexus-shared already built by setup-gitnexus, just
  install web deps without rebuilding

* fix(ci): use prepare script so gitnexus-shared builds during npm ci

Move typescript from devDependencies to dependencies in gitnexus-shared
so the prepare script (tsc) works when npm resolves file: deps during
npm ci. No GHA modifications needed — npm handles the build lifecycle
automatically.

Remove manual gitnexus-shared build steps from setup-gitnexus and
setup-gitnexus-web actions.

* fix(ci): build gitnexus-shared explicitly in setup actions

The file: dependency protocol doesn't reliably run prepare scripts
because devDependencies aren't installed first. Instead of fragile
lifecycle hacks, build gitnexus-shared explicitly in both setup actions:
- setup-gitnexus: npm install && npm run build in gitnexus-shared/
- setup-gitnexus-web: same, before npm ci in gitnexus-web/
- ci-tests: shared already built by setup-gitnexus, web just npm ci

No prepare script, no dist in git, no typescript as a prod dependency.

* fix: remove CALL from CYPHER_WRITE_RE — breaks FTS and vector search

CALL is used by read-only procedures: CALL QUERY_FTS_INDEX(...) and
CALL QUERY_VECTOR_INDEX(...). Adding it to the write guard blocked all
FTS search, causing 3 test failures. The database is opened in read-only
mode as defense-in-depth against write procedures via CALL.

Keep INSTALL and LOAD in the blocklist (genuinely dangerous).

* fix(web): update vercel.json for gitnexus-shared, remove COOP/COEP

- Add installCommand that builds gitnexus-shared before installing
  web deps (Vercel doesn't know about the monorepo file: dependency)
- Remove Cross-Origin-Opener-Policy and Cross-Origin-Embedder-Policy
  headers (no longer needed — WASM LadybugDB removed)

* fix(web): update tests for deleted modules

- Delete csv-generator.test.ts (tests deleted WASM-only csv-generator)
- Update security-guards.test.ts: import NODE_TABLES/REL_TYPES from
  gitnexus-shared instead of deleted src/core/lbug/schema
- Update server-connection.test.ts: import normalizeServerUrl from
  backend-client, remove extractFileContents tests (function deleted)

* fix(e2e): remove Server tab click — UI is now server-only

The DropZone no longer has ZIP/GitHub/Server tabs (browser ingestion
was removed). The server URL input is directly visible on the landing
page. Update e2e test to skip the tab click and go straight to input.

All 5 e2e tests pass locally.

* refactor: use gitnexus-shared for PipelinePhase/PipelineProgress types

CLI was duplicating PipelinePhase and PipelineProgress locally instead
of importing from gitnexus-shared. Updated all consumers to import
directly. Also removed dead code: SerializablePipelineResult,
serializePipelineResult(), deserializePipelineResult().

* fix(server): address PR #536 review — security, race conditions, dead code

- Fix path traversal in POST /api/analyze: split into isAbsolute + normalize check
- Add shared repo lock (activeRepoPaths) preventing concurrent analyze+embed on same repo
- Fix 202 response returning actual job.status instead of hardcoded 'queued'
- Add 30-minute timeout for embedding jobs (was missing unlike analyze jobs)
- Fix DropZone calling startAnalyze without setting backend URL first
- Add SSE reconnect with exponential backoff (3 retries) and Last-Event-ID
- Fix normalizeServerUrl to return base URL (no /api suffix) — clear contract
- Delete dead code: proxy.ts, server-graph-hydration.ts, pipeline.ts re-export barrel
- Update LoadingOverlay to import PipelineProgress directly from gitnexus-shared

* fix(server): fix repo lock key mismatch and embed cancel race

- Use getStoragePath(targetPath) as lock key in analyze handler to match
  embed handler's entry.storagePath — keys now always align
- Guard embed completion: don't overwrite 'failed' with 'complete' when
  job was cancelled while pipeline was still running
- Remove unused jobType parameter from acquireRepoLock
- Log backend.init() errors instead of silently swallowing

* fix: add gitnexus-shared as a local dependency in package-lock.json

* refactor: move language detection to gitnexus-shared, add syntax highlighting for all 15 languages

Move getLanguageFromFilename() from CLI to gitnexus-shared with COBOL
support added. Add getSyntaxLanguageFromFilename() for Prism-compatible
syntax highlighting covering all 15 code languages plus auxiliary
formats (json, yaml, markdown, html, css, bash, sql, xml).

Refactor CodeReferencesPanel to use shared function instead of a local
30-line switch. Delete dead gitnexus-web/src/config/supported-languages.ts
(web already imports SupportedLanguages from gitnexus-shared).

* feat(web): add first-time user onboarding with auto server detection

Replace the manual "Connect to Server" panel with an automatic onboarding
flow that guides first-time users through starting the GitNexus server.

Server detection:
- useBackend hook polls via setTimeout chain (3s, no overlap)
- Page Visibility API pauses polling when tab is hidden
- SSE heartbeat (/api/heartbeat) for instant disconnect detection

Onboarding UI (OnboardingGuide.tsx):
- Step-by-step flow: copy command → run → auto-connect
- Smart command: shows `gitnexus serve` in dev, `npx gitnexus@latest serve` in prod
- Node.js version auto-detected from package.json via Vite define
- Faux terminal windows with copy-to-clipboard, platform tabs, polling indicator

Transitions (DropZone.tsx):
- Crossfade wrapper with snapshot pattern for smooth phase transitions
- Three phases: onboarding → success (1.2s hold) → loading → graph
- Auto-recovery: falls back to onboarding if server dies or connect fails

Server changes:
- GET /api/heartbeat: SSE endpoint for liveness detection
- GET /api/info: version, launch context, Node.js version
- npm run serve script for local development
- app.disable('x-powered-by') hardening

* feat(web): add repo analysis UI, SSE heartbeat, and review fixes

Repo analysis:
- AnalyzeOnboarding: empty-state card when server has zero repos
- RepoAnalyzer: GitHub URL + Local Folder tabs with browse button
- Header repo dropdown: click project badge to switch repos or analyze new
- DropZone 'analyze' phase integrated into Crossfade transitions

Reliability fixes from 5-agent review:
- Polling: stop scheduling timers when tab hidden, restart on visibility return
- Heartbeat: exponential backoff (1s/2s/4s, 3 retries) prevents graph loss on blip
- RepoAnalyzer: completion timer tracked in ref, cleaned up on unmount
- DropZone: standardized card padding (p-7), heading sizes (text-lg)

Accessibility:
- prefers-reduced-motion global CSS rule (WCAG 2.3.3)
- focus-visible rings on CopyButton
- cursor-pointer on all Header buttons
- Consistent rounded-xl on all dropdowns

Cleanup:
- Deleted dead AnalyzeSheet.tsx (219 LOC) and BackendRepoSelector.tsx (89 LOC)
- Fixed AnalyzeProgress lucide import (lucide-react → @/lib/lucide-icons)

* fix(server): resolve analyze worker fork crash in dev mode

The forked analyze worker was crashing immediately with exit code 1
when running via `npm run serve` (tsx). Two issues:

1. Worker path resolved to `analyze-worker.js` but only `.ts` exists
   in the source directory — the `.js` file is only in `dist/`.

2. On Windows, bare `--import tsx` in execArgv fails because Node's
   ESM resolver for --import uses the child's CWD, not the parent's
   node_modules. Windows also rejects raw paths as `d:` is not a
   valid URL scheme.

Fix: detect dev vs prod via `import.meta.url` extension. In dev mode,
resolve `tsx/esm` to an absolute `file://` URL via `pathToFileURL()`
anchored to the parent's `createRequire` context. This works on all
platforms and doesn't depend on the child's CWD or PATH.

Also captures child stderr for better crash diagnostics.

Verified: `POST /api/analyze` with GitHub URL completes successfully
in dev mode (tsx) — status goes from cloning → analyzing → complete.

* fix(server): add worker auto-retry, error handling, and crash diagnostics

Worker resilience:
- Auto-retry up to 2 times with exponential backoff (1s, 2s) on crash
- SSE progress shows "Retrying after crash (1/2)..." during retry
- Captures child stderr for crash diagnostics in failure message
- AnalyzeJob tracks retryCount per job

Server error handling:
- app.listen wrapped in Promise so EADDRINUSE/EACCES propagate cleanly
- serve.ts catches startup errors with friendly messages and exit code 1
- EADDRINUSE gets actionable guidance (stop other process or --port flag)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- DEBUG=1 env var shows full stack traces

* feat: add e2e tests for onboarding flows, worker retry, and error handling

E2E tests (onboarding.spec.ts — 11 tests):
- Flow 1: OnboardingGuide shown when server unreachable (6 tests)
- Flow 2: Auto-connect with success card, analyze phase for zero repos
- Flow 3: Analyze form — GitHub URL validation, Local Folder tab, tab switching
- Flow 4: Repo dropdown in exploring view (skipped without live server)

Updated server-connect.spec.ts:
- Replaced manual Connect button flow with auto-connect waitForGraphLoaded

Server resilience:
- Worker auto-retry (2 attempts with exponential backoff) on crash
- Friendly error messages for serve startup failures (EADDRINUSE etc.)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- app.listen wrapped in Promise for proper error propagation

* refactor(shared): enforce exhaustive language coverage via Record types

Replace the if/else chain in getLanguageFromFilename with two exhaustive
Record<SupportedLanguages, ...> maps:

- EXTENSION_MAP: every language → its file extensions
- SYNTAX_MAP: every language → its Prism syntax identifier

Adding a new member to the SupportedLanguages enum without adding it to
both maps now produces a TypeScript compile error:

  Property '[SupportedLanguages.NewLang]' is missing in type...

This matches the existing pattern in languages/index.ts (providers table)
which already uses `satisfies Record<SupportedLanguages, LanguageProvider>`.

Three compile-time enforcement points now exist:
1. EXTENSION_MAP in language-detection.ts (file extensions)
2. SYNTAX_MAP in language-detection.ts (Prism syntax identifiers)
3. providers in languages/index.ts (LanguageProvider instances)

* feat(web): load source code from server and scroll to selected line

CodeReferencesPanel now fetches file content via GET /api/file when a
node is selected, instead of showing "Code not available in memory".

- Fetches via readFile() from backend-client when selectedFilePath changes
- Shows loading spinner while fetching
- After content loads, auto-scrolls to the selected node's startLine
- Highlights the selected line range with a cyan left border
- Cancels in-flight fetch if selection changes before it completes

Also: refactored language-detection.ts to use exhaustive Record types
(EXTENSION_MAP and SYNTAX_MAP) so adding a new SupportedLanguages enum
member without implementing extensions/syntax is a compile error.

* feat: buffered file reading for Code Inspector

Server: GET /api/file now supports ?startLine=N&endLine=M for reading
a line range instead of the entire file. Returns { content, startLine,
endLine, totalLines }.

Client: readFile() returns ReadFileResult with metadata. When selecting
a symbol (function, class, method), fetches only ±50 lines around the
symbol's startLine/endLine instead of the full file. File nodes still
fetch the entire file.

SyntaxHighlighter startingLineNumber set from the buffer offset so line
numbers are correct even for partial reads.

* fix: adapt readFile callers to new ReadFileResult return type

tools.ts: readFile comes from GraphRAGBackend interface which returns
Promise<string> (the adapter in useAppState extracts .content), so
revert the { content } destructuring back to plain string assignment.

useAppState.tsx: wrap backendReadFile with { repo } options object
and extract .content to satisfy the GraphRAGBackend interface.

* fix(web): ensure new repos appear in list immediately after analysis

Two fixes:

1. DropZone: handleAnalyzeComplete now passes the repoName through to
   connectToServer so the specific newly-analyzed repo loads — not the
   server's default first repo.

2. App.tsx: fetchRepos() is now awaited BEFORE handleServerConnect in
   both the DropZone and Header flows. This ensures the repo list is
   populated before the exploring view renders, so the new repo appears
   in the header dropdown immediately without a page reload.

* feat: delete repos, re-analyze with force, select after analysis

Server — DELETE /api/repo:
- Acquires repo lock first (409 if analyze/embed in flight)
- Closes LadybugDB, deletes index + clone dir, unregisters, re-inits
- Lock released in finally block

Server — analyze complete:
- backend.init() must succeed before SSE complete fires
- If backend.init() fails, job is marked failed (not complete)

Web — Header repo dropdown:
- Re-analyze: calls POST /api/analyze with force=true, shows spinning
  icon + inline progress bar via SSE
- Delete: acquires lock, aborts any running re-analysis SSE for same
  repo, refreshes list, switches to next repo
- After analysis completes: refreshes repo list, connects to the
  specific repo by name, loads graph, shows in explorer
- Retry with 1.5s backoff on 404 (server may still be reinitializing)

Type safety:
- err: any → err: unknown + instanceof BackendError in retry loop
- Added missing BackendRepo + BackendError imports in App.tsx
2026-03-28 14:07:11 +00:00
jreakin
83ec256cec fix(web): address React rendering review — stale closure, ref dep, O(1) lookups
- MarkdownRenderer: wrap handleLinkClick in useCallback, add to
  markdownComponents useMemo deps (fixes stale closure)
- GraphCanvas: remove sigmaRef from useEffect deps (ref identity
  never changes), extract handleToggleAIHighlights to useCallback
- CodeReferencesPanel: add nodeById Map for O(1) focus-in-graph
  lookup (was O(N) graph.nodes.find on every click)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:34 -05:00
jreakin
4199c5e4f3 perf(web): O(1) lookups, memoization, bundle optimizations, React fixes
Performance:
- nodeById Map in GraphCanvas for O(1) click/hover lookups
- fileNodeByPath Map in useAppState for O(1) file path lookups
- Set.has() for HIGHLIGHT_NODES/IMPACT matching (was O(N²))
- useMemo for primaryLanguage in StatusBar
- useCallback on toggleLabelVisibility/toggleEdgeVisibility
- Recursive FileTreePanel search (full subtree, not 1 level)

React fixes:
- Remove stale queryResult dep from clearAICodeReferences
- Cancel RAF chains in CodeReferencesPanel on cleanup
- Clean up timeouts in SettingsPanel and MarkdownRenderer on unmount
- try/catch on localStorage in DropZone (private browsing)
- Await handleServerConnect in App.tsx auto-connect
- pendingToolCalls counter replaces allToolsDone boolean in agent.ts
- JSON.parse try/catch in agent streaming
- sessionStorage JSDoc fix in settings-service

Bundle:
- Centralized lucide icon deep imports

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
Candido Sales Gomes
0999595444
feat(ruby): Add Ruby language support for CLI and web (#111) 2026-03-13 21:46:59 +00:00
abhigyanpatwari
c90576442e feat: merge gitnexus-mcp into gitnexus package - unified CLI+MCP 2026-02-04 01:12:41 +05:30
Renamed from gitnexus/src/components/CodeReferencesPanel.tsx (Browse further)