GitNexus/gitnexus/test/unit/analyze-api.test.ts
Felipe c2ca132620
fix: web citation/code panel bugs and serve analyze/route hardening (#3348)
* chore: ignore local Vercel link artifacts

Keep .vercel and env files out of the repo after a local preview link.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): repair citation chips, code panel line math, stale agent state

Audit findings in the web client, each verified against the source:

- RightPanel: citation chips ([[path:10-20]], [[Class:Foo]]) were wired to a
  stub resolver that always returned null, so clicking any citation did
  nothing. Expose resolveFilePath from useAppState and use it.
- CodeReferencesPanel: graph startLine/endLine are 1-based but were treated
  as 0-based, so the highlighted range and scroll target were off by one
  line; AI citation cards always rendered "code not available" because the
  snippet loader was a stub. Fetch per-citation snippets via /api/file.
- useAppState: sendChatMessage read llmSettings.activeProvider outside its
  deps (stale provider capabilities after switching provider);
  initializeAgent trapped projectName at '' for callers without an override
  (system prompt labelled the codebase "project"); the embeddings 409 dedup
  matched a message the server never sends for same-repo jobs.
- tools.ts impact: for path targets every symbol defined in the file shares
  the filePath, so the disambiguation always picked the first row and could
  analyze an arbitrary symbol while reporting a file impact. Prefer the File
  node.
- useSigma: the layout timeout called stop() but never kill(), leaking one
  ForceAtlas2 Web Worker plus four graph listeners per completed layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): release repo lock on cancel, IPC-first worker cancel, route hardening

Audit findings in gitnexus/src, each verified against the source:

- analyze-launch: cancelJob marks the job failed before the worker exits,
  so the exit handler's terminal early-return skipped releaseLockOnce and
  the repo stayed locked ("Another job is already active") until restart.
  Release on terminal exit and when forkWorker bails on a terminal job.
- analyze-job: cancellation now sends { type: 'cancel' } over IPC first and
  signals only after a 15s grace. On Windows child.kill('SIGTERM') is a
  forceful termination, so leading with it could kill the worker inside a
  LadybugDB write. Mirrors core/auto-sync/analysis-worker-launch.
- api resolveRepo: a job that FAILED during the hold-queue wait fell through
  to the { __timedOut } sentinel ("taking longer than expected") instead of
  404; only /api/repo checked the sentinel, so graph/query/search/file/grep/
  embed/delete crashed on entry.storagePath with a 500 after a 5 minute hang.
  Return null on failed jobs and check the sentinel in every consumer.
- api processes/process/clusters/cluster: resolve ?repo= through the HTTP
  resolver (documented policy on resolveRegisteredRepoEntry) and pass the
  registered absolute path to the backend; add the standard rate limiter.
  The raw param previously reached the MCP resolver, which runs a
  cwd-relative realpathSync probe + registry refresh on a bare-name miss
  and accepts unambiguous partial names.
- api body handling: Express 5 leaves req.body undefined without a JSON
  content type, turning "Missing X" 400s into TypeError 500s; body-parser
  4xx errors (malformed JSON, over-limit) were also reported as 500.
- /api/file: the lexical path.relative check cannot see symlinks; re-check
  containment on realpath so a cloned repo containing evil -> /etc/passwd
  cannot read outside the root.
- repo-manager unregisterRepo: used the lenient reader, so a transient read
  error (EBUSY/EPERM racing another process's atomic rename) turned into
  writing [] and deregistering every repo. Use the strict-if-present reader.
- clean --branch: compared registry paths with raw path.resolve instead of
  the canonical registryPathEquals used everywhere else (macOS /private/var,
  Windows short names / drive-letter case) and reported indexed branches as
  not indexed.

Tests: cancelJob IPC-before-signal contract; /api/file symlink escape (403)
and in-repo symlink (200), skipped where the host cannot create symlinks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: keep example env files visible and ignore local Cursor config

.env* also hid gitnexus/.env.example and eval/.env.example. .vercel was already ignored. The web app's .cursor/ stays local, including its MCP file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add realtime execution ops dashboard for Vercel monitoring.

Expose /api/ops snapshots over the local serve process and a ?view=ops SPA panel so analyze/embed jobs can be watched live from the hosted web UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: address gitnexus-check review on cite/embed/resolve paths

Pass req into resolveRepo for process/cluster routes, tighten same-repo embed 409 handling, guard empty citation paths, fix snippet retry races, and drop the lone-File ambiguity fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): harden CORS, redact paths, and bound ops streams

Restrict Vercel CORS to exact production hosts, omit raw repo paths/URLs from the unauthenticated ops feed, rate-limit and cap SSE connections, and fix dashboard SSE/poll edge cases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): stabilize citation fetches and file impact matching

Retry cancelled snippet loads without duplicate in-flight reads, cap range-less citation downloads, and make impact file matching unique-suffix-aware with a synthetic File target when LIMIT drops the File node.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web/ops): close gitnexus-check review threads on SSE and redaction

Cap citation reads, reconnect ops on applied server URL, skip SSE onError after abort, and strip URL query/fragment from public repoName.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops): abort SSE on poll fallback and harden repoName parsing

Prevent dual SSE+poll after a failed safety snapshot, skip overlapping poll ticks, ignore aborted streamSSE onError, and basename Windows drive-letter URLs so ops never leaks path segments.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ops/web): close gitnexus-check threads on SSE budget and credential leak

Keep finite SSE retries across short 200s, strip backend URL userinfo before ?server=, and redact progress messages on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Restore 0-based GraphNode line math, redact public job poll/error fields, fix omit-?repo= 400, hold the analyze lock across cancel-during-settle, and restore the Vercel shared compile.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): contain leftover-slot reclaim to the slot and keep --stale status honest

Preview and force now share one branches/ containment rule, nested junctions cannot walk a sibling index, and a mid-loop git failure no longer claims leftovers were not deleted after a successful rm.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(cli): share leftover-slot helpers without changing reclaim behavior

Pull the rolling I/O pool and owned-cwd storage lookup into one place so clean --stale/--branch and leftover listing stop restating the same ownership and concurrency paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

Keep the omitted-repo snapshot instead of re-listing, stop citation and ops races, redact full public repo URLs, and restore fake timers in teardown.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Store graph node citation lines as 0-based offsets and correct the default-port origin comment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Redact public SSE progress text, stop citation append retries from
cancelling in-flight reads, and keep the ops dashboard from showing a
stale snapshot or clearing a failed Connect.

Note: pre-existing failure in incremental-index-extension-dml-gate and other lbug/env unit tests not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Prune citation snippets when AI refs are cleared, and poll /api/ops at 2s so the fallback stays under the 60/min limit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): apply prettier class order for format check (#3348)

CI quality/format runs root-only npm ci, so prettier-plugin-tailwindcss
sorts scrollbar-thin without the web Tailwind catalog.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a unique suffix match for graph-backed citation paths so
  ambiguous names like index.ts no longer open the first graph file.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

- Require a path-component boundary so unique citation suffixes cannot match filename substrings like myindex.ts

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact filesystem paths in ops text

Unauthenticated /api/ops and poll replay worker errors. URLs were
scrubbed but home-directory and Windows paths still leaked.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): omit repoPath from public SSE frames

/api/ops lists job ids, so the unauthenticated progress stream
must not replay the analyzed filesystem path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): treat embed lock 409 as a busy error

Analyze and embed share the same lock string. Mapping that 409 to
embedding hid an in-flight analyze as a successful embed start.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): bound impact File path suffix matches

Unbounded endsWith let lib/foo.ts select src/mylib/foo.ts. Require
an exact path or a unique /suffix, matching resolveUniqueIndexedPath.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): keep canceled analyze slot until exit

Marking failed before the worker exited let a second POST start
cloneOrPull against a LadybugDB file still being written.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): skip Windows SIGTERM on job dispose

child.kill('SIGTERM') is TerminateProcess there. Ask over IPC first
and leave the 15s grace timer to SIGKILL.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): let process routes skip the analyze hold

GET /api/processes and /api/clusters always waited up to 300s.
?awaitAnalysis=false fails fast; default still waits like /api/repo.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): cover public ops and analyze SSE user flows

Lock the unauthenticated dashboard and analyze complete/fail/cancel paths so a leaked repoPath, token, or home path cannot ship unnoticed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Keep caller cancel reasons over the worker's generic IPC, skip publish while cancel is pending, and redact scp-style remotes on the public ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address remaining PR review feedback (#3348)

Release the analyze slot when a worker fails to spawn, and omit branch refs from the unauthenticated ops feed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): do not reuse an analyze job that is pending cancel

A dying same-repo job still occupies the single slot; 202-reuse would
attach a new client to a cancel in flight.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): hold the analyze lock until the worker exits after cancel

Cancel error IPC used to drop the repo lock while the child was still
checkpointing. Abort settle immediately on pending cancel so the slot
is not held for a 60s disk poll.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): redact known repo paths with spaces in public ops text

Known repoPath/repoUrl literals are replaced first so a clone dir with
spaces cannot leak past the whitespace-bounded path regex.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(server): return the live job on analyze and embed DELETE

Hard-coding failed made clients retry immediately and 409 while the
child still occupied the slot. resolveRepo now returns not-found as
soon as that job fails instead of waiting out the hold timeout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): drive public analyze and ops through a live gitnexus serve

Spawn the real backend and observe requests instead of intercepting
them, so slot occupancy, redaction, and reconnect stay honest.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(web): reuse code-panel helpers and drop dead UI aliases

Citation fetches already had selectedNodeFileRange and snippetRepoKey;
the impact File suffix filter already handled exact paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* Address PR review feedback (#3348)

- Skip createJob reuse after cancel IPC is consumed while the child remains
- Hold the repo lock until exit when complete IPC races a pending cancel
- Scrub full remote URLs before known repoUrl prefixes in public ops text
- Reject unique impact File suffix matches from a truncated LIMIT 10 page
- Make live e2e helpers bound probes, clean up failed startups, and wait out the cancel slot

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): open live e2e pages on the Vite host CI actually bound

Absolute 127.0.0.1:5173 navigation refused on Actions because wait-on
and Vite use localhost (often ::1). Honor FRONTEND_URL when set, else
pick the first of localhost / 127.0.0.1 that answers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3348)

Hold the analyze lock until worker exit when cancel aborts settle, and assert GitLab failure chrome does not leak host or path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): keep live analyze e2e under the analyze rate limit

POST /api/analyze allows 10 requests per minute per IP. The slot-free
helper re-POSTed every 400ms while a cancelled worker was exiting, spent
that budget, and the lock test's hold request got 429 instead of 202.

- postAnalyze waits out a 429 using the RateLimit reset and retries
- slot polling backs off to 2s and leaves a small POST budget for callers
- the slot probe is a clone that fails before any worker fork, so the
  probe itself no longer holds the slot after reporting failed
- the lock test holds the slot with a real local analyze
- token and GitLab tests wait for a free slot before posting from the UI
- request fetches carry a timeout; teardown signals the serve process group
- an empty FRONTEND_URL falls back to the default base URL

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- ops view: a `?server=` link no longer auto-connects to another origin
  while a deploy token is held; it prefills and waits for Connect
- analyze completion: a local-path run reconnects by the path this client
  submitted, so duplicate basenames stay collision-safe without repoPath
  on the public SSE frame
- stale slot cleanup: revalidate each nested directory (lstat + realpath)
  right before readdir, so a mid-cleanup junction swap aborts instead of
  walking an outside tree; list phases run sequentially so the slot-I/O
  cap is global
- e2e: 429 backoff honours the caller deadline; the unreachable-backend
  ops test navigates through the resolved frontend URL
- drop the unused `repo` member from the clean integration helper

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(server): reconnect after analyze by an opaque repo id

Public job views and the SSE terminal frame no longer carry repoPath, and
repoName is not unique, so a post-analyze reconnect by name could load a
same-named sibling. The server now issues `repoId`: an HMAC of the
canonical registry path under a per-process random key. It is set on
complete jobs (ops view, analyze poll, SSE terminal frame) and matches
the new `id` on `GET /api/repos` entries. The web client resolves it to
the exact entry path on completion; unknown ids fall back to the name.
This covers URL clones and folder uploads, and replaces the local-path
only fallback.

With reconnect off the label, public `repoName` for a branch-pinned URL
clone is the repository name, not the `<repo>__<branch slug>` registry
name, so the requested branch stays off the unauthenticated feed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- stale slot cleanup: a descendant that vanishes before its unlink is
  treated as removed instead of aborting the reclaim
- e2e teardown: escalate to SIGKILL on the process group when the live
  backend ignores SIGTERM for 5s
- ops view: drop the dead initial value CodeQL flagged

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- public redaction: a known repoPath now also consumes its descendant
  tail, so `<repoPath>/src/secret.ts` becomes `[path]` instead of
  `[path]/src/secret.ts`; a same-prefix sibling is left to the path scrub
- e2e: the cancel test waits for the analyze slot the previous failed
  local-path job still holds; `fetchOps` carries the request timeout
- docs: SSE terminal payload and RepoAnalyzer `onComplete` describe
  `repoId` resolution

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address PR review feedback (#3348)

- RepoAnalyzer: drop a completion that resolves after unmount, so a slow
  /api/repos lookup cannot switch repos after the sheet was dismissed
- e2e: select the local-path input by test id (the placeholder differs on
  Windows); the slot probe's job wait honours the caller's deadline

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(server): assert the public SSE terminal frame in analyze-api

The #2790 terminality tests still expected `repoPath` on the terminal
frame. This PR replaced it with the opaque `repoId`, so assert that shape
and that the analyzed path never appears in the stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(web): wait for the analyze slot between duplicate-repo setup runs

The server now keeps the single analyze slot until the worker exits, even
after its job reports complete. repo-path-identity posted the second
duplicate's analyze immediately and got 409 in CI. Use the shared
slot-aware POST, which also waits out a 429.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 13:27:22 +01:00

715 lines
30 KiB
TypeScript

import fs from 'node:fs/promises';
import os from 'node:os';
import path from 'node:path';
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { JobManager } from '../../src/server/analyze-job.js';
import {
startSSEHarness,
terminalFrame,
terminalFrameCount,
type SSEHarness,
} from '../helpers/sse-harness.js';
import { publicRepoId } from '../../src/server/public-repo-id.js';
import {
resolveEmbedRunOutcome,
withMeasuredEmbeddingCount,
type EmbeddingRunResult,
} from '../../src/server/embed-run-outcome.js';
import { mintInterruptedCheckpoint } from '../../src/core/embedding-checkpoint.js';
import {
measurePersistedEmbeddingCount,
persistedEmbeddingCountOrUndefined,
} from '../../src/core/embedding-count.js';
import { loadMeta, saveMeta, type RepoMeta } from '../../src/storage/repo-manager.js';
import { deriveEmbeddingMode } from '../../src/core/embedding-mode.js';
/**
* NOTHING in this file imports `src/server/api.ts` for behavior. That module
* pulls Express, cors, the LadybugDB native adapter and the whole MCP wiring:
* reaching three pure helpers through it cost one 30s TIMEOUT and ~20s/~22s on
* the runs that passed, against a 30s `testTimeout` (#2790 review, finding 9).
* The helpers now live in `src/server/{sse-progress,embed-run-outcome}.ts` and
* `src/core/embedding-{count,checkpoint}.ts`, none of which import a database
* or a server.
*/
describe('analyze API logic', () => {
let manager: JobManager;
beforeEach(() => {
manager = new JobManager();
});
afterEach(() => {
manager.dispose();
});
it('creates a job and returns 202 shape', () => {
const job = manager.createJob({ repoUrl: 'https://github.com/user/repo' });
const response = { jobId: job.id, status: job.status };
expect(response.jobId).toBeTruthy();
expect(response.status).toBe('queued');
});
it('rejects when job already active for different repo', () => {
const job1 = manager.createJob({ repoUrl: 'https://github.com/user/repo1' });
manager.updateJob(job1.id, { status: 'analyzing' });
expect(() => manager.createJob({ repoUrl: 'https://github.com/user/repo2' })).toThrow(
/already in progress/,
);
});
it('returns existing job for same repo URL', () => {
const job1 = manager.createJob({ repoUrl: 'https://github.com/user/repo' });
manager.updateJob(job1.id, { status: 'analyzing' });
const job2 = manager.createJob({ repoUrl: 'https://github.com/user/repo' });
expect(job2.id).toBe(job1.id);
});
it('SSE progress listener receives all events including terminal', () => {
const job = manager.createJob({ repoUrl: 'https://github.com/user/sse-test' });
const events: Array<{ phase: string; percent: number }> = [];
const unsub = manager.onProgress(job.id, (progress) => {
events.push({ phase: progress.phase, percent: progress.percent });
});
manager.updateJob(job.id, {
status: 'analyzing',
progress: { phase: 'parsing', percent: 30, message: 'Parsing' },
});
manager.updateJob(job.id, {
progress: { phase: 'calls', percent: 50, message: 'Tracing calls' },
});
manager.updateJob(job.id, { status: 'complete', repoName: 'sse-test' });
unsub();
expect(events).toEqual([
{ phase: 'parsing', percent: 30 },
{ phase: 'calls', percent: 50 },
{ phase: 'complete', percent: 100 },
]);
});
});
const IDENTITY = { model: 'test-model', dimensions: 384, provider: 'local' };
const CLEAN_RUN: EmbeddingRunResult = {
nodesProcessed: 412,
chunksProcessed: 900,
failedNodeIds: [],
};
/** Progress figures an in-flight checkpoint records. */
const PROGRESS = { nodesProcessed: 4, totalNodes: 12, chunksProcessed: 9 };
/**
* ── #2790: an SSE client must not be told a partial run succeeded ──────────
*
* `runEmbeddingPipeline` emits `phase: 'ready'` / 100% UNCONDITIONALLY before
* returning — including when it dropped nodes to endpoint failures — and
* /api/embed relayed that as a progress phase before it had measured anything
* or decided the outcome. The relay treated a progress PHASE STRING of
* 'complete'/'failed' as terminal, so it wrote `event: complete` with
* `error: undefined`, called `res.end()` and unsubscribed; the route's later
* `updateJob({status:'failed'})` went into a stream with no listener. The web
* client fired `onComplete` and showed "ready" while a `GET /api/embed/:jobId`
* poller saw `failed` — the two consumers of one job disagreeing about whether
* the data is complete, and a regression against the pre-#2790 behavior where
* the pipeline threw and the client received the failure.
*
* These tests drive the REAL relay over a REAL HTTP server (same harness as
* server-sse-payload.test.ts) and subscribe BEFORE the misleading event is
* emitted — subscribing after it is exactly why the previous version of this
* suite passed while the bug was live.
*/
describe('mountSSEProgress terminality (#2790)', () => {
let harness: SSEHarness;
let manager: JobManager;
let baseUrl = '';
beforeEach(async () => {
// Mirrors both production mounts in createServer().
harness = await startSSEHarness('/api/embed/:jobId/progress');
manager = harness.manager;
baseUrl = harness.baseUrl;
});
afterEach(() => harness.close());
it('a partial run reaches the client as a failure, not a success', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-partial' });
manager.updateJob(job.id, {
repoName: 'embed-partial',
status: 'analyzing',
progress: { phase: 'embedding', percent: 40, message: 'Embedding nodes (40%)...' },
});
// The client is connected and listening BEFORE anything terminal-looking is
// emitted. `fetch` resolves once headers arrive, and the handler subscribes
// synchronously before that (see server-sse-payload.test.ts).
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
// A progress event that CLAIMS to be terminal. Production now maps the
// pipeline's `ready` to 'finalizing' instead, but a phase string must not be
// able to end the stream no matter who sends it — that is the invariant.
manager.updateJob(job.id, {
progress: { phase: 'complete', percent: 100, message: 'Embeddings complete' },
});
// Only now does the route learn the run dropped nodes.
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 10,
chunksProcessed: 24,
failedNodeIds: ['node-a', 'node-b'],
});
manager.updateJob(job.id, {
status: 'failed',
error: outcome.error,
partial: outcome.partial,
progress: { phase: 'failed', percent: 100, message: String(outcome.error) },
});
const body = await response.text();
expect(body).not.toContain('event: complete');
expect(body).not.toContain('/ws/embed-partial');
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'failed')).toMatchObject({
repoName: 'embed-partial',
error: expect.stringContaining('finished partially') as unknown as string,
// The distinction a UI needs to offer "retry 2 nodes" instead of a bare
// red chip — carried without adding a `status` union member.
partial: { kind: 'embedding-partial', pendingNodeCount: 2, nodesProcessed: 10 },
});
});
it('a clean run produces exactly one terminal complete event', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-clean' });
manager.updateJob(job.id, {
repoName: 'embed-clean',
status: 'analyzing',
progress: { phase: 'embedding', percent: 40, message: 'Embedding nodes (40%)...' },
});
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
// What the route actually emits between the pipeline returning and the
// outcome being known.
manager.updateJob(job.id, {
progress: { phase: 'finalizing', percent: 100, message: 'Finalizing embeddings...' },
});
manager.updateJob(job.id, {
status: 'complete',
progress: { phase: 'complete', percent: 100, message: 'Embeddings complete' },
});
const body = await response.text();
expect(body).not.toContain('event: failed');
// Exactly one — the status update carries a `progress` too, and #2264's
// single-emit rule is what keeps that from double-writing the terminal frame.
expect(terminalFrameCount(body)).toBe(1);
// Public frame: display name + opaque repoId, never the analyzed path.
expect(terminalFrame(body, 'complete')).toEqual({
repoName: 'embed-clean',
repoId: publicRepoId('/ws/embed-clean'),
});
// The 'finalizing' frame was relayed as ordinary progress, not swallowed.
expect(body).toContain('"phase":"finalizing"');
});
it('the analyze path still closes on its own terminal update', async () => {
// /api/analyze mounts the same relay. Its worker reports phases like
// 'parsing' and 'done' (never 'complete'), so the fix must not leave that
// stream open — it closes when the job's STATUS becomes terminal.
const job = manager.createJob({ repoPath: '/ws/reels' });
manager.updateJob(job.id, {
status: 'analyzing',
progress: { phase: 'parsing', percent: 30, message: 'Parsing' },
});
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
manager.updateJob(job.id, {
progress: { phase: 'done', percent: 100, message: 'Done' },
});
manager.updateJob(job.id, { status: 'complete', repoName: 'reels' });
const body = await response.text();
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'complete')).toEqual({
repoName: 'reels',
repoId: publicRepoId('/ws/reels'),
});
});
it('a job that finished before the client connected replays its outcome', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-late' });
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 3,
chunksProcessed: 9,
failedNodeIds: ['node-a'],
});
manager.updateJob(job.id, {
status: 'failed',
repoName: 'embed-late',
error: outcome.error,
partial: outcome.partial,
});
const body = await (await fetch(`${baseUrl}/api/embed/${job.id}/progress`)).text();
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'failed')).toMatchObject({
error: expect.stringContaining('finished partially') as unknown as string,
partial: { kind: 'embedding-partial', pendingNodeCount: 1, nodesProcessed: 3 },
});
});
});
/**
* ── #2790: POST /api/embed must not report unqualified success ─────────
*
* The pipeline no longer throws when a sub-batch loses its endpoint — it
* deletes the affected nodes' rows and names them in `failedNodeIds`. The route
* discarded that receipt: it cleared `embeddingCheckpoint` and marked the job
* 'complete', so a partial run looked identical to a clean one and the dropped
* nodes were never retried (pre-#2790 the pipeline threw and the catch marked
* the job failed).
*/
describe('resolveEmbedRunOutcome (#2790)', () => {
it('clears the checkpoint and reports no error on a clean, measured run', () => {
const outcome = resolveEmbedRunOutcome(IDENTITY, CLEAN_RUN, { measuredEmbeddings: 412 });
expect(outcome.checkpoint).toBeUndefined();
expect(outcome.error).toBeUndefined();
expect(outcome.partial).toBeUndefined();
});
it('retains the checkpoint with the dropped ids and reports an error on a partial run', () => {
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 10,
chunksProcessed: 24,
failedNodeIds: ['node-a', 'node-b'],
});
// The record of what failed survives — this is the pending set the next
// run's `forceReembedNodeIds` re-embeds.
expect(outcome.checkpoint).toMatchObject({
pendingNodeIds: ['node-a', 'node-b'],
nodesProcessed: 10,
totalNodes: 12,
chunksProcessed: 24,
model: 'test-model',
dimensions: 384,
provider: 'local',
// The run COMPLETED: these nodes provably hold zero rows, so a later
// identity mismatch may drop the set with a warning instead of wedging
// every subsequent run (repo-manager.ts).
kind: 'partial',
});
expect(outcome.error).toMatch(/2 node\(s\)/);
expect(outcome.partial).toEqual({
kind: 'embedding-partial',
pendingNodeCount: 2,
nodesProcessed: 10,
});
});
it('stamps no attempt count on a fresh partial run', () => {
const outcome = resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 10, chunksProcessed: 24, failedNodeIds: ['node-a'] },
// Resumed from an in-flight marker, not a partial one.
{ resumedFrom: mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']) },
);
expect(outcome.checkpoint).toMatchObject({ kind: 'partial' });
expect(outcome.checkpoint?.attempts).toBeUndefined();
});
it('advances the attempt count only when a resumed pending node fails again', () => {
const resumedFrom: RepoMeta['embeddingCheckpoint'] = {
at: new Date(0).toISOString(),
nodesProcessed: 10,
totalNodes: 12,
chunksProcessed: 24,
...IDENTITY,
kind: 'partial',
attempts: 1,
pendingNodeIds: ['node-a', 'node-b'],
};
// Same node failed again → the retry is not converging; the budget advances.
expect(
resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 11, chunksProcessed: 26, failedNodeIds: ['node-a'] },
{ resumedFrom },
).checkpoint,
).toMatchObject({ kind: 'partial', attempts: 2 });
// The resumed set cleared and DIFFERENT nodes were lost → a fresh partial,
// so the budget resets. The bound exists for a node the endpoint rejects
// deterministically, not for an endpoint that is merely flaky.
expect(
resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 11, chunksProcessed: 26, failedNodeIds: ['node-z'] },
{ resumedFrom },
).checkpoint?.attempts,
).toBeUndefined();
});
});
describe('the mid-run marker /api/embed writes (mintInterruptedCheckpoint, #2790)', () => {
it('stamps interrupted, so resume regenerates a possibly half-written window', () => {
const checkpoint = mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a', 'node-b']);
expect(checkpoint).toMatchObject({
kind: 'interrupted',
nodesProcessed: 4,
totalNodes: 12,
chunksProcessed: 9,
model: 'test-model',
dimensions: 384,
provider: 'local',
pendingNodeIds: ['node-a', 'node-b'],
});
// `attempts` bounds retries of a 'partial' set; an in-flight marker has no
// such budget because its rows may exist.
expect(checkpoint.attempts).toBeUndefined();
});
});
/**
* ── The /api/embed count omission (silent embedding loss) ──────────────
*
* The route generated embeddings and wrote `embeddingCheckpoint`, but never
* `stats.embeddings`. A repo embedded purely through the server therefore kept
* whatever count the last CLI `analyze` stamped — `0` for a repo analyzed
* without embeddings. The next CLI run reads that as `existingEmbeddingCount`,
* `deriveEmbeddingMode` sees `hasExisting: false` → `shouldLoadCache: false`,
* and `gitnexus analyze --force` wipes the database with no cache load: every
* server-generated embedding is destroyed with no warning.
*
* The route body is an inline closure inside `createServer`, so its finalize
* sequence is replayed here over the SAME helpers the route calls, with real
* meta.json I/O and the real `deriveEmbeddingMode`. The consequence is what
* these tests pin, not the field.
*/
describe('POST /api/embed records the embedding count it measured', () => {
let metaDir: string;
let seeded: RepoMeta;
beforeEach(async () => {
metaDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-embed-count-'));
});
afterEach(async () => {
await fs.rm(metaDir, { recursive: true, force: true });
});
/** What a CLI `analyze` (plus any mid-run checkpoint) leaves on disk. */
const seedMeta = async (
embeddings: number | undefined,
embeddingCheckpoint?: RepoMeta['embeddingCheckpoint'],
): Promise<void> => {
seeded = {
repoPath: '/repo/embed-count',
lastCommit: 'abc123',
indexedAt: new Date(0).toISOString(),
stats: { nodes: 500, ...(embeddings === undefined ? {} : { embeddings }) },
embeddingCheckpoint,
};
await saveMeta(metaDir, seeded);
};
const rowsWith = (cnt: unknown) => async () => [{ cnt } as Record<string, unknown>];
/** The route's finalize sequence: measure → re-read meta → resolve → write. */
const finalizeEmbedRun = async (
runQuery: (cypher: string) => Promise<Array<Record<string, unknown>> | undefined>,
pipelineResult: EmbeddingRunResult,
): Promise<RepoMeta | null> => {
const measured = await measurePersistedEmbeddingCount(runQuery);
const finalMeta = (await loadMeta(metaDir)) ?? seeded;
const outcome = resolveEmbedRunOutcome(IDENTITY, pipelineResult, {
measuredEmbeddings: persistedEmbeddingCountOrUndefined(measured),
onDisk: finalMeta,
});
await saveMeta(
metaDir,
withMeasuredEmbeddingCount(
{ ...finalMeta, embeddingCheckpoint: outcome.checkpoint },
measured,
),
);
return loadMeta(metaDir);
};
const embeddingCountOf = (meta: RepoMeta | null): number => meta?.stats?.embeddings ?? 0;
it('writes the measured count into meta on a clean run, without disturbing the other stats', async () => {
await seedMeta(0);
const asked: string[] = [];
const written = await finalizeEmbedRun(async (cypher) => {
asked.push(cypher);
return [{ cnt: 412 }];
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { nodes: 500, embeddings: 412 } });
// A clean, MEASURED run clears the checkpoint (#2790 contract).
expect(written?.embeddingCheckpoint).toBeUndefined();
// Measured, not restated: the count comes from the live embedding table.
expect(asked).toEqual([expect.stringMatching(/MATCH \(e:\w+\) RETURN count\(e\) AS cnt/)]);
});
it('is what makes the next CLI run preserve instead of wipe', async () => {
await seedMeta(0);
// Pre-fix state: the server embedded 412 nodes but meta still says 0.
const stale = embeddingCountOf(await loadMeta(metaDir));
expect(stale).toBe(0);
expect(deriveEmbeddingMode({ force: true }, stale)).toMatchObject({
// `--force` rebuilds without loading the embedding cache → the 412
// server-generated vectors are destroyed.
shouldLoadCache: false,
preserveExistingEmbeddings: false,
});
const written = await finalizeEmbedRun(rowsWith(412), CLEAN_RUN);
const honest = embeddingCountOf(written);
expect(honest).toBe(412);
// Post-fix: `--force` loads the cache and regenerates on top of it rather
// than discarding the index. (`preserveExistingEmbeddings` is false here by
// design — `--force` upgrades to `forceRegenerateEmbeddings`; the wipe
// protection is `shouldLoadCache`.)
expect(deriveEmbeddingMode({ force: true }, honest)).toMatchObject({
shouldLoadCache: true,
forceRegenerateEmbeddings: true,
});
// A routine `analyze` preserves them outright.
expect(deriveEmbeddingMode({}, honest)).toMatchObject({
shouldLoadCache: true,
preserveExistingEmbeddings: true,
});
});
it('treats an unanswerable count query as unknown rather than 0', async () => {
// The query throws for reasons unrelated to how many rows were written.
await expect(
measurePersistedEmbeddingCount(async () => {
throw new Error('Connection closed');
}),
).resolves.toMatchObject({ kind: 'unknown', reason: 'Connection closed' });
// No row / no cell: an empty table would still answer with a 0.
await expect(measurePersistedEmbeddingCount(async () => [])).resolves.toMatchObject({
kind: 'unknown',
});
await expect(measurePersistedEmbeddingCount(async () => undefined)).resolves.toMatchObject({
kind: 'unknown',
});
// Non-numeric cell — same class of unknown.
await expect(measurePersistedEmbeddingCount(rowsWith('many'))).resolves.toMatchObject({
kind: 'unknown',
});
// A real zero is still a real answer.
await expect(measurePersistedEmbeddingCount(rowsWith(0))).resolves.toEqual({
kind: 'measured',
count: 0,
});
});
it('leaves the previous count alone when the measurement fails, never writing a fabricated 0', async () => {
await seedMeta(137);
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { embeddings: 137 } });
// The dangerous direction is wrong-LOW: a fabricated 0 here would arm the
// wipe the test above describes.
expect(deriveEmbeddingMode({ force: true }, embeddingCountOf(written))).toMatchObject({
shouldLoadCache: true,
});
});
it('keeps the recovery marker when a clean run cannot verify its own count', async () => {
// The state that arms the silent wipe: meta records 0 embeddings (a repo
// analyzed without them, embedded through the server), the run succeeded,
// and the count query cannot answer — so no honest count can be stamped.
const midRunMarker = mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']);
await seedMeta(0, midRunMarker);
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
// No fabricated value: neither a 0 nor a NaN/null lands in meta.
expect(written).toMatchObject({ stats: { nodes: 500, embeddings: 0 } });
// …and the marker this run wrote SURVIVES, so something on disk still
// records that embeddings were produced. Clearing it here would leave the
// index with zero evidence of its own embeddings.
expect(written?.embeddingCheckpoint).toMatchObject({
kind: 'interrupted',
pendingNodeIds: ['node-a'],
});
});
it('still clears the marker on an unverifiable run once meta records embeddings', async () => {
// Same unmeasurable run, but the recorded count already proves the index is
// accounted for — nothing needs preserving, so the clean-run contract wins.
await seedMeta(412, mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']));
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { embeddings: 412 } });
expect(written?.embeddingCheckpoint).toBeUndefined();
});
it('records the honest count on a partial run, alongside the pending checkpoint', async () => {
await seedMeta(0);
const written = await finalizeEmbedRun(rowsWith(300), {
nodesProcessed: 300,
chunksProcessed: 700,
failedNodeIds: ['node-a', 'node-b'],
});
// A partial index that is honest about itself survives the next run: the
// count keeps `--force` from wiping it, the checkpoint re-embeds the rest.
expect(written).toMatchObject({
stats: { embeddings: 300 },
embeddingCheckpoint: {
pendingNodeIds: ['node-a', 'node-b'],
nodesProcessed: 300,
kind: 'partial',
},
});
expect(deriveEmbeddingMode({ force: true }, embeddingCountOf(written))).toMatchObject({
shouldLoadCache: true,
});
});
});
/**
* Wiring guard for the route. Everything the helpers DECIDE is pinned
* behaviorally above; what remains is that the inline route closure inside
* `createServer` still asks them — the helper being right while the call site
* keeps writing `embeddingCheckpoint: undefined` is exactly the regression
* #2790 is about, and that closure cannot be reached without booting a server
* over a real repo + LadybugDB + embedding endpoint. Static-analysis layer of
* last resort, same precedent as api-readonly-wiring.test.ts.
*/
describe('POST /api/embed route wiring (#2790)', () => {
const readSource = () =>
fs.readFile(path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'), 'utf-8');
/**
* The body of the route's `withLbugDb` callback — everything that may only
* run while the database connection is open. Sliced rather than matched with
* a character-distance regex so a comment edit cannot silently un-assert it.
*/
const insideWithLbugDb = (source: string): string => {
// The open is a multi-line `withLbugDb(lbugPath, async () => {…}, opts)` call
// since #3091 added the FTS-mode options argument, so anchor on the call head
// and close on that options argument rather than a fixed-indent literal.
const head = source.match(/await withLbugDb\(\s*lbugPath,\s*async \(\) => \{/);
expect(head).not.toBeNull();
const start = head!.index!;
const end = source.indexOf('skipFtsOption(ftsSession.skipFts)', start);
expect(end).toBeGreaterThan(start);
return source.slice(start, end);
};
it('feeds the pipeline result through resolveEmbedRunOutcome into the finalize write', async () => {
const source = await readSource();
// The result is captured, not discarded…
expect(source).toContain('const pipelineResult = await runEmbeddingPipeline(');
// …handed to the helper with the finalize context…
expect(source).toMatch(
/resolveEmbedRunOutcome\(\s*embeddingIdentity,\s*pipelineResult,\s*finalizeContext,\s*\)/,
);
// …and its checkpoint is what the finalize meta write persists (pre-fix: a
// hardcoded `embeddingCheckpoint: undefined`).
expect(source).toContain('embeddingCheckpoint: outcome.checkpoint');
expect(source).toContain('partialRunError = outcome.error;');
// A partial run does not reach `status: 'complete'`, and carries its detail.
expect(source).toMatch(
/partialRunError === undefined[\s\S]{0,400}status: 'complete'[\s\S]{0,600}status: 'failed'/,
);
expect(source).toContain('partial: partialRunDetail,');
});
it('measures after the WAL flush, inside withLbugDb, and folds the result into the write', async () => {
const source = await readSource();
const region = insideWithLbugDb(source);
// Inside the open connection — this is the route's only chance to stamp
// `stats.embeddings`, and the next CLI run's preserve-or-wipe decision
// hangs on it.
expect(region).toContain('const measuredEmbeddings = await countPersistedEmbeddings();');
expect(region).toContain('await saveMeta(storagePath, embeddingMeta);');
// Ordering, without brittle character spans: flush → measure → decide →
// write. Counting before the flush would describe rows still in the WAL.
const flushed = region.lastIndexOf('await flushWAL();');
const measured = region.indexOf('const measuredEmbeddings = await countPersistedEmbeddings();');
const decided = region.indexOf('const outcome = resolveEmbedRunOutcome(');
const folded = region.indexOf('embeddingMeta = withMeasuredEmbeddingCount(', measured);
expect(flushed).toBeLessThan(measured);
expect(measured).toBeLessThan(decided);
expect(decided).toBeLessThan(folded);
expect(region.slice(folded)).toContain('measuredEmbeddings,');
});
it('measures in the post-flush checkpoint callback and nowhere else in the pipeline options', async () => {
const source = await readSource();
expect(source).toMatch(
/await saveEmbeddingCheckpoint\(\s*checkpoint,\s*\[\],\s*await countPersistedEmbeddings\(\),?\s*\)/,
);
// The window-start callback fires before any row exists — it must pass no
// count rather than restate a stale one.
expect(source).toMatch(
/onCheckpointWindowStart: async \(\{ nodeIds, \.\.\.checkpoint \}\) => \{\s*await saveEmbeddingCheckpoint\(checkpoint, nodeIds\);\s*\},/,
);
});
it('resolves a found checkpoint through the shared resume decision', async () => {
const source = await readSource();
const region = insideWithLbugDb(source);
// The route asks the SAME decider the CLI does, instead of hard-throwing on
// any identity mismatch and ignoring `attempts` — the disagreement that let
// a CLI-written `'partial'` marker wedge every later `POST /api/embed`.
expect(region).toMatch(/decideEmbeddingResume\(priorCheckpoint, embeddingIdentity\)/);
// Every action is routed: abort fails the run, abandon warns and proceeds
// with an empty pending set, resume hands the decision's ids to the pipeline.
expect(region).toContain("if (resume?.action === 'abort') throw new Error(resume.error);");
expect(region).toMatch(/resume\?\.action === 'resume'\s*\?\s*resume\.pendingNodeIds/);
// No second copy of the gate: the route no longer authors its own message.
expect(region).not.toContain('Cannot resume embedding checkpoint:');
});
it('never maps the pipeline ready phase to a phase a client can read as terminal', async () => {
const source = await readSource();
// `ready` fires unconditionally before the route knows the outcome (#2790).
expect(source).toMatch(/p\.phase === 'ready'\s*\?\s*'finalizing'/);
expect(source).not.toMatch(/p\.phase === 'ready' \? 'complete'/);
});
});
describe('HTTP repo catalog validation', () => {
const readSource = () =>
fs.readFile(path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'), 'utf-8');
it('lists and resolves repos with validate: true, and maps StorageRequirementError', async () => {
const source = await readSource();
expect(source).toMatch(/const repos = await listRegisteredRepos\(\{\s*validate:\s*true\s*\}\)/);
expect(source).toMatch(
/const freshRepos = await listRegisteredRepos\(\{\s*validate:\s*options\.validateStorage !== false,\s*\}\)/,
);
expect(source).toMatch(
/app\.get\('\/api\/repos'[\s\S]*listRegisteredRepos\(\{\s*validate:\s*true\s*\}\)/,
);
expect(source).toMatch(/sendStorageRequirementHttp\(err, res\)/);
expect(source).toMatch(/storageRequirementToHttp\(err\)/);
expect(source).toMatch(/code: 'index-unavailable'/);
});
});