mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-02 02:11:29 +00:00
* chore: ignore local Vercel link artifacts
Keep .vercel and env files out of the repo after a local preview link.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): repair citation chips, code panel line math, stale agent state
Audit findings in the web client, each verified against the source:
- RightPanel: citation chips ([[path:10-20]], [[Class:Foo]]) were wired to a
stub resolver that always returned null, so clicking any citation did
nothing. Expose resolveFilePath from useAppState and use it.
- CodeReferencesPanel: graph startLine/endLine are 1-based but were treated
as 0-based, so the highlighted range and scroll target were off by one
line; AI citation cards always rendered "code not available" because the
snippet loader was a stub. Fetch per-citation snippets via /api/file.
- useAppState: sendChatMessage read llmSettings.activeProvider outside its
deps (stale provider capabilities after switching provider);
initializeAgent trapped projectName at '' for callers without an override
(system prompt labelled the codebase "project"); the embeddings 409 dedup
matched a message the server never sends for same-repo jobs.
- tools.ts impact: for path targets every symbol defined in the file shares
the filePath, so the disambiguation always picked the first row and could
analyze an arbitrary symbol while reporting a file impact. Prefer the File
node.
- useSigma: the layout timeout called stop() but never kill(), leaking one
ForceAtlas2 Web Worker plus four graph listeners per completed layout.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): release repo lock on cancel, IPC-first worker cancel, route hardening
Audit findings in gitnexus/src, each verified against the source:
- analyze-launch: cancelJob marks the job failed before the worker exits,
so the exit handler's terminal early-return skipped releaseLockOnce and
the repo stayed locked ("Another job is already active") until restart.
Release on terminal exit and when forkWorker bails on a terminal job.
- analyze-job: cancellation now sends { type: 'cancel' } over IPC first and
signals only after a 15s grace. On Windows child.kill('SIGTERM') is a
forceful termination, so leading with it could kill the worker inside a
LadybugDB write. Mirrors core/auto-sync/analysis-worker-launch.
- api resolveRepo: a job that FAILED during the hold-queue wait fell through
to the { __timedOut } sentinel ("taking longer than expected") instead of
404; only /api/repo checked the sentinel, so graph/query/search/file/grep/
embed/delete crashed on entry.storagePath with a 500 after a 5 minute hang.
Return null on failed jobs and check the sentinel in every consumer.
- api processes/process/clusters/cluster: resolve ?repo= through the HTTP
resolver (documented policy on resolveRegisteredRepoEntry) and pass the
registered absolute path to the backend; add the standard rate limiter.
The raw param previously reached the MCP resolver, which runs a
cwd-relative realpathSync probe + registry refresh on a bare-name miss
and accepts unambiguous partial names.
- api body handling: Express 5 leaves req.body undefined without a JSON
content type, turning "Missing X" 400s into TypeError 500s; body-parser
4xx errors (malformed JSON, over-limit) were also reported as 500.
- /api/file: the lexical path.relative check cannot see symlinks; re-check
containment on realpath so a cloned repo containing evil -> /etc/passwd
cannot read outside the root.
- repo-manager unregisterRepo: used the lenient reader, so a transient read
error (EBUSY/EPERM racing another process's atomic rename) turned into
writing [] and deregistering every repo. Use the strict-if-present reader.
- clean --branch: compared registry paths with raw path.resolve instead of
the canonical registryPathEquals used everywhere else (macOS /private/var,
Windows short names / drive-letter case) and reported indexed branches as
not indexed.
Tests: cancelJob IPC-before-signal contract; /api/file symlink escape (403)
and in-repo symlink (200), skipped where the host cannot create symlinks.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore: keep example env files visible and ignore local Cursor config
.env* also hid gitnexus/.env.example and eval/.env.example. .vercel was already ignored. The web app's .cursor/ stays local, including its MCP file.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Add realtime execution ops dashboard for Vercel monitoring.
Expose /api/ops snapshots over the local serve process and a ?view=ops SPA panel so analyze/embed jobs can be watched live from the hosted web UI.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: address gitnexus-check review on cite/embed/resolve paths
Pass req into resolveRepo for process/cluster routes, tighten same-repo embed 409 handling, guard empty citation paths, fix snippet retry races, and drop the lone-File ambiguity fallback.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops): harden CORS, redact paths, and bound ops streams
Restrict Vercel CORS to exact production hosts, omit raw repo paths/URLs from the unauthenticated ops feed, rate-limit and cap SSE connections, and fix dashboard SSE/poll edge cases.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): stabilize citation fetches and file impact matching
Retry cancelled snippet loads without duplicate in-flight reads, cap range-less citation downloads, and make impact file matching unique-suffix-aware with a synthetic File target when LIMIT drops the File node.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web/ops): close gitnexus-check review threads on SSE and redaction
Cap citation reads, reconnect ops on applied server URL, skip SSE onError after abort, and strip URL query/fragment from public repoName.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops): abort SSE on poll fallback and harden repoName parsing
Prevent dual SSE+poll after a failed safety snapshot, skip overlapping poll ticks, ignore aborted streamSSE onError, and basename Windows drive-letter URLs so ops never leaks path segments.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ops/web): close gitnexus-check threads on SSE budget and credential leak
Keep finite SSE retries across short 200s, strip backend URL userinfo before ?server=, and redact progress messages on the public ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Restore 0-based GraphNode line math, redact public job poll/error fields, fix omit-?repo= 400, hold the analyze lock across cancel-during-settle, and restore the Vercel shared compile.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(cli): contain leftover-slot reclaim to the slot and keep --stale status honest
Preview and force now share one branches/ containment rule, nested junctions cannot walk a sibling index, and a mid-loop git failure no longer claims leftovers were not deleted after a successful rm.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(cli): share leftover-slot helpers without changing reclaim behavior
Pull the rolling I/O pool and owned-cwd storage lookup into one place so clean --stale/--branch and leftover listing stop restating the same ownership and concurrency paths.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* Address PR review feedback (#3348)
Keep the omitted-repo snapshot instead of re-listing, stop citation and ops races, redact full public repo URLs, and restore fake timers in teardown.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Store graph node citation lines as 0-based offsets and correct the default-port origin comment.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Redact public SSE progress text, stop citation append retries from
cancelling in-flight reads, and keep the ops dashboard from showing a
stale snapshot or clearing a failed Connect.
Note: pre-existing failure in incremental-index-extension-dml-gate and other lbug/env unit tests not addressed by this PR.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Prune citation snippets when AI refs are cleared, and poll /api/ops at 2s so the fallback stays under the 60/min limit.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ci): apply prettier class order for format check (#3348)
CI quality/format runs root-only npm ci, so prettier-plugin-tailwindcss
sorts scrollbar-thin without the web Tailwind catalog.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
- Require a unique suffix match for graph-backed citation paths so
ambiguous names like index.ts no longer open the first graph file.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
- Require a path-component boundary so unique citation suffixes cannot match filename substrings like myindex.ts
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): redact filesystem paths in ops text
Unauthenticated /api/ops and poll replay worker errors. URLs were
scrubbed but home-directory and Windows paths still leaked.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): omit repoPath from public SSE frames
/api/ops lists job ids, so the unauthenticated progress stream
must not replay the analyzed filesystem path.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): treat embed lock 409 as a busy error
Analyze and embed share the same lock string. Mapping that 409 to
embedding hid an in-flight analyze as a successful embed start.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): bound impact File path suffix matches
Unbounded endsWith let lib/foo.ts select src/mylib/foo.ts. Require
an exact path or a unique /suffix, matching resolveUniqueIndexedPath.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): keep canceled analyze slot until exit
Marking failed before the worker exited let a second POST start
cloneOrPull against a LadybugDB file still being written.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): skip Windows SIGTERM on job dispose
child.kill('SIGTERM') is TerminateProcess there. Ask over IPC first
and leave the 15s grace timer to SIGKILL.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): let process routes skip the analyze hold
GET /api/processes and /api/clusters always waited up to 300s.
?awaitAnalysis=false fails fast; default still waits like /api/repo.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(web): cover public ops and analyze SSE user flows
Lock the unauthenticated dashboard and analyze complete/fail/cancel paths so a leaked repoPath, token, or home path cannot ship unnoticed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Keep caller cancel reasons over the worker's generic IPC, skip publish while cancel is pending, and redact scp-style remotes on the public ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address remaining PR review feedback (#3348)
Release the analyze slot when a worker fails to spawn, and omit branch refs from the unauthenticated ops feed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): do not reuse an analyze job that is pending cancel
A dying same-repo job still occupies the single slot; 202-reuse would
attach a new client to a cancel in flight.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): hold the analyze lock until the worker exits after cancel
Cancel error IPC used to drop the repo lock while the child was still
checkpointing. Abort settle immediately on pending cancel so the slot
is not held for a 60s disk poll.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): redact known repo paths with spaces in public ops text
Known repoPath/repoUrl literals are replaced first so a clone dir with
spaces cannot leak past the whitespace-bounded path regex.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(server): return the live job on analyze and embed DELETE
Hard-coding failed made clients retry immediately and 409 while the
child still occupied the slot. resolveRepo now returns not-found as
soon as that job fails instead of waiting out the hold timeout.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(web): drive public analyze and ops through a live gitnexus serve
Spawn the real backend and observe requests instead of intercepting
them, so slot occupancy, redaction, and reconnect stay honest.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(web): reuse code-panel helpers and drop dead UI aliases
Citation fetches already had selectedNodeFileRange and snippetRepoKey;
the impact File suffix filter already handled exact paths.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* Address PR review feedback (#3348)
- Skip createJob reuse after cancel IPC is consumed while the child remains
- Hold the repo lock until exit when complete IPC races a pending cancel
- Scrub full remote URLs before known repoUrl prefixes in public ops text
- Reject unique impact File suffix matches from a truncated LIMIT 10 page
- Make live e2e helpers bound probes, clean up failed startups, and wait out the cancel slot
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): open live e2e pages on the Vite host CI actually bound
Absolute 127.0.0.1:5173 navigation refused on Actions because wait-on
and Vite use localhost (often ::1). Honor FRONTEND_URL when set, else
pick the first of localhost / 127.0.0.1 that answers.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3348)
Hold the analyze lock until worker exit when cancel aborts settle, and assert GitLab failure chrome does not leak host or path.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(web): keep live analyze e2e under the analyze rate limit
POST /api/analyze allows 10 requests per minute per IP. The slot-free
helper re-POSTed every 400ms while a cancelled worker was exiting, spent
that budget, and the lock test's hold request got 429 instead of 202.
- postAnalyze waits out a 429 using the RateLimit reset and retries
- slot polling backs off to 2s and leaves a small POST budget for callers
- the slot probe is a clone that fails before any worker fork, so the
probe itself no longer holds the slot after reporting failed
- the lock test holds the slot with a real local analyze
- token and GitLab tests wait for a free slot before posting from the UI
- request fetches carry a timeout; teardown signals the serve process group
- an empty FRONTEND_URL falls back to the default base URL
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- ops view: a `?server=` link no longer auto-connects to another origin
while a deploy token is held; it prefills and waits for Connect
- analyze completion: a local-path run reconnects by the path this client
submitted, so duplicate basenames stay collision-safe without repoPath
on the public SSE frame
- stale slot cleanup: revalidate each nested directory (lstat + realpath)
right before readdir, so a mid-cleanup junction swap aborts instead of
walking an outside tree; list phases run sequentially so the slot-I/O
cap is global
- e2e: 429 backoff honours the caller deadline; the unreachable-backend
ops test navigates through the resolved frontend URL
- drop the unused `repo` member from the clean integration helper
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(server): reconnect after analyze by an opaque repo id
Public job views and the SSE terminal frame no longer carry repoPath, and
repoName is not unique, so a post-analyze reconnect by name could load a
same-named sibling. The server now issues `repoId`: an HMAC of the
canonical registry path under a per-process random key. It is set on
complete jobs (ops view, analyze poll, SSE terminal frame) and matches
the new `id` on `GET /api/repos` entries. The web client resolves it to
the exact entry path on completion; unknown ids fall back to the name.
This covers URL clones and folder uploads, and replaces the local-path
only fallback.
With reconnect off the label, public `repoName` for a branch-pinned URL
clone is the repository name, not the `<repo>__<branch slug>` registry
name, so the requested branch stays off the unauthenticated feed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- stale slot cleanup: a descendant that vanishes before its unlink is
treated as removed instead of aborting the reclaim
- e2e teardown: escalate to SIGKILL on the process group when the live
backend ignores SIGTERM for 5s
- ops view: drop the dead initial value CodeQL flagged
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- public redaction: a known repoPath now also consumes its descendant
tail, so `<repoPath>/src/secret.ts` becomes `[path]` instead of
`[path]/src/secret.ts`; a same-prefix sibling is left to the path scrub
- e2e: the cancel test waits for the analyze slot the previous failed
local-path job still holds; `fetchOps` carries the request timeout
- docs: SSE terminal payload and RepoAnalyzer `onComplete` describe
`repoId` resolution
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address PR review feedback (#3348)
- RepoAnalyzer: drop a completion that resolves after unmount, so a slow
/api/repos lookup cannot switch repos after the sheet was dismissed
- e2e: select the local-path input by test id (the placeholder differs on
Windows); the slot probe's job wait honours the caller's deadline
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(server): assert the public SSE terminal frame in analyze-api
The #2790 terminality tests still expected `repoPath` on the terminal
frame. This PR replaced it with the opaque `repoId`, so assert that shape
and that the analyzed path never appears in the stream.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(web): wait for the analyze slot between duplicate-repo setup runs
The server now keeps the single analyze slot until the worker exits, even
after its job reports complete. repo-path-identity posted the second
duplicate's analyze immediately and got 409 in CI. Use the shared
slot-aware POST, which also waits out a 429.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
541 lines
20 KiB
TypeScript
541 lines
20 KiB
TypeScript
/**
|
|
* `createLaunchAnalysisWorker`'s collapsed-index guard — ORDERING, not just status.
|
|
*
|
|
* `backend.init()` is the PUBLISH step (it is `LocalBackend.refreshRepos()`,
|
|
* which swaps the freshly-registered repo into the in-memory map every MCP tool
|
|
* and HTTP route resolves through). The guard added in #2899 read
|
|
* `graphWriteCollapsed` only AFTER that call had already resolved, so a
|
|
* known-incomplete database was live and queryable before the job was ever
|
|
* marked `failed` — the job status was a label on a published index rather than
|
|
* a gate. These tests pin the order, because the order is the defect.
|
|
*
|
|
* `analyze-launch.ts` had ZERO test coverage before this file, which is why a
|
|
* field-name drift against `analyze-worker-ipc.ts`'s wire shape would have made
|
|
* the branch permanently dead and silently restored the pre-guard behaviour.
|
|
* The worker messages below are therefore built by calling the PRODUCTION
|
|
* projection `projectAnalyzeResultForIpc` rather than hand-rolling a literal, so
|
|
* a rename of `graphWriteCollapsed` breaks these tests instead of disabling the
|
|
* branch they cover.
|
|
*/
|
|
import { afterEach, beforeEach, describe, expect, it, vi, type Mock } from 'vitest';
|
|
import { EventEmitter } from 'node:events';
|
|
|
|
// `vi.mock` factories are hoisted above every top-level `const`, and this file
|
|
// imports the module under test statically — so anything a factory closes over
|
|
// must be hoisted with it.
|
|
const H = vi.hoisted(() => {
|
|
const STORAGE_PATH = '/tmp/gitnexus-test-storage';
|
|
return {
|
|
forkMock: vi.fn(),
|
|
STORAGE_PATH,
|
|
REPO_PATH: '/tmp/gitnexus-test-repo',
|
|
METADATA_FILE: 'gitnexus.json',
|
|
// When false, the finalization gate sees no fresh index (timeout / lock-hold tests).
|
|
settleOk: true,
|
|
requireStoragePath: vi.fn(async () => STORAGE_PATH),
|
|
};
|
|
});
|
|
const { forkMock, REPO_PATH } = H;
|
|
|
|
vi.mock('child_process', async () => {
|
|
const actual = await vi.importActual<typeof import('child_process')>('child_process');
|
|
return { ...actual, fork: H.forkMock };
|
|
});
|
|
|
|
// The launcher's finalization gate (`waitForSettledIndex`) probes the
|
|
// ownership-validated storage path. Pin the filesystem so the gate settles on
|
|
// its FIRST poll — the gate itself is not under test here and its 200ms poll
|
|
// would otherwise put a real timer between the worker message and the assertions.
|
|
vi.mock('../../src/storage/repo-manager.js', () => ({
|
|
INDEX_METADATA_FILE: H.METADATA_FILE,
|
|
}));
|
|
|
|
vi.mock('../../src/storage/storage-resolver.js', () => ({
|
|
ANALYZE_STORAGE_REQUIREMENTS: { allowedStates: ['missing', 'empty', 'owned'] },
|
|
ANALYZE_FORCE_STORAGE_REQUIREMENTS: {
|
|
allowedStates: ['missing', 'empty', 'owned', 'unowned', 'foreign'],
|
|
},
|
|
requireStoragePath: H.requireStoragePath,
|
|
}));
|
|
|
|
vi.mock('node:fs', async () => {
|
|
const actual = await vi.importActual<typeof import('node:fs')>('node:fs');
|
|
return {
|
|
...actual,
|
|
statSync: () => {
|
|
if (!H.settleOk) {
|
|
throw Object.assign(new Error('ENOENT'), { code: 'ENOENT' });
|
|
}
|
|
// Both index files were (re)written far in the future relative to jobStartMs.
|
|
return { mtimeMs: Number.MAX_SAFE_INTEGER };
|
|
},
|
|
// No WAL/shadow/checkpoint sidecar remains when the gate is allowed to settle.
|
|
existsSync: () => false,
|
|
};
|
|
});
|
|
|
|
import { createLaunchAnalysisWorker } from '../../src/server/analyze-launch.js';
|
|
import { JobManager } from '../../src/server/analyze-job.js';
|
|
import { projectAnalyzeResultForIpc } from '../../src/server/analyze-worker-ipc.js';
|
|
import type { AnalyzeResult } from '../../src/core/run-analyze.js';
|
|
import type { CompleteMessage } from '../../src/server/analyze-worker.js';
|
|
|
|
const REPO_NAME = 'collapse-fixture';
|
|
|
|
/**
|
|
* Build the exact `complete` message the worker puts on the wire, by running the
|
|
* production projection. The `graphWriteCollapsed` key is therefore whatever
|
|
* `analyze-worker-ipc.ts` actually sends — not a literal this test invented.
|
|
*/
|
|
const completeMessage = (graphWriteCollapsed?: { expected: number; persisted: number }) => {
|
|
const result = {
|
|
repoName: REPO_NAME,
|
|
repoPath: REPO_PATH,
|
|
storagePath: H.STORAGE_PATH,
|
|
stats: { files: 10, nodes: 100, edges: 500 },
|
|
...(graphWriteCollapsed ? { graphWriteCollapsed } : {}),
|
|
} satisfies Partial<AnalyzeResult> as AnalyzeResult;
|
|
return { type: 'complete', result: projectAnalyzeResultForIpc(result) } satisfies CompleteMessage;
|
|
};
|
|
|
|
interface FakeChild extends EventEmitter {
|
|
stderr: EventEmitter;
|
|
send: Mock<(msg: unknown) => boolean>;
|
|
kill: Mock<(signal?: NodeJS.Signals) => boolean>;
|
|
pid?: number;
|
|
}
|
|
|
|
const makeChild = (): FakeChild => {
|
|
const child = new EventEmitter() as FakeChild;
|
|
child.stderr = new EventEmitter();
|
|
child.send = vi.fn();
|
|
child.kill = vi.fn();
|
|
return child;
|
|
};
|
|
|
|
describe('createLaunchAnalysisWorker — collapsed index is never published', () => {
|
|
let jobManager: JobManager;
|
|
let child: FakeChild;
|
|
let calls: string[];
|
|
let backendInit: Mock<() => Promise<unknown>>;
|
|
let closeDbHandle: Mock<() => Promise<void>>;
|
|
|
|
/** Drive one analyze to its terminal state and return the observed call order. */
|
|
const runWorker = async (msg: CompleteMessage) => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {
|
|
calls.push('releaseRepoLock');
|
|
},
|
|
closeDbHandle,
|
|
});
|
|
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
child.emit('message', msg);
|
|
|
|
await vi.waitFor(() => expect(calls).toContain('updateJob:terminal'));
|
|
return jobManager.getJob(job.id);
|
|
};
|
|
|
|
beforeEach(() => {
|
|
calls = [];
|
|
H.settleOk = true;
|
|
H.requireStoragePath.mockClear();
|
|
jobManager = new JobManager();
|
|
child = makeChild();
|
|
forkMock.mockImplementation(() => child);
|
|
|
|
backendInit = vi.fn(async () => {
|
|
calls.push('backend.init');
|
|
return true;
|
|
});
|
|
closeDbHandle = vi.fn(async () => {
|
|
calls.push('closeDbHandle');
|
|
});
|
|
|
|
const realUpdate = jobManager.updateJob.bind(jobManager);
|
|
vi.spyOn(jobManager, 'updateJob').mockImplementation((id, update) => {
|
|
calls.push(`updateJob:${update.status ?? 'progress'}`);
|
|
realUpdate(id, update);
|
|
// Recorded after the real call so the marker only lands once the status is
|
|
// committed — `updateJob` drops any update to an already-terminal job.
|
|
calls.push(
|
|
...['complete', 'failed']
|
|
.filter((s) => s === update.status)
|
|
.map(() => 'updateJob:terminal'),
|
|
);
|
|
});
|
|
});
|
|
|
|
it('forwards the Spring Actuator snapshot path to the worker', async () => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {},
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
|
|
await launch(job, REPO_PATH, { springActuatorPath: 'runtime/actuator' });
|
|
|
|
expect(child.send).toHaveBeenCalledWith(
|
|
expect.objectContaining({
|
|
options: expect.objectContaining({
|
|
springActuatorPath: 'runtime/actuator',
|
|
}),
|
|
}),
|
|
);
|
|
});
|
|
|
|
it('uses the force storage set only when launch options request force', async () => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {},
|
|
closeDbHandle,
|
|
});
|
|
|
|
const ordinary = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(ordinary, REPO_PATH, {});
|
|
expect(H.requireStoragePath).toHaveBeenLastCalledWith(REPO_PATH, {
|
|
allowedStates: ['missing', 'empty', 'owned'],
|
|
});
|
|
|
|
const forced = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(forced, REPO_PATH, { force: true });
|
|
expect(H.requireStoragePath).toHaveBeenLastCalledWith(REPO_PATH, {
|
|
allowedStates: ['missing', 'empty', 'owned', 'unowned', 'foreign'],
|
|
});
|
|
});
|
|
|
|
it('forwards the index-branch selector to the worker', async () => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {},
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
|
|
await launch(job, REPO_PATH, { branch: 'development' });
|
|
|
|
// `StartMessage.options` is typed as `AnalyzeOptions`, so this key IS
|
|
// `AnalyzeOptions.branch` — the field `resolveWriteTarget` reads to choose
|
|
// the run's storage slot. (It does not always mean a `branches/<slug>/`
|
|
// sub-slot: `resolveBranchPlacement` keeps the flat slot when that slot has
|
|
// no owner, or when its owner is already this label.) A rename breaks this
|
|
// test.
|
|
expect(child.send).toHaveBeenCalledWith(
|
|
expect.objectContaining({
|
|
options: expect.objectContaining({ branch: 'development' }),
|
|
}),
|
|
);
|
|
});
|
|
|
|
it('omits branch entirely when the caller did not select one', async () => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {},
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
|
|
await launch(job, REPO_PATH, {});
|
|
|
|
// Not merely undefined: absent. `AnalyzeOptions.branch === undefined` is the
|
|
// documented signal for "target the flat workspace slot", so sending the key
|
|
// with an undefined value must not become the way that default is expressed.
|
|
const sent = child.send.mock.calls.at(0)?.[0] as { options: Record<string, unknown> };
|
|
expect(Object.hasOwn(sent.options, 'branch')).toBe(false);
|
|
});
|
|
|
|
afterEach(() => {
|
|
vi.useRealTimers();
|
|
jobManager.dispose();
|
|
vi.restoreAllMocks();
|
|
forkMock.mockReset();
|
|
H.settleOk = true;
|
|
});
|
|
|
|
it('does not publish the index — backend.init() is never called for a collapsed run', async () => {
|
|
await runWorker(completeMessage({ expected: 500, persisted: 3 }));
|
|
|
|
// The defect: init() resolved FIRST, so the incomplete graph was live and
|
|
// queryable by every MCP/API consumer before the job was marked failed.
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(calls).not.toContain('backend.init');
|
|
// The cached handle is still evicted — the worker rewrote the DB files on
|
|
// disk, so a pre-rewrite handle is stale whatever the outcome was. Eviction
|
|
// is not publication.
|
|
expect(closeDbHandle).toHaveBeenCalledTimes(1);
|
|
expect(calls.indexOf('closeDbHandle')).toBeLessThan(calls.indexOf('updateJob:failed'));
|
|
// Lock is held through the collapse decision and dropped afterwards, once.
|
|
expect(calls.indexOf('updateJob:failed')).toBeLessThan(calls.indexOf('releaseRepoLock'));
|
|
});
|
|
|
|
it('marks the collapsed run failed and still reports repoName', async () => {
|
|
const job = await runWorker(completeMessage({ expected: 500, persisted: 3 }));
|
|
|
|
expect(job?.status).toBe('failed');
|
|
// The success path sets repoName; api.ts's repo-resolution wait matches jobs
|
|
// on it first. Dropping it here cost one of three match keys for no reason.
|
|
expect(job?.repoName).toBe(REPO_NAME);
|
|
expect(job?.error).toContain('INCOMPLETELY');
|
|
expect(job?.error).toContain('3 of 500');
|
|
// The failure is explicit about the index being unreachable, not merely stale.
|
|
expect(job?.error).toContain('NOT published');
|
|
});
|
|
|
|
it('publishes and completes a healthy run, in that order', async () => {
|
|
const job = await runWorker(completeMessage());
|
|
|
|
expect(job?.status).toBe('complete');
|
|
expect(job?.repoName).toBe(REPO_NAME);
|
|
expect(backendInit).toHaveBeenCalledTimes(1);
|
|
// Publish strictly BEFORE the terminal complete, so the repo really is
|
|
// queryable when the client receives the SSE complete event.
|
|
expect(calls).toEqual([
|
|
'updateJob:analyzing',
|
|
'closeDbHandle',
|
|
'backend.init',
|
|
'updateJob:complete',
|
|
'updateJob:terminal',
|
|
'releaseRepoLock',
|
|
]);
|
|
});
|
|
|
|
it('reads the collapse flag under the name analyze-worker-ipc.ts actually sends', async () => {
|
|
const wire = completeMessage({ expected: 500, persisted: 3 });
|
|
|
|
// Guards against a silent rename: the branch under test keys off this exact
|
|
// field, and the message was produced by the production projection.
|
|
expect(Object.keys(wire.result)).toContain('graphWriteCollapsed');
|
|
expect(wire.result.graphWriteCollapsed).toEqual({ expected: 500, persisted: 3 });
|
|
|
|
// A projection that stopped carrying the field must not read as healthy.
|
|
const healthy = completeMessage();
|
|
expect(healthy.result.graphWriteCollapsed).toBeUndefined();
|
|
const job = await runWorker(healthy);
|
|
expect(job?.status).toBe('complete');
|
|
});
|
|
|
|
it('fails and does not publish when index finalization never becomes visible', async () => {
|
|
vi.useFakeTimers();
|
|
H.settleOk = false;
|
|
const releaseRepoLock = vi.fn(() => {
|
|
calls.push('releaseRepoLock');
|
|
});
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock,
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
child.emit('message', completeMessage());
|
|
|
|
expect(releaseRepoLock).not.toHaveBeenCalled();
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(jobManager.getJob(job.id)?.status).toBe('analyzing');
|
|
|
|
// Must match FINALIZE_SETTLE_TIMEOUT_MS + one poll in analyze-launch.ts.
|
|
await vi.advanceTimersByTimeAsync(61_000);
|
|
|
|
const done = jobManager.getJob(job.id);
|
|
expect(done?.status).toBe('failed');
|
|
expect(done?.error).toMatch(/finalization not visible after timeout/i);
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(closeDbHandle).not.toHaveBeenCalled();
|
|
expect(releaseRepoLock).toHaveBeenCalledTimes(1);
|
|
});
|
|
|
|
it('holds the write lock until settle resolves, then releases once after publish', async () => {
|
|
vi.useFakeTimers();
|
|
H.settleOk = false;
|
|
const releaseRepoLock = vi.fn(() => {
|
|
calls.push('releaseRepoLock');
|
|
});
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock,
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
child.emit('message', completeMessage());
|
|
|
|
// First poll failed; the gate is sleeping. Lock must still be held.
|
|
expect(releaseRepoLock).not.toHaveBeenCalled();
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(jobManager.getJob(job.id)?.status).toBe('analyzing');
|
|
|
|
H.settleOk = true;
|
|
await vi.advanceTimersByTimeAsync(200);
|
|
|
|
expect(jobManager.getJob(job.id)?.status).toBe('complete');
|
|
expect(backendInit).toHaveBeenCalledTimes(1);
|
|
expect(releaseRepoLock).toHaveBeenCalledTimes(1);
|
|
expect(calls.indexOf('backend.init')).toBeLessThan(calls.indexOf('releaseRepoLock'));
|
|
});
|
|
});
|
|
|
|
describe('createLaunchAnalysisWorker — pending cancel', () => {
|
|
let jobManager: JobManager;
|
|
let child: FakeChild;
|
|
let backendInit: Mock<() => Promise<unknown>>;
|
|
let closeDbHandle: Mock<() => Promise<void>>;
|
|
|
|
const launchOne = async () => {
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock: () => {},
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
return job;
|
|
};
|
|
|
|
beforeEach(() => {
|
|
H.settleOk = true;
|
|
jobManager = new JobManager();
|
|
child = makeChild();
|
|
forkMock.mockImplementation(() => child);
|
|
backendInit = vi.fn(async () => true);
|
|
closeDbHandle = vi.fn(async () => {});
|
|
});
|
|
|
|
afterEach(() => {
|
|
vi.useRealTimers();
|
|
jobManager.dispose();
|
|
vi.restoreAllMocks();
|
|
forkMock.mockReset();
|
|
H.settleOk = true;
|
|
});
|
|
|
|
it('keeps the caller cancel reason when the worker reports a generic cancel error', async () => {
|
|
const job = await launchOne();
|
|
expect(jobManager.cancelJob(job.id, 'Analysis timed out (30 minute limit)')).toBe(true);
|
|
child.emit('message', {
|
|
type: 'error',
|
|
message: 'Analysis cancelled (parent requested cancellation)',
|
|
});
|
|
const done = jobManager.getJob(job.id);
|
|
expect(done?.status).toBe('failed');
|
|
expect(done?.error).toBe('Analysis timed out (30 minute limit)');
|
|
});
|
|
|
|
it('does not publish after complete IPC if cancel is already pending', async () => {
|
|
const job = await launchOne();
|
|
expect(jobManager.cancelJob(job.id, 'Cancelled by user')).toBe(true);
|
|
child.emit('message', completeMessage());
|
|
await vi.waitFor(() => expect(jobManager.getJob(job.id)?.status).toBe('failed'));
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(jobManager.getJob(job.id)?.error).toBe('Cancelled by user');
|
|
});
|
|
|
|
it('does not publish if cancel arrives while settle is in flight', async () => {
|
|
vi.useFakeTimers();
|
|
H.settleOk = false;
|
|
const job = await launchOne();
|
|
child.emit('message', completeMessage());
|
|
expect(jobManager.getJob(job.id)?.status).toBe('analyzing');
|
|
expect(jobManager.cancelJob(job.id, 'Cancelled by user')).toBe(true);
|
|
H.settleOk = true;
|
|
await vi.advanceTimersByTimeAsync(200);
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(jobManager.getJob(job.id)?.status).toBe('failed');
|
|
expect(jobManager.getJob(job.id)?.error).toBe('Cancelled by user');
|
|
});
|
|
|
|
it('releases the analyze slot when the worker emits error without exit', async () => {
|
|
const job = await launchOne();
|
|
child.emit('error', new Error('spawn ENOENT'));
|
|
expect(jobManager.getJob(job.id)?.status).toBe('failed');
|
|
expect(jobManager.createJob({ repoPath: '/tmp/other' }).status).toBe('queued');
|
|
});
|
|
|
|
it('holds the repo lock until exit after complete IPC when cancel is already pending', async () => {
|
|
const releaseRepoLock = vi.fn();
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock,
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
|
|
expect(jobManager.cancelJob(job.id, 'Cancelled by user')).toBe(true);
|
|
child.emit('message', completeMessage());
|
|
|
|
expect(jobManager.getJob(job.id)?.status).toBe('failed');
|
|
expect(jobManager.getJob(job.id)?.error).toBe('Cancelled by user');
|
|
expect(backendInit).not.toHaveBeenCalled();
|
|
expect(releaseRepoLock).not.toHaveBeenCalled();
|
|
expect(() => jobManager.createJob({ repoPath: '/tmp/other' })).toThrow(/already in progress/);
|
|
|
|
child.emit('exit', 0);
|
|
|
|
expect(releaseRepoLock).toHaveBeenCalledTimes(1);
|
|
expect(jobManager.createJob({ repoPath: '/tmp/other' }).status).toBe('queued');
|
|
});
|
|
|
|
it('holds the repo lock until exit after cancel error IPC', async () => {
|
|
const releaseRepoLock = vi.fn();
|
|
const launch = createLaunchAnalysisWorker({
|
|
jobManager,
|
|
backend: { init: backendInit },
|
|
acquireRepoLock: () => null,
|
|
releaseRepoLock,
|
|
closeDbHandle,
|
|
});
|
|
const job = jobManager.createJob({ repoPath: REPO_PATH });
|
|
await launch(job, REPO_PATH, {});
|
|
|
|
expect(jobManager.cancelJob(job.id, 'Cancelled by user')).toBe(true);
|
|
child.emit('message', {
|
|
type: 'error',
|
|
message: 'Analysis cancelled (parent requested cancellation)',
|
|
});
|
|
|
|
expect(jobManager.getJob(job.id)?.status).toBe('failed');
|
|
expect(jobManager.getJob(job.id)?.error).toBe('Cancelled by user');
|
|
expect(releaseRepoLock).not.toHaveBeenCalled();
|
|
expect(() => jobManager.createJob({ repoPath: '/tmp/other' })).toThrow(/already in progress/);
|
|
|
|
child.emit('exit', 0);
|
|
|
|
expect(releaseRepoLock).toHaveBeenCalledTimes(1);
|
|
expect(jobManager.createJob({ repoPath: '/tmp/other' }).status).toBe('queued');
|
|
});
|
|
|
|
it('does not release the analyze slot on post-spawn child error until exit', async () => {
|
|
const job = await launchOne();
|
|
child.pid = 123;
|
|
child.emit('error', new Error('write EPIPE'));
|
|
|
|
expect(jobManager.getJob(job.id)?.status).toBe('failed');
|
|
expect(jobManager.getJob(job.id)?.error).toMatch(/write EPIPE/);
|
|
expect(() => jobManager.createJob({ repoPath: '/tmp/other' })).toThrow(/already in progress/);
|
|
|
|
child.emit('exit', 1);
|
|
|
|
expect(jobManager.createJob({ repoPath: '/tmp/other' }).status).toBe('queued');
|
|
});
|
|
});
|