GitNexus/gitnexus/test/unit/lbug-pool-forced-refusal-heal.test.ts
Nguyễn Đăng Minh Lực 5be80b1fc3
Some checks are pending
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Gitleaks / gitleaks (push) Waiting to run
Publish / Classify release event (push) Waiting to run
Publish / RC guard (marker + release-PR skip) (push) Blocked by required conditions
Publish / ci (push) Blocked by required conditions
Publish / Publish to npm (push) Blocked by required conditions
Publish / Build & Push RC Docker images (push) Blocked by required conditions
Scorecard / Scorecard analysis (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-cli) (push) Waiting to run
Trivy Image Scan / Trivy (gitnexus-web) (push) Waiting to run
fix(lbug): self-heal read-only opens refused by an interrupted checkpoint (#3340)
* fix(lbug): self-heal read-only opens refused by an interrupted checkpoint

Homelab repro 2026-09-19 (image 1.6.10-20260917, @ladybugdb/core 0.19.x):
a wiki pod killed mid-CHECKPOINT left the engine's checkpoint artifacts on
disk (lbug.wal.checkpoint / lbug.shadow / checkpoint intent+apply locks),
and every later READ-ONLY open refused with 'Cannot open database in
read-only mode while checkpoint is in progress' — permanently, until a
writable open (any gitnexus analyze) happened to run.

Two defects, both fixed:

1. The refusal was unclassified. ensureReadOnlyConnectionUsable (direct
   adapter) and openReadOnlyDatabase (pool) recovered missing-shadow and
   shadow-replay errors but rethrew this one raw, so 'LadybugDB
   unavailable for __wiki__' repeated forever. Add
   isReadOnlyCheckpointInProgressError (LADYBUGDB-CONTRACT, live-verified
   against 0.19.1) and route it through the same writable-open recovery —
   including at OPEN time, where the refusal fires before any probe can
   run (pool: init() moved inside the try; direct: doInitLbug catch).

2. The existing shadow-replay recovery was not durable. Reproduction
   matrix on 0.19.1: the writable probe replays the WAL in MEMORY only —
   without an explicit CHECKPOINT the engine drops the pages at close and
   the follow-up read-only open silently serves the pre-checkpoint state.
   Both recovery paths now CHECKPOINT after the probe, which applies the
   replay, consumes the sidecars, and clears the checkpoint locks.

Verified end-to-end: real engine 0.19.1, killed-mid-CHECKPOINT state →
exact refusal → pool adapter self-heals → 80,800 rows intact, sidecars
consumed. The planted-signature integration test runs on any engine
version (refusal asserted only where the engine emits it, 0.19+).

* docs(architecture): list checkpoint-in-flight artifacts and the read-path self-heal

* fix(lbug): address review findings — shared cursor closer, tighten structural guard

Both findings from the gitnexus-check review on this PR:

1. pool-adapter.ts: the recovery CHECKPOINT closed its cursor with a bare
   unawaited result.close?.(). Use the shared best-effort closer
   (closeQueryResults) the repo already funnels both adapters through, so a
   cursor-close failure stays cleanup and cannot escape as an unhandled
   rejection.

2. sidecar-recovery.test.ts: the structural regex omitted the leading
   negation, so as a substring match it also accepted the inverted
   predicate (recover ONLY shadow-replay, exclude checkpoint) — the exact
   regression the guard exists to prevent. Pin the full
   '!isReadOnlyShadowReplayError(err) && !isReadOnlyCheckpointInProgressError(err)'
   throw-through shape.

* test(lbug): wire the recovery plant into lbug-db/LBUG_NATIVE; force the refusal on any engine pin

Address the tri-review findings (all four):

P1 — the planted-signature integration test was collected by the parallel
'default' project and never by the serialized 'lbug-db' project, and
Windows/macOS CI never ran it: register it next to its sibling in
vitest.config.ts (lbug-db include + default exclude) and in
cross-platform-tests.ts LBUG_NATIVE, per TESTING.md's rule for native
@ladybugdb/core suites. Verified via 'vitest list --project lbug-db'.

P2 — on the committed 0.18.3 pin the plant passes as 'pool opens and
count(n)=300' without ever exercising the new classifier or recovery
CHECKPOINT. Add forced-refusal behavioral tests for BOTH adapters: a mocked
native layer whose first read-only Database refuses with the canonical
0.19 message, asserting the exact self-heal shape (ro-refused -> writable
open -> CHECKPOINT -> ro retry) and that a healthy db never triggers a
writable open. The direct adapter is lazy, so its refusal is scripted at
the first probe query rather than init().

P2 — the LADYBUGDB-CONTRACT header on isReadOnlyCheckpointInProgressError
claimed '^0.18.0' like its siblings; only 0.19.x emits this string (0.18.3
tolerates the state). State the first-observed version so a bump reviewer
validates the right binary.

P2 — the pool's writable replay recovery quarantined on missing-shadow
even when the error came from the post-probe CHECKPOINT, where the main
file has already changed: mirror the direct adapter's probeSucceeded guard
(replaySucceeded) and fail closed instead of parking a live sidecar.

Also assert walBuffer.byteLength > 0 in the plant so the fixture cannot
silently degrade into an empty shell on tolerant engines.

* test(lbug): fix hosted-CI failures — Windows handle release, version-gated plant, prettier

Address the CHANGES_REQUESTED review of the hosted run:

Windows blocker (Win32 Error 33, locked file region at the pooled reopen):
the fixture now makes handle release explicit — waitForFixtureRelease
probe-reads the db and its residual WAL with bounded retries after every
native close (plant, raw refusal probe, pool close), mirroring the
adapter's own Windows handle-release probing. The WAL is re-planted from
the captured bytes instead of renamed: a close-time auto-checkpoint can
consume the live .wal out from under the rename (ENOENT, second flake).

Version-gate the plant itself: on < 0.19 engines that tolerate the planted
signature, its synthetic sidecars are not a consistent staging state for
the old engine (double-apply replays surfaced as 'Person already exists in
catalog' — the third flake), and no refusal can be forced there anyway.
The plant now skips below 0.19 with that rationale in-file; the behavioral
contract on every pin stays with the forced-refusal units, and this native
suite remains registered in lbug-db + LBUG_NATIVE so Win/macOS exercise it
as soon as the pin moves off 0.18.3.

Also: prettier on the two flagged files (format gate).

* fix(lbug): keep failed checkpoint heals from re-entering CHECKPOINT

Wrap recovery failures without repeating native refusal text, skip already-wrapped errors in doInitLbug, and pin constructor-time heal plus the Windows reopen skip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3340)

- Clean each forced-heal tmpDir in afterEach so earlier cases do not leak
- Compare major.minor when version-gating the interrupted-checkpoint plant

Co-authored-by: Cursor <cursoragent@cursor.com>

* Address PR review feedback (#3340)

Reset the pool forced-refusal mock Database sequence in native.reset() so later tests can still script the first construction as the checkpoint victim.

* Address PR review feedback (#3340)

Count every MATCH on the pool forced-refusal mock and assert the writable replay probe so CHECKPOINT-without-probe cannot stay green.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 19:18:16 +01:00

151 lines
5.5 KiB
TypeScript

/**
* Behavioral recovery test — POOL adapter, refusal FORCED on any engine pin.
*
* The integration plant (`lbug-interrupted-checkpoint-recovery.test.ts`) only
* exercises the refusal on engines that emit it (0.19+); on the committed
* 0.18.3 pin CI can stay green with the new classifier and recovery CHECKPOINT
* deleted (review finding: "0.18.3 CI never forces the refusal"). This suite
* pins the BEHAVIOR behind a mocked native layer instead: the read-only open
* throws the canonical 0.19 refusal, and the pool must run the full
* self-heal — writable open, probe, CHECKPOINT, read-only retry — regardless
* of which engine binary is installed.
*/
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import fs from 'node:fs/promises';
import os from 'node:os';
import path from 'node:path';
const REFUSAL =
'Connection exception: Cannot open database in read-only mode while checkpoint is in progress. Please retry later.';
const { native } = vi.hoisted(() => {
const native = {
calls: {
/** Every Database construction in order — the self-heal shape. */
constructions: [] as Array<'ro' | 'rw'>,
checkpoints: 0,
/** Every MATCH — victim, writable replay probe, and post-heal retry. */
probes: 0,
},
/** Script the FIRST read-only Database (index 0) as the checkpoint victim. */
refuseFirst: true,
reset(options: { refuseFirst?: boolean } = {}) {
native.seq = 0;
native.calls.constructions = [];
native.calls.checkpoints = 0;
native.calls.probes = 0;
native.refuseFirst = options.refuseFirst ?? true;
},
seq: 0,
};
return { native };
});
// The Database at index 0 is the interrupted-checkpoint victim: its init()
// throws the canonical refusal (the pool calls init explicitly) and any query
// on it refuses. Every LATER Database is healthy — in particular the WRITABLE
// recovery open and the read-only retry must never refuse, or the recovery
// would kill itself.
vi.mock('@ladybugdb/core', () => {
class Database {
role: 'ro' | 'rw';
index: number;
constructor(
_path: unknown,
_bufferManagerSize: unknown,
_compression: unknown,
readOnly = false,
) {
this.role = readOnly ? 'ro' : 'rw';
this.index = native.seq++;
native.calls.constructions.push(this.role);
}
async init(): Promise<void> {
if (this.index === 0 && this.role === 'ro' && native.refuseFirst) {
throw new Error(REFUSAL);
}
}
async close(): Promise<void> {}
}
const refuseIfVictim = (db: Database, text: string): void => {
const t = text.trim().toUpperCase();
if (t.startsWith('MATCH')) native.calls.probes++;
if (db.index === 0 && db.role === 'ro' && native.refuseFirst && t.startsWith('MATCH')) {
throw new Error(REFUSAL);
}
if (t === 'CHECKPOINT') native.calls.checkpoints++;
};
const emptyResult = {
getAll: async () => [],
close: async () => {},
isSuccess: () => true,
getErrorMessage: async () => '',
};
class Connection {
constructor(private db: Database) {}
async query(text: string) {
refuseIfVictim(this.db, text);
return emptyResult;
}
async prepare(cypher: string) {
refuseIfVictim(this.db, cypher);
return { isSuccess: () => true, getErrorMessage: async () => '' };
}
async execute() {
return emptyResult;
}
async close(): Promise<void> {}
}
const mod = { Database, Connection };
return { ...mod, default: mod, lbug: mod };
});
import { closeLbug, executeQuery, initLbug } from '../../src/core/lbug/pool-adapter.js';
describe('pool adapter self-heals a refused read-only open (forced refusal)', () => {
let dbPath: string;
let tmpDir: string;
const REPO = 'test-pool-forced-refusal';
beforeEach(async () => {
native.reset();
tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gitnexus-pool-forced-cp-'));
dbPath = path.join(tmpDir, 'lbug');
// The native layer is mocked; the file only has to EXIST (doInitLbug
// stats it before opening).
await fs.writeFile(dbPath, '');
});
afterEach(async () => {
await closeLbug(REPO).catch(() => {});
if (tmpDir) await fs.rm(tmpDir, { recursive: true, force: true });
});
it('runs writable open + CHECKPOINT + read-only retry, then serves queries', async () => {
await initLbug(REPO, dbPath);
// Self-heal shape: refused read-only open → writable recovery open →
// read-only retry. Exactly one CHECKPOINT, issued by the recovery.
expect(native.calls.constructions).toEqual(['ro', 'rw', 'ro']);
expect(native.calls.checkpoints).toBe(1);
// Writable replay MATCH + post-heal read-only MATCH. A CHECKPOINT without
// its required replay probe would leave this at 1.
expect(native.calls.probes).toBe(2);
const rows = await executeQuery(REPO, 'MATCH (n:Person) RETURN count(n) AS c');
expect(rows).toEqual([]);
});
it('does not issue a CHECKPOINT when the read-only open succeeds outright', async () => {
// A healthy db: the first read-only construction is the only one, and the
// recovery machinery must stay out of the way (no writable open, no
// CHECKPOINT on the read path).
native.reset({ refuseFirst: false });
await initLbug(REPO, dbPath);
await executeQuery(REPO, 'MATCH (n:Person) RETURN count(n) AS c');
expect(native.calls.constructions).toEqual(['ro']);
expect(native.calls.checkpoints).toBe(0);
expect(native.calls.probes).toBe(2);
});
});