mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-11 03:38:07 +00:00
fix(indexing): abort on exhausted or collapsed graph writes (#3540)
* docs(indexing): document native pool exhaustion recovery (U3) Publish the reviewed recovery runbook and collapsed-write guardrails for #3526. * docs(indexing): explain native buffer pool sizing (U3) * fix(indexing): reject collapsed staged graph writes (U2) * fix(indexing): abort native buffer exhaustion (U1) * test(indexing): cover exhausted and collapsed graph writes * ci(tests): split ubuntu coverage into 4 shards Coverage shard 1/3 runs 20-25 minutes across branches and now hits the 25-minute job timeout (cancelled on this PR and on unrelated branches), which skips coverage-merge and fails "every test executed" plus every platform compatibility gate downstream. A fourth shard brings each slice well under the cap. Shard job names are not required checks; the merge job globs coverage-blob-*, so no other wiring changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(lbug): remove branching from buffer-exhaustion tests Split the collapsed-staged-graph case into first-time and existing-index tests sharing one helper, so every assertion runs unconditionally instead of behind `if (hasExistingIndex)`. Parameterize the first-COPY exhaustion case with its statement matcher (reusing isRelationshipCopy) rather than a ternary on the phase name. Assertions are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
ecb444b4c8
commit
114ae0b411
8 changed files with 363 additions and 32 deletions
2
.github/workflows/ci-tests.yml
vendored
2
.github/workflows/ci-tests.yml
vendored
|
|
@ -200,7 +200,7 @@ jobs:
|
|||
- id: gen
|
||||
run: |
|
||||
TOTAL=4 # cross-platform (windows/macOS) shards per OS
|
||||
COV_TOTAL=3 # ubuntu coverage shards (merged before thresholds)
|
||||
COV_TOTAL=4 # ubuntu coverage shards (merged before thresholds)
|
||||
if [ "$TOTAL" -lt 1 ] || [ "$COV_TOTAL" -lt 1 ]; then
|
||||
echo "shard totals must be >= 1" >&2; exit 1
|
||||
fi
|
||||
|
|
|
|||
|
|
@ -60,9 +60,9 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
|
||||
### Analyze reports INCOMPLETE with a collapsed graph write
|
||||
|
||||
- **Trigger:** `npx gitnexus status` reports `incompleteReasons: ["graph-write-collapsed"]`; the analyze summary printed `Repository indexed INCOMPLETELY` naming an expected and a persisted relationship count, and the CLI exited non-zero.
|
||||
- **Do:** Re-run `npx gitnexus analyze --force`. If it recurs, check free disk space on the volume holding `.gitnexus/`, confirm no second `analyze` is running against the same repo (both stage through `.gitnexus/csv`), then run `npx gitnexus doctor`.
|
||||
- **Why:** The run finished and wrote metadata, but far fewer relationships are readable back than the pipeline produced. Nothing throws: the DB holds rows and the metadata is valid, so every query answers with missing edges rather than an error — a confident empty answer, which is worse than a failure because it looks like a result. Unlike `incremental-in-progress` and `embedding-checkpoint-pending`, which describe a run that did what it said and left work for next time, this one means most of your edges are gone, so it is the one incomplete reason that also fails the exit code. The check compares in-memory totals (including rows streamed out of the heap) against the post-write count, refuses to answer when the count cannot be read, and is skipped on incremental runs where whole-scope counts are not comparable.
|
||||
- **Trigger:** Analysis rejects a staged graph write with expected and persisted relationship counts, or an in-place run prints `Repository indexed INCOMPLETELY`, exits nonzero, and records `incompleteReasons: ["graph-write-collapsed"]`.
|
||||
- **Do:** Inspect COPY errors, available memory, and free disk space on the volume holding `.gitnexus/`; confirm no second `analyze` is running against the same repo, then run `npx gitnexus doctor`. For native buffer-pool exhaustion, see [RUNBOOK.md](RUNBOOK.md#memory--analyze-crashes) for `GITNEXUS_LBUG_BUFFER_POOL_SIZE` and memory-budget guidance. Re-run `npx gitnexus analyze --force` after addressing the cause.
|
||||
- **Why:** Far fewer relationships are readable back than the pipeline produced. A staged build throws before publication and preserves the previous index, if any. In-place writes cannot roll back, so the incomplete stamp and nonzero exit prevent a partial graph from being presented as successful. Recognized buffer exhaustion also throws immediately during COPY or fallback inserts. The collapse check compares in-memory totals (including streamed rows) against the post-write count, stays unmeasurable when the count cannot be read, and is skipped on incremental runs where whole-scope counts are not comparable.
|
||||
|
||||
### MCP lists no repos
|
||||
|
||||
|
|
|
|||
24
RUNBOOK.md
24
RUNBOOK.md
|
|
@ -185,6 +185,30 @@ Orchestrator: `.github/workflows/ci.yml`.
|
|||
|
||||
Analyze re-execs Node with a **large old-space heap** when needed (`analyze.ts`). If you still OOM on huge repos, close other processes, avoid `--embeddings` for a first pass, or analyze a smaller path if supported by your workflow.
|
||||
|
||||
`Unable to allocate memory` / `The buffer pool is full` during graph writes
|
||||
refers to LadybugDB's native buffer pool. Increasing `NODE_OPTIONS`,
|
||||
`--memory-budget`, or worker heap limits does not resize that pool. Recognized
|
||||
pool exhaustion aborts analysis instead of retrying with row skipping.
|
||||
|
||||
If the host has enough memory, set an explicit pool size in **bytes**, then
|
||||
rebuild. For example, 3 GiB is `3221225472` bytes:
|
||||
|
||||
```bash
|
||||
GITNEXUS_LBUG_BUFFER_POOL_SIZE=3221225472 npx gitnexus analyze --force
|
||||
```
|
||||
|
||||
The override bypasses automatic pool sizing. Budget for the Node heap, parse
|
||||
workers, other native allocations, and the OS alongside it; a larger pool can
|
||||
exhaust the host. `GITNEXUS_LBUG_BUFFER_POOL_SIZE=0` restores LadybugDB's native
|
||||
80%-of-RAM default, not an unlimited supply of memory. If memory is insufficient,
|
||||
exclude generated/vendor directories or use a larger host.
|
||||
|
||||
A measured graph-write collapse aborts a staged build before publication,
|
||||
leaving the previous index, if any, intact. An in-place run cannot roll back:
|
||||
it records `graph-write-collapsed` and exits nonzero. Rebuild a previously
|
||||
collapsed index with `--force` after addressing memory, disk space, or the
|
||||
reported COPY error; do not rely on its incomplete query results.
|
||||
|
||||
---
|
||||
|
||||
## LadybugDB / lock errors
|
||||
|
|
|
|||
|
|
@ -723,7 +723,7 @@ Configure the behavior with these environment variables:
|
|||
| `GITNEXUS_STREAM_GRAPH_EMIT` | `0`, `1` | `1` (on) | **On by default** on a full rebuild (`--force`); incremental runs ignore it. Holds structural relationships (CALLS, IMPORTS, ACCESSES, CONTAINS, ...) as CSV-on-disk plus compact in-memory columns instead of as objects in three overlapping indexes, cutting peak in-memory graph heap by ~1.4x at no measurable CPU cost (measured A/B on a synthetic 400k-node / 1.08M-edge graph: 819 MB -> 584 MB, iteration at parity, scaling verified linear from 100k to 800k nodes, with every edge still visible through the graph interface; no end-to-end measurement on a real repository yet). Nothing is traded away — community detection, process extraction, PDG taint summaries and the local-symbol pruner all read a complete relationship set and behave identically. Set to `0` only to bisect a suspected streaming-related fault. |
|
||||
| `GITNEXUS_COMMUNITY_ENGINE` | `graphology`, `icebug`, `auto` | `graphology` | Community-detection engine used during analyze. `graphology` is the supported default. `icebug` and `auto` are **experimental** and currently behave identically: both try the optional `@ladybugmem/icebug` native Leiden over a CSR export and fall back to Graphology if it is not installed, cannot load, or lacks the deterministic thread/seed controls. Experimental engines partition differently, so community IDs are not comparable across engines. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
|
||||
| `GITNEXUS_LBUG_BUFFER_POOL_SIZE` | integer `>= 0` (bytes) | min(2 GiB, 80% RAM) | LadybugDB buffer-pool ceiling for every GitNexus database (analyze, MCP server, serve, group bridges). Bounded so a long-lived `gitnexus mcp` process or a large incremental `analyze` cannot grow toward LadybugDB's native 80%-of-RAM default and OOM the host (#2557). `0` restores that native unbounded default; invalid values warn and fall back to the default. During `analyze` the pool is right-sized to the graph and, on non-4 KiB-page hosts (Apple Silicon 16 KiB, Ascend/aarch64 64 KiB), scaled by the page-size granule ratio up to min(2 GiB × pageSize/4 KiB, 80% RAM) (#2631); this env var overrides all of that as an absolute value. |
|
||||
| `GITNEXUS_LBUG_BUFFER_POOL_SIZE` | integer `>= 0` (bytes) | min(2 GiB, 80% RAM) | LadybugDB buffer-pool ceiling for every GitNexus database (analyze, MCP server, serve, group bridges). Bounded so a long-lived `gitnexus mcp` process or a large incremental `analyze` cannot grow toward LadybugDB's native 80%-of-RAM default and OOM the host (#2557). `0` restores the native 80%-of-RAM default; invalid values warn and fall back to the default. During `analyze` the pool is right-sized to the graph and, on non-4 KiB-page hosts (Apple Silicon 16 KiB, Ascend/aarch64 64 KiB), scaled by the page-size granule ratio up to min(2 GiB × pageSize/4 KiB, 80% RAM) (#2631); this env var overrides all of that as an absolute value. |
|
||||
| `GITNEXUS_LBUG_MAX_DB_SIZE` | positive integer (bytes) | `17179869184` (16 GiB) | Upper bound for a single LadybugDB database file. This is an mmap/disk-address-space ceiling, not a memory limit — it does not constrain the buffer pool (use `GITNEXUS_LBUG_BUFFER_POOL_SIZE` for that). Raise it when indexing genuinely huge monorepos; invalid values silently fall back to the default. |
|
||||
|
||||
```bash
|
||||
|
|
@ -744,6 +744,27 @@ GITNEXUS_FTS_CJK_SEGMENTATION=bigram npx gitnexus analyze --force
|
|||
|
||||
### Analysis runs out of memory
|
||||
|
||||
If graph writes fail with `Unable to allocate memory` or `The buffer pool is
|
||||
full`, LadybugDB's **native buffer pool** is exhausted. Analysis stops on these
|
||||
errors rather than retrying with row skipping. This pool is separate from the
|
||||
Node/V8 heap and parse-worker budgets; raising `NODE_OPTIONS`,
|
||||
`--memory-budget`, or worker heap limits does not increase it.
|
||||
|
||||
On a host with memory to spare, override the pool in **bytes** and rebuild:
|
||||
|
||||
```bash
|
||||
# 3 GiB native pool; leave room for Node, workers, other native memory, and the OS
|
||||
GITNEXUS_LBUG_BUFFER_POOL_SIZE=3221225472 npx gitnexus analyze --force
|
||||
```
|
||||
|
||||
An explicit value bypasses automatic pool sizing and can exhaust the host if
|
||||
set too high. `0` restores LadybugDB's native 80%-of-RAM default; it does not
|
||||
provide unlimited memory. Reduce the indexed scope or use more RAM when the
|
||||
combined budgets do not fit. A collapsed staged build is discarded before
|
||||
publication, preserving the previous index if one exists. In-place collapse
|
||||
leaves an incomplete index and exits nonzero; recover with `--force` once the
|
||||
underlying problem is fixed.
|
||||
|
||||
Memory management is automatic: `analyze` sizes its heap to the machine
|
||||
(always below physical RAM), caps each parse worker, and — rather than
|
||||
grinding into a GC death spiral or crash — stops early with a message telling
|
||||
|
|
|
|||
|
|
@ -1099,12 +1099,19 @@ const doInitLbug = async (
|
|||
|
||||
export type LbugProgressCallback = (message: string) => void;
|
||||
|
||||
const throwIfBufferPoolExhausted = (error: unknown, context?: string): void => {
|
||||
const message = error instanceof Error ? error.message : String(error);
|
||||
const remedy = bufferPoolExhaustionRemedy(message);
|
||||
if (remedy) throw new Error(`${context ?? message} ${remedy}`, { cause: error });
|
||||
};
|
||||
|
||||
/**
|
||||
* Run a COPY, retrying once with IGNORE_ERRORS=true (which skips row-level
|
||||
* errors) on first failure. Log the original failure and native COPY/warning
|
||||
* receipts even when the retry succeeds. Node COPY requires every row; a
|
||||
* skipped-node receipt is a failure. Call sites retain their own message
|
||||
* limits and relationship fallback policy.
|
||||
* errors) on a non-resource first failure. Buffer exhaustion is fatal, including
|
||||
* retained retry warnings, and bypasses the recoverable-error callback. Log the
|
||||
* original failure and native COPY/warning receipts even when the retry succeeds.
|
||||
* Node COPY requires every row; a skipped-node receipt is a failure. Call sites
|
||||
* retain their own message limits and relationship fallback policy.
|
||||
*/
|
||||
const copyCsvWithRetry = async (
|
||||
targetConn: lbug.Connection,
|
||||
|
|
@ -1115,6 +1122,7 @@ const copyCsvWithRetry = async (
|
|||
try {
|
||||
await queryAndDrain(targetConn, copyQuery);
|
||||
} catch (firstError) {
|
||||
throwIfBufferPoolExhausted(firstError);
|
||||
logger.warn(
|
||||
{ err: firstError, copyQuery },
|
||||
'First COPY failure; retrying with IGNORE_ERRORS=true',
|
||||
|
|
@ -1157,6 +1165,13 @@ const copyCsvWithRetry = async (
|
|||
},
|
||||
'COPY retry completed; retained warnings describe skipped rows (a lower bound)',
|
||||
);
|
||||
// Inspect every retained warning, not just the diagnostic samples.
|
||||
// Exhaustion cannot be recovered by discarding rows or replaying a
|
||||
// partially committed relationship COPY through the fallback path.
|
||||
const resourceWarning = warnings.find((warning) =>
|
||||
bufferPoolExhaustionRemedy(String(warning.message ?? '')),
|
||||
);
|
||||
if (resourceWarning) throw new Error(String(resourceWarning.message));
|
||||
if (expectedRows !== undefined && (copiedRows !== expectedRows || warnings.length > 0)) {
|
||||
throw new Error(
|
||||
`COPY retry skipped rows or could not verify a complete node load ` +
|
||||
|
|
@ -1168,8 +1183,10 @@ const copyCsvWithRetry = async (
|
|||
} catch (retryErr) {
|
||||
const firstMessage = firstError instanceof Error ? firstError.message : String(firstError);
|
||||
const retryMessage = retryErr instanceof Error ? retryErr.message : String(retryErr);
|
||||
const message = `COPY retry failed: ${retryMessage}; first failure: ${firstMessage}`;
|
||||
throwIfBufferPoolExhausted(retryErr, message);
|
||||
onError(
|
||||
new Error(`COPY retry failed: ${retryMessage}; first failure: ${firstMessage}`, {
|
||||
new Error(message, {
|
||||
cause: retryErr,
|
||||
}),
|
||||
);
|
||||
|
|
@ -1247,14 +1264,7 @@ const copyNodeCSVs = async (
|
|||
copyQuery,
|
||||
(retryErr) => {
|
||||
const retryMsg = retryErr instanceof Error ? retryErr.message : String(retryErr);
|
||||
// Pool exhaustion gets a remedy (#2631): the raw binder text gives the
|
||||
// operator nothing to act on, and on non-4K-page hosts (Ascend aarch64,
|
||||
// Apple Silicon) the pool bills up to pageSize/4KiB x faster than the
|
||||
// sizing was calibrated for — name the knob and the mechanism.
|
||||
const remedy = bufferPoolExhaustionRemedy(retryMsg);
|
||||
throw new Error(
|
||||
`COPY failed for ${table}: ${retryMsg.slice(0, 200)}${remedy ? ` ${remedy}` : ''}`,
|
||||
);
|
||||
throw new Error(`COPY failed for ${table}: ${retryMsg.slice(0, 200)}`);
|
||||
},
|
||||
rows,
|
||||
);
|
||||
|
|
@ -1454,7 +1464,6 @@ export const loadGraphToLbug = async (
|
|||
|
||||
const insertedRels = totalValidRels + (graphEmitManifest?.totalRows ?? 0);
|
||||
const warnings: string[] = [];
|
||||
let poolRemedyIssued = false;
|
||||
if (insertedRels > 0) {
|
||||
log(`Loading edges: ${insertedRels.toLocaleString()} across ${copyJobs.length} CSV files`);
|
||||
|
||||
|
|
@ -1487,17 +1496,6 @@ export const loadGraphToLbug = async (
|
|||
await copyCsvWithRetry(writeConn, copyQuery, (retryErr) => {
|
||||
const retryMsg = retryErr instanceof Error ? retryErr.message : String(retryErr);
|
||||
warnings.push(`${fromLabel}->${toLabel} (${rows} edges): ${retryMsg.slice(0, 80)}`);
|
||||
// One remedy per bulk load, not per pair (#2631): pool exhaustion
|
||||
// repeats for every remaining pair once it starts. logger.warn, not
|
||||
// just warnings.push — the returned warnings array has no consumer at
|
||||
// any call site, so a push alone would leave the remedy invisible
|
||||
// while the row-by-row fallback quietly degrades the load.
|
||||
const remedy = poolRemedyIssued ? undefined : bufferPoolExhaustionRemedy(retryMsg);
|
||||
if (remedy) {
|
||||
poolRemedyIssued = true;
|
||||
warnings.push(remedy);
|
||||
logger.warn(remedy);
|
||||
}
|
||||
failedPairEdges += rows;
|
||||
failedPairCsvPaths.add(pairCsvPath);
|
||||
});
|
||||
|
|
@ -1696,8 +1694,9 @@ export const fallbackRelationshipInserts = async (
|
|||
CREATE (a)-[:${REL_TABLE_NAME} {type: ${formatCypherValue(relType)}, confidence: ${confidence}, reason: ${formatCypherValue(reason)}, step: ${step}, staticGated: ${staticGated}}]->(b)
|
||||
`,
|
||||
);
|
||||
} catch {
|
||||
// skip
|
||||
} catch (error) {
|
||||
throwIfBufferPoolExhausted(error);
|
||||
// Ordinary per-row failures remain skippable.
|
||||
}
|
||||
}
|
||||
};
|
||||
|
|
|
|||
|
|
@ -4489,6 +4489,18 @@ async function runFullAnalysisInner(
|
|||
existingMeta?.graphWriteCollapsed,
|
||||
);
|
||||
if (graphWriteCollapsed) {
|
||||
if (useAtomicSwap) {
|
||||
// A collapsed staging DB can be discarded before it replaces the live
|
||||
// index. In-place writes below still need their incomplete stamp.
|
||||
throw new Error(
|
||||
`Graph write incomplete — the pipeline produced ${expectedRelationships} ` +
|
||||
`relationships but only ${persistedRelationships} are readable from the staged index. ` +
|
||||
`Analysis aborted before publication; the previous index, if any, is left intact. ` +
|
||||
`Review COPY warnings and available memory and disk space. If the native buffer pool ` +
|
||||
`was exhausted, set GITNEXUS_LBUG_BUFFER_POOL_SIZE to a larger byte value that fits ` +
|
||||
`available memory, then re-run \`gitnexus analyze --force\`.`,
|
||||
);
|
||||
}
|
||||
log(
|
||||
`Warning: graph write incomplete — the pipeline produced ${expectedRelationships} ` +
|
||||
`relationships but only ${persistedRelationships} are readable from the index. Recording the ` +
|
||||
|
|
|
|||
|
|
@ -68,6 +68,7 @@ import {
|
|||
runFullAnalysis,
|
||||
} from '../../src/core/run-analyze.js';
|
||||
import { getStoragePaths, loadMeta, readRegistry } from '../../src/storage/repo-manager.js';
|
||||
import { executeQuery } from '../../src/core/lbug/lbug-adapter.js';
|
||||
import {
|
||||
initLbug as poolInit,
|
||||
executeQuery as poolQuery,
|
||||
|
|
@ -177,6 +178,89 @@ describe.skipIf(isWin)('atomic full-rebuild swap (#2)', () => {
|
|||
}
|
||||
}, 180_000);
|
||||
|
||||
// Commit a replacement graph, then force a rebuild whose staged load loses
|
||||
// every relationship, and assert the collapse guard rejects it unpublished.
|
||||
const expectCollapsedRebuildRejected = async (repo: string) => {
|
||||
const { lbugPath } = getStoragePaths(repo);
|
||||
const oldRegistry = await readRegistry();
|
||||
await fs.writeFile(
|
||||
path.join(repo, 'a.ts'),
|
||||
'export function replacement() { return "new graph"; }\n',
|
||||
);
|
||||
execSync('git -c user.name=t -c user.email=t@t commit -am replacement', {
|
||||
cwd: repo,
|
||||
stdio: 'pipe',
|
||||
});
|
||||
ctx.loadMock.mockImplementationOnce(
|
||||
async (...args: Parameters<LbugAdapter['loadGraphToLbug']>) => {
|
||||
const result = await ctx.realLoad!(...args);
|
||||
// Keep real nodes and a measurable DB, but lose every relationship
|
||||
// after the load so the actual collapse guard must reject it.
|
||||
await executeQuery('MATCH ()-[r:CodeRelation]->() DELETE r');
|
||||
return result;
|
||||
},
|
||||
);
|
||||
|
||||
const failure = await runFullAnalysis(repo, { force: true }, { onProgress: () => {} }).catch(
|
||||
(error: unknown) => error,
|
||||
);
|
||||
expect(failure).toBeInstanceOf(Error);
|
||||
expect(failure).toMatchObject({
|
||||
message: expect.stringMatching(/produced [1-9]\d* relationships but only 0/),
|
||||
});
|
||||
expect(failure).toMatchObject({
|
||||
message: expect.stringContaining('GITNEXUS_LBUG_BUFFER_POOL_SIZE'),
|
||||
});
|
||||
expect(analyzeFailureMayHaveMutatedLiveIndex(failure)).toBe(false);
|
||||
expect(await lingeringTemp(lbugPath)).toEqual([]);
|
||||
expect(await readRegistry()).toEqual(oldRegistry);
|
||||
};
|
||||
|
||||
it('rejects a collapsed first-time staged graph before publication', async () => {
|
||||
const { repo, cleanup } = await makeRepo();
|
||||
const { lbugPath, storagePath } = getStoragePaths(repo);
|
||||
try {
|
||||
await expectCollapsedRebuildRejected(repo);
|
||||
await expect(fs.stat(lbugPath)).rejects.toMatchObject({ code: 'ENOENT' });
|
||||
// Storage setup may leave its empty metadata shell, but no analysis
|
||||
// commit or completed graph statistics may be published.
|
||||
const failedMeta = await loadMeta(storagePath);
|
||||
expect(failedMeta?.lastCommit).toBe('');
|
||||
expect(failedMeta?.stats).toBeUndefined();
|
||||
} finally {
|
||||
await cleanup();
|
||||
}
|
||||
}, 180_000);
|
||||
|
||||
it('rejects a collapsed rebuild and keeps the previous index readable', async () => {
|
||||
const { repo, cleanup } = await makeRepo();
|
||||
const repoId = 'atomic-collapse-existing';
|
||||
const { lbugPath, storagePath } = getStoragePaths(repo);
|
||||
try {
|
||||
await runFullAnalysis(repo, {}, { onProgress: () => {} });
|
||||
const oldGraph = await fs.readFile(lbugPath);
|
||||
const oldMeta = await loadMeta(storagePath);
|
||||
|
||||
await expectCollapsedRebuildRejected(repo);
|
||||
|
||||
expect((await fs.readFile(lbugPath)).equals(oldGraph)).toBe(true);
|
||||
expect(await loadMeta(storagePath)).toMatchObject({
|
||||
lastCommit: oldMeta!.lastCommit,
|
||||
indexedAt: oldMeta!.indexedAt,
|
||||
});
|
||||
await poolInit(repoId, lbugPath);
|
||||
const names = (await poolQuery(repoId, 'MATCH (f:Function) RETURN f.name AS n')).flatMap(
|
||||
(row) => Object.values(row as Record<string, unknown>).map(String),
|
||||
);
|
||||
expect(names.sort()).toEqual(['caller', 'greet']);
|
||||
const edges = await poolQuery(repoId, 'MATCH ()-[r:CodeRelation]->() RETURN count(r) AS n');
|
||||
expect(Number((edges[0] as Record<string, unknown>).n)).toBeGreaterThan(0);
|
||||
} finally {
|
||||
await poolClose(repoId);
|
||||
await cleanup();
|
||||
}
|
||||
}, 180_000);
|
||||
|
||||
it.each(
|
||||
(['full', 'incremental'] as const).flatMap((writeMode) =>
|
||||
(['close', 'swap', 'metadata'] as const).map((failurePoint) => ({
|
||||
|
|
|
|||
|
|
@ -59,6 +59,7 @@ beforeAll(async () => {
|
|||
|
||||
afterEach(() => {
|
||||
emitMock.mockReset();
|
||||
vi.restoreAllMocks();
|
||||
});
|
||||
|
||||
afterAll(async () => {
|
||||
|
|
@ -258,3 +259,193 @@ describe('loadGraphToLbug overlap error paths (#2226 F2)', () => {
|
|||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('graph write buffer exhaustion (#3526)', () => {
|
||||
let fixtureId = 0;
|
||||
|
||||
const emitGraphCSVs = (malformedRelationship = false) => {
|
||||
const prefix = `pool-${++fixtureId}`;
|
||||
emitMock.mockImplementation(
|
||||
async (_g: unknown, _r: unknown, dir: string, onNodes?: (n: NodeFiles) => void) => {
|
||||
await fs.mkdir(dir, { recursive: true });
|
||||
const csvPath = path.join(dir, 'file.csv');
|
||||
await fs.writeFile(
|
||||
csvPath,
|
||||
'id,name,filePath,content\n' +
|
||||
`"File:${prefix}-a.ts","a.ts","${prefix}-a.ts",""\n` +
|
||||
`"File:${prefix}-b.ts","b.ts","${prefix}-b.ts",""\n`,
|
||||
);
|
||||
const nodeFiles = new Map([['File', { csvPath, rows: 2 }]]) as NodeFiles;
|
||||
onNodes?.(nodeFiles);
|
||||
const relPath = path.join(dir, 'rel_File_File.csv');
|
||||
await fs.writeFile(
|
||||
relPath,
|
||||
REL_HEADER +
|
||||
',staticGated\n' +
|
||||
`"File:${prefix}-a.ts","File:${prefix}-b.ts","IMPORTS",${malformedRelationship ? 'bad' : '1'},"first",0,0\n` +
|
||||
`"File:${prefix}-b.ts","File:${prefix}-a.ts","IMPORTS",1,"second",0,0\n`,
|
||||
);
|
||||
return {
|
||||
...emptyResult(),
|
||||
nodeFiles,
|
||||
relsByPair: new Map([['File|File', { csvPath: relPath, rows: 2 }]]),
|
||||
totalValidRels: 2,
|
||||
};
|
||||
},
|
||||
);
|
||||
return prefix;
|
||||
};
|
||||
|
||||
// Keep real native results (including warning cursors and their cleanup),
|
||||
// replacing only the particular statement that simulates the failure.
|
||||
const injectQueries = async (replace: (sql: string) => Error | string | undefined) => {
|
||||
const { default: lbug } = await import('@ladybugdb/core');
|
||||
const originalQuery = lbug.Connection.prototype.query;
|
||||
const seen: string[] = [];
|
||||
vi.spyOn(lbug.Connection.prototype, 'query').mockImplementation(function (
|
||||
this: unknown,
|
||||
sql: string,
|
||||
...rest: unknown[]
|
||||
) {
|
||||
seen.push(sql);
|
||||
const replacement = replace(sql);
|
||||
if (replacement instanceof Error) return Promise.reject(replacement);
|
||||
return originalQuery.call(this, replacement ?? sql, ...rest);
|
||||
});
|
||||
return seen;
|
||||
};
|
||||
|
||||
const isRelationshipCopy = (sql: string) => /^COPY CodeRelation\b/.test(sql);
|
||||
const isFallbackInsert = (sql: string) => /CREATE \(a\)-\[:CodeRelation/.test(sql);
|
||||
|
||||
it.each([
|
||||
{ phase: 'node', isCopy: (sql: string) => /^COPY File\(/.test(sql) },
|
||||
{ phase: 'relationship', isCopy: isRelationshipCopy },
|
||||
])(
|
||||
'aborts a first $phase COPY exhaustion without retrying or falling back',
|
||||
async ({ isCopy }) => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
emitGraphCSVs();
|
||||
const original = new Error('Unable to allocate memory! The buffer pool is full!');
|
||||
const seen = await injectQueries((sql) =>
|
||||
isCopy(sql) && !sql.includes('IGNORE_ERRORS') ? original : undefined,
|
||||
);
|
||||
|
||||
const error = await adapter
|
||||
.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath)
|
||||
.catch((err: unknown) => err);
|
||||
expect(error).toBeInstanceOf(Error);
|
||||
expect(error).toMatchObject({ cause: original });
|
||||
expect((error as Error).message).toContain(original.message);
|
||||
expect((error as Error).message).toContain('GITNEXUS_LBUG_BUFFER_POOL_SIZE');
|
||||
expect(seen.some((sql) => sql.includes('IGNORE_ERRORS'))).toBe(false);
|
||||
expect(seen.some(isFallbackInsert)).toBe(false);
|
||||
},
|
||||
);
|
||||
|
||||
it('aborts retry exhaustion before relationship fallback', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
emitGraphCSVs();
|
||||
const original = new Error('Unable to allocate memory during relationship retry');
|
||||
const seen = await injectQueries((sql) => {
|
||||
if (isRelationshipCopy(sql)) {
|
||||
return sql.includes('IGNORE_ERRORS') ? original : new Error('ordinary COPY failure');
|
||||
}
|
||||
});
|
||||
|
||||
const error = await adapter
|
||||
.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath)
|
||||
.catch((err: unknown) => err);
|
||||
expect(error).toBeInstanceOf(Error);
|
||||
expect(error).toMatchObject({ cause: original });
|
||||
expect((error as Error).message).toContain(original.message);
|
||||
expect((error as Error).message).toContain('ordinary COPY failure');
|
||||
expect((error as Error).message).toContain('GITNEXUS_LBUG_BUFFER_POOL_SIZE');
|
||||
expect(seen.filter(isRelationshipCopy)).toHaveLength(2);
|
||||
expect(seen.some(isFallbackInsert)).toBe(false);
|
||||
});
|
||||
|
||||
it('aborts when a retained resource warning is beyond the five logged samples', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
emitGraphCSVs();
|
||||
const seen = await injectQueries((sql) => {
|
||||
if (isRelationshipCopy(sql) && !sql.includes('IGNORE_ERRORS')) {
|
||||
return new Error('ordinary COPY failure');
|
||||
}
|
||||
if (sql.startsWith('CALL SHOW_WARNINGS()')) {
|
||||
return (
|
||||
"UNWIND ['row one', 'row two', 'row three', 'row four', 'row five', " +
|
||||
"'Unable to allocate memory in COPY'] AS message RETURN message, '' AS file_path, 1 AS line_number"
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
await expect(
|
||||
adapter.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath),
|
||||
).rejects.toThrow(/Unable to allocate memory.*GITNEXUS_LBUG_BUFFER_POOL_SIZE/);
|
||||
expect(seen.some(isFallbackInsert)).toBe(false);
|
||||
});
|
||||
|
||||
it('stops fallback on resource exhaustion before the next relationship', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
emitGraphCSVs();
|
||||
const original = new Error('The buffer pool is full during CREATE');
|
||||
const seen = await injectQueries((sql) => {
|
||||
if (isRelationshipCopy(sql)) return new Error('ordinary COPY failure');
|
||||
if (isFallbackInsert(sql)) return original;
|
||||
});
|
||||
|
||||
const error = await adapter
|
||||
.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath)
|
||||
.catch((err: unknown) => err);
|
||||
expect(error).toMatchObject({ cause: original });
|
||||
expect((error as Error).message).toContain('GITNEXUS_LBUG_BUFFER_POOL_SIZE');
|
||||
expect(seen.filter(isFallbackInsert)).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('still skips an ordinary fallback row error and inserts subsequent relationships', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
const prefix = emitGraphCSVs();
|
||||
let inserts = 0;
|
||||
const seen = await injectQueries((sql) => {
|
||||
if (isRelationshipCopy(sql)) return new Error('ordinary COPY failure');
|
||||
if (isFallbackInsert(sql) && ++inserts === 1) return new Error('ordinary row failure');
|
||||
});
|
||||
|
||||
await expect(
|
||||
adapter.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath),
|
||||
).resolves.toMatchObject({ success: true });
|
||||
expect(seen.filter(isFallbackInsert)).toHaveLength(2);
|
||||
const rows = await adapter.executeQuery(
|
||||
`MATCH (a:File {id: 'File:${prefix}-b.ts'})-[r:CodeRelation]->(b) RETURN r.reason AS reason`,
|
||||
);
|
||||
expect(rows).toEqual([{ reason: 'second' }]);
|
||||
});
|
||||
|
||||
it('loads a healthy graph through real COPY', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
const prefix = emitGraphCSVs();
|
||||
await expect(
|
||||
adapter.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath),
|
||||
).resolves.toMatchObject({ success: true, warnings: [] });
|
||||
const rows = await adapter.executeQuery(
|
||||
`MATCH (a:File)-[r:CodeRelation]->(b) WHERE a.id STARTS WITH 'File:${prefix}-' RETURN count(r) AS count`,
|
||||
);
|
||||
expect(Number(rows[0].count)).toBe(2);
|
||||
});
|
||||
|
||||
it('still skips an ordinary malformed relationship during COPY retry without replaying it', async () => {
|
||||
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
|
||||
const prefix = emitGraphCSVs(true);
|
||||
const seen = await injectQueries(() => undefined);
|
||||
await expect(
|
||||
adapter.loadGraphToLbug(buildTestGraph([], []), tmpBase, storagePath),
|
||||
).resolves.toMatchObject({ success: true });
|
||||
expect(seen.filter(isRelationshipCopy)).toHaveLength(2);
|
||||
expect(seen.some(isFallbackInsert)).toBe(false);
|
||||
const rows = await adapter.executeQuery(
|
||||
`MATCH (a:File)-[r:CodeRelation]->(b) WHERE a.id STARTS WITH 'File:${prefix}-' RETURN r.reason AS reason`,
|
||||
);
|
||||
expect(rows).toEqual([{ reason: 'second' }]);
|
||||
});
|
||||
});
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue