GitNexus/gitnexus/test/unit/incremental-orchestration.test.ts
glier b15ff2d888
feat(ingestion): mint Destination nodes from AsyncAPI 3.x documents (#3140)
* feat(ingestion): read AsyncAPI 3.x documents into broker addresses

Adds a format-driven reader that turns the `operations[]` entries of an
AsyncAPI 3.x document into (broker, address, direction) triples, plus the
protocol-to-broker map behind it. Nothing consumes it yet.

The reader lives outside `frameworks/spring/` on purpose, like
`destination-key.ts` and for the same reason: an AsyncAPI document is a
published artifact emitted by generators across several language toolchains
and written by hand as often as generated. The entry criterion is therefore
the document format -- a root `asyncapi` key -- and never the generator.

AsyncAPI 2.x is refused under its own countable reason rather than mapped.
Its `publish`/`subscribe` are inverted relative to 3.x `send`/`receive`, so
a naive mapping reverses every direction in the async graph while leaving it
connected: nothing fails, the arrows simply point the wrong way. A silent
skip would be indistinguishable from "this service publishes no document",
which is the one thing the refusal count has to be able to tell us.

The broker is read twice over -- from the operation's bindings and from its
channel's server protocol -- and the two readings must agree. A destination
keyed on the wrong broker joins a stranger, and with the document
contradicting itself there is no way to tell which reading is right, so the
operation is refused rather than decided by a coin flip.

An unmapped protocol passes through as its own literal instead of being
dropped, because `destinationNodeKey` takes a plain string precisely so a
non-Spring caller can attest to a broker Spring has no member for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(cli): add --asyncapi-spec, an explicit path to AsyncAPI documents

Threads an `asyncApiSpecPath` option from the CLI, the server analyze
endpoint, and the programmatic entry through to `PipelineOptions`. Nothing
reads it yet; the reader added in the previous commit is still unwired.

Shaped deliberately after `springActuatorPath`, the existing option for an
out-of-band artifact: an explicit local path, accepting a directory or a
single file, resolved against the repository root so a committed
`docs/asyncapi` and an absolute cache populated by something else are both
natural, and `undefined` keeping the feature entirely off. Mirroring that
option rather than inventing a mechanism is what lets a downstream consumer
point the reader at documents fetched out of band without patching a file
here.

`analyze --watch` REJECTS the flag, exactly as it rejects --spring-actuator.
The watcher reacts to source changes and nothing watches a document
directory, so honouring it there would read the documents once and then
serve a stale answer for the rest of the session -- worse than refusing,
because it looks like it worked.

Additive only: 49 inserted lines, no deletions and no modified lines. Every
new interface member is optional and every forward is an object-literal
spread of an undefined value, so with the option unset the analyzer takes
byte-identical paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ingestion): mint Destination nodes from AsyncAPI documents

Wires the reader into the destinations phase. With `--asyncapi-spec` set,
every `send` operation emits PUBLISHES_TO and every `receive` emits
CONSUMES_FROM against the ordinary resolved `Destination` node -- same key,
same `address` property -- so a document and a source site that name one
address on one broker land on ONE node and the two halves of a conversation
meet. Verified end to end: an address named only in a document is minted
with the right broker and direction, and an address a source site already
resolved stays a single node with its own `literal` provenance while the
same address on a different broker stays separate.

This claims only what a document states -- that the service talks to that
address, on that broker, in that direction -- and never which method does
it. The addresses in one document partition by (broker, action) into buckets
that usually hold more than one operation, so any assignment past a bucket
of size one is a heuristic, and a wrong one attaches a real address to the
wrong handler: a false connection wearing the clothes of a resolved one. The
edge therefore starts at the document, not at a callable. That is weaker
than a source-derived edge and worth having anyway, because it is available
where the source supplies nothing at all -- a programmatically registered
listener, a broker with no patterns here, a language whose messaging idiom
nobody has taught this codebase yet.

Documents are read even when the source pass found no messaging, which is
why the early return had to move: a repository whose brokers are invisible
to the patterns is precisely the case a published document covers, and an
early return keyed on source sites skipped the documents exactly there.

Their counters are kept in their own block rather than folded into the
existing ones. `refusalsByReason` is the denominator of the SOURCE
unresolved fraction, and a mistyped specification directory must not be able
to make the source look worse than it is. The block is absent -- not zeroed
-- when no path was configured, so "not asked for" stays distinguishable
from "asked for and found nothing"; those need different answers from an
operator and one zero cannot say which happened.

The direction assertion is the one that matters and it is pinned by type,
not by existence: inverting the mapping in the source tree fails exactly one
test, because every other assertion passes identically under both readings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document --asyncapi-spec and why step 4 of the cascade stays empty

Adds the flag to both READMEs and to the three byte-identical copies of the
CLI skill, which a sync test pins together.

Also rewrites the note on the `specification` seam in the address cascade.
It said "nothing supplies it today", which was true and is now misleading:
a reader exists, and the hook is still unsupplied because of a decision
rather than for want of one.

A document names addresses; it does not name the method that uses one. To
hand an address to a particular candidate something must choose which of the
document's operations belongs to it, and the only division both sides agree
on -- (broker, action) -- leaves buckets that usually hold more than one
operation. On a real generated document exactly one bucket of four was
unambiguous. Every assignment past a bucket of size one is a heuristic, and
a wrong one puts a REAL address on a joining node under the wrong site: a
false connection wearing the clothes of a resolved one, which is the outcome
the keying rule exists to prevent.

The note also records the two things that would change that and are not
heuristics -- a document carrying the implementing symbol, or a
configuration source answering the `${key}` the candidate already recorded
-- and that the second wants its own resolver, since what it needs is the
placeholder key rather than the candidate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ingestion): refuse the document shapes that would forge a join

Four ways a conformant AsyncAPI document could mint a Destination that
connects two services which have said nothing about each other. Every one is
reachable from ordinary 3.x vocabulary, not from malformed input, and each is
now a countable refusal.

A PARAMETERIZED ADDRESS is a pattern, not a place. Two services that publish
`{env}.orders` share a template; one deploys with env=prod and the other with
env=staging, and keying on the template text merges them into a single node
with a publisher on one side and a subscriber on the other. This is the
document-side twin of `overridable-config-default`, which argues the same
thing about `${key:default}` in source. A channel declaring `parameters` is
the specification's own statement that its address is a template, so the
detector is a reading rather than a guess; the `{` test catches generators
that template without declaring.

ANY BINDINGS KEY WAS TAKEN AS A PROTOCOL. AsyncAPI allows `bindings` to be a
Reference Object, so the map's own key can be `$ref` -- and passed through,
that becomes half of a join key carrying no broker information at all. Two
services that both reference shared bindings and both name `orders` then land
on one node, defeating the broker-in-key rule that keeps `kafka orders` and
`rabbit orders` apart. A broker must now be spelled like a protocol name.

A BROKER CONTAINING A SPACE COLLIDES, because the node key joins with one:
("kafka orders", "x") and ("kafka", "orders x") are the same key. That was
latent while every broker came from Spring's closed union. This module is the
first caller to feed the shared helper text that a document wrote, which is
exactly the condition under which it stops being latent, so it is closed here
-- at the producer -- rather than by changing an encoding that `routeNodeKey`
shares.

THE ADDRESS WAS TRIMMED, while the source cascade keeps an address exactly as
written so `" orders "` stays its own node. Two producers of one key held
opposite whitespace policies and the document side erred toward joining.

Also: fold the transport-security protocol variants (`kafka-secure`,
`secure-mqtt`, `wss`, `stomps`, `https`) onto their base protocol. The
`amqp`->`rabbit` argument already in this file demands it -- AsyncAPI's server
vocabulary distinguishes them and its bindings vocabulary does not, so a
secured cluster's own document was being read as self-contradictory. Treat a
channel with no `servers` as available on all of the document's servers, which
is the specification's default and was costing every single-server document
its destinations. Bound the address and operation-id lengths, because
`generateId` concatenates rather than hashes, and bound total operations
across the run rather than only per document.

Read each file through ONE handle for both the size gate and the read, as
`actuator-runtime.ts` does and for the reason its comment gives (CodeQL
js/file-system-race): re-resolving the path lets a swapped file bypass the
cap, and the out-of-band cache this option reads is written by other tooling
by definition. Open it with O_NONBLOCK: the type check that rejects a FIFO is
unreachable without it, because opening a FIFO for reading blocks in open(2)
until a writer appears -- found by writing the test first and watching it time
out rather than fail.

Count symlinked entries and walk truncation instead of dropping them in
silence. A symlinked cache and a wrong path were producing identical results.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): rebuild when documents are configured, and report what was read

An AsyncAPI document is external to git freshness in exactly the way an
Actuator snapshot is: replacing one moves no commit and dirties no file. The
option's own README paragraph advertises an absolute cache written by other
tooling, and on the second run of that workflow the already-up-to-date fast
path fired, no document was ever opened, and the previous run's addresses were
served as current. Measured, not reasoned: editing a document and re-running
printed "Already up to date" and left the old address in the graph; with this
change the same probe re-reads and the new address replaces it.

So an enabled run forces a rebuild and dropping the option forces one more, to
clear document-derived evidence -- the treatment `springActuatorPath` already
gets, for the same reason. Only the FLAG is recorded in index metadata, not
the path: Actuator retains its inputs so future scans keep excluding them,
whereas a committed document is deliberately NOT excluded (it wants its real
`File` node), so there is nothing to retain and recording the path would put
an operator's directory layout into metadata for no consumer.

That also settles a defect it would have been tempting to patch separately. A
synthetic `File` node for an out-of-tree document carries a path that is in no
write set and is not covered by `isGraphWideNode`, so an incremental writeback
dropped the node while keeping its edges -- which then COPY against a row that
was never written, and fail into an IGNORE_ERRORS retry that reports success.
A forced rebuild has no incremental subgraph to get that wrong.

Distinguish the two meanings of `resolution: 'specification'`. That value
belongs to the address cascade and means a CODE candidate was resolved through
the step-4 hook; a node minted from a document has no code site and now says
`asyncapi-document`. Reusing one string would leave a query that groups by
provenance unable to separate an address a document states from one a document
was used to resolve, and only the second is a claim about source.

Report what was read. The stats block was justified on the grounds that an
operator must be able to tell a mistyped directory from a repository with no
documents -- and nothing surfaced it, so the justification was aspirational. A
configured path that yields nothing, or a walk that hit a bound, now warns
unconditionally, as `spring-auto-configuration.ts` does for the same class of
input. The phase summary carries the refusal breakdown rather than only the
totals, because the unresolved fraction is the number this work is judged on
and a bare count says how big the gap is without saying what would close it.

Tests for the three wiring lines that were individually deletable with a green
suite, following the templates already in the repository: a row in the
`--watch` rejection table, the CLI-threading assertion beside the Actuator
one, and the shipped-skill fragment that pins the flag in all three copies.
Also pin `filePath: ''` on a spec-minted destination -- the half of the keying
rule that stops a shared node becoming collateral damage of one document's
next change -- and the in-repo `File` branch, which was dead-code-able.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ingestion): close the join-forging paths a second review round found

The `$ref` exclusion added last round fixed one instance of a class and left
the class open. A `bindings` map key of `x-scs-function` -- an ordinary
Specification Extension, which generators emit -- still became the broker, so
two unrelated services carrying one vendor annotation and one address landed
on ONE node whose broker half said nothing about any broker. And
`{ kafka: {}, x-internal: {} }` read as two brokers, losing a conformant
document and reporting it as self-contradictory, which also made any document
author a one-line saboteur of their own cross-service links.

So the two readers are now separate functions with opposite defaults, and the
header says why they must be. `servers[].protocol` is a FIELD DECLARED to hold
a protocol: an unrecognized value there is the document's own claim and passes
through, because refusing it would lose a destination the document states
plainly. A `bindings` MAP KEY is not that -- the specification puts `$ref` and
`x-` in the same namespace -- so a non-protocol key is the EXPECTED case and
only AsyncAPI's binding vocabulary may answer. The syntactic test that was
applied to both was the right rule for one of them.

The walk fix from last round introduced something worse than it reported. A
shared abort flag meant depth exhaustion in ONE branch terminated the whole
traversal, so ten good documents beside a twelve-deep unrelated subtree were
kept or lost depending on whether that subtree sorted before or after them.
Truncation and budget-exhaustion are now separate: depth returns from its own
branch, and only a genuinely global bound stops the walk.

Rewriting `protocol.ts` dropped the whitespace check from the server-protocol
path and made the node-key collision reachable again. The test written for
that collision last round caught it within the minute; the comment now records
that it was learned twice.

Everything else measured this round:

- The broker is the THIRD string that reaches a graph identifier, and it was
  unbounded while the header claimed there were two. A one-megabyte protocol in
  a document satisfying every other cap was measured producing a gigabyte of
  resident identifier strings, because `generateId` concatenates rather than
  hashes and the phase mints one id per node and per edge.
- The run-wide operation budget counted ACCEPTED operations, reproducing at the
  run level the exact defect the per-document cap was corrected for last round:
  a run whose every operation is refused never decrements it. Both now count
  operations EXAMINED, and `operation-cap` sets `truncated` -- it is a bound
  that stopped the operation count, which is what that flag is documented to
  mean.
- The channel-inherits-all-servers rule ran per operation. Hoisted: it depends
  only on the servers.
- A subdirectory that cannot be listed is counted rather than dropped, so a
  mixed-permission cache cannot report a clean, complete read.
- The read LOOPS, like the Actuator reader this claims to follow. A single read
  was never short across seven hundred probes on APFS, but POSIX permits it and
  FUSE mounts with `direct_io` -- the deployment this option targets -- return
  short counts. A document truncated at a line boundary still parses, so the
  failure is silent: operations vanish with `refusals: {}`.
- `parameters: {}` no longer refuses a literal address; an empty container
  states nothing and generators emit them.
- A channel that is itself a Reference Object gets its own reason instead of
  `no-address`, which was telling operators their documents omit addresses when
  the reader simply stops one hop short.
- A multi-protocol document resolves from its operation's own bindings; only
  when those are silent does an inherited multi-protocol server set refuse, and
  under `ambiguous-server-default` rather than a reason that says the document
  contradicts itself. It does not.
- HTTP and WebSocket are refused for destination minting. For a broker the
  topic is the namespace; for HTTP the host is, so keying on the path alone
  would make every service exposing `/events` one node. A `Route` already
  models an HTTP endpoint, with its method in the key.
- The sniff window is a parse gate, not a read gate, and four kilobytes refused
  a good document behind a licence header.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ingestion): pin the wiring and the reporting that were deletable

Three lines could be deleted with the whole suite green, and each of them
makes the feature partly or wholly inert: the forward from `run-analyze` into
`PipelineOptions`, the forced rebuild while documents are configured, and the
cleanup rebuild when the option is dropped. None is visible one layer up,
where the CLI test asserts on a mock's arguments.

One integration test closes all three. It drives the real `runFullAnalysis`
against a real repository and asserts, in order: the enabled run does not take
the up-to-date fast path and logs the rebuild; the destination reaches the
graph, which only happens if the option is forwarded; a document edited with
the tree clean and the commit unchanged is re-read; dropping the option
rebuilds once and removes the document-derived evidence; and the run after
that is up to date again -- the "rebuilds once" half, which rests on the
metadata being written as a fresh literal rather than merged, and which
nothing pinned.

The document lives OUTSIDE the repository on purpose. That is the workflow the
option is documented for, and it is the only one where the hazard exists:
editing a tracked file dirties the tree and forces a rebuild anyway, so an
in-repo fixture would pass with the freshness fix reverted. Verified by
reverting both: deleting the forward fails on the empty destination list,
deleting the forced rebuild fails on the missing log line.

Also pinned, each because deleting the code it covers left the suite green:
the `parameters` half of the templated-address refusal (its old test supplied
a braced address too, so the `{` half alone satisfied it); the phase actually
forwarding `symlinksSkipped`; the unconditional warning, whose whole argument
is that a tally nobody can see is not a tally -- captured through the
repository's own `_captureLogger`; a `.yml` document; a character device,
which is the case the `isFile` check exists for and which the FIFO test does
not reach; and the bound that stops a walk.

Reject an empty `--asyncapi-spec` at the CLI. It resolved to the repository
root and walked the whole tree, defeating this module's own rule that there is
no glob-based auto-discovery -- and the HTTP entry point already rejected the
identical value. Two doors onto one option must not hold different rules.

Surface the flag in the MCP context resource beside `spring_actuator`. It
matters more there than for its neighbour: Actuator annotates nodes the source
pass already found, while document reading mints destinations and edges with
no code site, and nothing said where they came from.

Log the configured path relative to the repository. The same change refuses to
persist that path to index metadata because it would record an operator's
directory layout; holding that rule for metadata and not for logs was holding
it in one place.

Both test files now clean up their temporary directories.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* style(ingestion): apply the repository's prettier contract

`quality / format` runs `npx prettier --check .`, and three files added by this
branch were not formatted to it. No behaviour changes: the reader's line
breaks and two test literals move, nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ingestion): release the mini-repo handle the document test allocated

`setupMiniRepo` documents that the caller owns cleanup, and every other test
in this file calls `repo.cleanup()` in its `finally`. The AsyncAPI document
test removed only the document directory, so each run left a temporary
repository behind.

The two owners are separate on purpose: the document directory is a SIBLING of
the repository, placed outside the working tree so that editing it cannot
dirty the tree and force a rebuild on its own. The repo's cleanup therefore
does not reach it, and both calls belong in the same block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ingestion): stop partial server evidence from reading as unanimous

Six review findings, every one a way this reader could name a broker the
document does not name. They share a shape: something is DROPPED rather than
refused, the remaining evidence agrees with itself, and an operation is
attributed with confidence to a broker its document never settled on. A wrong
broker is half a join key, so it does not produce a missing edge -- it produces
an edge to a stranger, reported as a fact.

CAPPED SERVER MAPS. A channel with no `servers` inherits all of them, and that
map is capped at 1,000. A document whose first thousand servers are Kafka and
whose thousand-and-first is JMS read as unanimously Kafka, because unanimity
was tested on the slice. Counted `server-cap` and set `truncated`, but neither
stopped the attribution. The inherited path now refuses under
`capped-server-default` -- checked BEFORE agreement, since a subset agrees with
itself for free.

ROOT SERVER REFERENCE OBJECTS. The Servers Object patterned field is
`Server Object | Reference Object`, so `{ $ref: '#/components/servers/prod' }`
is conformant. Reading `protocol` off the raw value dropped every one: an
all-reference document had no protocol at all, and -- worse -- a MIXED set lost
its disagreeing half and became unanimous. One hop is now followed, through
`#/servers` and `#/components/servers`; anything else is refused under
`unresolved-server-reference` rather than skipped.

CHANNEL BINDINGS. Only the operation's bindings were read. A conformant channel
carrying `bindings: { kafka: {} }` with no operation binding was dropped as
`protocol-unknown` while the document said plainly which broker it meant, and a
disagreement between the two levels was invisible. Both are read; a conflict is
`protocol-disagreement`.

EMPTY `servers`. "If `servers` is absent or empty, this channel MUST be
available on all the servers defined in the Servers Object" -- one sentence,
both cases. A zero-iteration loop returned `explicit: true`, which blocked the
inherited fallback and dropped valid operations.

POINTER DECODING ORDER. RFC 6901 percent-decodes the fragment BEFORE splitting
on `/`. The raw token was tested for a separator first, so `#/channels/orders%2Fv1`
passed a check it should have failed and then decoded into two segments -- a
pointer addressing `channels.orders.v1` was read as a channel named `orders/v1`,
inventing a channel the document never declared. A malformed escape is now
refused rather than resolved against its undecoded text. `~1` still resolves; it
is the pointer's own escape and belongs after segmentation.

THE SNIFF WINDOW. A fixed window decides by where the root key sits rather than
whether it is there, so every window is a false negative waiting for a longer
preamble -- 4 KiB was replaced by 64 KiB for that reason and inherited the same
defect. The whole text is scanned; it is already bounded and already in memory,
and the gate exists to skip the PARSE, which is the expensive half. A leading
UTF-8 BOM is stripped before both sniff and parse.

Twelve of the fourteen new tests were run against the unfixed reader and all
twelve failed; the other two are controls that must pass either way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(ingestion): simplify AsyncAPI pointer and binding resolution

Decode each $ref once, union binding evidence, and stop walking capped
server maps whose brokers are unused on the inherit path.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-02 22:01:15 +00:00

1944 lines
83 KiB
TypeScript

/**
* Integration coverage for the `runFullAnalysis` incremental-orchestration
* wiring (Claude PR-review Finding 2).
*
* These tests exercise the *real runtime path* — they call
* `runFullAnalysis` against a real on-disk git repo backed by a real
* LadybugDB at `<repo>/.gitnexus/`, and assert behaviours that pure
* unit tests on `diffFileHashes` / `extractChangedSubgraph` cannot
* catch:
*
* - the `isIncremental` decision (post-pipeline eligibility check)
* - `incrementalInProgress` dirty-flag set-before-mutation and
* clear-on-success
* - the importer-closure expansion (1-hop reached via the writable
* set, transitive reachable via bounded BFS)
* - the "forced full rebuild on dirty-flag-from-prior-crash" path
*
* Each test creates a temporary git repo, runs the analyzer, and asserts
* on the resulting `meta.json` and graph state. Cleanup is best-effort
* (Windows LadybugDB handle release can lag; `cleanupTempDir` retries).
*/
import { execSync } from 'child_process';
import { writeFile, readFile, mkdir, rm } from 'fs/promises';
import path from 'path';
import { afterEach, beforeAll, beforeEach, describe, it, expect, vi } from 'vitest';
import {
getStoragePaths,
saveMeta,
loadMeta,
type RepoMeta,
} from '../../src/storage/repo-manager.js';
import { setupMiniRepo as setupSharedMiniRepo } from '../helpers/mini-repo.js';
import { createTempDir } from '../helpers/test-db.js';
// Shared embedding-seed trio (this shipping review, FIX 8) — the KTD9
// zero-vector seeding pattern lives in one helper module now instead of two
// divergent copies here and in incremental-dirty-recovery.test.ts.
import {
readEmbeddingNodeIds,
seedEmbeddingForNodeId,
seedEmbeddingsForFiles,
stampEmbeddingCount,
} from '../helpers/embedding-seed.js';
import { CLASS_FRAMEWORK_ANNOTATIONS_FEATURE } from '../../src/core/analysis-features.js';
import { SCHEMA_FINGERPRINT } from '../../src/core/lbug/schema.js';
import {
SPRING_AOP_FEATURE,
SPRING_BEAN_INVENTORY_FEATURE,
SPRING_CONDITIONALS_FEATURE,
SPRING_NON_HTTP_HANDLERS_FEATURE,
SPRING_ROUTE_BINDINGS_FEATURE,
} from '../../src/core/ingestion/frameworks/spring/analysis-features.js';
import { springVendorPrefixesKey } from '../../src/core/ingestion/frameworks/spring/vendor-prefixes.js';
import {
decodeSpringAopReason,
SPRING_AOP_EVIDENCE_ID_PREFIX,
} from '../../src/core/ingestion/frameworks/spring/aop.js';
import { SPRING_AUTO_CONFIGURATION_SYNTHETIC_ID_PREFIX } from '../../src/core/ingestion/frameworks/spring/auto-configuration.js';
import {
JAVA_ENUM_INTERFACE_HERITAGE_FEATURE,
SPRING_CONFIG_BINDINGS_FEATURE,
} from '../../src/core/ingestion/languages/java/analysis-features.js';
const setupMiniRepo = () => setupSharedMiniRepo('gitnexus-incr-orch-');
/** Stage + commit everything in the temp repo (mirrors mini-repo.ts's git calls). */
const gitCommitAll = (cwd: string, message: string): void => {
execSync('git -c user.name=test -c user.email=t@t -c commit.gpgsign=false add -A', {
cwd,
stdio: 'pipe',
});
execSync(
`git -c user.name=test -c user.email=t@t -c commit.gpgsign=false commit -q -m "${message}"`,
{ cwd, stdio: 'pipe' },
);
};
const SPRING_SERVICE = 'org.springframework.stereotype.Service';
function withoutAnalysisFeature(meta: RepoMeta, featureId: string): RepoMeta {
return {
...meta,
analysisFeatures: Object.fromEntries(
Object.entries(meta.analysisFeatures ?? {}).filter(([id]) => id !== featureId),
),
};
}
async function setupSpringBeanIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-spring-bean-');
const src = path.join(repo.dbPath, 'src', 'com', 'other');
await mkdir(src, { recursive: true });
await writeFile(
path.join(src, 'WildcardService.java'),
'package com.other;\n' +
'import org.springframework.stereotype.*;\n\n' +
'@Service public class WildcardService {}\n',
'utf-8',
);
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial spring bean candidate');
return repo;
}
async function setupJavaEnumHeritageIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-java-enum-heritage-');
const src = path.join(repo.dbPath, 'src');
await mkdir(src, { recursive: true });
await writeFile(
path.join(src, 'Status.java'),
'interface Named { String label(); }\n' +
'enum Status implements Named {\n' +
' ACTIVE;\n' +
' public String label() { return "active"; }\n' +
'}\n',
'utf-8',
);
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial Java enum heritage');
return repo;
}
async function setupKotlinSpringBeanIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-spring-bean-kotlin-');
const src = path.join(repo.dbPath, 'src', 'com', 'other');
await mkdir(src, { recursive: true });
await writeFile(
path.join(src, 'WildcardService.kt'),
'package com.other\n' +
'import org.springframework.stereotype.*\n\n' +
'@Service class WildcardService\n',
'utf-8',
);
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial Kotlin spring bean candidate');
return repo;
}
const springAopAspectSource = (pointcut: string): string =>
`package com.example;\n` +
`import org.aspectj.lang.annotation.Aspect;\n` +
`import org.aspectj.lang.annotation.Before;\n\n` +
`@Aspect public class TraceAspect {\n` +
` @Before("${pointcut}") public void trace() {}\n` +
`}\n`;
async function setupSpringAopIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-spring-aop-');
const javaSrc = path.join(repo.dbPath, 'src', 'main', 'java', 'com', 'example');
const kotlinSrc = path.join(repo.dbPath, 'src', 'main', 'kotlin', 'com', 'example');
const resources = path.join(repo.dbPath, 'src', 'main', 'resources');
await Promise.all([
mkdir(javaSrc, { recursive: true }),
mkdir(kotlinSrc, { recursive: true }),
mkdir(resources, { recursive: true }),
]);
await Promise.all([
writeFile(path.join(repo.dbPath, '.gitignore'), '.gitnexus/\n', 'utf-8'),
writeFile(
path.join(javaSrc, 'FirstService.java'),
'package com.example;\n\n' +
'public class FirstService {\n' +
' public void first() {}\n' +
'}\n',
'utf-8',
),
writeFile(
path.join(javaSrc, 'TraceAspect.java'),
springAopAspectSource('within(com.example.FirstService)'),
'utf-8',
),
writeFile(
path.join(kotlinSrc, 'KotlinService.kt'),
'package com.example\n\n' +
'import org.springframework.transaction.annotation.Transactional as Tx\n\n' +
'class KotlinService {\n' +
' @Tx fun kotlinTx() {}\n' +
'}\n',
'utf-8',
),
writeFile(path.join(resources, 'application.properties'), 'feature.enabled=true\n', 'utf-8'),
]);
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial spring aop fixture');
return repo;
}
async function setupSpringBeanFactoryIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-spring-bean-factory-');
const src = path.join(repo.dbPath, 'src', 'com', 'other');
await mkdir(src, { recursive: true });
await writeFile(
path.join(src, 'WildcardConfiguration.java'),
'package com.other;\n' +
'import org.springframework.context.annotation.*;\n\n' +
'@Configuration class WildcardConfiguration {\n' +
' @Bean Gateway gateway() { return new DefaultGateway(); }\n' +
'}\n' +
'interface Gateway {}\n' +
'class DefaultGateway implements Gateway {}\n',
'utf-8',
);
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial spring bean factory');
return repo;
}
async function setupSpringConfigIncrementalRepo() {
const repo = await createTempDir('gitnexus-incr-spring-config-');
const resources = path.join(repo.dbPath, 'src', 'main', 'resources');
await mkdir(resources, { recursive: true });
await writeFile(path.join(resources, 'application.properties'), 'service.timeout=30\n', 'utf-8');
execSync('git init', { cwd: repo.dbPath, stdio: 'pipe' });
gitCommitAll(repo.dbPath, 'initial spring configuration');
return repo;
}
async function setupKotlinSpringConfigConsumerIncrementalRepo() {
const repo = await setupSpringConfigIncrementalRepo();
const kotlin = path.join(repo.dbPath, 'src', 'main', 'kotlin', 'com', 'example');
await mkdir(kotlin, { recursive: true });
await writeFile(
path.join(kotlin, 'ConfigConsumer.kt'),
'package com.example\n' +
'import org.springframework.beans.factory.annotation.Value\n\n' +
'class ConfigConsumer {\n' +
' @Value("\\${service.timeout}")\n' +
' var timeout: Int = 0\n' +
'}\n',
'utf-8',
);
gitCommitAll(repo.dbPath, 'add Kotlin Spring config consumer');
return repo;
}
async function readWildcardServiceAnnotations(repoPath: string): Promise<string[]> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
"MATCH (c:Class) WHERE c.name = 'WildcardService' " +
'RETURN c.frameworkAnnotations AS frameworkAnnotations LIMIT 1',
)) as Array<{ frameworkAnnotations?: unknown }>;
const value = rows[0]?.frameworkAnnotations;
return Array.isArray(value)
? value.filter((item): item is string => typeof item === 'string')
: [];
} finally {
await adapter.closeLbug();
}
}
async function readSpringConfigPropertyNames(repoPath: string): Promise<string[]> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
"MATCH (p:Property) WHERE p.filePath = 'src/main/resources/application.properties' " +
'RETURN p.name AS name ORDER BY p.name',
)) as Array<{ name?: unknown }>;
return rows.map((row) => String(row.name));
} finally {
await adapter.closeLbug();
}
}
async function readKotlinConfigConsumerState(repoPath: string): Promise<{
description: string;
bindingCount: number;
}> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
"MATCH (p:Property) WHERE p.name = 'timeout' " +
'RETURN p.description AS description LIMIT 1',
)) as Array<{ description?: unknown }>;
const bindings = (await adapter.executeQuery(
"MATCH (p:Property {name: 'timeout'})-[r:CodeRelation]->(c:Property) " +
"WHERE r.type = 'USES' AND r.reason STARTS WITH 'spring-config:' " +
'RETURN count(r) AS count',
)) as Array<{ count?: number | bigint }>;
return {
description: String(rows[0]?.description ?? ''),
bindingCount: Number(bindings[0]?.count ?? 0),
};
} finally {
await adapter.closeLbug();
}
}
async function readActuatorSnapshotLeakRows(
repoPath: string,
snapshotPath: string,
secretValue: string,
): Promise<Array<{ filePath?: unknown; content?: unknown }>> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH (f:File) RETURN f.filePath AS filePath, f.content AS content`,
)) as Array<{ filePath?: unknown; content?: unknown }>;
return rows.filter(
(row) =>
row.filePath === snapshotPath ||
(typeof row.content === 'string' && row.content.includes(secretValue)),
);
} finally {
await adapter.closeLbug();
}
}
async function readRuntimePropertyEvidence(repoPath: string): Promise<{
description: string;
reasons: string[];
}> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const propertyRows = (await adapter.executeQuery(
`MATCH (p:Property) WHERE p.name = 'runtime.secret' ` +
`RETURN p.description AS description LIMIT 1`,
)) as Array<{ description?: unknown }>;
const relationshipRows = (await adapter.executeQuery(
`MATCH (:File)-[r:CodeRelation]->(p:Property) WHERE p.name = 'runtime.secret' ` +
`AND r.type = 'DECLARES' RETURN r.reason AS reason ORDER BY reason`,
)) as Array<{ reason?: unknown }>;
return {
description: String(propertyRows[0]?.description ?? ''),
reasons: relationshipRows.map((row) => String(row.reason ?? '')),
};
} finally {
await adapter.closeLbug();
}
}
async function countStatusImplementsNamed(repoPath: string): Promise<number> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
"MATCH (e:Enum {name: 'Status'})-[r:CodeRelation]->(i:Interface {name: 'Named'}) " +
"WHERE r.type = 'IMPLEMENTS' RETURN count(r) AS c",
)) as Array<{ c: number | bigint }>;
return Number(rows[0]?.c ?? 0);
} finally {
await adapter.closeLbug();
}
}
async function deleteStatusImplementsNamed(repoPath: string): Promise<void> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
await adapter.executeQuery(
"MATCH (e:Enum {name: 'Status'})-[r:CodeRelation]->(i:Interface {name: 'Named'}) " +
"WHERE r.type = 'IMPLEMENTS' DELETE r",
);
} finally {
await adapter.closeLbug();
}
}
/**
* Direct count over INJECTS CodeRelation rows — mirrors pdg-mode-flip's
* countBasicBlocks: reopen the repo DB, count, close (runFullAnalysis closes
* the singleton on completion, so each count owns its own open/close).
*/
async function countInjects(repoPath: string): Promise<number> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH ()-[r:CodeRelation]->() WHERE r.type = 'INJECTS' RETURN count(r) AS c`,
)) as Array<{ c: number | bigint }>;
return Number(rows[0]?.c ?? 0);
} finally {
await adapter.closeLbug();
}
}
async function countSpringAutoConfigurationDeclarations(repoPath: string): Promise<number> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH ()-[r:CodeRelation]->() WHERE r.type = 'DECLARES' ` +
`AND (r.reason = 'spring-auto-configuration-import' ` +
`OR r.reason = 'spring-auto-configuration-factory') RETURN count(r) AS c`,
)) as Array<{ c: number | bigint }>;
return Number(rows[0]?.c ?? 0);
} finally {
await adapter.closeLbug();
}
}
async function countSpringBeanFactoryDeclarations(repoPath: string): Promise<number> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH ()-[r:CodeRelation]->() WHERE r.type = 'DECLARES' ` +
`AND r.reason STARTS WITH 'spring-bean-factory:' RETURN count(r) AS c`,
)) as Array<{ c: number | bigint }>;
return Number(rows[0]?.c ?? 0);
} finally {
await adapter.closeLbug();
}
}
async function countSpringAutoConfigurationSyntheticClasses(repoPath: string): Promise<number> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH (n:Class) WHERE n.id STARTS WITH ` +
`'${SPRING_AUTO_CONFIGURATION_SYNTHETIC_ID_PREFIX}' ` +
`RETURN count(n) AS c`,
)) as Array<{ c: number | bigint }>;
return Number(rows[0]?.c ?? 0);
} finally {
await adapter.closeLbug();
}
}
interface SpringAopPersistedRelationship {
readonly relType: string;
readonly sourceId: string;
readonly sourceName: string;
readonly targetId: string;
readonly targetName: string;
readonly reason: string;
}
interface SpringAopPersistedEvidence {
readonly id: string;
readonly description: string;
}
interface SpringAopPersistedSnapshot {
readonly relationships: readonly SpringAopPersistedRelationship[];
readonly evidence: readonly SpringAopPersistedEvidence[];
}
async function readSpringAopSnapshot(repoPath: string): Promise<SpringAopPersistedSnapshot> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const relationships = (await adapter.executeQuery(
`MATCH (s)-[r:CodeRelation]->(t) ` +
`WHERE (r.type = 'ADVISED_BY' OR r.type = 'DECLARES') ` +
`AND r.reason STARTS WITH 'spring-aop:v1:' ` +
`RETURN r.type AS relType, s.id AS sourceId, s.name AS sourceName, ` +
`t.id AS targetId, t.name AS targetName, r.reason AS reason ` +
`ORDER BY relType, sourceId, targetId, reason`,
)) as SpringAopPersistedRelationship[];
const evidence = (await adapter.executeQuery(
`MATCH (n:CodeElement) WHERE n.id STARTS WITH '${SPRING_AOP_EVIDENCE_ID_PREFIX}' ` +
`RETURN n.id AS id, n.description AS description ORDER BY id`,
)) as SpringAopPersistedEvidence[];
return { relationships, evidence };
} finally {
await adapter.closeLbug();
}
}
function assertSpringAopSnapshotShape(
snapshot: SpringAopPersistedSnapshot,
expectedAdviceSource: string,
): void {
expect(snapshot.relationships).toHaveLength(4);
expect(snapshot.evidence).toHaveLength(3);
const decoded = snapshot.relationships.map((relationship) => ({
relationship,
reason: decodeSpringAopReason(relationship.reason),
}));
expect(decoded.map(({ reason }) => reason?.kind).sort()).toEqual([
'advice',
'aspect',
'behavior',
'pointcut',
]);
expect(decoded.find(({ reason }) => reason?.kind === 'behavior')?.relationship.sourceName).toBe(
'kotlinTx',
);
const advice = decoded.find(({ reason }) => reason?.kind === 'advice')?.relationship;
expect(advice?.sourceName).toBe(expectedAdviceSource);
expect(advice?.targetName).toBe('trace');
expect(
new Set(
snapshot.relationships.map(
(relationship) =>
`${relationship.relType}\0${relationship.sourceId}\0${relationship.targetId}\0${relationship.reason}`,
),
).size,
).toBe(snapshot.relationships.length);
expect(new Set(snapshot.evidence.map(({ id }) => id)).size).toBe(snapshot.evidence.length);
}
/** Java DI fixture (#2200): `@Autowired List<IFoo>` + 2 implementers ⇒ exactly
* 2 INJECTS edges (Consumer→FooA, Consumer→FooB). Same shapes as the
* spring-di-pipeline integration fixture. */
const JAVA_DI_FIXTURE: ReadonlyArray<readonly [string, string]> = [
['IFoo.java', 'package com.example;\n\npublic interface IFoo {}\n'],
['FooA.java', 'package com.example;\n\npublic class FooA implements IFoo {}\n'],
['FooB.java', 'package com.example;\n\npublic class FooB implements IFoo {}\n'],
[
'Consumer.java',
'package com.example;\n' +
'import java.util.List;\n' +
'import org.springframework.beans.factory.annotation.Autowired;\n' +
'\n' +
'public class Consumer {\n' +
' @Autowired private List<IFoo> foos;\n' +
'}\n',
],
];
describe('runFullAnalysis — incremental orchestration', () => {
afterEach(() => {
vi.unstubAllEnvs();
});
it('first run populates fileHashes + schemaVersion and clears incrementalInProgress on success', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta).not.toBeNull();
expect(meta!.schemaFingerprint).toBe(SCHEMA_FINGERPRINT);
expect(meta!.fileHashes).toBeDefined();
expect(Object.keys(meta!.fileHashes ?? {}).length).toBeGreaterThan(0);
expect(meta!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
});
// Dirty flag MUST be cleared after a successful run.
expect(meta!.incrementalInProgress).toBeUndefined();
} finally {
await repo.cleanup();
}
}, 180_000);
it('second run on unchanged state takes the alreadyUpToDate fast path', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
const first = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(first.alreadyUpToDate).toBeUndefined();
const second = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
// lastCommit==HEAD && working tree clean (mod GitNexus output) →
// early-return fast path.
expect(second.alreadyUpToDate).toBe(true);
} finally {
await repo.cleanup();
}
}, 300_000);
it('rebuilds for Actuator snapshots and once more when runtime enrichment is disabled', async () => {
const repo = await setupMiniRepo();
const runtimeInput = 'runtime-actuator';
const runtimeInputDir = path.join(repo.dbPath, runtimeInput);
const secretValue = 'ACTUATOR_META_SECRET_2418';
try {
await mkdir(runtimeInputDir, { recursive: true });
await writeFile(
path.join(runtimeInputDir, 'env.json'),
JSON.stringify({
propertySources: [
{
name: 'systemEnvironment',
properties: { 'runtime.secret': { value: secretValue, origin: 'env' } },
},
],
}),
'utf-8',
);
gitCommitAll(repo.dbPath, 'add actuator runtime snapshot');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
const enabled = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, springActuatorPath: runtimeInput },
{ onProgress: () => {} },
);
expect(enabled.alreadyUpToDate).toBeUndefined();
const { storagePath } = getStoragePaths(repo.dbPath);
const enabledMeta = await loadMeta(storagePath);
if (enabledMeta === null)
throw new Error('Expected Actuator metadata after enabled analysis');
expect(enabledMeta.springActuator).toEqual({
enabled: true,
repoRelativeInputs: [runtimeInput],
});
expect(JSON.stringify(enabledMeta)).not.toContain(secretValue);
expect(Object.keys(enabledMeta.fileHashes ?? {})).not.toContain(`${runtimeInput}/env.json`);
expect(await readRuntimePropertyEvidence(repo.dbPath)).toEqual({
description: expect.stringContaining('Spring Actuator env runtime-confirmed'),
reasons: ['spring-actuator:env:runtime-confirmed'],
});
expect(
await readActuatorSnapshotLeakRows(repo.dbPath, `${runtimeInput}/env.json`, secretValue),
).toEqual([]);
await saveMeta(storagePath, {
...enabledMeta,
springActuator: {
enabled: true,
repoRelativeInputs: [runtimeInput, 42],
},
} as RepoMeta);
await expect(
runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} }),
).rejects.toThrow('Cannot safely disable Spring Actuator runtime enrichment');
await saveMeta(storagePath, enabledMeta);
const disableLogs: string[] = [];
const disabled = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => disableLogs.push(message) },
);
expect(disabled.alreadyUpToDate).toBeUndefined();
expect(disableLogs.join('\n')).toContain(
'Spring Actuator runtime enrichment disabled; rebuilding to remove runtime evidence.',
);
expect((await loadMeta(storagePath))?.springActuator).toEqual({
enabled: false,
repoRelativeInputs: [runtimeInput],
});
expect(
await readActuatorSnapshotLeakRows(repo.dbPath, `${runtimeInput}/env.json`, secretValue),
).toEqual([]);
const steady = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(steady.alreadyUpToDate).toBe(true);
const forcedSteady = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, force: true },
{ onProgress: () => {} },
);
expect(forcedSteady.alreadyUpToDate).toBeUndefined();
expect(
await readActuatorSnapshotLeakRows(repo.dbPath, `${runtimeInput}/env.json`, secretValue),
).toEqual([]);
expect((await loadMeta(storagePath))?.springActuator).toEqual({
enabled: false,
repoRelativeInputs: [runtimeInput],
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('a same-commit v8 index missing the global Class capability rebuilds before the fast path', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta!.schemaFingerprint).toBe(SCHEMA_FINGERPRINT);
await saveMeta(
storagePath,
withoutAnalysisFeature(meta!, CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id),
);
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain(`missing:${CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id}`);
expect((await loadMeta(storagePath))!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('a Java index missing enum heritage evidence rebuilds before the fast path (#2918)', async () => {
const repo = await setupJavaEnumHeritageIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta!.analysisFeatures).toMatchObject({
[JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.id]: JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.version,
});
expect(await countStatusImplementsNamed(repo.dbPath)).toBe(1);
await deleteStatusImplementsNamed(repo.dbPath);
expect(await countStatusImplementsNamed(repo.dbPath)).toBe(0);
await saveMeta(
storagePath,
withoutAnalysisFeature(meta!, JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.id),
);
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain(`missing:${JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.id}`);
expect((await loadMeta(storagePath))!.analysisFeatures).toMatchObject({
[JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.id]: JAVA_ENUM_INTERFACE_HERITAGE_FEATURE.version,
});
expect(await countStatusImplementsNamed(repo.dbPath)).toBe(1);
} finally {
await repo.cleanup();
}
}, 300_000);
it('a same-commit index with NO fingerprint (pre-#2798) rebuilds once, not grandfathered', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
// Every index built before the field existed. Grandfathering absence
// would stamp a fresh fingerprint onto a database whose DDL was never
// verified — permanently certifying the very index this guard catches.
await saveMeta(storagePath, { ...meta!, schemaFingerprint: undefined });
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
// An absent stamp is unattributable — this build cannot tell a pre-#2798
// index from a hand-cleared one — so the notice names no version.
expect(logs.join('\n')).toContain(
'index schema changed (built by an unidentified GitNexus build,',
);
// The extra "non-git repositories never record a schema fingerprint"
// sentence is conditional on the repo having no git dir. setupMiniRepo
// builds a real git repo, so appending it here would be a false
// explanation for an absence this build is genuinely responsible for.
expect(logs.join('\n')).not.toContain('Non-git repositories never record');
// One-time: the rebuild restamps it, so the next run is eligible again.
expect(await loadMeta(storagePath)).toMatchObject({
schemaFingerprint: SCHEMA_FINGERPRINT,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('a JVM index missing Bean inventory evidence rebuilds and restores the scoped stamp', async () => {
const repo = await setupKotlinSpringBeanIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
[SPRING_AOP_FEATURE.id]: SPRING_AOP_FEATURE.version,
[SPRING_BEAN_INVENTORY_FEATURE.id]: SPRING_BEAN_INVENTORY_FEATURE.version,
[SPRING_CONDITIONALS_FEATURE.id]: SPRING_CONDITIONALS_FEATURE.version,
[SPRING_CONFIG_BINDINGS_FEATURE.id]: SPRING_CONFIG_BINDINGS_FEATURE.version,
[SPRING_NON_HTTP_HANDLERS_FEATURE.id]: SPRING_NON_HTTP_HANDLERS_FEATURE.version,
[SPRING_ROUTE_BINDINGS_FEATURE.id]: SPRING_ROUTE_BINDINGS_FEATURE.version,
});
await saveMeta(storagePath, withoutAnalysisFeature(meta!, SPRING_BEAN_INVENTORY_FEATURE.id));
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain(`missing:${SPRING_BEAN_INVENTORY_FEATURE.id}`);
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([SPRING_SERVICE]);
expect((await loadMeta(storagePath))!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
[SPRING_AOP_FEATURE.id]: SPRING_AOP_FEATURE.version,
[SPRING_BEAN_INVENTORY_FEATURE.id]: SPRING_BEAN_INVENTORY_FEATURE.version,
[SPRING_CONDITIONALS_FEATURE.id]: SPRING_CONDITIONALS_FEATURE.version,
[SPRING_CONFIG_BINDINGS_FEATURE.id]: SPRING_CONFIG_BINDINGS_FEATURE.version,
[SPRING_NON_HTTP_HANDLERS_FEATURE.id]: SPRING_NON_HTTP_HANDLERS_FEATURE.version,
[SPRING_ROUTE_BINDINGS_FEATURE.id]: SPRING_ROUTE_BINDINGS_FEATURE.version,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('rebuilds a JVM index when the registered Spring vendor prefixes change', async () => {
const repo = await setupSpringBeanIncrementalRepo();
try {
vi.stubEnv('GITNEXUS_SPRING_VENDOR_PREFIXES', 'Win');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
expect((await loadMeta(storagePath))?.springVendorPrefixes).toBe(springVendorPrefixesKey());
vi.stubEnv('GITNEXUS_SPRING_VENDOR_PREFIXES', 'Acme,Win');
const logs: string[] = [];
const rebuilt = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(rebuilt.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain('Spring vendor mapping prefixes changed');
expect((await loadMeta(storagePath))?.springVendorPrefixes).toBe(springVendorPrefixesKey());
const steady = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(steady.alreadyUpToDate).toBe(true);
} finally {
await repo.cleanup();
}
}, 300_000);
it('a config-only index missing Spring config evidence rebuilds and restores the scoped stamp', async () => {
const repo = await setupSpringConfigIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
[SPRING_CONFIG_BINDINGS_FEATURE.id]: SPRING_CONFIG_BINDINGS_FEATURE.version,
});
await saveMeta(storagePath, withoutAnalysisFeature(meta!, SPRING_CONFIG_BINDINGS_FEATURE.id));
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain(`missing:${SPRING_CONFIG_BINDINGS_FEATURE.id}`);
expect(await readSpringConfigPropertyNames(repo.dbPath)).toEqual(['service.timeout']);
expect((await loadMeta(storagePath))!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
[SPRING_CONFIG_BINDINGS_FEATURE.id]: SPRING_CONFIG_BINDINGS_FEATURE.version,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('persists Kotlin unresolved markers when a Spring config key is deleted incrementally', async () => {
const repo = await setupKotlinSpringConfigConsumerIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readKotlinConfigConsumerState(repo.dbPath)).toEqual({
description: '',
bindingCount: 1,
});
const configPath = path.join(
repo.dbPath,
'src',
'main',
'resources',
'application.properties',
);
await writeFile(configPath, '', 'utf-8');
gitCommitAll(repo.dbPath, 'delete Spring config key');
const logs: string[] = [];
await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(logs.join('\n')).toContain('Spring config consumer property drift');
expect(await readKotlinConfigConsumerState(repo.dbPath)).toEqual({
description: 'Spring config unresolved: service.timeout',
bindingCount: 0,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('adding the first JVM file re-evaluates capabilities after the pipeline and avoids a top-up', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
expect((await loadMeta(storagePath))!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
});
await writeFile(
path.join(repo.dbPath, 'src', 'FirstBean.kt'),
'import org.springframework.stereotype.Service\n\n@Service class FirstBean\n',
'utf-8',
);
gitCommitAll(repo.dbPath, 'add first JVM source file');
const logs: string[] = [];
await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
expect(logs.join('\n')).toContain(`missing:${SPRING_BEAN_INVENTORY_FEATURE.id}`);
expect(logs.join('\n')).not.toContain('Incremental:');
expect((await loadMeta(storagePath))!.analysisFeatures).toEqual({
[CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.id]: CLASS_FRAMEWORK_ANNOTATIONS_FEATURE.version,
[SPRING_AOP_FEATURE.id]: SPRING_AOP_FEATURE.version,
[SPRING_BEAN_INVENTORY_FEATURE.id]: SPRING_BEAN_INVENTORY_FEATURE.version,
[SPRING_CONDITIONALS_FEATURE.id]: SPRING_CONDITIONALS_FEATURE.version,
[SPRING_CONFIG_BINDINGS_FEATURE.id]: SPRING_CONFIG_BINDINGS_FEATURE.version,
[SPRING_NON_HTTP_HANDLERS_FEATURE.id]: SPRING_NON_HTTP_HANDLERS_FEATURE.version,
[SPRING_ROUTE_BINDINGS_FEATURE.id]: SPRING_ROUTE_BINDINGS_FEATURE.version,
});
} finally {
await repo.cleanup();
}
}, 300_000);
it('replaces Spring AOP evidence across real incremental runs without duplicates', async () => {
const repo = await setupSpringAopIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const firstSnapshot = await readSpringAopSnapshot(repo.dbPath);
assertSpringAopSnapshotShape(firstSnapshot, 'first');
const firstPointcutEvidenceId = firstSnapshot.relationships.find(
({ reason }) => decodeSpringAopReason(reason)?.kind === 'pointcut',
)?.targetId;
expect(firstPointcutEvidenceId).toBeDefined();
const aspectPath = path.join(
repo.dbPath,
'src',
'main',
'java',
'com',
'example',
'TraceAspect.java',
);
await writeFile(
aspectPath,
springAopAspectSource(
'@annotation(org.springframework.transaction.annotation.Transactional)',
),
'utf-8',
);
gitCommitAll(repo.dbPath, 'retarget spring advice');
const retargetLogs: string[] = [];
const retargeted = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => retargetLogs.push(message) },
);
expect(retargeted.alreadyUpToDate).toBeUndefined();
expect(retargetLogs.join('\n')).toContain('Incremental: changed=1');
expect(retargetLogs.join('\n')).not.toContain('switching to a full DB write');
const secondSnapshot = await readSpringAopSnapshot(repo.dbPath);
assertSpringAopSnapshotShape(secondSnapshot, 'kotlinTx');
const secondPointcutEvidenceId = secondSnapshot.relationships.find(
({ reason }) => decodeSpringAopReason(reason)?.kind === 'pointcut',
)?.targetId;
expect(secondPointcutEvidenceId).toBeDefined();
expect(secondPointcutEvidenceId).not.toBe(firstPointcutEvidenceId);
expect(secondSnapshot.evidence.map(({ id }) => id)).not.toContain(firstPointcutEvidenceId);
const propertyPath = path.join(
repo.dbPath,
'src',
'main',
'resources',
'application.properties',
);
await writeFile(propertyPath, 'feature.enabled=false\n', 'utf-8');
gitCommitAll(repo.dbPath, 'change unrelated resource');
const replayLogs: string[] = [];
const replayed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{
onLog: (message) => replayLogs.push(message),
onProgress: () => {},
},
);
expect(replayed.alreadyUpToDate).toBeUndefined();
expect(replayLogs.join('\n')).toContain('Incremental: changed=1');
expect(replayLogs.join('\n')).not.toContain('switching to a full DB write');
expect(await readSpringAopSnapshot(repo.dbPath)).toEqual(secondSnapshot);
} finally {
await repo.cleanup();
}
}, 600_000);
it('second run after a comment-only edit takes the incremental path, clears the dirty flag, and preserves graph stats exactly', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const firstMeta = await loadMeta(storagePath);
// Modify a source file with a COMMENT-ONLY edit — by construction
// this changes the content hash (driving the incremental code path)
// without changing any symbol, scope binding, call edge, import,
// or community membership. Therefore every graph-stat invariant
// (files / nodes / edges / communities / processes) MUST be
// bit-identical to the first run. Anything else is a regression.
const target = path.join(repo.dbPath, 'src', 'logger.ts');
const before = await readFile(target, 'utf-8');
await writeFile(target, before + '\n// touched by test\n', 'utf-8');
const second = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
// The early-return alreadyUpToDate path must NOT fire (the dirty
// tree should kick the run through to incremental writeback).
expect(second.alreadyUpToDate).toBeUndefined();
const secondMeta = await loadMeta(storagePath);
expect(secondMeta).not.toBeNull();
// Dirty flag must be cleared on success.
expect(secondMeta!.incrementalInProgress).toBeUndefined();
// fileHashes[logger.ts] must have rotated to the new content.
expect(secondMeta!.fileHashes?.['src/logger.ts']).toBeDefined();
expect(secondMeta!.fileHashes?.['src/logger.ts']).not.toBe(
firstMeta!.fileHashes?.['src/logger.ts'],
);
// Exact-equality stats invariant. DoD §2.7: avoid bounds-only
// assertions that would mask a regression dropping half the graph.
expect(secondMeta!.stats?.files).toBe(firstMeta!.stats?.files);
expect(secondMeta!.stats?.nodes).toBe(firstMeta!.stats?.nodes);
expect(secondMeta!.stats?.edges).toBe(firstMeta!.stats?.edges);
expect(secondMeta!.stats?.communities).toBe(firstMeta!.stats?.communities);
expect(secondMeta!.stats?.processes).toBe(firstMeta!.stats?.processes);
} finally {
await repo.cleanup();
}
}, 300_000);
it('skips the framework annotation drift query when no Bean source changed', async () => {
const repo = await setupMiniRepo();
try {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const target = path.join(repo.dbPath, 'src', 'logger.ts');
const before = await readFile(target, 'utf-8');
await writeFile(target, before + '\n// non-bean-source incremental touch\n', 'utf-8');
const querySpy = vi.spyOn(adapter, 'executeQuery');
try {
const incremental = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(incremental.alreadyUpToDate).toBeUndefined();
expect(incremental.incrementalStats).toMatchObject({
changedFiles: 1,
affectedDependents: 2,
deletedFiles: 0,
writeMode: 'incremental',
});
expect(incremental.incrementalStats?.reparsedFiles).toBe(1);
expect(
querySpy.mock.calls.some(
([query]) =>
typeof query === 'string' &&
query.includes('RETURN c.id AS id, c.frameworkAnnotations AS frameworkAnnotations'),
),
).toBe(false);
} finally {
querySpy.mockRestore();
}
} finally {
await repo.cleanup();
}
}, 300_000);
it('incremental output is byte-equivalent to a full rebuild (incremental ≡ --force on the same repo state)', async () => {
// The central correctness contract of this PR: an incremental run
// and a full rebuild from the same repo state must produce identical
// graph stats. We exercise it end-to-end:
//
// 1. setup mini-repo + run analyze (populates the index)
// 2. edit one source file (comment-only — same graph)
// 3. run incremental analyze → record secondMeta
// 4. run analyze --force from the same state → record forceMeta
// 5. assert every stats invariant is exactly equal.
//
// Steps 3 and 4 share the same on-disk file contents, so any
// divergence is purely an artifact of the writeback strategy. If
// any invariant differs, the PR's load-bearing claim is violated.
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
// Step 1: initial index.
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
// Step 2: comment-only edit, same as the test above.
const target = path.join(repo.dbPath, 'src', 'logger.ts');
const original = await readFile(target, 'utf-8');
await writeFile(target, original + '\n// equivalence test touch\n', 'utf-8');
// Step 3: incremental writeback for the edited file.
const incremental = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(incremental.alreadyUpToDate).toBeUndefined();
const { storagePath } = getStoragePaths(repo.dbPath);
const secondMeta = await loadMeta(storagePath);
expect(secondMeta).not.toBeNull();
// Step 4: force a full rebuild from the SAME on-disk file state.
const forced = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, force: true },
{ onProgress: () => {} },
);
expect(forced.alreadyUpToDate).toBeUndefined();
const forceMeta = await loadMeta(storagePath);
expect(forceMeta).not.toBeNull();
// Step 5: exact-equality across every stat. `toEqual` would also
// work but `toBe` per-field makes a failure pinpoint the field.
expect(secondMeta!.stats?.files).toBe(forceMeta!.stats?.files);
expect(secondMeta!.stats?.nodes).toBe(forceMeta!.stats?.nodes);
expect(secondMeta!.stats?.edges).toBe(forceMeta!.stats?.edges);
expect(secondMeta!.stats?.communities).toBe(forceMeta!.stats?.communities);
expect(secondMeta!.stats?.processes).toBe(forceMeta!.stats?.processes);
} finally {
await repo.cleanup();
}
}, 600_000);
it('rewrites unchanged Spring bean metadata when same-package shadowing changes', async () => {
const repo = await setupSpringBeanIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([SPRING_SERVICE]);
const shadow = path.join(repo.dbPath, 'src', 'com', 'other', 'Service.java');
await writeFile(shadow, 'package com.other;\npublic @interface Service {}\n', 'utf-8');
gitCommitAll(repo.dbPath, 'add same-package annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([]);
await rm(shadow);
gitCommitAll(repo.dbPath, 'remove same-package annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([SPRING_SERVICE]);
} finally {
await repo.cleanup();
}
}, 600_000);
it('rewrites unchanged Kotlin Spring bean metadata when same-package shadowing changes', async () => {
const repo = await setupKotlinSpringBeanIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([SPRING_SERVICE]);
const shadow = path.join(repo.dbPath, 'src', 'com', 'other', 'Service.kt');
await writeFile(shadow, 'package com.other\nannotation class Service\n', 'utf-8');
gitCommitAll(repo.dbPath, 'add same-package Kotlin annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([]);
await rm(shadow);
gitCommitAll(repo.dbPath, 'remove same-package Kotlin annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await readWildcardServiceAnnotations(repo.dbPath)).toEqual([SPRING_SERVICE]);
} finally {
await repo.cleanup();
}
}, 600_000);
it('rewrites unchanged Bean factory declarations when same-package shadowing changes', async () => {
const repo = await setupSpringBeanFactoryIncrementalRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringBeanFactoryDeclarations(repo.dbPath)).toBe(1);
const shadow = path.join(repo.dbPath, 'src', 'com', 'other', 'Bean.java');
await writeFile(shadow, 'package com.other;\npublic @interface Bean {}\n', 'utf-8');
gitCommitAll(repo.dbPath, 'add same-package Bean annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringBeanFactoryDeclarations(repo.dbPath)).toBe(0);
await rm(shadow);
gitCommitAll(repo.dbPath, 'remove same-package Bean annotation shadow');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringBeanFactoryDeclarations(repo.dbPath)).toBe(1);
} finally {
await repo.cleanup();
}
}, 600_000);
// #2409: a large-fraction effective write set must escalate to the full DB
// write plan (wipe + bulk COPY of the already-built graph) instead of the
// surgical per-file writeback — at that size the surgical plan measured
// SLOWER than a full load and its delete storm is the write pattern behind
// the reported native mid-writeback deaths. The escalated result must be
// indistinguishable from a --force rebuild of the same state.
it('a hub edit whose write set covers most of a large repo escalates to the full DB write plan (#2409)', async () => {
const repo = await setupMiniRepo();
try {
// Grow the repo past INCREMENTAL_ESCALATION_MIN_FILES (50) with a hub
// imported by every generated file: touching the hub pulls the whole
// family into the importer closure → write fraction ≈ 100% > 50%.
const src = path.join(repo.dbPath, 'src');
await writeFile(
path.join(src, 'hub.ts'),
'export function hubValue(x: number): number {\n return x + 1;\n}\n',
'utf-8',
);
for (let i = 0; i < 60; i++) {
await writeFile(
path.join(src, `spoke-${String(i).padStart(3, '0')}.ts`),
`import { hubValue } from './hub';\n\nexport function spoke${i}(): number {\n return hubValue(${i});\n}\n`,
'utf-8',
);
}
gitCommitAll(repo.dbPath, 'add hub + spokes');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
// Touch the hub — comment-only, so graph stats must be preserved.
const hub = path.join(src, 'hub.ts');
await writeFile(hub, (await readFile(hub, 'utf-8')) + '// escalation touch\n', 'utf-8');
const logs: string[] = [];
const incremental = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (m) => logs.push(m) },
);
expect(incremental.alreadyUpToDate).toBeUndefined();
const joined = logs.join('\n');
// The importer expansion fired AND the valve rerouted the write plan.
expect(joined).toContain('importer(s) added to writable set');
expect(joined).toContain('switching to a full DB write');
const { storagePath } = getStoragePaths(repo.dbPath);
const escalatedMeta = await loadMeta(storagePath);
expect(escalatedMeta).not.toBeNull();
// Dirty flag cleared on success — the escalated plan converges on the
// same meta-save as every other successful run.
expect(escalatedMeta!.incrementalInProgress).toBeUndefined();
// The escalated write must be indistinguishable from --force on the
// same state: any stale surviving row would show up as a stats delta.
await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, force: true },
{ onProgress: () => {} },
);
const forcedMeta = await loadMeta(storagePath);
expect(escalatedMeta!.stats?.files).toBe(forcedMeta!.stats?.files);
expect(escalatedMeta!.stats?.nodes).toBe(forcedMeta!.stats?.nodes);
expect(escalatedMeta!.stats?.edges).toBe(forcedMeta!.stats?.edges);
expect(escalatedMeta!.stats?.communities).toBe(forcedMeta!.stats?.communities);
expect(escalatedMeta!.stats?.processes).toBe(forcedMeta!.stats?.processes);
} finally {
await repo.cleanup();
}
}, 600_000);
// U4 / KTD10 (tri-review 4669518496): a SURGICAL preserve-mode run (below
// both valve gates) must keep embedding rows in lockstep with their files
// now that deleteNodesForFiles really deletes embedding rows via the
// nodeId join:
// - changed-file rows are deleted with their nodes and RESTORED from the
// cache (the old insert-all restore lost them when a surviving row's
// PK conflict aborted the rest of the batch),
// - deleted-file rows are gone (join-delete) and NOT resurrected by the
// restore (live-graph filter),
// - unchanged rows are untouched (restore-scope filter, no conflicts),
// - a LEGACY ORPHAN row — stranded while the embedding delete was a
// no-op, unreachable by the node join forever — is swept by exact id
// (this shipping review, FIX 3).
it('surgical incremental run keeps embedding rows in lockstep: changed restored, deleted gone, unchanged intact, legacy orphan swept (tri-review 4669518496 KTD10 + FIX 3)', async () => {
const repo = await setupMiniRepo();
try {
const CHANGED_FILE = 'src/logger.ts';
const UNCHANGED_FILE = 'src/db.ts';
const DELETED_FILE = 'src/formatter.ts';
// A fabricated nodeId no graph will ever contain: the P2-1-era no-op
// delete left rows like this stranded in real DBs (schema version
// stays 6, so they are still out there).
const LEGACY_ORPHAN_NODE_ID = 'Function:src/ghost.ts:ghost:1';
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const idsByFile = await seedEmbeddingsForFiles(
repo.dbPath,
[CHANGED_FILE, UNCHANGED_FILE, DELETED_FILE],
3,
);
for (const fp of [CHANGED_FILE, UNCHANGED_FILE, DELETED_FILE]) {
expect((idsByFile.get(fp) ?? []).length).toBeGreaterThan(0);
}
await seedEmbeddingForNodeId(repo.dbPath, LEGACY_ORPHAN_NODE_ID);
const seededTotal = [...idsByFile.values()].flat().length + 1;
await stampEmbeddingCount(storagePath, seededTotal);
// One file modified (comment-only, appended at EOF so node ids keep
// their line numbers), one file deleted — committed so lastCommit moves.
const target = path.join(repo.dbPath, CHANGED_FILE);
await writeFile(
target,
(await readFile(target, 'utf-8')) + '\n// embeddings parity touch\n',
'utf-8',
);
await rm(path.join(repo.dbPath, DELETED_FILE));
gitCommitAll(repo.dbPath, 'modify logger + delete formatter');
const logs: string[] = [];
const run = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (m) => logs.push(m) },
);
expect(run.alreadyUpToDate).toBeUndefined();
// 7-file repo — far below the 50-file valve floor: this MUST have been
// the surgical write plan, or every assertion below is vacuously about
// the escalated path instead.
expect(logs.join('\n')).not.toContain('switching to a full DB write');
// The surgical run swept the fabricated legacy orphan by exact id
// (FIX 3). The logged count also includes DELETED_FILE's cached rows —
// live-graph rejects whose DB rows were already join-deleted with the
// file, so their exact-id DELETEs match nothing (documented no-op).
const expectedSweepCount = 1 + (idsByFile.get(DELETED_FILE) ?? []).length;
expect(logs.join('\n')).toContain(
`Swept ${expectedSweepCount} cached embedding row(s) with no live owning node`,
);
const after = await loadMeta(storagePath);
expect(after!.incrementalInProgress).toBeUndefined();
const expectedSurvivors = [
...(idsByFile.get(CHANGED_FILE) ?? []),
...(idsByFile.get(UNCHANGED_FILE) ?? []),
];
// stats.embeddings excludes both the deleted-file rows AND the swept
// legacy orphan.
expect(after!.stats?.embeddings).toBe(expectedSurvivors.length);
// Exact surviving nodeId set — pins all four behaviors at once (a
// batch-abort loss, a leaked deleted-file row, a dropped unchanged
// row, or a lingering legacy orphan each break set equality).
expect((await readEmbeddingNodeIds(repo.dbPath)).sort()).toEqual(
[...expectedSurvivors].sort(),
);
} finally {
await repo.cleanup();
}
}, 600_000);
// #2409 defect 2 (dirty-flag recovery parks WAL/shadow sidecars before any
// open) is covered in incremental-dirty-recovery.test.ts — its own file so
// the cross-platform CI matrix runs it on windows-latest without pulling in
// this whole suite.
it('a stale incrementalInProgress flag at startup forces a full rebuild that clears it', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
// First run lays down a normal index.
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
// Manually corrupt meta.json with a stale dirty flag — simulates
// a crashed previous incremental run.
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta).not.toBeNull();
const tampered: RepoMeta = {
...meta!,
incrementalInProgress: {
startedAt: Date.now() - 60_000,
toWriteCount: 3,
phase: 'load-graph',
importerExpansion: 153,
effectiveWriteCount: 167,
deleteCount: 169,
},
};
await saveMeta(storagePath, tampered);
const logs: string[] = [];
// Next run must detect the flag, force a full rebuild (which
// overwrites meta), and clear the flag.
const recovered = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
// A full rebuild was taken — the alreadyUpToDate fast path
// explicitly cannot fire because the dirty-flag check rewrote
// `options.force` to true.
expect(recovered.alreadyUpToDate).toBeUndefined();
const after = await loadMeta(storagePath);
expect(after!.incrementalInProgress).toBeUndefined();
expect(logs.join('\n')).toContain(
'last dirty state: phase=load-graph, toWrite=3, importerExpansion=153, effectiveWrite=167, deleteCount=169',
);
} finally {
await repo.cleanup();
}
}, 300_000);
// An index carrying a schema stamp that is not this build's must not take the
// alreadyUpToDate fast path. The schema mismatch guard runs before lastCommit
// equality can short-circuit the pipeline, so node-identity migrations receive
// a full rebuild. Pinned on the RESULT (no fast path, restamped meta) rather
// than the log line, so the ordering invariant survives a reworded notice.
it('a foreign schema fingerprint forces a full rebuild on an unchanged-commit re-analyze', async () => {
const repo = await setupMiniRepo();
try {
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
// First run stamps the digest of the DDL this build creates.
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta).not.toBeNull();
expect(meta!.schemaFingerprint).toBe(SCHEMA_FINGERPRINT);
// Simulate an index whose tables were created from a different DDL, at
// the same commit with a clean tree. Well-formed (12 lowercase hex, so it
// clears the echo-shape gate) but not this build's — every other fast-path
// condition holds, so only the schema guard can stop the early return.
const downgraded: RepoMeta = { ...meta!, schemaFingerprint: 'b1c2d3e4f5a6' };
await saveMeta(storagePath, downgraded);
const logs: string[] = [];
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => logs.push(message) },
);
// Pipeline actually ran (schemaFingerprint mismatch → force=true), and the
// notice names the stamp it rejected rather than a generic placeholder.
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
expect(logs.join('\n')).toContain('index schema changed (built by b1c2d3e4f5a6,');
// And the rebuild restamped this build's digest (that path runs saveMeta).
const restamped = await loadMeta(storagePath);
expect(restamped!.schemaFingerprint).toBe(SCHEMA_FINGERPRINT);
} finally {
await repo.cleanup();
}
}, 300_000);
// #2331/#2339: mirrors the schema-fingerprint mismatch test above, but for
// the CJK segmentation mode stamp. Uses a non-default mode ('bigram') rather
// than 'none' — with the default, (undefined ?? 'none') !== 'none' is
// false regardless of whether the stamp was ever actually written, so a
// dropped-stamp bug would pass this test vacuously. 'bigram' makes an
// omitted stamp manifest as a real comparator mismatch instead.
it('a stale cjkSegmentation stamp forces a full rebuild on an unchanged-commit re-analyze', async () => {
const repo = await setupMiniRepo();
try {
vi.stubEnv('GITNEXUS_FTS_CJK_SEGMENTATION', 'bigram');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
const { storagePath } = getStoragePaths(repo.dbPath);
const meta = await loadMeta(storagePath);
expect(meta).not.toBeNull();
expect(meta!.cjkSegmentation).toBe('bigram');
// Simulate a repo indexed under 'none' (or a pre-#2339 build with no
// stamp at all) that's now being served/re-analyzed with bigram mode.
const downgraded: RepoMeta = { ...meta!, cjkSegmentation: 'none' };
await saveMeta(storagePath, downgraded);
const reanalyzed = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
// Pipeline actually ran (cjkSegmentation mismatch → force=true).
expect(reanalyzed.alreadyUpToDate).toBeUndefined();
// And the meta is restamped to the live resolved mode.
const restamped = await loadMeta(storagePath);
expect(restamped!.cjkSegmentation).toBe('bigram');
} finally {
await repo.cleanup();
}
}, 300_000);
it('first-ever analyze of a brand-new repo proceeds without a spurious CJK mode force-rebuild', async () => {
const repo = await setupMiniRepo();
try {
const { storagePath } = getStoragePaths(repo.dbPath);
// No meta.json exists yet — existingMeta is falsy, so the
// cjkSegmentationModeMismatch guard is skipped entirely (never calls
// the comparator), same as the pdg/schemaVersion guards above it.
expect(await loadMeta(storagePath)).toBeNull();
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
const result = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(result.alreadyUpToDate).toBeUndefined();
const meta = await loadMeta(storagePath);
expect(meta!.cjkSegmentation).toBe('none');
} finally {
await repo.cleanup();
}
}, 300_000);
// U7 (#2200): the INJECTS delete-before-writeback must be UNCONDITIONAL.
// extractChangedSubgraph re-includes ALL INJECTS edges from the fresh graph
// on every incremental run (isGraphWideRelType), and CodeRelation has no PK
// and no read-side dedup — so a pdg-gated delete (literal TAINT_PATH
// mirroring) would append without deleting on every non-pdg incremental
// run: N runs = N copies of every INJECTS row. This test is the assertion
// that catches exactly that mistake.
it('incremental runs neither strand nor duplicate INJECTS edges (delete-all is not pdg-gated) (#2200)', async () => {
const repo = await setupMiniRepo();
try {
const src = path.join(repo.dbPath, 'src');
for (const [name, content] of JAVA_DI_FIXTURE) {
await writeFile(path.join(src, name), content, 'utf-8');
}
gitCommitAll(repo.dbPath, 'add java di fixture');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
// Full index: Consumer.foos fans out to the two IFoo implementers.
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countInjects(repo.dbPath)).toBe(2);
// Incremental run 1: comment-only touch of an UNRELATED file (none of
// the Java DI files change), committed so lastCommit moves.
const target = path.join(src, 'logger.ts');
const beforeFirstTouch = await readFile(target, 'utf-8');
await writeFile(target, beforeFirstTouch + '\n// di idempotency touch 1\n', 'utf-8');
gitCommitAll(repo.dbPath, 'unrelated touch 1');
const run1 = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(run1.alreadyUpToDate).toBeUndefined();
expect(await countInjects(repo.dbPath)).toBe(2);
// Incremental run 2: second unrelated touch. A gated delete would have
// appended two more rows per writeback (4 by now) — must still be 2.
const beforeSecondTouch = await readFile(target, 'utf-8');
await writeFile(target, beforeSecondTouch + '\n// di idempotency touch 2\n', 'utf-8');
gitCommitAll(repo.dbPath, 'unrelated touch 2');
const run2 = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(run2.alreadyUpToDate).toBeUndefined();
expect(await countInjects(repo.dbPath)).toBe(2);
} finally {
await repo.cleanup();
}
}, 600_000);
it('incremental runs do not duplicate repository-wide Spring DECLARES edges (#2415)', async () => {
const repo = await setupMiniRepo();
try {
const sourceDir = path.join(repo.dbPath, 'src', 'main', 'java', 'com', 'example');
const metadataDir = path.join(repo.dbPath, 'src', 'main', 'resources', 'META-INF', 'spring');
await mkdir(sourceDir, { recursive: true });
await mkdir(metadataDir, { recursive: true });
await writeFile(
path.join(metadataDir, 'org.springframework.boot.autoconfigure.AutoConfiguration.imports'),
'com.example.ExampleAutoConfiguration\n',
'utf-8',
);
gitCommitAll(repo.dbPath, 'add auto configuration metadata');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringAutoConfigurationDeclarations(repo.dbPath)).toBe(1);
expect(await countSpringAutoConfigurationSyntheticClasses(repo.dbPath)).toBe(1);
const target = path.join(repo.dbPath, 'src', 'logger.ts');
for (const run of [1, 2]) {
const before = await readFile(target, 'utf-8');
await writeFile(target, `${before}\n// auto-register idempotency touch ${run}\n`, 'utf-8');
gitCommitAll(repo.dbPath, `unrelated auto-register touch ${run}`);
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringAutoConfigurationDeclarations(repo.dbPath)).toBe(1);
expect(await countSpringAutoConfigurationSyntheticClasses(repo.dbPath)).toBe(1);
}
await writeFile(
path.join(sourceDir, 'ExampleAutoConfiguration.java'),
'package com.example;\npublic class ExampleAutoConfiguration {}\n',
'utf-8',
);
gitCommitAll(repo.dbPath, 'add source for metadata-only auto configuration');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
expect(await countSpringAutoConfigurationDeclarations(repo.dbPath)).toBe(1);
expect(await countSpringAutoConfigurationSyntheticClasses(repo.dbPath)).toBe(0);
} finally {
await repo.cleanup();
}
}, 600_000);
});
/**
* U3 (tri-review 4669518496 P1): the #2409 escalation valve wipes the DB
* files — HNSW vector index included. The Phase 3.5 restore brought the
* embedding ROWS back, but nothing recreated the index and meta still
* stamped `vector-index`: semantic search on a >10k-embedding repo silently
* lost its vector lane while meta certified otherwise. This suite pins the
* fix end-to-end: escalated preserve-mode run → index recreated → meta honest.
*
* SEPARATE from the `--force` escalation parity test above (KTD9): seeding
* embeddings there would make its force leg derive forceRegenerateEmbeddings
* and boot a real embedder in CI. This run stays preserve-only (no force).
*
* Skip-gated on VECTOR availability (the lbug-vector-extension.test.ts
* pattern): skipped only where the extension genuinely cannot load —
* no platform is categorically excluded any more (#2623 follow-up) — and
* where it cannot, the honest stamp is
* 'exact-scan', which the unit-level wiring pin in
* run-analyze-fts-repair.test.ts covers platform-independently.
*/
describe('runFullAnalysis — escalated wipe recreates the vector index (#2409, tri-review 4669518496 P1)', () => {
let vectorAvailable = false;
let skipWarned = false;
beforeAll(async () => {
// Probe VECTOR the way the analyze write path loads it. loadVectorExtension
// needs an open connection, and this suite (unlike the withTestLbugDB
// vector suites) has no ambient DB — probe against a scratch one.
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { resolveAnalyzeInstallPolicy } = await import('../../src/core/lbug/extension-loader.js');
const tmp = await createTempDir('gitnexus-incr-orch-vector-probe-');
try {
await adapter.initLbug(path.join(tmp.dbPath, 'probe-lbug'));
vectorAvailable = await adapter.loadVectorExtension(undefined, {
policy: resolveAnalyzeInstallPolicy(),
});
} finally {
await adapter.closeLbug();
await tmp.cleanup();
}
}, 120_000);
beforeEach((ctx) => {
if (!vectorAvailable) {
if (!skipWarned) {
skipWarned = true;
console.warn(
'[incremental-orchestration] Skipping vector-index recreation test — the ' +
'LadybugDB VECTOR extension is unavailable (unsupported platform or ' +
'could not be installed).',
);
}
ctx.skip();
}
});
it('recreates the HNSW index after an escalated wipe-and-restore and stamps meta honestly (tri-review 4669518496 P1)', async () => {
const repo = await setupMiniRepo();
try {
// Hub+spokes repo shape VERBATIM from the escalation parity test above:
// the escalated run must clear BOTH valve gates (deleteCount ≥ 50 AND
// fraction > 0.5). A smaller fixture would silently take the surgical
// path, whose surviving index makes every assertion below pass
// vacuously.
const src = path.join(repo.dbPath, 'src');
await writeFile(
path.join(src, 'hub.ts'),
'export function hubValue(x: number): number {\n return x + 1;\n}\n',
'utf-8',
);
for (let i = 0; i < 60; i++) {
await writeFile(
path.join(src, `spoke-${String(i).padStart(3, '0')}.ts`),
`import { hubValue } from './hub';\n\nexport function spoke${i}(): number {\n return hubValue(${i});\n}\n`,
'utf-8',
);
}
gitCommitAll(repo.dbPath, 'add hub + spokes');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
await runFullAnalysis(repo.dbPath, { skipAgentsMd: true }, { onProgress: () => {} });
// 20 zero-vector embeddings on real Function nodes (one per spoke —
// fabricated ids would be dropped by the Phase 3.5 live-graph filter)
// + a stats stamp so deriveEmbeddingMode sees an embedded repo
// (preserve mode — the run below passes NO force flag, so no embedder
// ever fires; KTD9).
const SEED_COUNT = 20;
const { storagePath, lbugPath } = getStoragePaths(repo.dbPath);
const seedFiles: string[] = [];
for (let i = 0; i < SEED_COUNT; i++) {
seedFiles.push(`src/spoke-${String(i).padStart(3, '0')}.ts`);
}
const idsByFile = await seedEmbeddingsForFiles(repo.dbPath, seedFiles, 1);
expect([...idsByFile.values()].flat().length).toBe(SEED_COUNT);
await stampEmbeddingCount(storagePath, SEED_COUNT);
// Touch the hub — the importer closure covers the whole family, the
// valve fires, and the DB (index included) is wiped mid-run.
const hub = path.join(src, 'hub.ts');
await writeFile(hub, (await readFile(hub, 'utf-8')) + '// index recreation touch\n', 'utf-8');
const logs: string[] = [];
const escalated = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (m) => logs.push(m) },
);
expect(escalated.alreadyUpToDate).toBeUndefined();
// The valve rerouted the write plan — without this the index assertions
// below test the surgical path's surviving index, not the recreation.
expect(logs.join('\n')).toContain('switching to a full DB write');
// Every cached row was restored across the wipe…
const after = await loadMeta(storagePath);
expect(after!.incrementalInProgress).toBeUndefined();
expect(after!.stats?.embeddings).toBe(SEED_COUNT);
// …meta stamps what the DB actually holds…
expect(after!.capabilities?.vectorSearch.status).toBe('vector-index');
// …and the DB really does hold a recreated HNSW index (SHOW_INDEXES
// straight off the reopened store — the assertion that fails when the
// wipe destroys the index and nothing rebuilds it).
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
await adapter.initLbug(lbugPath);
try {
const idxRows = (await adapter.executeQuery('CALL SHOW_INDEXES() RETURN *')) as Array<{
index_name?: string;
index_type?: string;
}>;
const idx = idxRows.find((r) => r.index_name === 'code_embedding_idx');
expect(idx).toBeDefined();
expect(idx!.index_type).toBe('HNSW');
} finally {
await adapter.closeLbug();
}
} finally {
await repo.cleanup();
}
}, 600_000);
});
/** Document-derived destinations currently in the graph, address + provenance. */
async function readDocumentDestinations(
repoPath: string,
): Promise<Array<{ address: string; broker: string; resolution: string }>> {
const adapter = await import('../../src/core/lbug/lbug-adapter.js');
const { lbugPath } = getStoragePaths(repoPath);
await adapter.initLbug(lbugPath);
try {
const rows = (await adapter.executeQuery(
`MATCH (d:Destination) WHERE d.resolution = 'asyncapi-document' ` +
`RETURN d.address AS address, d.broker AS broker, d.resolution AS resolution ` +
`ORDER BY address`,
)) as Array<{ address?: unknown; broker?: unknown; resolution?: unknown }>;
return rows.map((row) => ({
address: String(row.address ?? ''),
broker: String(row.broker ?? ''),
resolution: String(row.resolution ?? ''),
}));
} finally {
await adapter.closeLbug();
}
}
describe('runFullAnalysis — AsyncAPI document reading', () => {
/**
* Drives the REAL `runFullAnalysis` with `asyncApiSpecPath` set.
*
* Three separate things were individually deletable with the whole suite
* green before this existed: the forward from `run-analyze` into
* `PipelineOptions`, the forced rebuild while the option is enabled, and the
* cleanup rebuild when it is dropped. Each of them makes the feature
* partially or wholly inert, and none of them is visible one layer up, where
* the CLI test asserts on a mock's arguments.
*
* The document lives OUTSIDE the repository on purpose. That is the workflow
* the option is documented for — a cache written by other tooling — and it is
* the only one where the hazard is real: editing a tracked file dirties the
* tree and would force a rebuild anyway, so an in-repo fixture would pass
* even with the freshness fix reverted.
*/
it('forces a rebuild while enabled, re-reads a changed document, and cleans up once on disable', async () => {
const repo = await setupMiniRepo();
// A SIBLING of the repository, not a child: the document must be outside
// the working tree, or editing it would dirty the tree and force a rebuild
// on its own, and the test would pass with the freshness fix reverted.
const specDir = path.join(repo.dbPath, '..', 'gnx-asyncapi-spec');
await mkdir(specDir, { recursive: true });
const specFile = path.join(specDir, 'orders.yaml');
const documentFor = (address: string): string =>
[
'asyncapi: 3.0.0',
'info: { title: Order Service, version: 1.0.0 }',
'servers: { broker: { host: "example:9092", protocol: kafka } }',
`channels: { c: { address: ${address}, servers: [{ $ref: "#/servers/broker" }] } }`,
'operations: { publishOrder: { action: send, channel: { $ref: "#/channels/c" } } }',
'',
].join('\n');
try {
await writeFile(specFile, documentFor('orders.v1'), 'utf-8');
const { runFullAnalysis } = await import('../../src/core/run-analyze.js');
const { storagePath } = getStoragePaths(repo.dbPath);
const enabledLogs: string[] = [];
const enabled = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, asyncApiSpecPath: specDir },
{ onProgress: () => {}, onLog: (message) => enabledLogs.push(message) },
);
expect(enabled.alreadyUpToDate).toBeUndefined();
expect(enabledLogs.join('\n')).toContain(
'AsyncAPI document reading requested; forcing a full rebuild.',
);
// The forward into PipelineOptions is what puts this node in the graph;
// without it the flag parses and nothing else happens.
expect(await readDocumentDestinations(repo.dbPath)).toEqual([
{ address: 'orders.v1', broker: 'kafka', resolution: 'asyncapi-document' },
]);
expect((await loadMeta(storagePath))?.asyncApiSpec).toEqual({ enabled: true });
// Same commit, clean tree, only the out-of-tree document changed. Git
// freshness cannot see this, so without the forced rebuild the run is a
// no-op that reports success and serves the previous address.
await writeFile(specFile, documentFor('orders.v2'), 'utf-8');
const rereadRun = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true, asyncApiSpecPath: specDir },
{ onProgress: () => {} },
);
expect(rereadRun.alreadyUpToDate).toBeUndefined();
expect(await readDocumentDestinations(repo.dbPath)).toEqual([
{ address: 'orders.v2', broker: 'kafka', resolution: 'asyncapi-document' },
]);
const disableLogs: string[] = [];
const disabled = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {}, onLog: (message) => disableLogs.push(message) },
);
expect(disabled.alreadyUpToDate).toBeUndefined();
expect(disableLogs.join('\n')).toContain(
'AsyncAPI document reading disabled; rebuilding to remove document-derived evidence.',
);
expect(await readDocumentDestinations(repo.dbPath)).toEqual([]);
expect((await loadMeta(storagePath))?.asyncApiSpec).toBeUndefined();
// Exactly ONE cleanup rebuild: the metadata write drops the flag, so the
// next run has nothing to react to. A merge-over-previous write here
// would rebuild forever with nothing failing.
const steady = await runFullAnalysis(
repo.dbPath,
{ skipAgentsMd: true },
{ onProgress: () => {} },
);
expect(steady.alreadyUpToDate).toBe(true);
} finally {
// The document directory is a sibling of the repository, so the repo's
// own cleanup does not reach it — both owners have to be called here.
await rm(specDir, { recursive: true, force: true });
await repo.cleanup();
}
}, 180_000);
});